12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Large Language Models for Evidence-Based Planning: Evaluating an SLR-RAG Framework for Knowledge Synthesis of Urban Vacant Land

Xinyu Wang · Figshare · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
C
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.6084/m9.figshare.31144189

Methodology & findings

Study design

Systematic literature review (SLR) combined with computational evaluation using RAG-enhanced LLMs.

Primary method

Structured automatic evaluation metrics and expert rating assessment (qualitative analysis); specific statistical methods not detailed in abstract.

Main result

The study found that "RAG significantly improves the accuracy of LLMs under structured automatic evaluation." However, the authors note that "although retrieval augmentation provides models with access to domain-specific evidence, its benefits for open-ended planning questions are not consistently reflected in expert ratings, particularly for consistency and creativity."

Reports effect sizes.

Research paradigm

Pragmatist/mixed-methods (combining systematic review with computational evaluation)

Author conclusions

The authors conclude that "while text-only retrieval is insufficient for context-rich analysis, future advancements in spatially aware hybrid retrieval offer a promising pathway forward." They also state that the study "contributes to the field by elucidating the capabilities and limitations of LLMs and RAG in urban studies."

Risk of bias

Potential selection bias in the literature review if search strategy was not comprehensive or reproducible; Expert rater bias in qualitative evaluation of consistency and creativity; Mismatch between evaluation metrics (automated vs. expert ratings) may introduce measurement bias; Limited domain expertise representation unclear in expert panel composition; Expert rating subjectivity; Limited domain expertise representation in evaluation; Potential mismatch between automated and human evaluation criteria; Single case study (urban vacant land) limiting generalizability; Expert rater subjectivity in evaluating open-ended planning responses; Potential selection bias in the literature included in the SLR; Mismatch between evaluation metrics and actual planning domain requirements

Limitations

  • The authors identify several limitations: "the limited improvement can be attributed to the characteristics of the planning questions, the mismatch between textual information and the spatial data required for urban planning in current RAG pipelines, and potentially ineffective prompting that fails to elicit deeper reasoning." Additionally, "text-only retrieval is insufficient for context-rich analysis."

Open questions raised

  • The authors identify the need for spatially aware hybrid retrieval systems that integrate spatial data with textual evidence. They highlight that current RAG pipelines lack adequate spatial grounding for urban planning applications, and suggest that future advancements should address the mismatch between textual information and spatial data requirements.
  • Future research should address: (1) the integration of spatial data with text-based retrieval in RAG pipelines, (2) development of spatially aware hybrid retrieval systems for urban planning applications, (3) improved prompting strategies to elicit deeper reasoning from LLMs on planning questions, and (4) evaluation frameworks that better align automated metrics with expert domain assessments.
  • The paper identifies the need for spatially aware hybrid retrieval systems that integrate spatial data with textual information, improved prompting strategies to elicit deeper reasoning in planning contexts, and further research on applying LLMs and RAG to context-rich urban planning problems
Data: not_statedCode: not_statedExtracted from: pdfAgreement 67%

Explore related topics

Related papers