Large Language Models for Evidence-Based Planning: Evaluating an SLR-RAG Framework for Knowledge Synthesis of Urban Vacant Land
Xinyu Wang · Figshare · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.6084/m9.figshare.31144189
Methodology & findings
Study design
Systematic literature review (SLR) combined with computational evaluation using RAG-enhanced LLMs.
Primary method
Structured automatic evaluation metrics and expert rating assessment (qualitative analysis); specific statistical methods not detailed in abstract.
Main result
The study found that "RAG significantly improves the accuracy of LLMs under structured automatic evaluation." However, the authors note that "although retrieval augmentation provides models with access to domain-specific evidence, its benefits for open-ended planning questions are not consistently reflected in expert ratings, particularly for consistency and creativity."
Reports effect sizes.
Research paradigm
Pragmatist/mixed-methods (combining systematic review with computational evaluation)
Author conclusions
The authors conclude that "while text-only retrieval is insufficient for context-rich analysis, future advancements in spatially aware hybrid retrieval offer a promising pathway forward." They also state that the study "contributes to the field by elucidating the capabilities and limitations of LLMs and RAG in urban studies."
Risk of bias
Potential selection bias in the literature review if search strategy was not comprehensive or reproducible; Expert rater bias in qualitative evaluation of consistency and creativity; Mismatch between evaluation metrics (automated vs. expert ratings) may introduce measurement bias; Limited domain expertise representation unclear in expert panel composition; Expert rating subjectivity; Limited domain expertise representation in evaluation; Potential mismatch between automated and human evaluation criteria; Single case study (urban vacant land) limiting generalizability; Expert rater subjectivity in evaluating open-ended planning responses; Potential selection bias in the literature included in the SLR; Mismatch between evaluation metrics and actual planning domain requirements
Limitations
- The authors identify several limitations: "the limited improvement can be attributed to the characteristics of the planning questions, the mismatch between textual information and the spatial data required for urban planning in current RAG pipelines, and potentially ineffective prompting that fails to elicit deeper reasoning." Additionally, "text-only retrieval is insufficient for context-rich analysis."
Open questions raised
- The authors identify the need for spatially aware hybrid retrieval systems that integrate spatial data with textual evidence. They highlight that current RAG pipelines lack adequate spatial grounding for urban planning applications, and suggest that future advancements should address the mismatch between textual information and spatial data requirements.
- Future research should address: (1) the integration of spatial data with text-based retrieval in RAG pipelines, (2) development of spatially aware hybrid retrieval systems for urban planning applications, (3) improved prompting strategies to elicit deeper reasoning from LLMs on planning questions, and (4) evaluation frameworks that better align automated metrics with expert domain assessments.
- The paper identifies the need for spatially aware hybrid retrieval systems that integrate spatial data with textual information, improved prompting strategies to elicit deeper reasoning in planning contexts, and further research on applying LLMs and RAG to context-rich urban planning problems
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Exploring Students’ Perceptions of ChatGPT: Thematic Analysis and Follow-Up SurveyAbdulhadi Shoufan · 2023 · 464 citations