LR-Robot: An Human-in-the-Loop LLM Framework for Systematic Literature Reviews with Applications in Financial Research
Wei Wei, Jin Zheng, Zining Wang, Weibin Feng · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods design science approach combining: (1) unsupervised topic modeling baseline evaluation (BERTopic on 12,666 papers), (2) expert-designed taxonomy development with four classification dimensions, (3) systematic evaluation of eleven LLM models on manually annotated samples (1,000 papers for dimension 1; 417 papers for dimensions 2-4), with each model run three times to measure consistency, and (4) full-dataset application with label-enhanced citation network analysis.
Primary method
Design science with iterative human-in-the-loop refinement; expert-guided taxonomy design combined with LLM-based automation
Main result
The framework demonstrates that "expert-designed prompt constraints significantly enhance classification performance" with gains ranging from moderate to dramatic across eleven LLM models. With constraints, best-performing models achieve F1 scores above 0.81 and self-consistency above 0.94 on binary classification tasks. For multi-label classification on well-defined dimensions (asset types and option types), all five evaluated LLMs achieve Sample F1 above 0.82, with the error analysis revealing that "most misclassifications arise from genuinely ambiguous cases rather than systematic model deficiencies, suggesting that performance limits are driven by the inherent difficulty of the task."
Research paradigm
Pragmatist/Design Science
Author conclusions
"LR-Robot resolves this tension through a clear division of labor: domain experts define classification taxonomies and prompt constraints that encode conceptual boundaries, while LLMs execute scalable classification under systematic human evaluation. The framework thus occupies a methodological niche that neither manual review nor fully automated methods can fill alone." The authors further conclude that "By accelerating labor-intensive review stages while preserving interpretive accuracy, LR-Robot provides a practical, customizable, and high-quality approach for AI-assisted SLRs."
Risk of bias
Selection bias: Limited to English-language articles in Scopus database only; Sampling bias: Manual evaluation only on 1,000 randomly selected papers (417 for multi-label); unclear if representative of full corpus; Expert bias: Single domain expert labeling without inter-rater reliability reporting; Methodological bias: BERTopic comparison uses hyperparameter tuning but LLM evaluation uses fixed prompts without full tuning exploration; Generalization risk: Results demonstrated only on option pricing literature; unknown applicability to other domains; Selection bias: Dataset limited to English-language publications in Scopus only; Annotation bias: Manual labeling by domain experts without inter-rater reliability reported; Temporal bias: Dataset spans 1970s-2026 with heavily skewed distribution toward recent years (700+ papers in 2024-2025 vs. fewer than 50 in early 1990s); Corpus bias: Option pricing literature may not be representative of all financial research domains; Selection bias: Manual annotation of representative sample may not capture all edge cases in full corpus; Corpus bias: Limited to Scopus database and English-language publications; option pricing domain-specific findings may not generalize; Labeler bias: Domain experts performed all manual labeling; no inter-rater reliability reported; Publication bias: Search strategy limited to published articles; may not capture working papers or grey literature; Abstract-based limitation: Reliance on abstracts may bias against papers with vague or non-descriptive abstracts; Model selection bias: Results reflect available LLM APIs at time of study; newer models may perform differently
Limitations
- The current application "relies on abstracts, which provide an efficient and scalable foundation for large-scale classification
- However, abstracts may omit important methodological details and contextual nuances, which can limit the ability to distinguish between closely related research approaches." Additionally, "the generalisability of the expert-guided approach" remains uncertain, and "cross-disciplinary validation in other fields characterised by high terminological overlap would help establish the generalisability of the expert-guided approach."
Open questions raised
- Need to extend classification beyond abstracts to introduction sections or full-text for improved accuracy on complex dimensions
- Cross-disciplinary validation needed in other fields with high terminological overlap to establish generalizability
- Integration with dynamic knowledge graphs could support identification of emerging research frontiers and predictive tracking of thematic shifts
- Investigation of performance on emerging or underrepresented research paradigms
- Extension to full-text analysis beyond abstracts to capture methodological details and contextual nuances for improved classification accuracy on complex dimensions
- Cross-disciplinary validation in other fields with high terminological overlap to establish generalisability of expert-guided approach
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations