LR-Robot: A Unified Supervised Intelligent Framework for Real-Time Systematic Literature Reviews with Large Language Models
Wei Wei, Jin Zheng, Zining Wang · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods case study combining machine learning classification with human-in-the-loop validation.
Primary method
Design science research with human-in-the-loop supervised learning approach
Main result
The framework demonstrates that "incorporating human-in-the-loop guidance enables AI systems to better understand domain-specific academic concepts and make more accurate and consistent classification decisions." In the option pricing application, Gemini Flash 2.0 achieved an average accuracy of 0.8327 and F1 score of 0.8152 with human-in-the-loop constraints, compared to 0.7281 and 0.7419 without constraints. The analysis identified 5,942 of 11,916 papers (49.86%) as focusing on pricing or volatility model development and comparison, revealing that "Analytical Models and Numerical Methods overwhelmingly dominate the literature, highlighting the enduring influence of classical mathematical and computational approaches that form the methodological core of option pricing research."
Research paradigm
positivist/computational
Author conclusions
"LR-Robot provides a practical, flexible, and high-quality approach for AI-assisted systematic reviews, bridging the gap between efficiency and methodological rigor." The authors conclude that "by integrating structured knowledge and human-in-the-loop evaluation, the framework can support multidimensional analyses, uncover key contributions, and facilitate historical and thematic mapping in a wide range of disciplines." They further state the framework "enables end-users to automatically generate customized literature syntheses and targeted document selection, producing their own systematic literature reviews tailored to specific research needs."
Risk of bias
Selection bias: Sample of 1,000 papers for manual labeling may not be representative of the full 11,916-paper corpus; Labeling bias: Manual annotations performed by researchers with domain expertise may reflect specific perspectives or interpretations; Model selection bias: Choice of best-performing model (Gemini Flash 2.0) based on sample performance may not generalize to full dataset; Data quality bias: 34% of initial records (15,424 to 11,916) removed due to missing abstracts or incomplete metadata; Parameter sensitivity: Classification results depend on carefully crafted prompts that required iterative refinement; Language bias: Dataset restricted to English-language publications only; Selection bias: Dataset limited to Scopus database and English-language publications only; Language bias: Exclusion of non-English research; Model-selection bias: Choice of Gemini Flash 2.0 based on specific evaluation metrics may not optimize for other downstream tasks; Incomplete data: 11,916 valid samples retained from initial 15,424 due to missing abstracts; Annotation bias: Manual labeling by unspecified number of human annotators on sample dataset; Selection bias in data collection: Initial search yielded 15,424 articles but 3,508 were excluded due to missing/incomplete metadata; Algorithmic bias: Different LLM models show varying performance; reliance on Gemini Flash 2.0 may not generalize across domains; Annotation bias: Manual labeling of sample dataset (n=1,000) by researchers may introduce subjective classification errors; Language bias: Dataset limited to English-language publications only; Database bias: Single source (Scopus) used for literature collection, potentially missing non-indexed publications
Limitations
- The authors note that "while the model demonstrates excellent internal consistency and strong recall performance, there remains a moderate precision gap, reflecting the inherent complexity and overlap among option pricing model categories." Additionally, for the fourth dimension classification, "the mean Jaccard similarity was 0.5545, suggesting a moderate level of agreement between AI-generated and human-labeled classifications but difficult to perfectly match the human annotations." The framework's applicability is demonstrated only through a single domain case study (option pricing), and the authors acknowledge that "Future research could extend the application of LR-Robot through longitudinal and cross-disciplinary studies to evaluate its generalizability."
Open questions raised
- The authors identify several future research directions: (1) longitudinal and cross-disciplinary studies to evaluate generalizability and refine best practices for human-AI collaboration; (2) expanding the framework to multiple research domains to explore interconnections between disciplines and evolution of scientific knowledge across fields; (3) integrating LR-Robot with dynamic knowledge graphs or citation-based forecasting models to enable identification of emerging research frontiers and prediction of future thematic shifts.
- Future research should include: (1) longitudinal and cross-disciplinary studies to evaluate generalizability and refine best practices for human-AI collaboration; (2) expansion to multiple research domains to explore interconnections between disciplines and evolution of scientific knowledge; (3) integration of LR-Robot with dynamic knowledge graphs or citation-based forecasting models to identify emerging research frontiers and predict future thematic shifts.
- Future research could extend LR-Robot through: (1) longitudinal and cross-disciplinary studies to evaluate generalizability and refine human-AI collaboration practices, (2) expansion to multiple research domains to explore interconnections between disciplines and evolution of scientific knowledge, and (3) integration with dynamic knowledge graphs or citation-based forecasting models to identify emerging research frontiers and predict future thematic shifts.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations