12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LR-Robot: An Human-in-the-Loop LLM Framework for Systematic Literature Reviews with Applications in Financial Research

Wei Wei, Jin Zheng, Zining Wang, Weibin Feng · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods design science approach combining: (1) unsupervised topic modeling baseline evaluation (BERTopic on 12,666 papers), (2) expert-designed taxonomy development with four classification dimensions, (3) systematic evaluation of eleven LLM models on manually annotated samples (1,000 papers for dimension 1; 417 papers for dimensions 2-4), with each model run three times to measure consistency, and (4) full-dataset application with label-enhanced citation network analysis.

Primary method

Design science with iterative human-in-the-loop refinement; expert-guided taxonomy design combined with LLM-based automation

Main result

The framework demonstrates that "expert-designed prompt constraints significantly enhance classification performance" with gains ranging from moderate to dramatic across eleven LLM models. With constraints, best-performing models achieve F1 scores above 0.81 and self-consistency above 0.94 on binary classification tasks. For multi-label classification on well-defined dimensions (asset types and option types), all five evaluated LLMs achieve Sample F1 above 0.82, with the error analysis revealing that "most misclassifications arise from genuinely ambiguous cases rather than systematic model deficiencies, suggesting that performance limits are driven by the inherent difficulty of the task."

Research paradigm

Pragmatist/Design Science

Author conclusions

"LR-Robot resolves this tension through a clear division of labor: domain experts define classification taxonomies and prompt constraints that encode conceptual boundaries, while LLMs execute scalable classification under systematic human evaluation. The framework thus occupies a methodological niche that neither manual review nor fully automated methods can fill alone." The authors further conclude that "By accelerating labor-intensive review stages while preserving interpretive accuracy, LR-Robot provides a practical, customizable, and high-quality approach for AI-assisted SLRs."

Risk of bias

Selection bias: Limited to English-language articles in Scopus database only; Sampling bias: Manual evaluation only on 1,000 randomly selected papers (417 for multi-label); unclear if representative of full corpus; Expert bias: Single domain expert labeling without inter-rater reliability reporting; Methodological bias: BERTopic comparison uses hyperparameter tuning but LLM evaluation uses fixed prompts without full tuning exploration; Generalization risk: Results demonstrated only on option pricing literature; unknown applicability to other domains; Selection bias: Dataset limited to English-language publications in Scopus only; Annotation bias: Manual labeling by domain experts without inter-rater reliability reported; Temporal bias: Dataset spans 1970s-2026 with heavily skewed distribution toward recent years (700+ papers in 2024-2025 vs. fewer than 50 in early 1990s); Corpus bias: Option pricing literature may not be representative of all financial research domains; Selection bias: Manual annotation of representative sample may not capture all edge cases in full corpus; Corpus bias: Limited to Scopus database and English-language publications; option pricing domain-specific findings may not generalize; Labeler bias: Domain experts performed all manual labeling; no inter-rater reliability reported; Publication bias: Search strategy limited to published articles; may not capture working papers or grey literature; Abstract-based limitation: Reliance on abstracts may bias against papers with vague or non-descriptive abstracts; Model selection bias: Results reflect available LLM APIs at time of study; newer models may perform differently

Limitations

  • The current application "relies on abstracts, which provide an efficient and scalable foundation for large-scale classification
  • However, abstracts may omit important methodological details and contextual nuances, which can limit the ability to distinguish between closely related research approaches." Additionally, "the generalisability of the expert-guided approach" remains uncertain, and "cross-disciplinary validation in other fields characterised by high terminological overlap would help establish the generalisability of the expert-guided approach."

Open questions raised

  • Need to extend classification beyond abstracts to introduction sections or full-text for improved accuracy on complex dimensions
  • Cross-disciplinary validation needed in other fields with high terminological overlap to establish generalizability
  • Integration with dynamic knowledge graphs could support identification of emerging research frontiers and predictive tracking of thematic shifts
  • Investigation of performance on emerging or underrepresented research paradigms
  • Extension to full-text analysis beyond abstracts to capture methodological details and contextual nuances for improved classification accuracy on complex dimensions
  • Cross-disciplinary validation in other fields with high terminological overlap to establish generalisability of expert-guided approach
Data: Data will be available upon request (stated as "Data and code will be available upon request" but no permanent repository or DOI provided); Data and code will be available upon request (not currently publicly available); Data will be available upon request (as stated: "Data and code will be available upon request")Extracted from: pdfAgreement 58%

Explore related topics

Related papers