LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation
Gregory Hok Tjoan Go, Khang Ly, Anders Søgaard, Seyed Amin Tabatabaei, Maarten de Rijke, Xinyi Chen · Proceedings of the AAAI Conference on Artificial Intelligence · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i47.41489
Methodology & findings
Study design
Comparative evaluation study using benchmark datasets (SciReviewGen with 125 sampled reviews; internal ScienceDirect dataset with 125 expert-written reviews across 23 subject areas).
Sample
N = 250, 2 groups
Primary method
Statistical tests applied for retrieval experiment comparison: "We examined if the results when using retrieval differed significantly compared to the baseline researcher setting, which was tested using the appropriate statistical tests." Metrics reported with mean ± standard deviation notation. Specific statistical test names and software not explicitly named in main text (referenced as 'extended version' for details).
Main result
LiRA achieves the highest ROUGE scores on SciReviewGen (0.13 ± 0.0) and ScienceDirect (0.13 ± 0.0), indicating stronger lexical alignment with human-written reviews. Most importantly, "LiRA demonstrates the largest gains in citation reliability, achieving the highest Citation Quality F1 (CQF1) scores across both datasets (0.76 on SciReviewGen, 0.73 on ScienceDirect) and substantially outperforming AutoSurvey (≤0.63) and all other baselines." The framework generates concise reviews averaging 22,000 tokens compared to AutoSurvey's 50,000 tokens while maintaining information density and expert-preferred structural coherence.
Reports effect sizes and confidence intervals.
Research paradigm
Empirical / Computational
Author conclusions
"From this, it can be seen that LiRA overall produces literature reviews that are concise, structurally coherent, and citation-faithful, while maintaining competitive coverage. This balance between quality and reliability highlights LiRA as a more trustworthy and practically useful framework for automated survey writing." The authors note that "the results obtained show that LiRA is capable of performing the task of automated literature review quite well, outperforming all tested open-source methods when accounting for the varying output lengths, indicating a positive result for essentially every research question proposed. Moreover, it reduces hallucination through improved citation behavior and can demonstrably be used in real-world settings."
Risk of bias
Self-bias amplification from using same LLM (gpt-4o-mini) across all agents; Potential verbosity bias in recall-based metrics favoring longer outputs; Limited domain generalization (computer science primary dataset); SME annotation bias: different annotation procedures across datasets; Selection bias in reference sampling (capped at 50 references for outline, 25% max per subsection); Self-bias amplification when using same LLM (gpt-4o-mini) for all agents in the pipeline; Potential evaluator bias in human expert evaluation across different datasets (different annotation procedures for SciReviewGen vs. AutoSurvey/ScienceDirect); Presentation order bias mitigated by randomization in SME evaluation for SciReviewGen; Length bias in recall-based metrics (AutoSurvey produces 50,000 tokens vs. LiRA's 22,000); Dataset selection bias: primary dataset limited to computer science; secondary dataset limited to 23 subject areas; Self-bias amplification potential when using the same LLM (gpt-4o-mini) for all agents in the pipeline; Potential annotation bias in human expert evaluation (order of samples was randomized to mitigate); Limited dataset diversity: evaluation primarily on computer science and cross-domain internal dataset; lacks broader scientific field representation; Metric bias toward length: recall-based metrics naturally favor longer outputs (AutoSurvey produces 50,000 tokens vs LiRA's 22,000); Irreproducibility risk due to non-seedable LLM results
Limitations
- "Several improvements could be made, mainly regarding the irreproducibility of results due to the usage of gpt-4o-mini for all experiments." Additionally, "there is a lack of open-source datasets for this task specifically, which hinders the generalizability of all results to other scientific fields." The authors further note that "the current project does not take into account factors such as primary studies and risk of bias in randomized trials (i.e., the implementation of automated tools based on Higgins et al
Open questions raised
- Irreproducibility of results due to gpt-4o-mini usage - authors recommend seedable models
- Lack of open-source datasets for literature review generation task specifically
- Limited generalizability to non-computer science domains
- Need for end-to-end pipelines incorporating primary studies screening and risk of bias assessment
- Integration of search criteria definition and screening steps for better paper reproducibility
- Implementation of automated tools for formal systematic review guidelines (Higgins et al. 2024)
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations