12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation

Gregory Hok Tjoan Go, Khang Ly, Anders Søgaard, Seyed Amin Tabatabaei, Maarten de Rijke, Xinyi Chen · Proceedings of the AAAI Conference on Artificial Intelligence · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
1
Citations
15.05
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i47.41489

Methodology & findings

Study design

Comparative evaluation study using benchmark datasets (SciReviewGen with 125 sampled reviews; internal ScienceDirect dataset with 125 expert-written reviews across 23 subject areas).

Sample

N = 250, 2 groups

Primary method

Statistical tests applied for retrieval experiment comparison: "We examined if the results when using retrieval differed significantly compared to the baseline researcher setting, which was tested using the appropriate statistical tests." Metrics reported with mean ± standard deviation notation. Specific statistical test names and software not explicitly named in main text (referenced as 'extended version' for details).

Main result

LiRA achieves the highest ROUGE scores on SciReviewGen (0.13 ± 0.0) and ScienceDirect (0.13 ± 0.0), indicating stronger lexical alignment with human-written reviews. Most importantly, "LiRA demonstrates the largest gains in citation reliability, achieving the highest Citation Quality F1 (CQF1) scores across both datasets (0.76 on SciReviewGen, 0.73 on ScienceDirect) and substantially outperforming AutoSurvey (≤0.63) and all other baselines." The framework generates concise reviews averaging 22,000 tokens compared to AutoSurvey's 50,000 tokens while maintaining information density and expert-preferred structural coherence.

Reports effect sizes and confidence intervals.

Research paradigm

Empirical / Computational

Author conclusions

"From this, it can be seen that LiRA overall produces literature reviews that are concise, structurally coherent, and citation-faithful, while maintaining competitive coverage. This balance between quality and reliability highlights LiRA as a more trustworthy and practically useful framework for automated survey writing." The authors note that "the results obtained show that LiRA is capable of performing the task of automated literature review quite well, outperforming all tested open-source methods when accounting for the varying output lengths, indicating a positive result for essentially every research question proposed. Moreover, it reduces hallucination through improved citation behavior and can demonstrably be used in real-world settings."

Risk of bias

Self-bias amplification from using same LLM (gpt-4o-mini) across all agents; Potential verbosity bias in recall-based metrics favoring longer outputs; Limited domain generalization (computer science primary dataset); SME annotation bias: different annotation procedures across datasets; Selection bias in reference sampling (capped at 50 references for outline, 25% max per subsection); Self-bias amplification when using same LLM (gpt-4o-mini) for all agents in the pipeline; Potential evaluator bias in human expert evaluation across different datasets (different annotation procedures for SciReviewGen vs. AutoSurvey/ScienceDirect); Presentation order bias mitigated by randomization in SME evaluation for SciReviewGen; Length bias in recall-based metrics (AutoSurvey produces 50,000 tokens vs. LiRA's 22,000); Dataset selection bias: primary dataset limited to computer science; secondary dataset limited to 23 subject areas; Self-bias amplification potential when using the same LLM (gpt-4o-mini) for all agents in the pipeline; Potential annotation bias in human expert evaluation (order of samples was randomized to mitigate); Limited dataset diversity: evaluation primarily on computer science and cross-domain internal dataset; lacks broader scientific field representation; Metric bias toward length: recall-based metrics naturally favor longer outputs (AutoSurvey produces 50,000 tokens vs LiRA's 22,000); Irreproducibility risk due to non-seedable LLM results

Limitations

  • "Several improvements could be made, mainly regarding the irreproducibility of results due to the usage of gpt-4o-mini for all experiments." Additionally, "there is a lack of open-source datasets for this task specifically, which hinders the generalizability of all results to other scientific fields." The authors further note that "the current project does not take into account factors such as primary studies and risk of bias in randomized trials (i.e., the implementation of automated tools based on Higgins et al

Open questions raised

  • Irreproducibility of results due to gpt-4o-mini usage - authors recommend seedable models
  • Lack of open-source datasets for literature review generation task specifically
  • Limited generalizability to non-computer science domains
  • Need for end-to-end pipelines incorporating primary studies screening and risk of bias assessment
  • Integration of search criteria definition and screening steps for better paper reproducibility
  • Implementation of automated tools for formal systematic review guidelines (Higgins et al. 2024)
Data: SciReviewGen (Kasanishi et al. 2023) - publicly available benchmark from Semantic Scholar Open Research Corpus (S2ORC), 10,000 review articles in computer science (125 sampled); Internal ScienceDirect dataset - 125 expert-written reviews covering 23 subject areas (business, microbiology, materials science, etc.); SciReviewGen; ScienceDirect internal dataset; SciReviewGen: publicly available benchmark from Semantic Scholar Open Research Corpus (S2ORC), containing 10,000 review articles in computer science; 125 reviews randomly sampled for evaluation; Internal ScienceDirect dataset: 125 expert-written reviews covering 23 subject areas (business, microbiology, materials science, etc.)Code: LangGraph - open-source and production-ready agentic framework used for implementation; AutoSurvey (Wang et al. 2024) - modified for fair comparison with retrieval restricted to target review references; MASS-Survey (Tian et al. 2025) - baseline from automated survey-writing challenge; LiRA; AutoSurvey; MASS-Survey; LangGraph: open-source agentic framework used for implementing all agents (Python version mentioned for deployment)Extracted from: pdfAgreement 53%

Explore related topics

Related papers