12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Leveraging LLMs for semi-automatic corpus filtration in systematic literature reviews

Lucas Joos, Daniel A. Keim, Maximilian T. Fischer · Computers & Graphics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
3/4
Quality (LMQS)
C
Evidence
3
Citations
34.97
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.cag.2026.104537

Methodology & findings

Study design

Computational pipeline evaluation using ground-truth data from a real systematic literature review.

Main result

Results demonstrate that our pipeline significantly reduces manual effort while achieving lower error rates than single human annotators. Furthermore, modern open-source models prove sufficient for this task, making the method accessible and cost-effective. The consensus approach combining GPT-5, Claude Sonnet 4.5, and Llama 3.3 (70B) achieved "only 166 false positives and no false negatives," substantially improving upon single-model results and reducing manual review workload from weeks to minutes.

Research paradigm

Pragmatist/Empiricist - testing LLM-based automation against ground truth data

Author conclusions

"This work presents a semi-automated pipeline that leverages large language models to accelerate and enhance literature filtering for systematic literature reviews. By combining multiple LLMs in a consensus scheme and integrating human supervision through our open-source tool LLMSurver, researchers can efficiently reduce large corpora while maintaining high recall and transparency." The authors emphasize that "the rapid evolution of open models over the past year highlights how accessible, cost-effective, and privacy-preserving open AI tools can now match proprietary systems in quality for this task," and stress that "LLM-assisted and consensus-based workflows controlled through human-AI collaboration can streamline and facilitate academic work."

Risk of bias

Selection bias: Ground-truth data comes from a single recent survey (2025), limiting generalizability to other domains; Domain specificity: Evaluation dataset is specific to visual network analysis in immersive environments; Model selection bias: Choice of which LLMs to test may not represent full state-of-the-art; Prompt design bias: Prompt variations tested only on one model (Llama 3.1 8B), limiting insights on prompt optimization; Consensus voting bias: Conservative consensus scheme (inclusion if any model votes yes) may not optimize precision for all use cases; Domain-specific bias: Evaluation limited to single research topic (visual network analysis in immersive environments); Ground-truth validation bias: Only one domain's manually-labeled data used; generalization to other research areas not empirically tested; Model selection bias: Limited to models available in mid-2024 and fall 2025; earlier or alternative LLM architectures not evaluated; Prompt engineering bias: Prompts iteratively refined based on performance feedback, potentially introducing optimization bias toward specific models; Consensus scheme bias: Conservative strategy (include if any model recommends inclusion) may systematically favor higher recall over precision; Selection bias: Evaluation limited to single research domain (Visual Network Analysis in Immersive Environments), may not generalize to other fields; Model bias: Different LLM architectures show systematic differences in inclusion/exclusion patterns (e.g., open models more conservative with inclusions); Ground truth bias: Human-generated ground truth may contain inconsistencies, as authors note one ambiguous paper was disputed even by human evaluators; Domain-specific bias: Keywords and search strategy specific to computer science may not transfer to other domains

Limitations

  • The authors state that "we did not conduct a formal user study of LLMSurver
  • Such an evaluation could provide valuable insights into usability and guide further development." Additionally, the consensus scheme's conservative inclusion strategy ("includes a paper if any of the participating LLMs recommends its inclusion") may still result in false exclusions despite the multi-model approach, though such papers "can typically be recovered in a subsequent snowballing step." The evaluation was limited to a single domain (visual network analysis in immersive environments), and the corpus, while large (8,323 papers), may not represent all research areas with similar characteristics.

Open questions raised

  • Formal user study of LLMSurver tool for usability evaluation
  • Integration of automatic access to online libraries for corpus retrieval
  • Adaptive consensus methods that adjust to model confidence rather than fixed voting thresholds
  • Expansion of LLMSurver into collaborative platform for multiple reviewers
  • Adaptation of similar pipelines for related academic tasks (content screening, snowballing)
  • Formal user study of LLMSurver usability and adoption patterns
Data: Ground-truth dataset from Visual Network Analysis in Immersive Environments survey (8,323 papers with human classification labels) - referenced but not explicitly stated as publicly available; Ground-truth dataset: 8,323 papers from 'Visual Network Analysis in Immersive Environments: A Survey' (available via https://doi.org/10.48550/arXiv.2501.08500); Papers manually classified: 88 relevant, 8,235 irrelevant; Ground-truth dataset from Visual Network Analysis survey (8,323 papers with human classifications) - availability status not explicitly stated but used in evaluationCode: https://github.com/dbvis-ukon/LLMSurver (MIT License); LLMSurver web application: https://llmsurver.dbvis.de; LLMSurver application: https://github.com/dbvis-ukon/LLMSurver (MIT License); LLMSurver web demo: https://llmsurver.dbvis.de; https://github.com/dbvis-ukon/LLMSurver (MIT License, open-source implementation of LLMSurver); Web deployment: https://llmsurver.dbvis.de (freely available deployed version)Extracted from: pdfAgreement 36%

Explore related topics

Related papers