Leveraging LLMs for semi-automatic corpus filtration in systematic literature reviews
Lucas Joos, Daniel A. Keim, Maximilian T. Fischer · Computers & Graphics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.cag.2026.104537
Methodology & findings
Study design
Computational pipeline evaluation using ground-truth data from a real systematic literature review.
Main result
Results demonstrate that our pipeline significantly reduces manual effort while achieving lower error rates than single human annotators. Furthermore, modern open-source models prove sufficient for this task, making the method accessible and cost-effective. The consensus approach combining GPT-5, Claude Sonnet 4.5, and Llama 3.3 (70B) achieved "only 166 false positives and no false negatives," substantially improving upon single-model results and reducing manual review workload from weeks to minutes.
Research paradigm
Pragmatist/Empiricist - testing LLM-based automation against ground truth data
Author conclusions
"This work presents a semi-automated pipeline that leverages large language models to accelerate and enhance literature filtering for systematic literature reviews. By combining multiple LLMs in a consensus scheme and integrating human supervision through our open-source tool LLMSurver, researchers can efficiently reduce large corpora while maintaining high recall and transparency." The authors emphasize that "the rapid evolution of open models over the past year highlights how accessible, cost-effective, and privacy-preserving open AI tools can now match proprietary systems in quality for this task," and stress that "LLM-assisted and consensus-based workflows controlled through human-AI collaboration can streamline and facilitate academic work."
Risk of bias
Selection bias: Ground-truth data comes from a single recent survey (2025), limiting generalizability to other domains; Domain specificity: Evaluation dataset is specific to visual network analysis in immersive environments; Model selection bias: Choice of which LLMs to test may not represent full state-of-the-art; Prompt design bias: Prompt variations tested only on one model (Llama 3.1 8B), limiting insights on prompt optimization; Consensus voting bias: Conservative consensus scheme (inclusion if any model votes yes) may not optimize precision for all use cases; Domain-specific bias: Evaluation limited to single research topic (visual network analysis in immersive environments); Ground-truth validation bias: Only one domain's manually-labeled data used; generalization to other research areas not empirically tested; Model selection bias: Limited to models available in mid-2024 and fall 2025; earlier or alternative LLM architectures not evaluated; Prompt engineering bias: Prompts iteratively refined based on performance feedback, potentially introducing optimization bias toward specific models; Consensus scheme bias: Conservative strategy (include if any model recommends inclusion) may systematically favor higher recall over precision; Selection bias: Evaluation limited to single research domain (Visual Network Analysis in Immersive Environments), may not generalize to other fields; Model bias: Different LLM architectures show systematic differences in inclusion/exclusion patterns (e.g., open models more conservative with inclusions); Ground truth bias: Human-generated ground truth may contain inconsistencies, as authors note one ambiguous paper was disputed even by human evaluators; Domain-specific bias: Keywords and search strategy specific to computer science may not transfer to other domains
Limitations
- The authors state that "we did not conduct a formal user study of LLMSurver
- Such an evaluation could provide valuable insights into usability and guide further development." Additionally, the consensus scheme's conservative inclusion strategy ("includes a paper if any of the participating LLMs recommends its inclusion") may still result in false exclusions despite the multi-model approach, though such papers "can typically be recovered in a subsequent snowballing step." The evaluation was limited to a single domain (visual network analysis in immersive environments), and the corpus, while large (8,323 papers), may not represent all research areas with similar characteristics.
Open questions raised
- Formal user study of LLMSurver tool for usability evaluation
- Integration of automatic access to online libraries for corpus retrieval
- Adaptive consensus methods that adjust to model confidence rather than fixed voting thresholds
- Expansion of LLMSurver into collaborative platform for multiple reviewers
- Adaptation of similar pipelines for related academic tasks (content screening, snowballing)
- Formal user study of LLMSurver usability and adoption patterns
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations