Cutting Through the Clutter: The Potential of LLMs for Efficient Filtration in Systematic Literature Reviews
Lucas Joos, Daniel A. Keim, Maximilian T. Fischer · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2312/eurova.20251105
Methodology & findings
Study design
Empirical evaluation study using case study methodology.
Sample
N = 8323, 6 groups
Primary method
Classification accuracy evaluation using standard information retrieval metrics (precision, recall, accuracy, F1 score). Comparison of individual LLM performance versus consensus voting schemes. Manual comparison against ground-truth human classification. No formal statistical hypothesis tests reported.
Main result
The study found that "LLMs performed well with an accuracy above 90% across all models" and that "consensus voting only one paper would be discarded that should be part of the survey (based on the human ground-truth data)." The consensus approach achieved "Consensus (Best) approach comes with a lower FP rate (see Figure 4), reducing the manual filtering by 695 papers, and only requires three instead of five LLMs, reducing time and cost." Most significantly, "this entire process was completed in under 10 minutes for just $28.81 (as of July 2024), demonstrating the scalability of LLMs" compared to "a minimum of 69 hours of concentrated human effort."
Reports effect sizes.
Research paradigm
Empirical-pragmatist (mixed-methods evaluation of AI system performance)
Author conclusions
The authors conclude: "Our method addresses the limitations of traditional keyword-based filtering, which often struggles with semantic ambiguities and inconsistent terminology, requiring time-consuming manual checks. With our opensource tool LLMSurver, users can iteratively test different prompts and LLMs while interactively evaluating the results." They state that findings "show that LLMs can drastically accelerate the review process, shrinking search space by an order of magnitude and reducing weeks of effort to minutes, while maintaining recall (> 98%), even below typical human error rates. This efficiency not only enhances SLR but also holds promise for broader academic applications. Overall, this study highlights the effective use of LLMs to streamline academic research."
Risk of bias
Selection bias: single corpus may not generalize to other research areas; Training data bias: LLMs inherit biases from training data and RLHF process; Prompt design bias: performance influenced by specific prompt formulation; Model bias: notable differences between open-source (conservative) and commercial models (exclusion-focused); Corpus characteristics bias: writing style and terminology could influence performance; Database selection bias: initial bibliographical entries and source database selection validity risk; Selection bias: Single corpus and research domain limits generalizability; Prompt design bias: Results may be specific to the particular prompt formulation used; Model training data bias: LLMs exhibit inherent biases from training data and RLHF process; Researcher bias: Manual ground-truth classification performed by multiple researchers over weeks may introduce inconsistency (authors note 34 papers were reclassified after initial human filtration); Model-specific bias: Open-source models showed bias towards inclusion (high FP rate), commercial models towards exclusion (high FN rate); Model bias: Open-source models (Llama3 8B) showed bias towards inclusion with high false positive rates; Training data bias: Inherent biases from LLM training data and RLHF process mentioned as potential source of skewed results; Single corpus limitation: Study based on one research domain may not generalize; Prompt design effects: Performance influenced by prompt formulation and contextual understanding; Selection bias in corpus: Careful selection of initial bibliographical entries and source databases noted as validity risk
Limitations
- The authors state: "Our study is based on a single large corpus and prompt, which may not generalize to other research areas
- Also, our tool has not yet been evaluated in a controlled user study." Additionally, "A validity risk, in particular when avoiding snowballing, is a careful selection of the initial set of bibliographical entries and source databases
- Other factors, such as prompt design, corpus characteristics, or writing style, could also influence performance." The authors also note that "these models can still produce misleading outputs, with performance influenced by model quality, prompt formulation, and contextual understanding."
Open questions raised
- Generalizability to other research areas beyond the single tested corpus
- Controlled user study evaluation of the LLMSurver tool
- Prompt engineering optimization strategies
- Few-shot learning approaches to enhance accuracy
- Interactive literature review platforms with collaborative LLM-human feedback mechanisms
- LLMs with search engine access for automated corpus retrieval
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations