Einsatz KI-gestützter Systeme für Literaturreviews Explorative Analyse und kritische Reflexion
Michael Fellmann, Henrik Bongertmann, Niklas Götz, Carl Pommerencke · Ars digitalis · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/978-3-658-45839-3_7
Methodology & findings
Study design
Exploratory practical case study analysis combining hermeneutic literature analysis framework with hands-on testing of multiple AI-assisted tools (BingChat, Perplexity.ai, Phind, ConnectedPapers, ResearchRabbit, Iris.ai, elicit.org, ChatGPT, HeyGPT).
Primary method
Qualitative comparative analysis. Tools were evaluated against Scopus database results (62 results for search phase). Domain experts compared tool-generated summaries with reference summaries. No formal statistical testing reported.
Main result
The study found that "die automatische Extraktion ein zwiespältiges Bild hinterlässt" (automatic extraction leaves a mixed picture), with significant problems in that "nicht alle wichtigen Aussagen in dem zugrundeliegenden Text gefunden wurde" (not all important statements in the underlying text were found) and "inhaltliche Verfälschungen auftreten" (content distortions occur). Additionally, the research revealed that "KI-gestützte Werkzeuge gegenwärtig somit darin liegen, auf Basis einer einfachen, natürlichsprachlichen Anfrage einige wenige, relevante Forschungsarbeiten zu identifizieren" (the strengths of AI-supported tools currently lie in identifying a few highly relevant research works on the basis of a simple, natural language query).
Reports effect sizes.
Research paradigm
Interpretivist/Critical hermeneutic
Author conclusions
The central recommendation states that "der Einsatz solcher Werkzeuge bzw. deren erzeugte Ausgaben immer kritisch hinterfragt und die Ergebnisse überprüft werden müssen" (the use of such tools and their generated outputs must always be critically questioned and results verified). The authors propose a hybrid three-phase approach: an experimentation phase with LLM-based tools to identify keywords, initial search in AI-supported tools to identify highly relevant papers, and expansion using traditional tools that produce reproducible results. They conclude that "mit Hilfe KI-gestützter Tools können zwar Zusammenhänge in großen Datenmengen aufgespürt werden und für eine Frage relevante Fakten extrahiert werden. Dennoch bleibt das menschliche Urteilsvermögen am Ende entscheidend" (AI-supported tools can detect patterns in large datasets and extract relevant facts, but human judgment remains decisive at the end).
Risk of bias
Possible bias in AI system data foundations unclear; LLMs may amplify biases present in training data; Hallucinations and false statement generation in LLM-based tools; Non-deterministic character of AI tools producing inconsistent results; Tools may misattribute or fabricate sources; Prompt dependency—minor modifications can significantly alter results; Non-deterministic nature of LLM-based tools leading to non-reproducible results; Hallucinations and fabrication of sources in AI tools; Tool design bias and underlying dataset bias not transparent; Prompt dependency - minor modifications can significantly alter results; Index size limitations compared to traditional databases; Potential for amplification of existing biases in training data; Homogenization effects reducing diverse perspectives; Non-deterministic behavior of LLM-based tools producing non-reproducible results; Hallucinations in language models producing fabricated information; Undisclosed training data bias in AI systems; Prompt-dependency leading to inconsistent search results; Index size limitations in AI tools relative to traditional databases; Gender and occupational bias amplification in LLMs
Limitations
- The authors state that "die Extraktion kann damit nur als Grundgerüst oder Inspiration dienen, jeglicher Inhalt muss noch einmal manuell geprüft werden" (extraction can only serve as a framework or inspiration, all content must be manually reviewed again), and note that "die gleiche Eingabe beziehungsweise Anfrage mehrmals hintereinander unterschiedliche Ergebnisse liefern und nicht immer hinreichend reproduziert werden" (the same input or query can produce different results multiple times in succession and cannot always be sufficiently reproduced)
- Additionally, "die Indexgröße..
- ist bei KI-Tools in der Regel noch geringer als die der traditionellen Datenbanken, insbesondere GoogleScholar, sodass nicht alle relevanten Arbeiten gefunden werden können" (the index size of AI tools is usually still smaller than that of traditional databases, particularly Google Scholar, so not all relevant works can be found).
Open questions raised
- Need for improved maturity level of AI tools for independent use in literature reviews
- Lack of transparency and reproducibility in AI-supported search tools
- Gap between speed of tool development and assessment of actual utility
- Lack of technology assessment structures and resources in academia
- Need for didactic guidance on appropriate contexts for high-automation review creation
- Potential for future research to develop AI tools that pose critical questions to authors and support reflection processes
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations