12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Writing literature reviews with AI: principles, hurdles and some lessons learned

Saadi Lahlou, Annabelle Gouttebroze, Atrina Oraee, Julian Madera · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Qualitative comparative case study.

Sample

N = 280, 7 groups

Primary method

The study reports qualitative comparison methods rather than inferential statistics. Descriptive statistics are mentioned: "only 20% overlap between the paper selections by humans and the LLM." The paper mentions inter-rater reliability assessment through comparison of human versus LLM ratings (Appendix 4 references comparison of ratings by humans and LLM) but does not report formal statistical tests (e.g., Cohen's kappa) for agreement. The analysis is primarily qualitative, comparing the content, framing, and gaps across six review versions through close reading and thematic analysis.

Main result

The study found that "The same LLM, with same corpus of 280 papers but different selections produced dramatically different reviews -from mainstream and politically neutral to critical and post-colonial- though neither orientation was intended." The authors identified five main pitfalls: "1. The bias of ignorance (you don't know what you don't get) in the selection of relevant papers. 2. Alignment and digital sycophancy: commercial AI models slavishly take you further in the direction they understand you give. This reinforces biases. 3. Mainstreaming. Because of their statistical nature, LLM production will tend to favor mainstream perspective and content. As a result, there was only 20% overlap between the paper selections by humans and the LLM."

Reports effect sizes.

Research paradigm

Critical interpretive; qualitative comparison with epistemological reflexivity

Author conclusions

"Overall, AI can improve the span and quality of the review, but the gain of time is not as massive as one would expect, and 'press button' strategy leaving the AI to do the work is recipe for disaster." The authors conclude that "literature reviews, far from being neutral descriptions of the 'state of the art', are in fact deeply influenced by how the papers are selected, interpreted and presented. Using AI may shift the perspective in one or another direction, without the user being aware." They emphasize that "a proper use of these tools requires the user to have domain knowledge, and to do so, to actually read enough papers to be able to detect the biases and correct them." Most importantly: "The review itself becomes an act of scholarship rather than merely preparation for scholarship" and literature reviews can and should intervene in debates, challenge orthodoxies, and advance theoretical understanding.

Risk of bias

Selection bias in AI-selected papers (only 20% overlap with human selection); Bias of ignorance: missing relevant papers without researcher awareness; Mainstreaming bias: LLM favors mainstream perspectives due to statistical nature; Paywall bias: AI excludes papers behind paywalls, books, and older seminal works; Recency bias: exclusive reliance on online accessible papers biases toward recent literature; Alignment bias: commercial LLMs optimized to please user, reinforcing existing biases; Token window limitations: LLMs omitted papers from consideration without explanation; Political correctness bias: LLMs avoid polarizing conclusions and critical perspectives; Training data bias: LLM training dominated by mainstream academic perspectives; Ignorance bias - papers not included in consideration set are not evaluated, researcher remains unaware of what was excluded; Selection bias - human researchers' domain expertise influenced paper selection, potentially different from AI selection; Mainstreaming bias - LLM statistical nature favors mainstream perspectives and content; Alignment bias - commercial AI models trained to please users, reinforcing user biases; Digital sycophancy - LLM provides flattering rather than critical feedback; Availability bias - AI can only access online, openly available papers, excluding paywalled content, books, and older seminal works; Token window bias - LLM omitted 11 of specified papers likely due to token limitations; Confirmation bias - LLM alignment leads to users being fed-forward their own views; Researcher expertise bias - team had solid knowledge of food systems domain, potentially biasing what they considered relevant; Language and format bias in information retrieval - papers with different language/format availability; Recency bias - exclusion of older literature in favor of recent accessible papers; Selection bias: AI tends to select mainstream papers; only 20% overlap between human and LLM selections; Ignorance bias: Papers may be omitted from consideration set without researcher awareness; authors found 11 papers were entirely omitted by LLM in Review B; Mainstreaming bias: LLM production favors mainstream perspectives and content due to statistical nature of training; Paywall bias: AI relies on accessible online papers, excluding books, chapters, and older literature typically behind paywalls; Language/format bias: Papers tagged differently on the web may be missed; non-OCR'd papers lose information; Alignment bias: Commercial models are trained to please users and avoid controversy, reinforcing user biases; Recency bias: Tendency to favor recent literature over seminal older works; Political correctness bias: Models avoid polarizing conclusions and controversial perspectives; Confirmation bias: Continued user interaction with LLM in self-constructed bubble reinforces user views; Opacity in IR: Researchers have little insight into how LLM selects and rates papers

Limitations

  • The authors note that "Most pitfalls can be addressed by prompting, but only if the user knows the domain well enough to detect them
  • There is a paradox: producing a good AI-assisted review requires expertise that comes from reading the literature, which is precisely what AI was meant to reduce." They also acknowledge that the "bias of ignorance (you don't know what you don't get) in the selection of relevant papers" creates a "completely silent bias: you have no cognition of this ignorance." Additionally, "the review omitted eleven of the specified papers entirely
  • Rectifying this oversight required Annabelle to invest eight hours identifying gaps and strategically incorporating the missing literature—time that exceeded what would typically be required to draft a literature review from pre-existing notes." The study is limited to one knowledge domain (food system transitions) and one team's expertise.

Open questions raised

  • The paper identifies that existing AI literature review studies "do not explore in detail how using AI changes the content of the literature reviews themselves: are they better? Biased? How?" The authors found that "A fundamental challenge in assessing the quality of AI reviews is the absence of standard evaluation frameworks and established benchmarks, making it difficult to assess reliability across different research contexts."
  • The authors identify that existing quantitative studies of AI-augmented reviews show that "AI models enable surveying a larger number of papers, they are efficient but human supervision remains necessary" but fail to explore "how using AI changes the content of the literature reviews themselves: are they better? Biased? How?" They note the absence of "standard evaluation frameworks and established benchmarks" and identify unexplored areas including: how AI affects literature prioritization compared to humans, whether AI systems identify papers with equivalent relevance to research questions, comparative time investments when accounting for validation, and how AI use affects final review quality and comprehensiveness. They recommend further research on the mechanisms through which AI shapes review content and the conditions under which these mechanisms operate.
  • The authors identify that existing studies on AI-augmented literature reviews have focused primarily on efficiency metrics (time savings, screening accuracy) but have not adequately addressed fundamental questions about literature review quality. They note that "Key unexplored areas include: how AI affects literature prioritization decisions compared to human researchers; whether AI systems identify papers with equivalent relevance to research questions; the comparative time investments required for AI versus human approaches when accounting for validation and oversight; and most importantly, how the use of AI ultimately affects the quality and comprehensiveness of the final literature review produced." They call for investigation of how AI changes the content of literature reviews themselves and document five mechanisms through which AI shapes review content.
Data: 280 papers compiled in consideration set (detailed listings provided in Appendices 2-5); full review texts and comparison data provided in Appendices 6-8; The paper references a corpus of 280 papers compiled during the research. Appendices list the papers considered (Appendix 2) and those used in each review version (Appendix 5). Full lists of references are provided but the original dataset of papers is not stated as publicly available.Code: Not mentioned.Extracted from: pdfAgreement 51%

Explore related topics

Related papers