Auditing GenAI Literature Search Workflows: A Replicable Protocol for Traceable, Accountable Retrieval in Student-Facing Inquiry
Cristo León, Michelle Kudelka · AI in Education · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/aieduc2020008
Methodology & findings
Study design
Exploratory, autoethnographic audit design using Constructivist Grounded Theory (CGT).
Primary method
Constructivist Grounded Theory (CGT) with iterative protocol refinement. The study employed constant comparison across tool outputs and runs, memo-writing to refine analytic categories, and targeted theoretical sampling using LIS expert records to strengthen coverage of foundational information retrieval and evidence synthesis literature.
Main result
The audit found that "natural-language retrieval frequently produced non-verifiable or incorrectly matched citations, undermining traceability and metadata integrity under student-facing conditions." More specifically, "Aggregated across all natural-language runs (n = 100 items), only 58% met the DOI_Correct criterion, with the remainder distributed across non-resolving DOI, wrong-match DOI, and no-DOI outcomes." Boolean translation improved outcomes: "Across tools, Boolean translation increased yield consistency relative to natural-language prompting, including for tools that under-produced in natural-language mode." Additionally, "run-to-run drift persisted across tools and postures, indicating that per-item citation correctness does not guarantee reproducible evidence sets."
Research paradigm
constructivist grounded theory with pragmatist governance orientation
Author conclusions
"In conclusion, this study shows that AI-assisted literature search workflows are best evaluated as accountable inquiry rather than as generic tool performance. The audit protocol offers instructors, librarians, and research administrators a practical mechanism to govern AI use in student-facing work by foregrounding reproducibility and transparency at the level of citations and evidence trails." Furthermore, "Tool choice and retrieval posture materially shape whether students can demonstrate verifiable sourcing. Natural-language outputs frequently produced integrity failures that are difficult to detect without structured auditing, including DOI wrong-match errors that mimic verification while redirecting users to the wrong record. Boolean translation and database execution increased run completion and shifted retrieval toward DOI-bearing records, improving auditability." The authors recommend: "Decision rule. In student-facing inquiry, natural-language prompting should be limited to exploratory term generation and early topic orientation, whereas the evidence set used for screening, coding, and citation should be derived from database-executed Boolean retrieval with documented query logic and DOI-level verification."
Risk of bias
Tool output bias via DOI wrong-match, DOI non-resolving, bibliographic field instability; Author-name parsing errors affecting communities with multiple given/family names; Source-type substitution (non-peer-reviewed content presented in journal-like form); Reporting bias as documentation and provenance gaps; Selection bias in librarian benchmark construction (expert judgment dependent); Ranking opacity in tool outputs limiting reproducibility assessment; Under-yield runs forcing analyses to rely on uneven top-k sets; Tool-output bias: systematic integrity failures including DOI wrong-match, DOI non-resolving, bibliographic field instability, and source-type substitution; Author-name parsing errors, particularly for authors with compound given names or surnames; Reporting bias introduced by tool mediation, including missing rationales, opaque browsing processes, and substitution of non-peer-reviewed sources; Contingency on librarian benchmark selection and database access conditions; Fixed canonical prompt may not represent full range of student prompting behavior; Tool-output bias: Systematic integrity risks from DOI wrong-match, non-resolving DOIs, bibliographic field instability, and source-type substitution; Author-name parsing errors affecting communities with compound given names and/or surnames; Reporting bias: Documentation and provenance gaps introduced by tool mediation, including missing rationales for query construction, opaque browsing and ranking, and non-peer-reviewed source substitution; Librarian benchmark contingency: Overlap estimates depend on expert judgment and database access conditions; Top-k truncation effects: Drift metrics sensitive to under-yield runs and tool-specific ranking volatility
Limitations
- Two primary limitations bound interpretation
- "First, the audit evaluates tool outputs under a fixed canonical prompt and a top-k capture rule
- Alternative prompts, model settings, database filters, or ranking choices may materially alter yield, DOI traceability, metadata integrity, and run-to-run drift." "Second, overlap and alignment estimates depend on the librarian benchmark, which is contingent on expert judgment and database access conditions
- While librarian augmentation strengthens the methodological baseline by anchoring the study in established information-retrieval and evidence-synthesis practice, the benchmark remains contingent on expert judgment, search system access (for example, availability of Scopus and Web of Science), and the librarian's domain framing."
Open questions raised
- Most discussions of AI in reviews focus on efficiency and ethics while fewer studies operationalize traceability and metadata integrity as auditable outcomes at the record level
- Limited direct evaluation against expert information-retrieval baselines in educationally realistic settings with repeated runs to expose drift
- Need for expansion of overlap analyses at DOI-level identity across all AI runs with bounded precision and recall relative to librarian target set
- Testing of prompt variants to evaluate how prompt controls affect drift and integrity outcomes
- Replication across additional databases (Web of Science, ERIC) to test stability under alternative ranking, indexing, and filtering regimes
- Standardization of benchmark construction (e.g., multi-librarian adjudication or consensus procedures)
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations