12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Auditing GenAI Literature Search Workflows: A Replicable Protocol for Traceable, Accountable Retrieval in Student-Facing Inquiry

Cristo León, Michelle Kudelka · AI in Education · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/aieduc2020008

Methodology & findings

Study design

Exploratory, autoethnographic audit design using Constructivist Grounded Theory (CGT).

Primary method

Constructivist Grounded Theory (CGT) with iterative protocol refinement. The study employed constant comparison across tool outputs and runs, memo-writing to refine analytic categories, and targeted theoretical sampling using LIS expert records to strengthen coverage of foundational information retrieval and evidence synthesis literature.

Main result

The audit found that "natural-language retrieval frequently produced non-verifiable or incorrectly matched citations, undermining traceability and metadata integrity under student-facing conditions." More specifically, "Aggregated across all natural-language runs (n = 100 items), only 58% met the DOI_Correct criterion, with the remainder distributed across non-resolving DOI, wrong-match DOI, and no-DOI outcomes." Boolean translation improved outcomes: "Across tools, Boolean translation increased yield consistency relative to natural-language prompting, including for tools that under-produced in natural-language mode." Additionally, "run-to-run drift persisted across tools and postures, indicating that per-item citation correctness does not guarantee reproducible evidence sets."

Research paradigm

constructivist grounded theory with pragmatist governance orientation

Author conclusions

"In conclusion, this study shows that AI-assisted literature search workflows are best evaluated as accountable inquiry rather than as generic tool performance. The audit protocol offers instructors, librarians, and research administrators a practical mechanism to govern AI use in student-facing work by foregrounding reproducibility and transparency at the level of citations and evidence trails." Furthermore, "Tool choice and retrieval posture materially shape whether students can demonstrate verifiable sourcing. Natural-language outputs frequently produced integrity failures that are difficult to detect without structured auditing, including DOI wrong-match errors that mimic verification while redirecting users to the wrong record. Boolean translation and database execution increased run completion and shifted retrieval toward DOI-bearing records, improving auditability." The authors recommend: "Decision rule. In student-facing inquiry, natural-language prompting should be limited to exploratory term generation and early topic orientation, whereas the evidence set used for screening, coding, and citation should be derived from database-executed Boolean retrieval with documented query logic and DOI-level verification."

Risk of bias

Tool output bias via DOI wrong-match, DOI non-resolving, bibliographic field instability; Author-name parsing errors affecting communities with multiple given/family names; Source-type substitution (non-peer-reviewed content presented in journal-like form); Reporting bias as documentation and provenance gaps; Selection bias in librarian benchmark construction (expert judgment dependent); Ranking opacity in tool outputs limiting reproducibility assessment; Under-yield runs forcing analyses to rely on uneven top-k sets; Tool-output bias: systematic integrity failures including DOI wrong-match, DOI non-resolving, bibliographic field instability, and source-type substitution; Author-name parsing errors, particularly for authors with compound given names or surnames; Reporting bias introduced by tool mediation, including missing rationales, opaque browsing processes, and substitution of non-peer-reviewed sources; Contingency on librarian benchmark selection and database access conditions; Fixed canonical prompt may not represent full range of student prompting behavior; Tool-output bias: Systematic integrity risks from DOI wrong-match, non-resolving DOIs, bibliographic field instability, and source-type substitution; Author-name parsing errors affecting communities with compound given names and/or surnames; Reporting bias: Documentation and provenance gaps introduced by tool mediation, including missing rationales for query construction, opaque browsing and ranking, and non-peer-reviewed source substitution; Librarian benchmark contingency: Overlap estimates depend on expert judgment and database access conditions; Top-k truncation effects: Drift metrics sensitive to under-yield runs and tool-specific ranking volatility

Limitations

  • Two primary limitations bound interpretation
  • "First, the audit evaluates tool outputs under a fixed canonical prompt and a top-k capture rule
  • Alternative prompts, model settings, database filters, or ranking choices may materially alter yield, DOI traceability, metadata integrity, and run-to-run drift." "Second, overlap and alignment estimates depend on the librarian benchmark, which is contingent on expert judgment and database access conditions
  • While librarian augmentation strengthens the methodological baseline by anchoring the study in established information-retrieval and evidence-synthesis practice, the benchmark remains contingent on expert judgment, search system access (for example, availability of Scopus and Web of Science), and the librarian's domain framing."

Open questions raised

  • Most discussions of AI in reviews focus on efficiency and ethics while fewer studies operationalize traceability and metadata integrity as auditable outcomes at the record level
  • Limited direct evaluation against expert information-retrieval baselines in educationally realistic settings with repeated runs to expose drift
  • Need for expansion of overlap analyses at DOI-level identity across all AI runs with bounded precision and recall relative to librarian target set
  • Testing of prompt variants to evaluate how prompt controls affect drift and integrity outcomes
  • Replication across additional databases (Web of Science, ERIC) to test stability under alternative ranking, indexing, and filtering regimes
  • Standardization of benchmark construction (e.g., multi-librarian adjudication or consensus procedures)
Data: Dataset S3: Identification (All Runs, n=212) - CSV export via OSF; Dataset S4: Screening Cleaned (n=170) - CSV export via OSF; Dataset S5: Full-text eligibility (n=100) - CSV export via OSF; Dataset S6: Included Articles (n=37) - CSV export via OSF; Dataset S7: Expert Sources (n=27) - Librarian baseline set; Zotero group library (shared group planned for public access upon publication); Materials will be deposited via Open Science Framework; Dataset S3. Identification (All Runs, n = 212); Dataset S4. Screening Cleaned (n = 170); Dataset S5. Full-text eligibility (n = 100); Dataset S6. Included Articles (n = 37); Dataset S7. Expert Sources (n = 27); Materials to be deposited via Open Science Framework and institutional repository (Digital Commons STEM for Success); Dataset S3: Identification (All Runs, n=212) - to be released as Supplementary Materials in CSV format; Dataset S4: Screening Cleaned (n=170) - to be released via Open Science Framework; Dataset S5: Full-text eligibility (n=100) - to be released via Open Science Framework; Dataset S6: Included Articles (n=37) - to be released via Open Science Framework; Dataset S7: LIS Expert Sources baseline (n=27); Shared Zotero group (public accessibility upon publication planned)Code: Open Science Framework (OSF) - planned deposition location; Institutional repository (Digital Commons STEM for Success) - derivative presentations to be disseminatedExtracted from: pdfAgreement 59%

Explore related topics

Related papers