PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
Anna Martin, Cara A. C. Leckey, Martha C. Brown, Harmanpreet Kaur · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3772318.3791101
Methodology & findings
Study design
Within-subjects user study with 26 domain experts from a research organization comparing PaperTrail interface against baseline citation-based interface across two scholarly tasks (multi-paper synthesis and devil's advocate paper review).
Primary method
Design science research approach informed by human-centered AI, argumentation theory, and information foraging/sensemaking frameworks. The design is grounded in four explicit design goals (DG1-DG4) derived from scholarly QA literature and informed by prior work on evidence-based text generation, trust calibration, and coordination of multiple views.
Main result
The study found that "granular claim-evidence provenance information encourages more caution towards LLM outputs in scholarly settings" and that "People trust in LLMs is significantly lower after using PaperTrail compared to baseline." However, "this change does not translate to differences in perceived confidence or, importantly, changes in reliance behaviors," revealing what the authors characterize as a "trust-behavior gap in scholarly LLM use, showing that reduced trust alone is insufficient to change reliance behaviors without addressing systemic constraints of time, usability, and cognitive resources."
Research paradigm
pragmatist/mixed-methods (combining human-centered AI research with system design)
Author conclusions
The authors conclude that "while argument-based provenance can encourage healthy skepticism toward LLM outputs, translating this into changed reliance behavior requires overcoming substantial barriers related to time, usability, and ingrained patterns of tool use." They state that "The gap between recognizing verification needs and performing verification actions remains a challenge" and that "granular provenance information alone is insufficient to change behavior when users face the time pressures and cognitive constraints typical of research settings." They also emphasize: "Our findings indicate that the relationship between attitudes and actions in human-AI collaboration is complex, particularly in time-constrained, cognitively demanding contexts like scholarly tasks."
Risk of bias
Selection bias: participants from single organization (NASA research centers) may not represent broader scholar population; Task authenticity bias: artificial time constraints (20 minutes) differed from real scholarly workflows; Attrition bias: 12 participants excluded post-hoc based on task duration and engagement metrics; Learning effects: interface condition order was counterbalanced but single-session study prevented long-term adaptation; Familiarity bias: exclusion of participants with >3/7 familiarity with source papers may have introduced knowledge bias; Selection bias: Participants recruited from single NASA organization (narrow disciplinary and institutional sample); Task artificiality bias: 20-30 minute time constraints and unfamiliar papers may not represent authentic scholarly practices; Attrition: 12 of 38 participants (31.6%) excluded due to insufficient task engagement or system latency issues; Confounding variables: System latency (average 90 seconds per query) explicitly acknowledged as barrier to engagement, potentially conflating interface design effects with technical performance; Learning/order effects: Interface condition was randomized but task order was fixed (Task 1 then Task 2), potentially introducing fatigue or learning effects; Measurement confound: Edit distance metric conflates passive acceptance with informed delegation, per authors; Hawthorne effect: Participants in study context may modify behavior due to awareness of observation; Selection bias: 26 of 38 participants who started the study completed it; 12 were excluded post-hoc based on task completion time and data quality without reference to outcomes, but this ex-post exclusion could introduce bias; Attrition: 78 screened, 74 met inclusion, 38 participated, 26 completed—substantial dropout; Confounding: Time pressure and system latency (90 second average end-to-end response time) likely suppressed reliance behavior change independent of interface condition; Order effects: Task order was fixed (Task 1 then Task 2) to control for learning but interface condition was counterbalanced; Hawthorne effect: Artificial study context may have altered engagement patterns compared to authentic scholarly work; Sample homogeneity: All participants from single organization (NASA), predominantly STEM researchers with advanced degrees
Limitations
- "This took two forms: (1) participants engaged with unfamiliar papers within artificial time constraints, rather than conducting authentic literature searches or working with materials from their own research
- and (2) the 20-30 minute task duration guidance, that prevented the deep engagement that characterizes scholarly work in practice." Additionally, "our operationalization of reliance through edit distance may not fully capture the nuanced ways participants engaged with LLM assistance
- The measure conflates various behaviors (from wholesale acceptance to strategic delegation)," and "we did not evaluate the quality of participants' edited texts."
Open questions raised
- Need for field studies examining PaperTrail effectiveness in authentic research contexts with familiar literature and self-directed questions
- Development of more sophisticated behavioral measures distinguishing passive acceptance from informed delegation
- Component-specific trust measures to separate interface trust from LLM trust
- Quality assessment of edited texts to determine if unsupported claims were removed and overall argumentation improved
- Alternative design approaches beyond three-panel layout (inline annotations, progressive disclosure, conversational verification)
- Examination of claim-evidence provenance across diverse academic disciplines with varying argumentation conventions
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- PaperQA: Retrieval-Augmented Generative Agent for Scientific ResearchJakub Lála · 2023 · 52 citations
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations