12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A

Anna Martin, Cara A. C. Leckey, Martha C. Brown, Harmanpreet Kaur · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
1
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3772318.3791101

Methodology & findings

Study design

Within-subjects user study with 26 domain experts from a research organization comparing PaperTrail interface against baseline citation-based interface across two scholarly tasks (multi-paper synthesis and devil's advocate paper review).

Primary method

Design science research approach informed by human-centered AI, argumentation theory, and information foraging/sensemaking frameworks. The design is grounded in four explicit design goals (DG1-DG4) derived from scholarly QA literature and informed by prior work on evidence-based text generation, trust calibration, and coordination of multiple views.

Main result

The study found that "granular claim-evidence provenance information encourages more caution towards LLM outputs in scholarly settings" and that "People trust in LLMs is significantly lower after using PaperTrail compared to baseline." However, "this change does not translate to differences in perceived confidence or, importantly, changes in reliance behaviors," revealing what the authors characterize as a "trust-behavior gap in scholarly LLM use, showing that reduced trust alone is insufficient to change reliance behaviors without addressing systemic constraints of time, usability, and cognitive resources."

Research paradigm

pragmatist/mixed-methods (combining human-centered AI research with system design)

Author conclusions

The authors conclude that "while argument-based provenance can encourage healthy skepticism toward LLM outputs, translating this into changed reliance behavior requires overcoming substantial barriers related to time, usability, and ingrained patterns of tool use." They state that "The gap between recognizing verification needs and performing verification actions remains a challenge" and that "granular provenance information alone is insufficient to change behavior when users face the time pressures and cognitive constraints typical of research settings." They also emphasize: "Our findings indicate that the relationship between attitudes and actions in human-AI collaboration is complex, particularly in time-constrained, cognitively demanding contexts like scholarly tasks."

Risk of bias

Selection bias: participants from single organization (NASA research centers) may not represent broader scholar population; Task authenticity bias: artificial time constraints (20 minutes) differed from real scholarly workflows; Attrition bias: 12 participants excluded post-hoc based on task duration and engagement metrics; Learning effects: interface condition order was counterbalanced but single-session study prevented long-term adaptation; Familiarity bias: exclusion of participants with >3/7 familiarity with source papers may have introduced knowledge bias; Selection bias: Participants recruited from single NASA organization (narrow disciplinary and institutional sample); Task artificiality bias: 20-30 minute time constraints and unfamiliar papers may not represent authentic scholarly practices; Attrition: 12 of 38 participants (31.6%) excluded due to insufficient task engagement or system latency issues; Confounding variables: System latency (average 90 seconds per query) explicitly acknowledged as barrier to engagement, potentially conflating interface design effects with technical performance; Learning/order effects: Interface condition was randomized but task order was fixed (Task 1 then Task 2), potentially introducing fatigue or learning effects; Measurement confound: Edit distance metric conflates passive acceptance with informed delegation, per authors; Hawthorne effect: Participants in study context may modify behavior due to awareness of observation; Selection bias: 26 of 38 participants who started the study completed it; 12 were excluded post-hoc based on task completion time and data quality without reference to outcomes, but this ex-post exclusion could introduce bias; Attrition: 78 screened, 74 met inclusion, 38 participated, 26 completed—substantial dropout; Confounding: Time pressure and system latency (90 second average end-to-end response time) likely suppressed reliance behavior change independent of interface condition; Order effects: Task order was fixed (Task 1 then Task 2) to control for learning but interface condition was counterbalanced; Hawthorne effect: Artificial study context may have altered engagement patterns compared to authentic scholarly work; Sample homogeneity: All participants from single organization (NASA), predominantly STEM researchers with advanced degrees

Limitations

  • "This took two forms: (1) participants engaged with unfamiliar papers within artificial time constraints, rather than conducting authentic literature searches or working with materials from their own research
  • and (2) the 20-30 minute task duration guidance, that prevented the deep engagement that characterizes scholarly work in practice." Additionally, "our operationalization of reliance through edit distance may not fully capture the nuanced ways participants engaged with LLM assistance
  • The measure conflates various behaviors (from wholesale acceptance to strategic delegation)," and "we did not evaluate the quality of participants' edited texts."

Open questions raised

  • Need for field studies examining PaperTrail effectiveness in authentic research contexts with familiar literature and self-directed questions
  • Development of more sophisticated behavioral measures distinguishing passive acceptance from informed delegation
  • Component-specific trust measures to separate interface trust from LLM trust
  • Quality assessment of edited texts to determine if unsupported claims were removed and overall argumentation improved
  • Alternative design approaches beyond three-panel layout (inline annotations, progressive disclosure, conversational verification)
  • Examination of claim-evidence provenance across diverse academic disciplines with varying argumentation conventions
Data: SciClaimHunt dataset (Kumar et al.) - 50 paragraphs sampled with 50-sample holdout set; BioClaimDetect dataset (Achakulvisut et al.) - 50 abstracts randomly sampled from test set; SciClaimHunt dataset (Kumar et al., 2024) - used for offline evaluation; BioClaimDetect dataset (Achakulvisut et al., 2020) - used for offline evaluation; Four Mars exploration peer-reviewed papers used as corpus (full citations provided in source document corpus section); SciClaimHunt dataset (Kumar et al., 2024) - 50 paragraphs with claims sampled for evaluation; BioClaimDetect dataset (Achakulvisut et al.) - 50 abstracts from test set sampled for evaluationCode: Not mentioned in paper; Not mentioned in the paperExtracted from: pdfAgreement 45%

Explore related topics

Related papers