12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How Researchers Navigate Accountability, Transparency, and Trust When Using AI Tools in Early-Stage Research: A Think-Aloud Study

Sanjana Gautam, Houjiang Liu, Yujin Choi, Matthew Lease · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Think-aloud study (TAP) with 15 researchers who verbalized their thoughts while completing five structured research tasks using LLM-based tools (Research Rabbit and Elicit AI).

Sample

N = 15, 1 group

Primary method

Inductive coding for preliminary thematic identification from think-aloud interviews. Two-step analytical approach: initial inductive analysis followed by iterative refinement across additional data until thematic saturation was achieved. Iterative axial coding and theoretical grouping were employed to narrow 23 sub-themes to central themes consistent with core RAI principles. No quantitative statistical tests were conducted.

Main result

The study found that researchers experience systematic tensions around accountability, transparency, and trust when using AI tools in early-stage research. Specifically, "the confident tone of AI outputs obscured the epistemic uncertainty inherent in scholarly judgment, leading them to restrict AI use on core evaluative tasks such as assessing literature relevance and correctness." Additionally, "the opacity of AI retrieval and generation processes made establishing provenance impossible, prompting participants to rely on social credibility heuristics, manual verification, and constrained prompting." Rather than developing stable trust in AI tools, "participant confidence in AI was fragile and context-dependent, built incrementally through logical auditing and deliberately bounded to lower-stakes tasks."

Reports effect sizes.

Research paradigm

Qualitative interpretive/constructivist

Author conclusions

"Drawing on empirical evidence captured during task performance rather than retrospective reflection, we surface the concerns researchers experience when engaging with AI-mediated research, as well as the strategies they currently employ to address them. Together, these findings motivate the need for a slow, deliberate, and well-regulated adoption of AI tools in scientific practice to safeguard scientific integrity." The authors further emphasize that "advancing AI-mediated research cannot be reduced to improving models alone and it requires treating AI as part of a collective sociotechnical system in which transparency, accountability, and trust are actively designed for and socially negotiated."

Risk of bias

Selection bias: Participants recruited through convenience sampling via faculty mailing lists and personal contacts; may not represent broader researcher populations; Hawthorne effect: Participants' behavior may have been altered by the think-aloud protocol and knowledge of being observed; Tool-specific bias: Only two commercially available tools examined; findings may not generalize to other AI research platforms; Temporal bias: Tools underwent updates since data collection, limiting generalizability of findings; Task structure bias: Researchers used specific task prompts provided by experimenters rather than naturalistic, self-directed use; Selection bias: Criterion-based sampling through faculty mailing lists, personal contacts, and online channels may favor researchers already engaged with AI tools; Participant homogeneity: Sample consisted primarily of advanced PhD candidates and academic researchers; limited diversity across career stages and disciplinary representation; Task structure bias: Interactions with AI tools were structured by researcher-provided prompts, not representative of naturalistic open-ended use; Verbalization effect: Think-aloud methodology may alter cognitive processes under observation; Tool-specific bias: Findings tied to two specific tools that may not generalize to other AI research platforms; Researcher experience: Most advanced PhD participants had developed research proficiency before widespread AI availability, potentially affecting perceptions of AI utility; Selection bias: Participants self-selected through recruitment surveys; criterion-based sampling may not represent all researcher populations; Hawthorne effect: Participants' behavior may have been altered by awareness of being observed during think-aloud sessions; Task structure bias: Artificially designed tasks may not reflect naturalistic AI tool use; Verbalization bias: Think-aloud protocol requires explicit articulation which may be incomplete or alter cognitive processes; Tool update bias: Tools (Research Rabbit and Elicit) were updated since data collection, limiting generalizability; Expertise bias: Sample composed primarily of advanced PhD candidates and postdocs may not represent early-career or junior researchers

Limitations

  • The authors state that "Our participant pool draws primarily from academic researchers, excluding industry contexts where epistemic norms and accountability structures around AI may differ substantially." Additionally, "think-aloud methodology, while well-suited for capturing real-time reasoning, provides only a partial window into researcher cognition, as verbalization is effortful and incomplete, and the act of articulating judgment may itself alter the processes under observation." They also note that "Participant interactions with AI tools during the study were necessarily shaped by the task structure we provided, meaning that the prompts they issued and the outputs they encountered were not fully representative of the open-ended, self-directed AI use characteristic of naturalistic research practice."

Open questions raised

  • Need for longitudinal, naturalistic studies tracking how researchers' epistemic negotiations with AI tools evolve as both tools and norms surrounding them mature
  • Investigation of domain-specific requirements, as research practices and workflows vary widely across disciplines
  • Larger and more diverse participant pools beyond academic researchers to understand adoption in industry contexts
  • Integration of behavioral logs, interaction data, and quantitative metrics to provide objective measures of tool efficacy
  • Combining qualitative insights with quantitative analyses for comprehensive evaluation of AI tools
  • Lack of empirical investigation into how accountability, transparency, and trust concerns manifest in real-time scholarly judgments of active researchers
Data: Not mentionedCode: Not mentionedExtracted from: pdfAgreement 63%

Explore related topics

Related papers