12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

"Although Powerful, it's not Infallible": Investigating Academic Researchers' Verification Challenges with LLMs

Monica Visani Scozzi, Stephann Makri, Pranava Madhyastha · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3786304.3787865

Methodology & findings

Study design

Naturalistic observation study with think-aloud protocol.

Main result

The study found that academic researchers employ multiple verification strategies when interacting with LLMs, including "locating cited sources (either through the LLM tool or independently), verifying source authenticity and/or authority, examining original sources to verify specific claims, re-prompting to obtain more accurate responses, and comparing the prompt with the response." Additionally, "task complexity, prior knowledge and problem structure interact to influence academics' decisions on whether or not to verify LLM responses," with researchers consistently prioritizing "the need to rely on original sources versus the limitations of current LLMs in supporting users in doing so."

Research paradigm

Qualitative empirical (interpretive/phenomenological)

Author conclusions

The authors conclude that "Our naturalistic study highlighted a critical tension in academic information seeking with LLMs: the need to rely on original sources versus the limitations of current LLMs in supporting users in doing so." They further state: "Our findings illustrate that transparent source selection, improved source faithfulness, accurate alignment between responses and their sources and support for verification and prompting are not merely desirable technical enhancements, but essential design requirements." Additionally, they argue that "conversational information-seeking tools should engage in a meaningful exchange by empowering and supporting users to articulate their information needs and verify that the information provided is meeting them, rather than making assumptions about the user's intent and making it difficult to verify their outputs."

Risk of bias

Selection bias: Participants self-selected tasks and were recruited through university mailing lists, potentially introducing selection bias toward more engaged researchers; Observer effect: Think-aloud protocol may alter natural verification behavior; researchers may have been more deliberate in stating reasoning due to observation; Tool bias: Participants used different LLM tools (ChatGPT, Gemini, Copilot, Perplexity, ConsensusGPT) with varying features, potentially confounding results; Expertise bias: All 16 participants had prior LLM experience, limiting generalizability to novice users; Attrition/missing data: Fast thinking decisions were invisible to researchers, potentially missing verification decisions made unconsciously; Environmental confounding: Some sessions were disrupted by tool malfunctions or unexpected interface changes, altering natural verification patterns; Selection bias: Participants self-selected tasks and had prior LLM experience, potentially not representing all academic users; Observation effect: Think-aloud protocol may have altered natural verification behavior; Incomplete data capture: Fast thinking decisions to not verify were not observable; Tool-induced confounds: Unexpected events and tool unpredictability altered verification strategies during observation; Researcher bias in coding: First author conducted initial coding, though all authors reviewed for consistency; Selection bias: Participants self-selected and were recruited through university mailing lists, potentially representing researchers with greater LLM engagement or technical comfort; Observer effect: Think-aloud protocol may influence natural verification behaviour; Visibility bias in observational data: Fast, unconscious verification decisions were not observable; Tool-induced bias: Unexpected technical events (broken links, permission prompts) altered some participants' verification strategies during observation; Sample homogeneity: All 16 participants had prior LLM experience and represented academic researchers only

Limitations

  • The authors acknowledge that "with the think-aloud method, the majority of fast thinking decisions not to verify were invisible to the researcher
  • Only more conscious, stated decisions to verify (or not) were observable." Additionally, "due to the naturalistic nature of the study, some participants experienced unexpected events during the session, introduced either by design (choice between answers, new functionality like generating charts) or by programming error (e.g
  • links in ConsensusGPT intermittently broken)...this may have meant that their verification strategies were altered due to tool unpredictability." The study was also limited by its focus on 16 participants and did not isolate research expertise as a distinct factor.

Open questions raised

  • Limited empirical evidence on verification approaches when academic researchers interact with LLMs, especially for scholarly tasks
  • Need for user-centered design approaches supporting source verification in LLM tools
  • Further investigation into how domain knowledge, research expertise, and tool understanding interplay when interacting with opaque systems
  • Research on how conversational interaction influences misconceptions about how LLMs operate, beyond anthropomorphism
  • Investigation of cognitive processing and thought patterns during verification and how cognitive dissonance arises in LLM interactions
  • Limited empirical evidence on verification approaches when interacting with LLMs in scholarly contexts
Extracted from: pdfAgreement 78%

Explore related topics

Related papers