12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading

Yutao Wu, Jiale Ding, Xingjun Ma · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3774904.3792597

Methodology & findings

Study design

Controlled benchmark evaluation study.

Sample

not–specified, 3 groups

Main result

The study found "citation retrieval fails in 48–98% of multi-reference queries, section-specific content extraction fails in 72–91% of cases, and topical paper discovery yields F1 scores below 0.32, missing over 60% of relevant literature." These results demonstrate consistent reliability failures across multiple scholarly tasks when evaluating GPT-4o, GPT-5, and Gemini-2.5-Flash.

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "PaperAsk provides a reproducible and diagnostic framework for advancing the reliability evaluation of LLM-based scholarly assistance systems." They further note that distinct LLM failure patterns exist, with "ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers," and developed "lightweight reliability classifiers trained on PaperAsk data to identify unreliable outputs" as a solution.

Risk of bias

Selection bias: LLMs chosen may not represent all available models; Opacity of web interfaces may limit reproducibility and diagnostic capability; Limited documentation of human analysis methodology for failure attribution; Selection bias in choice of LLM models evaluated; Potential confounding from web interface opacity affecting search operation transparency; Human analysis subjectivity in attributing failure causes; Selection of specific LLM models may not represent broader LLM populations; Web interface dependency may introduce confounding factors unrelated to model capability; Human analysis of failure modes is subject to analyst interpretation bias; Lack of transparency in search operations prevents full causal attribution

Limitations

  • The paper indicates limitations in its scope and methodology, though specific limitation statements are not explicitly detailed in the abstract provided
  • The work focuses on "realistic usage conditions, using web interfaces where search operations are opaque to the user," which inherently constrains the ability to control or fully understand underlying system behaviors.

Open questions raised

  • The paper identifies the need for improved reliability evaluation of LLMs in scholarly tasks and proposes advancing methods for detecting unreliable outputs in LLM-based research assistance systems.
  • The paper identifies the need for improved reliability evaluation frameworks for LLMs in scholarly tasks and proposes developing better mechanisms to identify and mitigate unreliable outputs in LLM-based research assistance systems.
  • The paper identifies the need for improved reliability evaluation frameworks for LLMs in scholarly tasks, development of methods to address uncontrolled expansion of retrieved context, and techniques to help LLMs prioritize task instructions over semantic relevance.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 55%

Explore related topics

Related papers