12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery

Mengze Hong, Di Jiang, Zichang Guo, Yawen Li, Jun song Chen, Shaobo Cui et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

System design and evaluation using 40 sentence statements sampled from publicly available research papers across multiple disciplines, each accompanied by human-annotated search queries and ground-truth references.

Primary method

Design science approach; artifact-centric research with integrated user interface design within LaTeX editor environment

Main result

The proposed CiteLLM system demonstrates superior performance in reference discovery compared to baseline methods. The results show that "the proposed context-aware approach consistently outperforms both baselines across all three dimensions, demonstrating its ability to produce clearer, more specific, and human-aligned search queries." Additionally, "Table 1 summarizes the performance of three methods on the validation set, showing 100% valid references with high precision and usability in the proposed system," with CiteLLM achieving 100% validity, 84.0% human-evaluated precision, and 87.5% usability across 40 test sentences.

Research paradigm

Design science / pragmatism

Author conclusions

"The CiteLLM system represents a significant step toward the trustworthy, context-aware, and privacy-preserving integration of LLMs in academic writing." The authors conclude that "by automating reference discovery through context-aware query construction, validity verification, and a novel interaction paradigm, researchers can efficiently identify, assess, and cite relevant works without exposing sensitive manuscripts to external servers or fragmented third-party platforms." They emphasize that "Ultimately, CiteLLM demonstrates how responsible AI design can enhance academic workflows, enabling researchers to engage with scientific knowledge more efficiently while maintaining full control over their intellectual contributions."

Risk of bias

Evaluation reliance on small sample size (40 sentences); potential evaluator bias in human expert assessment; discrepancy noted between human and GPT-5 ratings suggesting standardization issues in evaluation metrics.; Small sample size (n=40 sentences) limits generalizability; Selection bias: sentences sampled from publicly available research papers only; Evaluator bias: only three experienced researchers scored query quality; LLM evaluation bias: documented misalignment between GPT-5 and human evaluator ratings; Limited disciplinary representation: evaluation across 'multiple disciplines' but specific distribution not detailed; Small evaluation set (n=40 sentences) may not represent full diversity of academic writing styles; Selection of papers from public repositories may introduce publication bias; Evaluation metrics reliance on GPT-5 as judge introduces potential algorithmic bias; Human evaluators may have limited disciplinary expertise across all 40 test sentences

Limitations

  • The paper states "This discrepancy highlights the need for caution when employing fully autonomous LLM agents in research, as they may compromise scientific rigor." Additionally, the authors note that "the current implementation relies on public preprint repositories: arXiv (computer science, physics, and mathematics), bioRxiv (biology-related claims), and medRxiv (clinical and medical research)" with the caveat that "The architecture naturally extends to restricted conference or journal corpora, provided appropriate database or API access is available." The evaluation set was limited to 40 sentences, and no discussion of computational latency or cost metrics is provided.

Open questions raised

  • Developing more user-friendly LLM utilities; optimizing integration to reduce latency and costs; enabling seamless and impactful AI automation in academic workflows. The discrepancy between human and GPT-5 evaluation standards highlights the need for caution when employing fully autonomous LLM agents in research.
  • "Future work should focus on developing more user-friendly LLM utilities and optimizing their integration to reduce latency and costs, enabling seamless and impactful AI automation in academic workflows." The architecture is noted to naturally extend to "restricted conference or journal corpora, provided appropriate database or API access is available," suggesting this as a future direction.
  • The authors identify that "Future work should focus on developing more user-friendly LLM utilities and optimizing their integration to reduce latency and costs, enabling seamless and impactful AI automation in academic workflows." Additionally, the need for more sophisticated LLM evaluation standards is highlighted by the misalignment between human and GPT-5 evaluation ratings.
Data: 40 sentence statements sampled from publicly available research papers (specific dataset link not provided); code repository mentioned but not fully detailedCode: https://github.com/kermitt2/grobid; https://github.com/kermitt2/grobid (GROBID tool for scholarly PDF processing, referenced but not authored by this team)Extracted from: pdfAgreement 57%

Explore related topics

Related papers