CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery
Mengze Hong, Di Jiang, Zichang Guo, Yawen Li, Jun song Chen, Shaobo Cui et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
System design and evaluation using 40 sentence statements sampled from publicly available research papers across multiple disciplines, each accompanied by human-annotated search queries and ground-truth references.
Primary method
Design science approach; artifact-centric research with integrated user interface design within LaTeX editor environment
Main result
The proposed CiteLLM system demonstrates superior performance in reference discovery compared to baseline methods. The results show that "the proposed context-aware approach consistently outperforms both baselines across all three dimensions, demonstrating its ability to produce clearer, more specific, and human-aligned search queries." Additionally, "Table 1 summarizes the performance of three methods on the validation set, showing 100% valid references with high precision and usability in the proposed system," with CiteLLM achieving 100% validity, 84.0% human-evaluated precision, and 87.5% usability across 40 test sentences.
Research paradigm
Design science / pragmatism
Author conclusions
"The CiteLLM system represents a significant step toward the trustworthy, context-aware, and privacy-preserving integration of LLMs in academic writing." The authors conclude that "by automating reference discovery through context-aware query construction, validity verification, and a novel interaction paradigm, researchers can efficiently identify, assess, and cite relevant works without exposing sensitive manuscripts to external servers or fragmented third-party platforms." They emphasize that "Ultimately, CiteLLM demonstrates how responsible AI design can enhance academic workflows, enabling researchers to engage with scientific knowledge more efficiently while maintaining full control over their intellectual contributions."
Risk of bias
Evaluation reliance on small sample size (40 sentences); potential evaluator bias in human expert assessment; discrepancy noted between human and GPT-5 ratings suggesting standardization issues in evaluation metrics.; Small sample size (n=40 sentences) limits generalizability; Selection bias: sentences sampled from publicly available research papers only; Evaluator bias: only three experienced researchers scored query quality; LLM evaluation bias: documented misalignment between GPT-5 and human evaluator ratings; Limited disciplinary representation: evaluation across 'multiple disciplines' but specific distribution not detailed; Small evaluation set (n=40 sentences) may not represent full diversity of academic writing styles; Selection of papers from public repositories may introduce publication bias; Evaluation metrics reliance on GPT-5 as judge introduces potential algorithmic bias; Human evaluators may have limited disciplinary expertise across all 40 test sentences
Limitations
- The paper states "This discrepancy highlights the need for caution when employing fully autonomous LLM agents in research, as they may compromise scientific rigor." Additionally, the authors note that "the current implementation relies on public preprint repositories: arXiv (computer science, physics, and mathematics), bioRxiv (biology-related claims), and medRxiv (clinical and medical research)" with the caveat that "The architecture naturally extends to restricted conference or journal corpora, provided appropriate database or API access is available." The evaluation set was limited to 40 sentences, and no discussion of computational latency or cost metrics is provided.
Open questions raised
- Developing more user-friendly LLM utilities; optimizing integration to reduce latency and costs; enabling seamless and impactful AI automation in academic workflows. The discrepancy between human and GPT-5 evaluation standards highlights the need for caution when employing fully autonomous LLM agents in research.
- "Future work should focus on developing more user-friendly LLM utilities and optimizing their integration to reduce latency and costs, enabling seamless and impactful AI automation in academic workflows." The architecture is noted to naturally extend to "restricted conference or journal corpora, provided appropriate database or API access is available," suggesting this as a future direction.
- The authors identify that "Future work should focus on developing more user-friendly LLM utilities and optimizing their integration to reduce latency and costs, enabling seamless and impactful AI automation in academic workflows." Additionally, the need for more sophisticated LLM evaluation standards is highlighted by the misalignment between human and GPT-5 evaluation ratings.
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations