12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary

Yutong Li, Lu Chen, Aiwei Liu, Kai Yu, Lijie Wen · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
3
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2403.02574

Methodology & findings

Study design

Design science research with comparative evaluation.

Primary method

Design science research with human workflow guidance as design principle

Main result

ChatCite outperformed other LLM-based literature summarization methods across quality dimensions. The study found that "ChatCite outperforms other models in various dimensions in the experiments" and "ChatCite performs best among LLM-based literature summarization methods, and the approach following the human workflow guidance is superior to the results obtained by the Chain of Thought (CoT) method." Specifically, ChatCite achieved a G-Score of 4.0642 compared to GPT-4 zero-shot at 3.5076 and LitLLM at 3.5448, with 51% human preference over other baselines.

Research paradigm

pragmatist

Author conclusions

The authors conclude: "LLMs are powerful tools in generating literature summaries, however, it poses the challenges of information omission, lack of comparative summaries and organizational deficiencies. In ChatCite, the Key Element Extractor contributes to improving content consistency, and the Comparative Incremental Generator effectively enhances the organizational structure, comparative analysis, and citation accuracy of the generated summary. Additionally, the literature summaries generated by ChatCite can be directly used for drafting literature reviews. Our study also demonstrated that the approach following the human workflow guidance is superior to the results obtained by the Chain of Thought (CoT) method."

Risk of bias

Limited dataset domain: only computer science papers (63 citations average), not generalizable to other fields; Small human study sample: only 10 papers evaluated by human annotators; Evaluator bias: human evaluators were computer science researchers, potentially biased toward technical language; Model dependency: evaluation relies heavily on GPT-4 as the automatic evaluator, introducing potential circular reasoning; Selection bias: test papers are 'well-received' papers with high citations, not representative of all research; Confounding variables: different context windows used for different models (16K for GPT-3.5 vs 128K for GPT-4); Dataset homogeneity: Only computer science papers included, no cross-domain validation; Model selection bias: Only GPT-3.5 and GPT-4 tested; no exploration of other LLM architectures; Evaluator bias: Human evaluation conducted on only 10 samples with unspecified number of annotators; Hyperparameter exploration: Limited investigation of different settings for GPT-3.5 model; Baseline implementation: LitLLM baseline implemented by authors based on paper description rather than official implementation; Dataset bias: Only 50 papers from computer science field, lacks diversity across disciplines; Model selection bias: Limited to GPT-3.5 and GPT-4 models for validation; Evaluator bias: Human study conducted with only 10 samples and researchers from single discipline; Evaluation metric bias: G-Score is LLM-based and may reflect LLM preferences rather than true quality

Limitations

  • "In this work, we focused mainly on the summarization of specific topics based on the selected literatures instead of the collection and the filtering of the literatures themselves
  • The datasets primarily consist of research articles in the area of computer science and lack research articles from other fields of study to validate our model
  • Our experimentation used Chat GPT 3.5 as the tool for validating the quality of the generated content and the functionalities of the various components of the agent
  • We did not explore any additional spec that can influence the result of the GPT3.5 model nor the possibility of using other models as the validation tool
  • The evaluation of the generated content poses a great challenge
  • We evaluated the generated results from multiple dimensions using G-Score as the performance metric, but there is still room for improvements over the accuracy of the automatic evaluation process

Open questions raised

  • Future work should extend validation beyond computer science to other research fields. The paper identifies randomness and instability in generated results as requiring further research. The authors note potential for improvements in automatic evaluation accuracy and suggest exploring other LLMs beyond GPT-3.5 as validation tools. They suggest broader application to complex inferential writing tasks.
  • Future work should address literature collection and filtering steps, not just summarization
  • Validation needed across research domains beyond computer science
  • Need to explore additional hyperparameters and LLM models beyond GPT-3.5 and GPT-4
  • Improvement of automatic evaluation accuracy for generated content
  • Enhancement of output stability and quality
Data: NudtRwG-Citation dataset (Wang et al., 2020) - test set of 50 academic research papers in Computer Science; NudtRwG-Citation dataset (Wang et al., 2020) - 50 academic research papers in computer science with ground truth related work sections and reference annotations; NudtRwG-Citation dataset (Wang et al., 2020) - 50 academic research papers in Computer Science used for experimentsCode: Code to be released after review process (as of publication date); Authors state: "Our code will be released after the review process."; Code availability stated: "Our code will be released after the review process."Extracted from: pdfAgreement 61%

Explore related topics

Related papers