CoAuthorAI: A Human in the Loop System For Scientific Book Writing
Yangjie Tian, Xungang Gu, Yun Zhao, Jiale Yang, Lin Yang, Ning Li et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods design combining system design and evaluation.
Primary method
Design Science Research with human-in-the-loop paradigm; iterative design with expert feedback
Main result
In evaluations of 500 multi-domain literature review chapters, CoAuthorAI achieved a maximum soft-heading recall of 98%; in a human evaluation of 100 articles, the generated content reached a satisfaction rate of 82%. The book AI for Rock Dynamics generated with CoAuthorAI and Kexin Technology's LUFFA AI model has been published with Springer Nature. These results show that "systematic human–AI collaboration can extend LLMs' capabilities from articles to full-length books, enabling faster and more reliable scientific publishing." Additionally, the system achieved an average citation accuracy of 77.4% with an average manual correction rate of 15.4% across chapters.
Research paradigm
Design Science Research / Human-Computer Interaction
Author conclusions
"The CoAuthorAI system offers a novel and effective approach to scientific book writing by integrating human expertise with the capabilities of large language models. By leveraging human guidance through expert-crafted outlines and iterative feedback loops, the system ensures the quality and precision of the generated content. The CoAuthorAI system has the potential to significantly streamline the scientific writing process and contribute to the production of high-quality scientific books."
Risk of bias
Selection bias: Only 20 of 500 literature reviews received human evaluation; unclear if representative sample; Evaluator bias: Human evaluation team composition and expertise not specified; single primary evaluator (Prim-eval) scored all articles; Publication bias: Book case study (AI for Rock Dynamics) was co-produced by system developers, not independent validation; Model selection bias: Evaluation focused on specific LLMs; unclear if other models were tested equivalently; Confounding: Citation accuracy metric depends on embedding model (bge-m3) similarity threshold, which is not validated; Selection bias in literature review dataset (500 English scientific reviews from unspecified sources); Human evaluation team size may be small (5 members) for inter-rater reliability; Evaluation on primarily domestic/international LLMs with potential performance disparities; Limited diversity in book-writing evaluation (single domain: Rock Dynamics); Potential evaluator bias in human assessment; limited to English scientific reviews; evaluation team composition not fully described; selection of reference materials for book writing may introduce content bias.
Limitations
- "Despite continuous adjustments to the prompts in Table 5, CoAuthorAI still retains a typical machine-generated format in the final books
- The lack of visual elements such as images and tables reduces the readability and appeal of the books
- Compared with traditional books, the content generated from existing literature lacks innovative materials and mainly consists of summaries and syntheses of past knowledge."
Open questions raised
- Future work may focus on further refining the system's capabilities and exploring its applications in other scientific writing tasks
- Future work may focus on further refining the system's capabilities and exploring its applications in other scientific writing tasks. The paper identifies limitations in visual element generation (images, tables) and the lack of innovative content generation beyond synthesis of existing literature.
- Future work may focus on further refining the system's capabilities and exploring its applications in other scientific writing tasks.
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations