Prêt-à-Submit: Desarrollo y evaluación de una herramienta basada en LLMs para la revisión automática de listas de verificación de conferencias académicas
Adrián Puyet Muñoz · UPM Digital Archive (Technical University of Madrid) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical evaluation study using corpus-based measurement.
Sample
N = 3748, 2 groups
Primary method
Comparative accuracy evaluation across nine open-source LLMs using author-provided responses as ground truth; performance metrics include accuracy percentages.
Main result
The study found that "lightweight models, such as qwen3.5:9b, achieve an accuracy of up to 83.3 %, while smaller models, such as gemma3:4b, yield acceptable results on consumer‐grade hardware." Additionally, "The analysis also reveals a systematic bias toward negative responses and inconsistencies across models, highlighting the necessity of human oversight in any practical implementation."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "Prêt-à-Submit: an open-source web application that enables researchers to perform verification locally or on their own servers. This tool ensures manuscript privacy by removing the reliance on third-party services, preventing unpublished works from being processed in external cloud environments. Furthermore, it optimizes the author's workflow by automating a repetitive and error-prone task, streamlining self-assessment through a simple and intuitive interface."
Risk of bias
Systematic bias toward negative responses in model outputs; Inconsistencies between models; Ground truth limitation: authors' responses may not represent objective correctness; Corpus limited to NeurIPS 2024 papers only, potential domain-specific bias; Systematic bias toward negative responses in model predictions; Inconsistencies across different LLM models; Potential dataset-specific bias (NeurIPS 2024 checklist may not generalize to other conferences); Inconsistencies between different models; Potential reference bias from using only author-provided responses as ground truth
Limitations
- The paper states that "The analysis also reveals a systematic bias toward negative responses and inconsistencies across models, highlighting the necessity of human oversight in any practical implementation." This indicates limitations in model consistency and systematic biases that restrict practical applicability without human review.
Open questions raised
- The authors identify the need for solutions that balance automation with human oversight, privacy concerns with accessibility, and the necessity of practical tools for researchers with limited resources. They highlight inconsistencies across models as an area requiring further attention.
- The authors identify the need for human oversight in practical implementations and highlight concerns about the accessibility of existing proprietary tools for researchers with limited resources, motivating the development of an open-source solution.
- The authors identify the need for solutions that address manuscript privacy concerns in existing proprietary tools and the challenge of providing accessible automated checklist verification for researchers with limited resources.
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations