The Erosion of LLM Signatures: CanWe Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee · International conference Recent advances in natural language processing · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.26615/978-954-452-098-4-129
Methodology & findings
Study design
Empirical experiment with systematic data generation and evaluation.
Sample
N = 846, 6 groups
Primary method
Macro F1-score for classification performance evaluation; Fisher's Discriminant Ratio (FDR) for measuring feature discriminability; Word Mover's Distance (WMD) using GloVe-wiki-gigaword-50 embeddings; Integrated Gradients (IG) for feature importance attribution; Training-validation loss curve analysis with slope computation; Three random train-test splits with averaged results; t-tests for performance comparison across conditions; Early stopping and hyperparameter tuning (batch size, epochs, dropout, early stopping criteria)
Main result
The study found that "even a simple logistic regression model attains 77%" F1-score in Stage 1, but detection performance significantly declines across stages. "We observe, BigBird consistently outperforms all other models, leveraging its superior context-length capability to capture both the research problem (RP) and idea representation effectively." Most critically, "an average decline of 25.4% from Stage 1 to Stage 5" reveals that "the 'LLM signature' becomes increasingly elusive, making it more challenging to establish a clear boundary between human and LLM-generated ideas."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "Unlike direct text-based detection, idea detection is significantly harder as paraphrasing progressively erodes distinctive LLM signatures, making idea attribution increasingly unreliable." They further state that "existing classifiers rely heavily on surface-level linguistic features rather than deeply understanding the underlying idea structures, leading to substantial performance declines as paraphrasing progresses." They advocate for future work to "extend this study beyond CS to other scientific disciplines" and to incorporate "the reasoning trajectory of LLMs during idea generation, as tracing the thought process may provide a more robust signal for detection."
Risk of bias
Topic bias potential due to selection from specific CS conferences (ACL, EMNLP, ICLR, ICML, NeurIPS); Model selection bias: mixture of OpenAI (63%) and Anthropic (37%) models with different capabilities; Temporal bias: papers limited to pre-2022 to exclude ChatGPT era, potentially not representative of current LLM characteristics; Sampling bias: 846 papers is limited sample size acknowledged by authors as constrained by computational resources; Selection bias: Sample limited to 846 papers from five top CS conferences (ACL, EMNLP, ICLR, ICML, NeurIPS) published 2017-2021, not representative of all scientific domains or publication venues; Domain bias: Exclusively CS-focused, limiting generalizability to other scientific disciplines; LLM selection bias: Data generation split 63% OpenAI vs 37% Anthropic models; unequal distribution may introduce model-specific artifacts; Temporal bias: Papers published only up to 2021 to ensure pre-ChatGPT generation; may not reflect contemporary LLM-human idea distinctions; Confounding: LLM signatures may reflect training data composition rather than inherent LLM properties; Manual verification bias: Only 1% of samples manually verified for consistency across all paraphrasing stages; Selection bias: Limited to five A*-rated CS conferences; may not generalize to other disciplines; Domain bias: All papers from 2017-2021 CS conferences; excludes other scientific domains; LLM selection bias: 63% of data generated by OpenAI models vs. 37% by Anthropic; unequal distribution; Temporal bias: Papers limited to pre-2022 to ensure human origin, excluding post-ChatGPT era; Paraphrasing bias: Cascade paraphrasing using same LLMs may introduce systematic artifacts
Limitations
- The sample size of 846 papers was limited
- as stated: "This sample size was chosen primarily due to the substantial computational and financial resources required for large-scale generation and extensive cascading paraphrasing using SOTA LLM APIs." Additionally, the study was restricted to computer science conferences (ACL, EMNLP, ICLR, ICML, NeurIPS) published up to 2021, and the authors note that "existing classifiers rely heavily on surface-level linguistic features rather than deeply understanding the underlying idea structures."
Open questions raised
- Extension to other scientific disciplines beyond CS
- Incorporation of LLM reasoning trajectory during idea generation
- Integration of structured knowledge-based embeddings
- Development of cross-attention modeling structures for better problem-idea interaction capture
- Need for deeper conceptual understanding beyond linguistic artifacts
- Extension beyond computer science to other scientific disciplines
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations