12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma

Ömer Aydın, Enis Karaarslan, Fatih Safa Erenay, Nebojša Bačanin · SSRN Electronic Journal · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
C
Evidence
17
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2139/ssrn.5133368

Methodology & findings

Study design

Comparative evaluation study using multiple large language models (DeepSeek v3, Qwen 2.5 Max, Qwen3 235B, ChatGPT, Gemini, Llama, Mistral, and Gemma) to generate academic content based on 40 papers on Digital Twin and Healthcare.

Main result

The study found that "paraphrased abstracts showed higher similarity rates, while question-based responses also exceeded acceptable levels" and that "AI detection tools consistently identified all outputs as AI-generated." Additionally, "all chatbots produced a sufficient volume of content" and "semantic similarity tests showed a strong overlap between generated and original texts," though "readability assessments indicated that the texts were insufficient in terms of clarity and accessibility."

Research paradigm

empirical/comparative analysis

Author conclusions

The authors conclude that "While these models generate substantial and semantically accurate content, concerns regarding plagiarism, AI detection, and readability must be addressed for their effective use in scholarly work." They also note that the study "comparatively highlights the potential and limitations of popular and latest large language models for academic writing."

Risk of bias

Selection bias in choice of 40 papers (all from Digital Twin and Healthcare domains); Potential bias in tool selection (comparing proprietary models like ChatGPT and Gemini with open-source models); Lack of human expert evaluation panel mentioned; No discussion of inter-rater reliability for qualitative assessments; Selection bias in choice of 40 papers on specific topics (Digital Twin and Healthcare); Potential variation in model versions and configurations not detailed; AI detection tool reliability and potential biases in detection algorithms; Limited specification of prompt engineering and generation parameters; Selection bias: Papers limited to Digital Twin and Healthcare domain, not representative of all academic writing; Sample size: Only 40 papers used, potentially insufficient for robust comparative conclusions; Tool selection bias: Choice of specific AI detection and readability assessment tools may influence results; No specification of paper selection criteria or randomization process; Potential publication bias as this is a preprint from SSRN

Limitations

  • The authors note that "While these models generate substantial and semantically accurate content, concerns regarding plagiarism, AI detection, and readability must be addressed for their effective use in scholarly work." The abstract indicates there is "a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generate original content in terms of quality, readability, and effectiveness."

Open questions raised

  • The authors identify "a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generate original content in terms of quality, readability, and effectiveness."
  • The authors identify a critical gap: "There is a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generate original content in terms of quality, readability, and effectiveness."
  • The abstract identifies "a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generate original content in terms of quality, readability, and effectiveness."
Data: not_statedCode: not_statedExtracted from: pdfAgreement 57%

Explore related topics

Related papers