On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence
Kay Smarsly · Publications · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/publications14020023
Methodology & findings
Study design
Mixed-methods approach combining: (1) Modified PRISMA-compliant systematic literature review with four phases (identification, screening, eligibility, snowballing) across four research questions examining citation metric manipulation, formal evaluation procedures, civil engineering domain specifics, and LLM usage in scientific papers.
Main result
The study found that "under the platform-specific conditions, indexing of LLM-assisted, strategically cited documents on Google Scholar is associated with increases in author-level citation metrics." Specifically, after Google Scholar indexed paper A on 19 June 2025, the citation count increased from 2555 to 2615 (+60 citations, 50 of which were attributed to paper A). Upon indexing of paper B on 15 July 2025, the total citation count rose from 2629 to 2715 (+86 citations, including 75 from paper B). The systematic review also documented that "manipulation of citation metrics is perpetrated through different methods" including "excessive self-citation, where authors deliberately cite their own publications to artificially inflate personal citation counts" and "citation cartels, i.e., groups of researchers who systematically and disproportionately cite each other's work."
Research paradigm
Critical realism with empirical validation
Author conclusions
The authors conclude: "Generative artificial intelligence, particularly large-language-model-based chat bots, substantially increases the vulnerability of citation metrics to manipulation in academic evaluation. While citation-based indicators, such as the h-index, i10-index and total citation counts, have historically been susceptible to manipulation through self-citation and strategic behaviors, generative AI introduces qualitatively new vulnerabilities due to its unprecedented scale and automation capabilities." They further state that "a critical vulnerability identified in this study is the possibility for researchers to manipulate their own citation metrics directly, which becomes particularly critical when evaluation committees rely on non-curated bibliographic databases (such as Google Scholar) that index all content without or with limited quality control." Additionally, they emphasize that "bibliographic databases might fail to detect manipulation patterns, as continuous uploads of fabricated papers do not cause suspicious spikes in the citation trajectory, rendering such manipulations difficult to detect."
Risk of bias
Selection bias in systematic review due to database restriction (Scopus only, searches June 21-23, 2026); Language bias: only English-language publications included; Publication bias: only articles, reviews, and conference papers; editorials and commentaries excluded; Single-platform limitation in case study (Google Scholar only); Single author profile tested (n=1), limiting generalizability; Temporal specificity: case study conducted in specific time period (May-July 2025); Confounding variables in case study: concurrent indexing of multiple papers, simultaneous shifts in citation practices; Potential attrition in snowballing phase: unclear how many records were identified and excluded at each iteration; Selection bias in systematic review: articles from English-language sources only and from specific document types (articles, reviews, conference papers); Platform specificity: findings limited to Google Scholar, which may not generalize to curated databases like Scopus or Web of Science; Single author profile in case study: n=1 limits generalizability; Temporal specificity: searches conducted June 2026; literature landscape may shift; Citation practices bias: both papers contained 'numerous citations to the author's prior work for scholarly reasons', creating confounding between legitimate scholarly citation and potential metric manipulation; No control group in proof-of-concept: no comparison with equivalent documents not containing concentrated citations; Publication venues: papers presented at non-peer-reviewed workshop before indexing; Selection bias: The systematic review relied on Scopus database only, potentially missing relevant studies in other databases or non-indexed sources; Screening bias: Single-researcher methodology not clearly stated for all phases (inter-rater reliability not mentioned); Publication bias: The review captures only published/indexed research, not gray literature or failed replication attempts; Case study confounding: The proof-of-concept involved papers by 'a different individual' to avoid self-citations, but the author of the main study created the experimental papers, introducing potential conflict of interest; Temporal bias: The case study was conducted during a specific 2-month period (June-July 2025) with Google Scholar's evolving indexing algorithm potentially affecting generalizability; Selection of target: Using papers containing concentrated citations to the study author creates non-representative conditions for typical paper dissemination patterns
Limitations
- The authors explicitly state: "The proof-of-concept has been designed as a demonstrative case study on a single platform (Google Scholar) and a single target profile (n = 1 author)
- It does not support claims about prevalence, representativeness, or generalizability beyond these conditions." Furthermore, "the results do not support broader causal claims beyond this platform-specific demonstration." Additionally, "the case study is intentionally limited in scope and should be interpreted as a proof-of-concept demonstration rather than as evidence of prevalence or system-wide effects." The authors also note that "the observations should not be generalized beyond Google Scholar or beyond the specific proof-of-concept setup examined in this study."
Open questions raised
- The authors identify the following gaps and future research directions: The need for research on detection mechanisms for AI-generated fraudulent papers; investigation of broader platform-level vulnerabilities beyond Google Scholar; study of system-wide effects across multiple evaluation systems; examination of coordinated manipulation attempts; research on disciplinary and geographical variations in citation metric vulnerabilities; and evaluation of the effectiveness of proposed governance interventions and recommendations.
- Limited understanding of generative AI's role in scaling citation manipulation across multiple platforms beyond Google Scholar
- Insufficient knowledge of detection mechanisms in different bibliographic databases (Scopus, Web of Science)
- Gap in understanding how citation metric manipulation affects downstream decisions in different disciplinary contexts
- Need for longitudinal studies examining sustained effects of AI-generated content on citation metrics
- Lack of empirical data on prevalence of undisclosed LLM use in manuscripts across different journal types and impact levels
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations