A Multidisciplinary Bibliometric Analysis of Differences and Commonalities Between GenAI in Science
Kacper Sieciński, Marian Oliński · Publications · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/publications13040067
Methodology & findings
Study design
Bibliometric analysis using Web of Science Core Collection database (2023-2025).
Sample
N = 14418, 10 groups
Primary method
Descriptive statistics (percentages, shares, growth rates); Co-citation network analysis in VOSviewer; Co-authorship network analysis with 75 countries; Keyword co-occurrence analysis with minimum occurrence thresholds adjusted to yield approximately 300 input terms per tool line; Jaccard index similarity calculations: |A ∩ B| / |A ∪ B| for pairwise tool comparisons; Sensitivity analysis of Jaccard matrices at Top-50, Top-100, and Top-200 keywords; h-index and i10-index calculations for impact assessment; Logarithmic function fitting for publication trend projection: y = 4050ln(x) + 2387.1
Main result
The study found that "ChatGPT accounts for approximately 80.6% of the total pool" of publications analyzed, with "Gemini ranks second with 17.0%, followed by Claude (7.5%) and LLaMA (7.4%)." The analysis reveals "a substantial increase in the share of publications in 2025 can be observed for all analyzed GenAI lines, making 2025 a clear turning point." Additionally, "the Jaccard matrices reveal two clear vectors of thematic convergence (general-purpose chatbots and open-source lines), intermediate search-oriented tools, and a handful of outliers."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
"The analysis shows differences in volume and citations, with ChatGPT receiving the most attention, as well as accelerated growth among later entrants. It also points to a split in the ecosystem into two modes of development, which distinguishes general purpose tools from open-source lines with a more engineering and methodological profile." Furthermore, "By combining publication dynamics, citation concentration, subject area profiles, and keyword-based Jaccard matrices, it describes a structured way to group tools into clusters, to identify bridge and outlier positions, and to assess the extent of topical overlap across different model lines. This framework can be applied to domain specific specialized models and to future generations of GenAI systems."
Risk of bias
Single database limitation (Web of Science only); Indexing delays affecting recency of data; Citation window bias favoring earlier-introduced tools; Domain imbalance from uneven subject area profiles; Temporal bias: shorter citation windows for newer tools (DeepSeek, Qwen, Grok); Keyword-based analysis may not capture full article contents; Lack of differentiation between model versions and configurations; Single database limitation (Web of Science only) may miss articles in other indexing systems; Indexing delays in Web of Science may underrepresent recent 2025 publications; Citation window bias: newer tools (DeepSeek, Qwen, Grok) have shorter citation windows, making citation metrics non-comparable; Disciplinary profile imbalances across tool corpora affect cross-tool comparisons; Keyword filtering applied post-hoc may introduce subjective bias in normalization and cleaning; Model versions not differentiated—aggregating across versions may obscure important differences; Ambiguity in tool names (e.g., 'Gemini' also astrological/astronomical term; 'Llama/Alpaca' zoological terms) despite exclusion efforts; Single database dependency (Web of Science only) may exclude relevant publications indexed elsewhere; Indexing delays and coverage gaps for 2025 publications; Citation window bias: later-introduced tools have shorter citation windows, artificially depressing their citation metrics; Subject area profile imbalances across tool corpora may confound cross-tool comparisons; Keyword-based thematic analysis may not capture full article content or nuanced topics; Model version heterogeneity within tool lines not differentiated in analysis; Recency effects for tools introduced late (DeepSeek, Qwen, Grok in 2025); Geographic concentration bias toward English-language publications indexed in Web of Science
Limitations
- "This study relies on a single database (Web of Science) and a 2023-2025 horizon, which limits coverage, exposes the results to indexing delays, and constrains the normalization of citation indicators
- The thematic analyses use keywords rather than full texts and therefore may not capture entire article contents
- The comparisons do not differentiate between model versions or detailed configurations, and some differences may arise from mixed disciplinary profiles within individual corpora
- Specialized tools (e.g., Med-PaLM 2, BioGPT, ClinicalGPT) were not examined in depth."
Open questions raised
- Expansion to additional databases beyond Web of Science
- Integration of preprint data sources
- Full-text models for topic discovery and semantic similarity
- Citation normalization by field and year
- Tracking of model versioning in publications
- Study of co-use of multiple tools within single publications
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- AI chatbots in programming education: Students’ use in a scientific computing course and consequences for learningS.E.A. Groothuijsen · 2024 · 65 citations