Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts
Joost de Winter · Scientometrics · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-024-04939-y
Methodology & findings
Study design
Observational empirical study using automated computational analysis.
Sample
N = 2222, 2 groups
Primary method
Principal Component Analysis (PCA) with Varimax rotation on standardized 2222 × 60 matrix of ChatGPT-4 item scores; Spearman rank-order correlation coefficients for associations with citations and altmetrics (chosen for robustness against outliers); Scree plot analysis to determine number of components to retain; Split-half reliability approach: component scores calculated from runs 1-2 and runs 3-4 separately; MATLAB R2021b for custom scripting and data submission to OpenAI API; Readability API by Ipeirotis (2023) for computing classical readability indices; Software: Scopus, Altmetric, Dimensions, Publish or Perish (for citation extraction)
Main result
The study found that "the Accessibility and Understandability component demonstrated weak to moderate positive correlations with several altmetrics scores, including the number of news mentions (ρ = 0.10), blog mentions (ρ = 0.07), Twitter mentions (ρ = 0.18), and Reddit mentions (ρ = 0.06), as well as a large correlation with the number of Mendeley readers (ρ = 0.40)." Additionally, "the predictive correlations of ChatGPT-based assessments surpassed traditional readability metrics," with the Novelty and Engagement component showing moderate positive correlations with citation counts (ρ = 0.18).
Reports effect sizes.
Research paradigm
Quantitative empiricism with computational methods
Author conclusions
The authors conclude: "This study demonstrated the potential application of large language models in scientometrics. Our study may also stimulate reflection from a broader perspective." They further note that "the current study pioneered a language-based evaluation of article abstracts," and observe that "in contrast to conventional scientometrics methods—which may be prone to overfitting and inconsistent results, as discussed in the Introduction—the present approach represents an innovative direction." However, they also caution: "It is important to recognize that these assessments arise from a model trained on previously generated human data," and emphasize that "as scientists strive to maintain academic integrity and autonomy, understanding who determines the quality of a scientific work remains a critical topic."
Risk of bias
Selection bias: Articles limited to PLOS ONE (open-access mega-journal), limiting generalizability to other journals; Temporal confounding: Papers from early 2022 chosen to avoid ChatGPT-4 knowledge cutoff contamination, but short follow-up period (22-24 months); Response bias: ChatGPT-4 ratings showed low reliability for individual abstracts (average r=0.27), though reliability was high at population level (r>0.998); Abstract representativeness: Unclear whether abstract characteristics reflect full paper quality or impact; Order effects: ChatGPT-4 output sensitive to item presentation order despite random ordering attempt; Order effects in item presentation (mitigated by randomizing item order); Potential model contamination from training data (mitigated by selecting articles published after knowledge cutoff); Limited reliability at individual abstract level (r = 0.27 average for individual abstract scores); Temporal bias: short follow-up period (22-24 months only); Potential unrepresentativeness of abstract for entire paper quality; Selection bias: Only PLOS ONE articles included (megajournal, open-access); Temporal bias: Articles from only January-February 2022 analyzed; Outcome measurement bias: Different citation databases (Dimensions, Scopus, Google Scholar) may have different coverage; Low reliability for individual abstracts: ChatGPT scores unreliable at individual abstract level (r=0.27 average), only reliable at population level; Order effects: ChatGPT-4 output sensitive to order of items presented despite randomization attempts; Potential 'contamination' risk minimized but not eliminated: Knowledge cutoff September 2021, but analysis conducted April 2023; Attrition: Some articles missing Mendeley reader data (manually inserted from Mendeley database)
Limitations
- The authors state that "while 2222 abstracts were analyzed, it would have been preferable to process a larger dataset" and note that "Our current method was not particularly cost-effective, amounting to approximately 0.04 USD per abstract." They also acknowledge that "our analysis was based on abstracts from January and February 2022, just after the knowledge cut-off of ChatGPT-4
- Consequently, the papers have had only 22–24 months to accumulate citations and altmetrics scores." Additionally, "the strength of the correlations was found to be only weak to moderate, with Spearman's ρ values ranging between 0.08 and 0.18."
Open questions raised
- Need for larger datasets to assess scalability beyond 2222 abstracts
- Future research could benefit from analyzing full-text articles instead of abstracts alone
- Need to clarify whether abstract characteristics are causal factors or epiphenomena of overall paper quality
- Exploration of different types of prompts and more precise ChatGPT evaluations
- Need for reassessment of abstracts as ChatGPT models improve
- Investigation of domain-specific variations in readability effects across disciplines
Explore related topics
Related papers
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Unlocking the Power of ChatGPT: A Framework for Applying Generative AI in EducationJiahong Su · 2023 · 550 citations
- Generative AI and the future of higher education: a threat to academic integrity or reformation? Evidence from multicultural perspectivesAbdullahi Yusuf · 2024 · 399 citations