12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts

Joost de Winter · Scientometrics · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
29
Citations
2.86
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-024-04939-y

Methodology & findings

Study design

Observational empirical study using automated computational analysis.

Sample

N = 2222, 2 groups

Primary method

Principal Component Analysis (PCA) with Varimax rotation on standardized 2222 × 60 matrix of ChatGPT-4 item scores; Spearman rank-order correlation coefficients for associations with citations and altmetrics (chosen for robustness against outliers); Scree plot analysis to determine number of components to retain; Split-half reliability approach: component scores calculated from runs 1-2 and runs 3-4 separately; MATLAB R2021b for custom scripting and data submission to OpenAI API; Readability API by Ipeirotis (2023) for computing classical readability indices; Software: Scopus, Altmetric, Dimensions, Publish or Perish (for citation extraction)

Main result

The study found that "the Accessibility and Understandability component demonstrated weak to moderate positive correlations with several altmetrics scores, including the number of news mentions (ρ = 0.10), blog mentions (ρ = 0.07), Twitter mentions (ρ = 0.18), and Reddit mentions (ρ = 0.06), as well as a large correlation with the number of Mendeley readers (ρ = 0.40)." Additionally, "the predictive correlations of ChatGPT-based assessments surpassed traditional readability metrics," with the Novelty and Engagement component showing moderate positive correlations with citation counts (ρ = 0.18).

Reports effect sizes.

Research paradigm

Quantitative empiricism with computational methods

Author conclusions

The authors conclude: "This study demonstrated the potential application of large language models in scientometrics. Our study may also stimulate reflection from a broader perspective." They further note that "the current study pioneered a language-based evaluation of article abstracts," and observe that "in contrast to conventional scientometrics methods—which may be prone to overfitting and inconsistent results, as discussed in the Introduction—the present approach represents an innovative direction." However, they also caution: "It is important to recognize that these assessments arise from a model trained on previously generated human data," and emphasize that "as scientists strive to maintain academic integrity and autonomy, understanding who determines the quality of a scientific work remains a critical topic."

Risk of bias

Selection bias: Articles limited to PLOS ONE (open-access mega-journal), limiting generalizability to other journals; Temporal confounding: Papers from early 2022 chosen to avoid ChatGPT-4 knowledge cutoff contamination, but short follow-up period (22-24 months); Response bias: ChatGPT-4 ratings showed low reliability for individual abstracts (average r=0.27), though reliability was high at population level (r>0.998); Abstract representativeness: Unclear whether abstract characteristics reflect full paper quality or impact; Order effects: ChatGPT-4 output sensitive to item presentation order despite random ordering attempt; Order effects in item presentation (mitigated by randomizing item order); Potential model contamination from training data (mitigated by selecting articles published after knowledge cutoff); Limited reliability at individual abstract level (r = 0.27 average for individual abstract scores); Temporal bias: short follow-up period (22-24 months only); Potential unrepresentativeness of abstract for entire paper quality; Selection bias: Only PLOS ONE articles included (megajournal, open-access); Temporal bias: Articles from only January-February 2022 analyzed; Outcome measurement bias: Different citation databases (Dimensions, Scopus, Google Scholar) may have different coverage; Low reliability for individual abstracts: ChatGPT scores unreliable at individual abstract level (r=0.27 average), only reliable at population level; Order effects: ChatGPT-4 output sensitive to order of items presented despite randomization attempts; Potential 'contamination' risk minimized but not eliminated: Knowledge cutoff September 2021, but analysis conducted April 2023; Attrition: Some articles missing Mendeley reader data (manually inserted from Mendeley database)

Limitations

  • The authors state that "while 2222 abstracts were analyzed, it would have been preferable to process a larger dataset" and note that "Our current method was not particularly cost-effective, amounting to approximately 0.04 USD per abstract." They also acknowledge that "our analysis was based on abstracts from January and February 2022, just after the knowledge cut-off of ChatGPT-4
  • Consequently, the papers have had only 22–24 months to accumulate citations and altmetrics scores." Additionally, "the strength of the correlations was found to be only weak to moderate, with Spearman's ρ values ranging between 0.08 and 0.18."

Open questions raised

  • Need for larger datasets to assess scalability beyond 2222 abstracts
  • Future research could benefit from analyzing full-text articles instead of abstracts alone
  • Need to clarify whether abstract characteristics are causal factors or epiphenomena of overall paper quality
  • Exploration of different types of prompts and more precise ChatGPT evaluations
  • Need for reassessment of abstracts as ChatGPT models improve
  • Investigation of domain-specific variations in readability effects across disciplines
Data: Raw data and MATLAB scripts available at https://doi.org/10.4121/71058da-ed2e-4d36-b8e4-ad02c3af1e65; Raw data and MATLAB scriptsCode: MATLAB R2021b script used for analysis available at https://doi.org/10.4121/71058da-ed2e-4d36-b8e4-ad02c3af1e65; MATLAB scripts for analysisExtracted from: pdfAgreement 49%

Explore related topics

Related papers