Assessing the sustainable development of a national research ecosystem: A generative AI-based evaluation of empirical educational research in China (2004–2023)
Sen Wang · PLoS ONE · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1371/journal.pone.0341620
Methodology & findings
Study design
Generative AI-based longitudinal ecosystem assessment using GPT-4o to score 2,145 empirical educational research papers published in four leading Chinese CSSCI-indexed journals between 2004-2023.
Sample
N = 2145, 3 groups
Primary method
Bland-Altman analysis for stability validation (comparing two pilot assessments); Kendall's W consistency test for inter-rater and inter-model reliability; Spearman correlation analysis for validity comparison with expert scores; Fuzzy comprehensive evaluation method; CRITIC (Criteria Importance Through Intercriteria Correlation) method for objective weighting; Fuzzy relation matrix construction; Chi-square testing (reported: χ² = 115.461)
Main result
The overall sustainability index of China's empirical educational research ecosystem over the 20-year period is 75.77 on a 100-point scale, with membership degrees concentrated at quality levels 7 (0.328) and 8 (0.435), indicating a generally robust and maturing system. "The ecosystem shows strong responsiveness to real-world educational problems, with high average scores for the relevance (8.45) and social significance (8.23) of research questions, as well as generally solid research design and data analysis practices." However, "relatively lower scores for transparency of data analysis (7.08) and accessibility of raw data (6.46) highlight persistent challenges for reproducibility, open science, and methodological innovation."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist with systems theory foundations
Author conclusions
"This study used a generative AI-driven methodology to conduct a 20-year longitudinal sustainability assessment of China's empirical educational research ecosystem (2004–2023)." The results show that "the ecosystem has reached a relatively high and stable level of performance but faces critical tasks in strengthening data openness, methodological renewal, and AI-augmented governance." The authors conclude that China's ecosystem "is at a pivotal juncture in its transition from a policy-driven to a quality-first, open-science paradigm." They assert that "the proposed generative AI–based evaluation framework may offer a scalable tool for continuous monitoring and governance of national research ecosystems, while its results should be interpreted as an auxiliary input rather than a substitute for expert peer assessment."
Risk of bias
Latent biases in GPT-4o training data; Selection bias: sample limited to four domestic Chinese journals, excluding international publications and non-journal outlets; Potential algorithmic bias in AI scoring; Black-box nature of AI evaluation limits transparency regarding scoring mechanisms; AI model bias from training data latent biases; Selection bias from focus on four domestic CSSCI-indexed journals only; Exclusion of international publication venues and book publications; Framework bias: omission of academic ethics dimensions assumed handled by pre-publication processes; Potential algorithmic bias in GAI scoring mechanisms; Single-model dependency despite multi-model validation checks; Black box nature of GAI evaluation with potential latent biases in training data; Selection bias from sampling only four domestic leading journals (excludes international publications and books); Potential algorithmic bias in GPT-4o scoring affecting reproducibility; Framework construction bias: omission of academic ethics dimension; Assumption that journal gatekeeping processes adequately screened for ethical issues
Limitations
- "First, the GAI evaluation, while validated against expert scores and multiple models, operates as a 'black box' to some extent
- its scoring may be influenced by latent biases within its training data
- Second, our sample, while representative of high-level research in China, was drawn from four leading domestic journals
- This focus necessarily excludes research published in international venues, books, and other outlets, meaning our findings are representative of the core of the domestic ecosystem but may not be generalizable to its entire intellectual output
- Third, the evaluation framework itself, though comprehensive, is a constructed representation of research quality and omits certain dimensions, like academic ethics, which were presumed to have been handled by the journals' pre-publication processes."
Open questions raised
- Comparative cross-national analyses assessing developmental trajectories of educational research ecosystems in other countries or different academic disciplines
- Fine-tuning large language models on expert-annotated papers to enhance nuance of AI assessments
- Continued longitudinal tracking beyond the 20-year baseline for real-time monitoring of ecosystem health and assessment of future policy interventions
- Integration of research published in international venues and non-journal outlets to achieve fuller picture of entire intellectual ecosystem
- Development of more robust methodological innovation indicators
- Limited integration of Chinese empirical educational research into mainstream international journals
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations