AI Research Is Not Magic, It Has to Be Reproducible and Responsible: Challenges in the AI Field from the Perspective of Its PhD Students
Andrea Hrčková, Jennifer Renoux, Rafael Tolosana‐Calasanz, Daniela Chudá, Martin Tamajka, Jakub Šimko · Lecture notes in computer science · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/978-3-032-11108-1_34
Methodology & findings
Study design
Exploratory qualitative research design using semi-structured, in-depth focus group interviews.
Sample
N = 28, 7 groups
Primary method
Qualitative content analysis with inductive coding. No inferential statistics, frequentist testing, or Bayesian methods employed. Findings presented as frequency counts (e.g., 'X groups out of 11') and thematic categories visualized in mind maps using Canva.com.
Main result
The study identified three critical areas where current AI research practices fall short: "(1) the findability and quality of AI resources such as datasets, models, and experiments; (2) the difficulties in replicating the experiments in AI papers; (3) and the lack of trustworthiness and interdisciplinarity." Among the major challenges, PhD students reported that "Sometimes you don't understand how they gained such results in the paper and how to replicate it" (P1), highlighting fundamental reproducibility obstacles. Additionally, respondents noted that "Papers do not mention data drifts, there is usually no information about the data preparation" (P3), indicating systematic gaps in documentation quality.
Reports effect sizes.
Research paradigm
Qualitative empiricism; interpretive/constructivist approach to understanding PhD student experiences
Author conclusions
The authors conclude: "Good research is collective and multidisciplinary, building on prior work to advance knowledge. In order to produce good research efficiently, this prior work must be usable - discoverable, reproducible, and responsible for researchers. Our study of AI PhD students shows this is not yet the case. Despite growing awareness and initiatives, more effort is needed to improve AI research practices. Efforts must be both technological—developing platforms for discoverability and reproducibility—and societal—encouraging multidisciplinarity and trustworthiness."
Risk of bias
Selection bias: Participants recruited through European project partners without random sampling; Sample size: 28 PhD students from only 13 European countries limits generalizability; Geographic limitation: European-only sample may not represent global AI research challenges; Potential response bias: Participants self-selected through project partner networks; Interviewer bias: Semi-structured format with selective question posing based on expertise; Coding bias: Initial coding by first author with validation by subset of authors; Selection bias: Participants were recruited through European project partners without random sampling, potentially limiting representativeness; Sample bias: 28 PhD students from 13 European countries—geographically concentrated and non-representative globally; Response bias: Participants self-selected through project partnerships; may represent researchers already engaged with reproducibility concerns; Interviewer bias: Only first and fourth authors conducted interviews; potential for subjective question selection and interpretation; Coding bias: While validation was performed by fourth and fifth authors, inductive coding is inherently subjective with risk of category drift; Small sample size (n=28) limiting generalizability; Geographic limitation: Only European PhD students included; Potential social desirability bias in focus group settings; Interviewer bias: Different authors conducted interviews with selective question presentation based on expertise
Limitations
- "Our study is limited to European PhD students and a moderate sample size, restricting quantification
- Broader validation is needed to assess these challenges globally." The authors also acknowledge that "addressing [systemic barriers to adoption of recommendations] lies beyond this paper's scope" and note that "interview data cannot be openly published" due to ethical protections, limiting external validation of coding and findings.
Open questions raised
- Global validation of identified challenges beyond European context
- Effects of irreproducibility on the AI community (noted as underexplored)
- Impact of reproducibility initiatives on researcher practices and outcomes
- Development of AI-specific tools for version control and experiment management
- Standardization of ethical assessment frameworks in AI research
- Mechanisms to incentivize researchers to create high-quality documentation and curated resources
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Challenges and Opportunities of Generative AI for Higher Education as Explained by ChatGPTRosario Michel‐Villarreal · 2023 · 749 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- What ChatGPT means for universities: Perceptions of scholars and studentsMehmet Fırat · 2023 · 405 citations