12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI Research Is Not Magic, It Has to Be Reproducible and Responsible: Challenges in the AI Field from the Perspective of Its PhD Students

Andrea Hrčková, Jennifer Renoux, Rafael Tolosana‐Calasanz, Daniela Chudá, Martin Tamajka, Jakub Šimko · Lecture notes in computer science · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/978-3-032-11108-1_34

Methodology & findings

Study design

Exploratory qualitative research design using semi-structured, in-depth focus group interviews.

Sample

N = 28, 7 groups

Primary method

Qualitative content analysis with inductive coding. No inferential statistics, frequentist testing, or Bayesian methods employed. Findings presented as frequency counts (e.g., 'X groups out of 11') and thematic categories visualized in mind maps using Canva.com.

Main result

The study identified three critical areas where current AI research practices fall short: "(1) the findability and quality of AI resources such as datasets, models, and experiments; (2) the difficulties in replicating the experiments in AI papers; (3) and the lack of trustworthiness and interdisciplinarity." Among the major challenges, PhD students reported that "Sometimes you don't understand how they gained such results in the paper and how to replicate it" (P1), highlighting fundamental reproducibility obstacles. Additionally, respondents noted that "Papers do not mention data drifts, there is usually no information about the data preparation" (P3), indicating systematic gaps in documentation quality.

Reports effect sizes.

Research paradigm

Qualitative empiricism; interpretive/constructivist approach to understanding PhD student experiences

Author conclusions

The authors conclude: "Good research is collective and multidisciplinary, building on prior work to advance knowledge. In order to produce good research efficiently, this prior work must be usable - discoverable, reproducible, and responsible for researchers. Our study of AI PhD students shows this is not yet the case. Despite growing awareness and initiatives, more effort is needed to improve AI research practices. Efforts must be both technological—developing platforms for discoverability and reproducibility—and societal—encouraging multidisciplinarity and trustworthiness."

Risk of bias

Selection bias: Participants recruited through European project partners without random sampling; Sample size: 28 PhD students from only 13 European countries limits generalizability; Geographic limitation: European-only sample may not represent global AI research challenges; Potential response bias: Participants self-selected through project partner networks; Interviewer bias: Semi-structured format with selective question posing based on expertise; Coding bias: Initial coding by first author with validation by subset of authors; Selection bias: Participants were recruited through European project partners without random sampling, potentially limiting representativeness; Sample bias: 28 PhD students from 13 European countries—geographically concentrated and non-representative globally; Response bias: Participants self-selected through project partnerships; may represent researchers already engaged with reproducibility concerns; Interviewer bias: Only first and fourth authors conducted interviews; potential for subjective question selection and interpretation; Coding bias: While validation was performed by fourth and fifth authors, inductive coding is inherently subjective with risk of category drift; Small sample size (n=28) limiting generalizability; Geographic limitation: Only European PhD students included; Potential social desirability bias in focus group settings; Interviewer bias: Different authors conducted interviews with selective question presentation based on expertise

Limitations

  • "Our study is limited to European PhD students and a moderate sample size, restricting quantification
  • Broader validation is needed to assess these challenges globally." The authors also acknowledge that "addressing [systemic barriers to adoption of recommendations] lies beyond this paper's scope" and note that "interview data cannot be openly published" due to ethical protections, limiting external validation of coding and findings.

Open questions raised

  • Global validation of identified challenges beyond European context
  • Effects of irreproducibility on the AI community (noted as underexplored)
  • Impact of reproducibility initiatives on researcher practices and outcomes
  • Development of AI-specific tools for version control and experiment management
  • Standardization of ethical assessment frameworks in AI research
  • Mechanisms to incentivize researchers to create high-quality documentation and curated resources
Data: Interview questions available at: https://zenodo.org/records/15920758; Interview data not openly published per ethical protections. Open-ended interview questions made available at: https://zenodo.org/records/15920758Code: None explicitly mentioned. Authors used AI tools (ChatGPT and Gemini) for language refinement only.; Interview questions available at https://zenodo.org/records/15920758Extracted from: pdfAgreement 58%

Explore related topics

Related papers