Can LLMs Produce Original Astronomy Research in a Semester? A Graduate Class Experiment
Ann Zabludoff, Chen-Yu Chuang, Parker Thomas Johnson, Yichen Liu, Brina Bianca Martinez, Neev Shah et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Qualitative case study of a graduate seminar course (ASTR 540 Structure and Dynamics of Galaxies) conducted in Fall 2025.
Main result
The study found that students successfully completed original research projects with LLM assistance, though results were mixed. Key findings include: "The LLMs allowed them to efficiently synthesize relevant literature, identify key papers, and generate research questions and methodologies, which significantly sped up those initial stages of the project." However, "the failures of the LLMs fell into several general categories," including struggles with "generating complex functional code for data analysis and physically accurate code for simulations, requiring extensive manual correction," and frequent citation errors where "the LLMs frequently-about 20% of the time-created links to the wrong sources or misattributed papers, requiring extensive manual review." By semester's end, all students completed research projects with novel results, but students reported they "would employ LLMs for research again, but in limited applications."
Research paradigm
Qualitative/Interpretive - Educational case study with reflective analysis
Author conclusions
By semester's end, "each student had completed a draft research paper on Overleaf...All papers included the results of novel data analyses or simulations and addressed an astronomy question not yet answered." Students concluded they would "employ LLMs for research again, but in limited applications and, for one student, only as a last resort," specifically for "1) to find papers in the literature if they did not know where to start, 2) to fix minor (e.g., syntax) errors or expand on code...and 3) in cases where they had experience iterating over LLM outputs, to generate plotting or data analysis scripts." The authors recommend that "students are trained to succeed in any career they wish to pursue. Students that adapt to use LLMs expertly, creatively, and efficiently will excel." They note that "In retrospect, students would have benefited from one class period and/or several early homework problems devoted to understanding LLMs, their limitations, and the best current practices for their use in research."
Risk of bias
Selection bias: Sample limited to 7 students from single institution with strong physics/math backgrounds; Recall bias: Students relied on memory for retrospective reporting of LLM usage patterns and time spent; Social desirability bias: Student responses may reflect perceived instructor expectations; Temporal specificity: LLM capabilities change rapidly; findings specific to models available Fall 2025; Survivorship bias: Only completed projects reported; no data on students who abandoned LLM-assisted approaches; Expertise heterogeneity: Prior LLM experience varied from 'limited to moderate'; Field heterogeneity: 70% of students worked in non-galaxy astronomy, limiting domain-specific conclusions; Selection bias: Seven volunteers in a single institution, possibly those more amenable to LLM use; Instructor bias: Assignment design by instructors may influence expectations and outcomes; Self-report bias: Student reflections on time spent and productivity may be subjective; Temporal bias: Rapid LLM landscape evolution makes findings potentially outdated; Generalization limitation: Results from astronomy students may not apply to other disciplines; Survivor bias: Only students who completed projects are represented; Selection bias: All students were first-year PhD students from a single institution with pre-existing strong physics/mathematics backgrounds; Self-report bias: Data derived entirely from student reflections; no independent observation of LLM usage; Novelty bias: Student experiences may reflect initial learning curve rather than stable LLM performance; Temporal bias: Fall 2025 represents a specific moment in LLM development; models may have improved since; Instructor bias: Course design and grading expectations not discussed; may influence student LLM usage patterns; Publication bias: Only successful or notable examples documented; unclear if all student experiences captured equally; Expertise confound: Most students (5/7) were new to galaxy research, confounding LLM effect with domain learning
Limitations
- This was "the first implementation of using LLMs to complete original research over a semester in a graduate course in the University of Arizona Astronomy department." The sample is small (n=7 students), all from a single institution and cohort, limiting generalizability
- Most students (5/7, ~70%) were pursuing PhD-track research outside galaxy science, potentially affecting relevance of findings to other astronomy subfields
- The paper relies on student self-report of LLM usage and time spent
- Students reported that "some of the problems encountered by students in Fall 2025 and discussed here, such as frequent coding and citation inaccuracies, may be resolved before the next academic cycle," indicating results may be time-specific to Fall 2025 LLM capabilities
- The course did not include "one class period and/or several early homework problems devoted to understanding LLMs, their limitations, and the best current practices for their use," which may have affected outcomes.
Open questions raised
- Need for LLM training for future graduate students: "students will be encouraged early in the term to review relevant and free online resources"
- Improvement needed in LLM code generation for complex analysis and simulations
- Need for LLM capability to access and query online astronomical data archives and APIs autonomously
- Improved handling of citations and elimination of hallucinated references in scientific contexts
- Better LLM ability to identify truly novel research questions versus replicating prior work
- Broader classroom discussion needed on "ethical and practical considerations that accompany the heavy use of LLMs," including "water and energy requirements for data centers, the environmental impacts of those centers, implicit biases included in training LLMs, corporate conflicts of interest, and lagging government regulation"
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations