Virtuous Machines: Towards Artificial General Science
Gabrielle Wehr, Reuben Rideaux, Jason M. Tangen, Jason B. Mattingley, Shane E. Ehrhardt · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.13421
Methodology & findings
Study design
Multi-component empirical system: (1) Autonomous hypothesis generation through literature search and validation; (2) Pre-registered experimental protocol design with power analysis; (3) Online cognitive experiments with human participants (N=288 for Study 1, 277 analyzable); (4) Automated data analysis with mixed-effects modeling; (5) Systematic result interpretation and theory refinement; (6) Automated visualization generation; (7) Manuscript drafting with citation validation; (8) AI-based peer review; (9) Document formatting in LaTeX/Word.
Sample
N = 277, 9 groups
Primary method
Mixed-effects linear and logistic regression models; split-half reliability analysis with Spearman-Brown correction; Pearson and Spearman correlations; circular statistics (von Mises distribution modeling, circular standard deviation); False Discovery Rate correction (Benjamini-Hochberg procedure, q=0.05); model comparison using Akaike Information Criterion (AIC); likelihood ratio tests; bootstrap confidence interval estimation (1000 resamples); multilevel modeling; quadratic vs. linear model comparison; sensitivity analyses using alternative precision measures. Software: R (pwr package for power analysis, rpy2 interface for Python), Python, mixed-effects modeling frameworks, custom code for circular statistics.
Main result
The system successfully conducted end-to-end empirical research with human participants, executing three complete studies "from initial conception to manuscript" in approximately 17 hours of runtime per study at an average cost of "∼$114 USD per research project (not including the human participant payments of ∼$4,500 USD for the current experiment)". Study 1 found "no correlation between individual performance patterns" in visual working memory and mental rotation tasks "despite both tasks showing expected difficulty effects; attributed to the established 'reliability paradox'". Study 2 found that "individuals with stronger imagery showed no greater carryover effects between trials, challenging theories that imagery and perception rely on common processing mechanisms". Study 3 found "negligible relationships and suggesting that apparent connections between visual-spatial tasks reflect general cognitive factors rather than specific shared processes".
Reports effect sizes and confidence intervals.
Research paradigm
Computational empiricism with autonomous agent-based scientific discovery
Author conclusions
"While capable of independent operation with minimal human intervention, the balance between autonomy and human collaboration offers distinct advantages depending on the research goal. Currently, humans provide most value to these systems in creative problem formulation, conceptual innovation, and ethical oversight—though the focal points for human contribution may shift as the technology advances." Furthermore: "This process mirrors cognitive development in children, who learn primarily through physical manipulation of objects and progressive refinement of their understanding. While current LLMs excel at pattern recognition within training data, they remain limited by their inability to autonomously expand beyond those boundaries. Here the system coupled internal representations—the exploration of latent connections between concepts—with external measurement through experimentation."
Risk of bias
High exclusion rates in online cognitive tasks (28.6% for Mental Rotation Task); Attrition bias from 30% anticipated online data collection dropout; Selection bias from restricted sample of highly engaged online participants; Measurement reliability issues: three of four slope parameters showed poor reliability; Model specification bias: mixture model assumptions validated across distributions; Multiple comparisons with False Discovery Rate correction applied; Anchoring bias in LLM systems: early-stage inaccuracies propagate downstream; Potential funding bias: system trained on corporate LLMs (Anthropic, OpenAI, xAI, Mistral, Google); Attrition: 23% initial exclusion from cognitive task battery (277 to 288 final sample); Selection bias: Online Prolific participant platform may not represent broader population; Measurement bias: Poor reliability of slope parameters (VWM delay slope r=-0.136, MRT accuracy slope r=0.160); Multiple comparisons: Multiple hypothesis testing across three studies without family-wise correction across studies; Model misspecification: Mixture model parameter validity concerns noted by prior literature; High exclusion rates: Mental Rotation Task 28.6% failure rate on attention checks suggests possible participant selection effects; Anchoring bias: System showed tendency to maintain early conceptualizations despite contradictory evidence; Anchoring bias in LLMs affecting early hypothesis formulation; Sample selection bias from stringent online participant screening; High exclusion rates in mental rotation task (28.6% of participants); Online testing environment variability affecting measurement precision; Potential restriction of range from exclusion criteria; Manual publication step introducing human oversight bias; Reliance on frontier LLMs with inherent training data biases
Limitations
- The paper explicitly states: "the system's ability to selectively filter relevant information while preserving focused knowledge representations of the necessary context enabled it to navigate the entire scientific workflow without conceptual drift
- Interestingly, many reasoning models generally degrade in their performance over extremely long chains, losing focus and coherence across iterations
- Sensitivity to early-stage accuracy also emerged as a challenge
- Poor question formulation or any conceptual errors introduced during hypothesis generation and methodological design propagate downstream -persisting through multiple verification cycles -at a detriment to research outcomes." Additionally: "The high exclusion rate in the Mental Rotation Task, where attention check failures eliminated a substantial proportion of participants, suggests that online cognitive assessment may introduce systematic biases that differentially affect individual differences measurement across tasks."
Open questions raised
- "Future studies should prioritize developing reliable individual difference measures through adaptive testing approaches that can dynamically adjust task difficulty to optimize between-subject variance while maintaining experimental control." "Longitudinal designs may prove particularly valuable for establishing the stability of individual differences patterns across time." "The integration of neural measures, such as event-related potentials or fMRI activation patterns, may offer complementary individual differences metrics." "Meta-analytic approaches are urgently needed to quantify true effect sizes across the literature, accounting for publication bias." "Future research should implement preregistered analytical plans with systematic comparison of alternative measurement approaches."
- Future research should: (1) Develop reliable individual difference measures through adaptive testing; (2) Implement longitudinal designs for stability assessment; (3) Integrate neural measures (ERP, fMRI) as complementary metrics; (4) Develop hybrid experimental paradigms combining VWM and mental rotation elements; (5) Conduct registered replication protocols; (6) Implement meta-analytic approaches to quantify true effect sizes accounting for publication bias; (7) Extend system beyond cognitive psychology to chemistry, biology, materials science through domain-specific toolset development.
- Need for extension of system to other scientific domains beyond cognitive psychology
- Integration with physical laboratory automation and robotics for chemistry/biology/materials science
- Enhancement of system's capacity for autonomous theory refinement
- Development of mechanisms for truly novel experimental methodology generation
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations