12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Virtuous Machines: Towards Artificial General Science

Gabrielle Wehr, Reuben Rideaux, Jason M. Tangen, Jason B. Mattingley, Shane E. Ehrhardt · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.13421

Methodology & findings

Study design

Multi-component empirical system: (1) Autonomous hypothesis generation through literature search and validation; (2) Pre-registered experimental protocol design with power analysis; (3) Online cognitive experiments with human participants (N=288 for Study 1, 277 analyzable); (4) Automated data analysis with mixed-effects modeling; (5) Systematic result interpretation and theory refinement; (6) Automated visualization generation; (7) Manuscript drafting with citation validation; (8) AI-based peer review; (9) Document formatting in LaTeX/Word.

Sample

N = 277, 9 groups

Primary method

Mixed-effects linear and logistic regression models; split-half reliability analysis with Spearman-Brown correction; Pearson and Spearman correlations; circular statistics (von Mises distribution modeling, circular standard deviation); False Discovery Rate correction (Benjamini-Hochberg procedure, q=0.05); model comparison using Akaike Information Criterion (AIC); likelihood ratio tests; bootstrap confidence interval estimation (1000 resamples); multilevel modeling; quadratic vs. linear model comparison; sensitivity analyses using alternative precision measures. Software: R (pwr package for power analysis, rpy2 interface for Python), Python, mixed-effects modeling frameworks, custom code for circular statistics.

Main result

The system successfully conducted end-to-end empirical research with human participants, executing three complete studies "from initial conception to manuscript" in approximately 17 hours of runtime per study at an average cost of "∼$114 USD per research project (not including the human participant payments of ∼$4,500 USD for the current experiment)". Study 1 found "no correlation between individual performance patterns" in visual working memory and mental rotation tasks "despite both tasks showing expected difficulty effects; attributed to the established 'reliability paradox'". Study 2 found that "individuals with stronger imagery showed no greater carryover effects between trials, challenging theories that imagery and perception rely on common processing mechanisms". Study 3 found "negligible relationships and suggesting that apparent connections between visual-spatial tasks reflect general cognitive factors rather than specific shared processes".

Reports effect sizes and confidence intervals.

Research paradigm

Computational empiricism with autonomous agent-based scientific discovery

Author conclusions

"While capable of independent operation with minimal human intervention, the balance between autonomy and human collaboration offers distinct advantages depending on the research goal. Currently, humans provide most value to these systems in creative problem formulation, conceptual innovation, and ethical oversight—though the focal points for human contribution may shift as the technology advances." Furthermore: "This process mirrors cognitive development in children, who learn primarily through physical manipulation of objects and progressive refinement of their understanding. While current LLMs excel at pattern recognition within training data, they remain limited by their inability to autonomously expand beyond those boundaries. Here the system coupled internal representations—the exploration of latent connections between concepts—with external measurement through experimentation."

Risk of bias

High exclusion rates in online cognitive tasks (28.6% for Mental Rotation Task); Attrition bias from 30% anticipated online data collection dropout; Selection bias from restricted sample of highly engaged online participants; Measurement reliability issues: three of four slope parameters showed poor reliability; Model specification bias: mixture model assumptions validated across distributions; Multiple comparisons with False Discovery Rate correction applied; Anchoring bias in LLM systems: early-stage inaccuracies propagate downstream; Potential funding bias: system trained on corporate LLMs (Anthropic, OpenAI, xAI, Mistral, Google); Attrition: 23% initial exclusion from cognitive task battery (277 to 288 final sample); Selection bias: Online Prolific participant platform may not represent broader population; Measurement bias: Poor reliability of slope parameters (VWM delay slope r=-0.136, MRT accuracy slope r=0.160); Multiple comparisons: Multiple hypothesis testing across three studies without family-wise correction across studies; Model misspecification: Mixture model parameter validity concerns noted by prior literature; High exclusion rates: Mental Rotation Task 28.6% failure rate on attention checks suggests possible participant selection effects; Anchoring bias: System showed tendency to maintain early conceptualizations despite contradictory evidence; Anchoring bias in LLMs affecting early hypothesis formulation; Sample selection bias from stringent online participant screening; High exclusion rates in mental rotation task (28.6% of participants); Online testing environment variability affecting measurement precision; Potential restriction of range from exclusion criteria; Manual publication step introducing human oversight bias; Reliance on frontier LLMs with inherent training data biases

Limitations

  • The paper explicitly states: "the system's ability to selectively filter relevant information while preserving focused knowledge representations of the necessary context enabled it to navigate the entire scientific workflow without conceptual drift
  • Interestingly, many reasoning models generally degrade in their performance over extremely long chains, losing focus and coherence across iterations
  • Sensitivity to early-stage accuracy also emerged as a challenge
  • Poor question formulation or any conceptual errors introduced during hypothesis generation and methodological design propagate downstream -persisting through multiple verification cycles -at a detriment to research outcomes." Additionally: "The high exclusion rate in the Mental Rotation Task, where attention check failures eliminated a substantial proportion of participants, suggests that online cognitive assessment may introduce systematic biases that differentially affect individual differences measurement across tasks."

Open questions raised

  • "Future studies should prioritize developing reliable individual difference measures through adaptive testing approaches that can dynamically adjust task difficulty to optimize between-subject variance while maintaining experimental control." "Longitudinal designs may prove particularly valuable for establishing the stability of individual differences patterns across time." "The integration of neural measures, such as event-related potentials or fMRI activation patterns, may offer complementary individual differences metrics." "Meta-analytic approaches are urgently needed to quantify true effect sizes across the literature, accounting for publication bias." "Future research should implement preregistered analytical plans with systematic comparison of alternative measurement approaches."
  • Future research should: (1) Develop reliable individual difference measures through adaptive testing; (2) Implement longitudinal designs for stability assessment; (3) Integrate neural measures (ERP, fMRI) as complementary metrics; (4) Develop hybrid experimental paradigms combining VWM and mental rotation elements; (5) Conduct registered replication protocols; (6) Implement meta-analytic approaches to quantify true effect sizes accounting for publication bias; (7) Extend system beyond cognitive psychology to chemistry, biology, materials science through domain-specific toolset development.
  • Need for extension of system to other scientific domains beyond cognitive psychology
  • Integration with physical laboratory automation and robotics for chemistry/biology/materials science
  • Enhancement of system's capacity for autonomous theory refinement
  • Development of mechanisms for truly novel experimental methodology generation
Data: "The complete dataset, analysis code, and materials are openly available on GitHub to facilitate replication and extend reproducibility standards in cognitive research." Specific repository URLs not provided in text.; "The complete dataset, analysis code, and materials are openly available on GitHub to facilitate replication." GitHub repository mentioned but specific URL not provided in text. Study materials and data collection platform integration: Pavlovia (hosted on GitLab) and Prolific participant management platform.; Complete datasets, analysis code, and materials stated as "openly available on GitHub" but specific URL not provided in excerpt. Prolific and GitLab repositories mentioned but links not explicitly given.Code: GitHub repositories mentioned but specific URLs not explicitly provided in extracted text. All analysis code and generated statistical code (mean 7696 lines per study, SD=2426) archived.; GitHub repository mentioned for complete dataset and analysis code (specific URL not provided in document). Code implementation: "All content was autonomously generated and typeset in LaTeX." Data analysis code totalled "mean of 7696 lines (SD = 2426) per study."; GitHub (materials and analysis code); GitLab/Pavlovia (experiment hosting); SurveyMonkey (questionnaire platform)Extracted from: pdfAgreement 41%

Explore related topics

Related papers