12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Pipette: Encoding scientific literature into an executable Skill Graph for multi-agent bioinformatics

Chirag Gupta, Ananya Sharma · bioRxiv (Cold Spring Harbor Laboratory) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.04.08.717332

Methodology & findings

Study design

Multi-case-study evaluation with real-world datasets and benchmarking against published reference analyses.

Primary method

Design science research; artifact-centric evaluation with comparative analysis against baseline approaches (unguided LLMs)

Main result

Pipette successfully recapitulated established biological and clinical findings across four technical domains. For single-cell RNA-seq analysis of the PBMC 68K dataset, "Pipette's automated annotations exhibited striking concordance with the established reference baseline" with a Pearson correlation of r=0.959 (p = 1.63 × 10⁻⁴). In bulk RNA-seq differential expression analysis, "Pearson correlations between the Pipette and the original study were exceptionally high across all experimental conditions. The correlations ranged from r=0.976 (p<0.0001, n=2,249) for the isolated drought contrast to r = 0.991 (p<0.0001, n=4,160) for the heat shock contrast." In clinical variant classification, the system achieved "7/7 sensitivity" on spike-in pathogenic variants with "100% specificity within the ACMG SF v3.2 gene panel." Across production deployments, "the Reviewer Agent functions as a substantive quality gate, systematically catching the same classes of methodological errors -QC thresholds, statistical rigour, and figure-text consistency -that human peer reviewers routinely flag in bioinformatics manuscripts."

Research paradigm

Design science / computational artifact evaluation

Author conclusions

The authors conclude: "By reducing the computational expertise required to execute standard genomic analyses, Pipette represents a step toward closing the gap between sequencing data generation and biological interpretation. As multi-omics datasets continue to grow in scale and complexity, literature-grounded, self-correcting agentic frameworks offer a promising approach to ensuring that analytical capacity keeps pace with data generation." They further note that "the central methodological contribution, however, is not improved NER but the finding that document-level positional ordering validated against EDAM-derived data type constraints recovers an order of magnitude more pipeline connections (60.5% vs. 6.2% recall) than sentence-level RE." The authors also emphasize the value of independent review: "The independent Reviewer Agent provides an additional layer of quality control through automated methodological auditing... the agent autonomously recognized and remediated missing physiological pH protonation in the ABL1 kinase docking workflow, flagged the absence of chromosome X in a clinical reference VCF, and correctly demoted a spurious PRKAG2 variant."

Risk of bias

Selection bias in literature corpus: ~20,000 PubMed Central papers may not represent all bioinformatics practices, particularly emerging or niche methodologies; Temporal bias: Skill Graph is static and represents a snapshot; no temporal weighting of recent methods documented; Tool annotation bias: NER/RE models trained on 800 documents with domain-specific tools may have uneven coverage across bioinformatics domains; Benchmarking bias: All benchmark datasets are published with known biological outcomes, not prospective validation on novel data; Literature retrieval bias: Automated module favored recent keyword-matched publications over foundational references; LLM vendor dependency: System uses Claude Opus 4.5 and GPT 5.4, introducing proprietary model biases and cost dependencies; Reference annotation artifacts: Original studies left ~10% of cells unclassified, creating apparent compositional differences with agent output; Dataset selection bias: all benchmark datasets were published with known outcomes, limiting generalizability to novel data; Literature bias: Skill Graph constructed from static corpus of ~20,000 papers representing snapshot at extraction time; risk of outdated workflows as tools deprecate; Vendor dependency: reliance on proprietary LLM (Claude Opus 4.5, GPT 5.4) introduces cost, latency, and service availability risks; Recall limitation in literature extraction: 65.1% recovery of ground-truth transitions means 34.9% of valid analytical pathways not captured; Annotation bias: 800-document training corpus may not represent full diversity of bioinformatics practices; Training corpus bias: Skill Graph constructed from static corpus of ~20,000 papers representing only published, peer-reviewed literature, potentially excluding novel/unpublished methods; Publication bias: Document-level ordering weighted by paper co-occurrence may bias toward widely-cited methods over emerging approaches; Vendor bias: System dependency on commercial LLMs (Claude Opus 4.5, GPT 5.4) with proprietary training data; Recency bias: Automated literature retrieval module favored recent keyword-matched publications over foundational references; Dataset selection bias: All benchmarks used established, published datasets with known biological outcomes; Annotation bias: Gold-standard knowledge graph derived from union consensus of two LLM reviewers (Claude Sonnet, Gemini) without human expert validation reported

Limitations

  • Several limitations should be noted
  • "First, Pipette's analytical repertoire is bounded by the Skill Graph
  • While the current graph covers the majority of standard bioinformatics workflows, novel algorithmic approaches that have not yet established a co-occurrence footprint in the published literature cannot be inferred." Second, "all benchmarks in this study used published datasets with known biological outcomes
  • This enables direct comparison with established findings but does not constitute a prospective evaluation of the system on novel, unpublished data where ground truth is unavailable." Third, "the automated literature retrieval module consistently favored recent keyword-matched publications over foundational references." Fourth, "while the Reviewer Agent caught several meaningful errors across our benchmarks, we did not systematically quantify its false negative rate (errors that pass review) or false positive rate (correct steps flagged as errors)." Fifth, "the system depends on a commercial large language model, introducing vendor dependency and associated cost and latency considerations that we have not formally characterized." Finally, "while the agent demonstrated standards-compliant application of ACMG/AMP classification guidelines, its outputs in clinical contexts must remain strictly investigational
  • All AI-generated clinical variant classifications require review and sign-off by a board-certified molecular geneticist before informing patient care."

Open questions raised

  • Mechanism for incorporating new tools and transitions at scale remains an area of active development
  • Need for prospective evaluation on novel, unpublished data where ground truth is unavailable
  • Systematic quantification of Reviewer Agent false negative and false positive rates across workflow types
  • Formal characterization of cost and latency considerations for commercial LLM dependencies
  • Automated detection of tool deprecation through version tracking and repository activity monitoring
  • Temporal weighting schemes that prioritize recent publications in Skill Graph construction
Data: PBMC 68K dataset (10X Genomics Fresh 68K PBMCs) - public; Human pancreas CEL-seq2 dataset (GSE85241, 3,072 cells, 4 donors) - publicly available at GEO; Rice leaf RNA-seq dataset (GEO accession GSE295637, 80 samples) - publicly available; ABL1 kinase crystal structure (PDB: 2HYY) - publicly available; p53-MDM2 co-crystal structure (PDB: 1YCR) - publicly available; HG002 Genome in a Bottle reference (GRCh38_1_22_v4.2.1_benchmark.vcf.gz) - publicly available from GIAB Consortium; All case study outputs at https://github.com/variomeanalytics/pipette_benchmark; Large data files at Zenodo (10.5281/zenodo.19433635); PBMC 68K dataset: 10X Genomics public dataset (Fresh 68K PBMCs from healthy human donor); Human pancreas CEL-seq2: GEO accession GSE85241 (3,072 cells, 4 donors); Rice leaf bulk RNA-seq: GEO accession GSE295637 (80 samples, 4 stress contrasts); Genome in a Bottle HG002 reference: https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/AshkenazimTrio/HG002_NA24385_son/NISTv4.2.1/GRCh38/; All case study outputs, conversation logs, provenance records: https://github.com/variomeanalytics/pipette_benchmark; Large data files (DESeq2 RDS, spiked-in VCF): https://zenodo.org/doi/10.5281/zenodo.19433635; 68K PBMC single-cell dataset (10X Genomics public dataset, Fresh 68K PBMCs); Human pancreas CEL-seq2 dataset (GSE85241, 3,072 cells, 4 donors); Rice leaf bulk RNA-seq dataset (GEO accession GSE295637, 80 samples); HG002 Genome in a Bottle reference standard (GRCh38, chr1-22, benchmark.vcf.gz); All case study outputs and provenance records: https://github.com/variomeanalytics/pipette_benchmark; Large data files (DESeq2 RDS object, spiked-in VCF) at Zenodo (10.5281/zenodo.19433635)Code: https://github.com/variomeanalytics/pipette_benchmark (case study code, conversation logs, provenance records); https://github.com/variomeanalytics/pipette_benchmark (complete conversation logs, provenance records, output files for all seven case studies); Skill Graph hosted at: http://skillgraph.pipette.bio (interactive navigation and workflow extraction); https://github.com/variomeanalytics/pipette_benchmark; Pipette.bio platform: https://pipette.bio; Skill Graph interactive interface: http://skillgraph.pipette.bioExtracted from: pdfAgreement 42%

Explore related topics

Related papers