12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Fine-tuning and structured prompting strategies for question answering over full-text biomedical research articles

Kaiming Tao, Rohit Satija, Jie Zhou, Zachary A. Osman, Vineet Ahluwalia, Chiara Sabatti et al. · PLoS ONE · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1371/journal.pone.0351631

Methodology & findings

Study design

Empirical evaluation study comparing three base language models (GPT-4o-mini-2024-07-18, Meta Llama-3.1-70B-Instruct, and Meta Llama-3.1-8B-Instruct) across four conditions: baseline, fine-tuning, question-specific prompting, and combined fine-tuning with question-specific prompting.

Sample

N = 400, 6 groups

Primary method

Wilcoxon signed-rank tests for pairwise comparisons between conditions, Fisher's exact test for pooled analyses comparing odds ratios (OR) between precision and recall effects. Metrics calculated: accuracy, precision, recall, and F1 score, averaged over 150 held-out studies.

Main result

Fine-tuning increased precision by 5% for GPT-4o, 16% for Llama-3.1-70B, and 8% for Llama-3.1-8B, and "fine-tuning also significantly increased recall for GPT-4o by 11%". Question-specific prompting increased recall for all three models (6% for GPT-4o, 7% for Llama-3.1-70B, and 18% for Llama-3.1-8B). "When pooled across the three models, fine-tuning was associated with a greater effect on precision than recall (OR = 4.35; p = 0.001; Fisher's exact test), whereas question-specific prompting led to a greater effect on recall than on precision (OR= 7.09; p = 0.0001; Fisher's exact test)."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-quantitative with computational methods

Author conclusions

"In this domain-focused proof-of-concept study, fine-tuning and question-specific prompting each led to improvement in one or more metrics for each of the three models. Pooled analyses indicated that fine-tuning improved precision, whereas question specific prompting preferentially improved recall."

Risk of bias

Potential selection bias in the choice of 250 HIV drug resistance studies used for fine-tuning instruction set construction. No mention of blinding or allocation concealment. The study focuses on a specific biomedical domain which may limit representation of other scientific domains. No cross-validation or external validation on independent datasets outside the held-out test set is mentioned.; Selection bias: Training set composition (250 HIV drug resistance studies) may not represent broader biomedical literature; Potential overfitting: Fine-tuning on domain-specific subset may limit generalizability to other research domains; Evaluation limitation: Performance assessed only on held-out studies from the same domain; Selection bias in study sample (250 HIV drug resistance studies used for fine-tuning); Potential data leakage if held-out test set was not properly isolated from fine-tuning; Limited to specific domain (HIV drug resistance) and question types; Only 150 held-out studies for evaluation

Data: not_statedCode: not_statedExtracted from: pdfAgreement 61%

Explore related topics

Related papers