Which Sections of a Research Paper Best Reveal Its Research Methods? Evidence from Library and Information Science
Qiuyu Fang, Jiayi Hao, Chengzhi Zhang · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational empirical study using supervised machine learning classification.
Sample
N = 1954, 6 groups
Primary method
Micro-averaged Precision, Recall, and F1-score (multi-label classification metrics); Approximate randomization test (Monte Carlo paired permutation test); 5,000 random swaps for statistical significance evaluation; Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) for encoder models; Quantized fine-tuning (QLoRA) for LLMs; Fair truncation strategy for token allocation across segments
Main result
The study found that "methodological features are non-uniformly distributed across the physical structure of academic papers" and that "segments in the middle-to-late and concluding intervals—which are more likely to carry implementation details and concluding remarks—exhibit higher discriminative value." By constructing TA-augmented dual-segment combinations (C4), the researchers "integrated global context from titles and abstracts with local methodological details, achieving superior performance over the baseline across most models."
Reports effect sizes.
Research paradigm
Positivist/empiricist - computational and quantitative evaluation of machine learning models on text classification tasks
Author conclusions
The authors conclude: "The results indicate that methodological features are non-uniformly distributed across the physical structure of academic papers. Segments in the middle-to-late and concluding intervals—which are more likely to carry implementation details and concluding remarks—exhibit higher discriminative value. By constructing TA-augmented dual-segment combinations (C4), we integrated global context from titles and abstracts with local methodological details, achieving superior performance over the baseline across most models. The distribution of optimal combinations reveals that high-performing pairs tend to cluster in the 'middle+middle-to-late' or 'middle-to-late+final' intervals, suggesting that the methodological evidence across different sections is inherently complementary."
Risk of bias
Limited corpus scope: only 3 LIS journals (Journal of Documentation, JASIST, LISR) from 2001-2010; Discipline-specific bias: findings may not generalize to other academic fields with different writing conventions; Annotation bias: original intercoder agreement at 86.7% (above 80% threshold but not perfect); Model selection bias: comparison limited to specific encoder and LLM architectures; Data cleaning bias: articles with non-standard structures (missing abstracts/headings) were excluded; Limited to three LIS journals (Journal of Documentation, JASIST, LISR), potentially not representative of LIS field broadly; Original annotation by Chu and Ke used 86.7% intercoder agreement, which while above 80% threshold, still leaves room for labeling error; Linear partitioning strategy may not align with semantic section boundaries, introducing segmentation bias; Manual auditing and cleaning process could introduce subjective bias; Dataset temporally limited to 2001-2010 publications; Limited to three LIS journals (Journal of Documentation, JASIST, LISR); disciplinary bias; Articles from 2001-2010 only; temporal bias; Excluded non-standard articles (missing abstracts or level-1 headings); selection bias; Manual cleaning and auditing introduces potential subjective bias; Use of DeepSeek-V3.2 for methodological summarization; reliance on single generative model introduces model-specific bias
Limitations
- The authors identified several limitations: "First, this study utilized articles from only three journals in the field of LIS
- Given the variations in writing norms and argumentative structures across different disciplines, the generalizability of these findings requires further validation across a broader range of fields
- Second, while our reliance on physical structure ensures universality and reproducibility, linear partitioning does not always align perfectly with logical sections, which may compromise semantic integrity at segment boundaries."
Open questions raised
- Generalizability across disciplines: findings limited to LIS; variations in writing norms and argumentative structures across fields require validation
- Semantic integrity at segment boundaries: linear partitioning may not align perfectly with logical sections
- Methodological summarization optimization: future research could further optimize mechanisms to provide more interpretable segment selection criteria
- Logic-based segmentation: incorporation of section-function-based segmentation as alternative to physical structure approach
- Generalizability beyond LIS field - variations in writing norms and argumentative structures across different disciplines require further validation
- Alignment of linear partitioning with logical sections - current approach may compromise semantic integrity at segment boundaries
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations