Optimal Multi-Model Embedding Combinations for Bibliometric Screening in Systematic Literature Reviews
Sebastian Matysik, Joanna Wiśniewska, Paweł Karol Frankowski · Journal of the Association for Information Systems · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Exhaustive computational search evaluating all 2^m - 1 = 1,023 non-empty subsets of 10 embedding models across a case study of 597 Scopus documents on corporate social responsibility (CSR) and consumer behavior.
Main result
The study found that "the best coherence-coverage trade-off depends on the analysis perspective: under the union perspective, the peak occurs at k = 5 (Score = 2.232, 95 articles); under the mean perspective, the best pair is distilroberta + Nomic (Score = 1.947, 58 articles); under the consensus perspective, distilroberta + Cohere preserves the most consensus articles (23), with cross-provider diversity being essential for effective combinations." Additionally, "multi-model consensus acts as a bibliometric coherence filter: keyword coherence of consensus articles rises from B = 0.471 at k = 1 to B = 1.333 at k = 10, confirming that multi-model agreement identifies the most relevant publications in the corpus."
Research paradigm
Positivist/Empiricist
Author conclusions
"This study conducted an exhaustive search across all 1,023 non-empty subsets of ten embedding models from five providers, evaluated on 597 Scopus publications using ten bibliometric indicators. The results establish three key findings. First, the all-distilroberta-v1 model emerges as the dominant core of best-performing configurations within the studied setting across union, mean, and consensus perspectives, appearing in all 15 of the top 15 union-ranked combinations. Second, the best coherence-coverage trade-off depends on the analysis perspective: under the union perspective, the peak occurs at k = 5 (Score = 2.232, 95 articles); under the mean perspective, the best pair is distilroberta + Nomic (Score = 1.947, 58 articles); under the consensus perspective, distilroberta + Cohere preserves the most consensus articles (23), with cross-provider diversity being essential for effective combinations. Third, multi-model consensus acts as a bibliometric coherence filter: keyword coherence of consensus articles rises from B = 0.471 at k = 1 to B = 1.333 at k = 10, confirming that multi-model agreement identifies the most relevant publications in the corpus."
Risk of bias
Single-domain case study bias (CSR and consumer behavior only); Selection bias in baseline configuration (top-n = 40, equal weighting); Construct validity risk: bibliometric indicators may not capture true relevance in interdisciplinary contexts; Query formulation sensitivity (shown by ρ = 0.38-0.50 between short and baseline queries); Corpus-specific findings that may not generalize across research fields; Single-domain evaluation (CSR/consumer behavior only) limits generalizability; Construct validity concern: bibliometric indicators (shared references, keywords, citations) may not reflect true relevance to human reviewers, particularly in interdisciplinary domains; Query formulation sensitivity: rank correlation drops to ρ=0.38-0.50 with concise query reformulation, indicating substantial sensitivity to how research question is phrased; Selection bias in bibliometric indicators: framework may under-reward cross-tradition papers in interdisciplinary settings; Potential bias introduced by equal-weighting of three indicator groups (references, keywords, citations) which may not reflect researcher priorities; Top-n saturation: at top-n=55, Spearman correlation drops to ρ=0.60, indicating union perspective becomes uninformative near corpus coverage; Single domain evaluation (CSR/consumer behavior only) limits generalizability; Bibliometric indicators (shared references, keywords, mutual citations) serve as proxies for relevance rather than measuring true human-reviewer relevance; Query formulation sensitivity: concise 13-word version yields ρ=0.38-0.50, extended 95-word version yields ρ=0.76-0.80; Potential under-reward of cross-disciplinary papers in interdisciplinary settings; Fixed top-n=40 configuration may not be optimal across all domains
Limitations
- "The study evaluates a single domain (CSR and consumer behavior)
- generalization requires additional empirical validation across diverse research fields." Additionally, "the baseline configuration used top-n = 40 with equal-group weighting
- the Sensitivity and Robustness Analysis demonstrates that the ranking is highly stable across top-n ∈ {35, 40, 45} (ρ ≥ 0.96) and across alternative weighting schemes (ρ = 0.94-0.98), but the identification of a single best configuration is conditional on weight allocation and query verbalisation." Furthermore, "a second category of limitation concerns construct validity
- Shared references, keyword overlap, and mutual citations measure the internal coherence of a selected set rather than its true relevance to a human reviewer's information need
- In domains organised around a coherent disciplinary tradition these indicators serve as a defensible proxy for screening relevance
- in interdisciplinary settings, however, highly relevant papers may be drawn from distinct citation communities and share few references or keywords, in which case the framework would under-reward cross-tradition relevance."
Open questions raised
- Generalization across diverse research fields beyond CSR and consumer behavior
- Validation on interdisciplinary corpora where highly relevant papers may be drawn from distinct citation communities
- Extension of exhaustive analysis to multiple domains
- Investigation of dynamic top-n selection in which n varies per model based on confidence scores
- Development of consensus-based voting mechanisms that leverage the coherence-enhancing properties of multi-model agreement
- Application across IS subdomains such as IT governance or digital innovation
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations