12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Optimal Multi-Model Embedding Combinations for Bibliometric Screening in Systematic Literature Reviews

Sebastian Matysik, Joanna Wiśniewska, Paweł Karol Frankowski · Journal of the Association for Information Systems · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Exhaustive computational search evaluating all 2^m - 1 = 1,023 non-empty subsets of 10 embedding models across a case study of 597 Scopus documents on corporate social responsibility (CSR) and consumer behavior.

Main result

The study found that "the best coherence-coverage trade-off depends on the analysis perspective: under the union perspective, the peak occurs at k = 5 (Score = 2.232, 95 articles); under the mean perspective, the best pair is distilroberta + Nomic (Score = 1.947, 58 articles); under the consensus perspective, distilroberta + Cohere preserves the most consensus articles (23), with cross-provider diversity being essential for effective combinations." Additionally, "multi-model consensus acts as a bibliometric coherence filter: keyword coherence of consensus articles rises from B = 0.471 at k = 1 to B = 1.333 at k = 10, confirming that multi-model agreement identifies the most relevant publications in the corpus."

Research paradigm

Positivist/Empiricist

Author conclusions

"This study conducted an exhaustive search across all 1,023 non-empty subsets of ten embedding models from five providers, evaluated on 597 Scopus publications using ten bibliometric indicators. The results establish three key findings. First, the all-distilroberta-v1 model emerges as the dominant core of best-performing configurations within the studied setting across union, mean, and consensus perspectives, appearing in all 15 of the top 15 union-ranked combinations. Second, the best coherence-coverage trade-off depends on the analysis perspective: under the union perspective, the peak occurs at k = 5 (Score = 2.232, 95 articles); under the mean perspective, the best pair is distilroberta + Nomic (Score = 1.947, 58 articles); under the consensus perspective, distilroberta + Cohere preserves the most consensus articles (23), with cross-provider diversity being essential for effective combinations. Third, multi-model consensus acts as a bibliometric coherence filter: keyword coherence of consensus articles rises from B = 0.471 at k = 1 to B = 1.333 at k = 10, confirming that multi-model agreement identifies the most relevant publications in the corpus."

Risk of bias

Single-domain case study bias (CSR and consumer behavior only); Selection bias in baseline configuration (top-n = 40, equal weighting); Construct validity risk: bibliometric indicators may not capture true relevance in interdisciplinary contexts; Query formulation sensitivity (shown by ρ = 0.38-0.50 between short and baseline queries); Corpus-specific findings that may not generalize across research fields; Single-domain evaluation (CSR/consumer behavior only) limits generalizability; Construct validity concern: bibliometric indicators (shared references, keywords, citations) may not reflect true relevance to human reviewers, particularly in interdisciplinary domains; Query formulation sensitivity: rank correlation drops to ρ=0.38-0.50 with concise query reformulation, indicating substantial sensitivity to how research question is phrased; Selection bias in bibliometric indicators: framework may under-reward cross-tradition papers in interdisciplinary settings; Potential bias introduced by equal-weighting of three indicator groups (references, keywords, citations) which may not reflect researcher priorities; Top-n saturation: at top-n=55, Spearman correlation drops to ρ=0.60, indicating union perspective becomes uninformative near corpus coverage; Single domain evaluation (CSR/consumer behavior only) limits generalizability; Bibliometric indicators (shared references, keywords, mutual citations) serve as proxies for relevance rather than measuring true human-reviewer relevance; Query formulation sensitivity: concise 13-word version yields ρ=0.38-0.50, extended 95-word version yields ρ=0.76-0.80; Potential under-reward of cross-disciplinary papers in interdisciplinary settings; Fixed top-n=40 configuration may not be optimal across all domains

Limitations

  • "The study evaluates a single domain (CSR and consumer behavior)
  • generalization requires additional empirical validation across diverse research fields." Additionally, "the baseline configuration used top-n = 40 with equal-group weighting
  • the Sensitivity and Robustness Analysis demonstrates that the ranking is highly stable across top-n ∈ {35, 40, 45} (ρ ≥ 0.96) and across alternative weighting schemes (ρ = 0.94-0.98), but the identification of a single best configuration is conditional on weight allocation and query verbalisation." Furthermore, "a second category of limitation concerns construct validity
  • Shared references, keyword overlap, and mutual citations measure the internal coherence of a selected set rather than its true relevance to a human reviewer's information need
  • In domains organised around a coherent disciplinary tradition these indicators serve as a defensible proxy for screening relevance
  • in interdisciplinary settings, however, highly relevant papers may be drawn from distinct citation communities and share few references or keywords, in which case the framework would under-reward cross-tradition relevance."

Open questions raised

  • Generalization across diverse research fields beyond CSR and consumer behavior
  • Validation on interdisciplinary corpora where highly relevant papers may be drawn from distinct citation communities
  • Extension of exhaustive analysis to multiple domains
  • Investigation of dynamic top-n selection in which n varies per model based on confidence scores
  • Development of consensus-based voting mechanisms that leverage the coherence-enhancing properties of multi-model agreement
  • Application across IS subdomains such as IT governance or digital innovation
Data: CSR and consumer behavior corpus; 597 Scopus documents retrieved September 15, 2025, using query: TITLE-ABS-KEY(("CSR influence" OR "Corporate Social Responsibility influence") AND "consumer behavior"); Dataset includes titles, abstracts, author keywords, references, and citation counts; 597 Scopus documents retrieved September 15, 2025, using query: TITLE-ABS-KEY(("CSR influence" OR "Corporate Social Responsibility influence") AND "consumer behavior"). Dataset includes titles, abstracts, author keywords, references, and citation counts.Code: EmbedSLR; EmbedSLR project (open-source extension) - Matysik et al. (2025a); EmbedSLR project (open-source extension), GitHub link not explicitly provided but referenced as: 'All algorithms are released as an extension of the open-source EmbedSLR project (Matysik et al., 2025a).'Extracted from: pdfAgreement 51%

Explore related topics

Related papers