How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description
Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark construction from 100 published bibliometric review papers.
Sample
N = 100, 3 groups
Primary method
Semantic similarity: BERTScore F1 with one-to-one optimal assignment matching. Coverage evaluation: Average maximum cosine similarity between corpus sentences and description atoms in shared embedding space. Clustering evaluation: Adjusted Rand Index (ARI) and silhouette score computed on induced paper assignments versus Louvain clustering. Graph evaluation: Modularity computed on the underlying paper-relation network. Reference validation: For Blind pipeline, fuzzy string similarity matching (threshold 0.8) against OpenAlex for title, title+year, and title+year+first-author. For other pipelines: reference classification as in-corpus valid, out-of-corpus valid, or invalid.
Main result
The study found that "LLMs can generate cluster descriptions that are semantically close to human-written ones, but are unreliable when asked to infer bibliometric structure from scratch" and "when given appropriate bibliometric structure, LLM-generated descriptions can score higher than human descriptions on our corpus-level, clustering, and graph-based metrics, with the strongest performance occurring when bibliometric algorithms first define the clusters and the LLM is used to interpret them."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-computational
Author conclusions
"Overall, LLM-based bibliometric analysis is most promising as a hybrid workflow in which algorithms provide auditable structure and LLMs translate that structure into readable descriptions. This finding reflects a broader pattern in LLM system design. Across retrieval-augmented generation, tool use, code generation, and data analysis, LLMs are most reliable when external systems provide structure and the model performs synthesis, explanation, or translation. In bibliometric analysis, algorithms should handle precise structural tasks such as clustering papers by citations or shared references, while LLMs should handle the interpretive task of writing coherent descriptions."
Risk of bias
Selection bias: Benchmark limited to 100 peer-reviewed bibliometric papers, which may not represent all research domains equally; Query reconstruction bias: Manual extraction of queries from published papers may introduce interpretation errors; Evaluation metric bias: Embedding-based metrics (BERTScore, cosine similarity) may not capture scholarly nuance or domain expertise; Model-specific bias: Results may differ across different LLM versions; only two models tested in ablation (GPT-5.4, Gemini-2.5-Flash); Corpus reconstruction bias: Rerunning queries in Scopus may yield different results than original studies due to database updates; Number of clusters fixed: Louvain resolution tuned only to match human cluster count, not optimized for other objectives; Selection bias in source papers: benchmark constructed only from peer-reviewed published bibliometric reviews; Corpus reconstruction mismatch: reconstructed Scopus corpora may differ from original author corpora; Embedding model dependency: evaluation metrics depend on choice of embedding model (partially mitigated by ablation); Human description optimization: human descriptions not optimized for metrics used in evaluation; LLM model selection: results may be specific to GPT-5.4 and Gemini-2.5-Flash; Selection bias in benchmark construction: 100 bibliometric papers manually selected from peer-reviewed sources may not represent the full diversity of bibliometric analysis practices; Reconstruction bias: Scopus corpus reconstruction may differ from original author queries due to database changes or field interpretation differences; Evaluation metric bias: Embedding-based metrics (BERTScore, coverage) may favor certain styles of description writing; LLM knowledge cutoff bias: Model performance depends on training data temporal boundaries; Ablation scope bias: Ablations conducted on only 20 studies due to computational cost, not full 100-paper benchmark
Limitations
- "These results should be interpreted as evidence of structural fidelity, not as a general comparison between LLMs and human experts
- Human descriptions were written for scholarly interpretation, not to optimize embedding-based coverage, ARI, or modularity
- Thus, when structured LLM pipelines score higher on these metrics, it means they better preserve the reconstructed bibliometric or citation structure used in our evaluation, not that they offer greater domain insight, nuance, or scholarly value." Additionally, "human-authored citation-cluster descriptions were not consistently available, so we leave human comparison for citation analysis outside our scope."
Open questions raised
- Broader evaluation needed: Results positioned as evidence of structural fidelity rather than direct replacement for human bibliometric expertise
- Need for expert judgment capture: Proposed metrics assess semantic alignment, corpus coverage, clustering fidelity, and reference grounding, but do not fully capture expert judgment, conceptual nuance, or scholarly usefulness
- Citation-analysis human comparisons: Human-authored citation-cluster descriptions were not consistently available in source papers, leaving human comparison for citation analysis outside the study scope
- Scalability and generalization: Ablation studies conducted on only 20 of 100 papers due to computational and financial constraints
- The paper identifies that recent work has begun to use LLMs around bibliometric analysis "mainly for auxiliary tasks such as search support, summarization, topic classification, and thematic mapping" but notes that "these studies do not systematically evaluate how different levels of LLM responsibility affect the quality of the bibliometric workflow itself." The authors position their work as contributing to a broader shift from "isolated single-task applications toward structured, multi-stage workflows."
- Need for broader evaluation of LLM responsibility in multi-stage workflows beyond isolated single-task applications
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations