12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark construction from 100 published bibliometric review papers.

Sample

N = 100, 3 groups

Primary method

Semantic similarity: BERTScore F1 with one-to-one optimal assignment matching. Coverage evaluation: Average maximum cosine similarity between corpus sentences and description atoms in shared embedding space. Clustering evaluation: Adjusted Rand Index (ARI) and silhouette score computed on induced paper assignments versus Louvain clustering. Graph evaluation: Modularity computed on the underlying paper-relation network. Reference validation: For Blind pipeline, fuzzy string similarity matching (threshold 0.8) against OpenAlex for title, title+year, and title+year+first-author. For other pipelines: reference classification as in-corpus valid, out-of-corpus valid, or invalid.

Main result

The study found that "LLMs can generate cluster descriptions that are semantically close to human-written ones, but are unreliable when asked to infer bibliometric structure from scratch" and "when given appropriate bibliometric structure, LLM-generated descriptions can score higher than human descriptions on our corpus-level, clustering, and graph-based metrics, with the strongest performance occurring when bibliometric algorithms first define the clusters and the LLM is used to interpret them."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-computational

Author conclusions

"Overall, LLM-based bibliometric analysis is most promising as a hybrid workflow in which algorithms provide auditable structure and LLMs translate that structure into readable descriptions. This finding reflects a broader pattern in LLM system design. Across retrieval-augmented generation, tool use, code generation, and data analysis, LLMs are most reliable when external systems provide structure and the model performs synthesis, explanation, or translation. In bibliometric analysis, algorithms should handle precise structural tasks such as clustering papers by citations or shared references, while LLMs should handle the interpretive task of writing coherent descriptions."

Risk of bias

Selection bias: Benchmark limited to 100 peer-reviewed bibliometric papers, which may not represent all research domains equally; Query reconstruction bias: Manual extraction of queries from published papers may introduce interpretation errors; Evaluation metric bias: Embedding-based metrics (BERTScore, cosine similarity) may not capture scholarly nuance or domain expertise; Model-specific bias: Results may differ across different LLM versions; only two models tested in ablation (GPT-5.4, Gemini-2.5-Flash); Corpus reconstruction bias: Rerunning queries in Scopus may yield different results than original studies due to database updates; Number of clusters fixed: Louvain resolution tuned only to match human cluster count, not optimized for other objectives; Selection bias in source papers: benchmark constructed only from peer-reviewed published bibliometric reviews; Corpus reconstruction mismatch: reconstructed Scopus corpora may differ from original author corpora; Embedding model dependency: evaluation metrics depend on choice of embedding model (partially mitigated by ablation); Human description optimization: human descriptions not optimized for metrics used in evaluation; LLM model selection: results may be specific to GPT-5.4 and Gemini-2.5-Flash; Selection bias in benchmark construction: 100 bibliometric papers manually selected from peer-reviewed sources may not represent the full diversity of bibliometric analysis practices; Reconstruction bias: Scopus corpus reconstruction may differ from original author queries due to database changes or field interpretation differences; Evaluation metric bias: Embedding-based metrics (BERTScore, coverage) may favor certain styles of description writing; LLM knowledge cutoff bias: Model performance depends on training data temporal boundaries; Ablation scope bias: Ablations conducted on only 20 studies due to computational cost, not full 100-paper benchmark

Limitations

  • "These results should be interpreted as evidence of structural fidelity, not as a general comparison between LLMs and human experts
  • Human descriptions were written for scholarly interpretation, not to optimize embedding-based coverage, ARI, or modularity
  • Thus, when structured LLM pipelines score higher on these metrics, it means they better preserve the reconstructed bibliometric or citation structure used in our evaluation, not that they offer greater domain insight, nuance, or scholarly value." Additionally, "human-authored citation-cluster descriptions were not consistently available, so we leave human comparison for citation analysis outside our scope."

Open questions raised

  • Broader evaluation needed: Results positioned as evidence of structural fidelity rather than direct replacement for human bibliometric expertise
  • Need for expert judgment capture: Proposed metrics assess semantic alignment, corpus coverage, clustering fidelity, and reference grounding, but do not fully capture expert judgment, conceptual nuance, or scholarly usefulness
  • Citation-analysis human comparisons: Human-authored citation-cluster descriptions were not consistently available in source papers, leaving human comparison for citation analysis outside the study scope
  • Scalability and generalization: Ablation studies conducted on only 20 of 100 papers due to computational and financial constraints
  • The paper identifies that recent work has begun to use LLMs around bibliometric analysis "mainly for auxiliary tasks such as search support, summarization, topic classification, and thematic mapping" but notes that "these studies do not systematically evaluate how different levels of LLM responsibility affect the quality of the bibliometric workflow itself." The authors position their work as contributing to a broader shift from "isolated single-task applications toward structured, multi-stage workflows."
  • Need for broader evaluation of LLM responsibility in multi-stage workflows beyond isolated single-task applications
Data: Code and reproducibility materials: https://anonymous.4open.science/r/How-Much-Structure-Do-LLMs-Need-EF56/; Benchmark construction code and evaluation scripts available in anonymized repository at https://anonymous.4open.science/r/How-Much-Structure-Do-LLMs-Need-EF56/; Raw Scopus data: not directly available due to licensing constraints; prompts, code, derived identifiers, and evaluation scripts provided instead; Benchmark code, prompts, derived identifiers, and evaluation scripts available at: https://anonymous.4open.science/r/How-Much-Structure-Do-LLMs-Need-EF56/Code: https://anonymous.4open.science/r/How-Much-Structure-Do-LLMs-Need-EF56/Extracted from: pdfAgreement 62%

Explore related topics

Related papers