12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Literature Review of Literature Reviews in Pattern Analysis and Machine Intelligence

Penghai Zhao, Xin Zhang, Ming‐Ming Cheng, Jian Yang, Xiang Li · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
1
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2402.12928

Methodology & findings

Study design

Tertiary analysis combining narrative synthesis and statistical analysis.

Sample

N = 3099, 4 groups

Primary method

Descriptive statistics (mean, median, mode, min, max, standard deviation implied in distributions); Maximum likelihood estimation for exponential distribution fitting (TNCSI calculation); Bézier curve analysis for citation trend modeling (IEI calculation); Gompertz function modification for reference quality measurement (RQM); Pearson correlation coefficient analysis; Spearman rank correlation coefficient analysis; Log-normal and power-law distribution fitting; LLM-based automated information extraction; Manual content coding and annotation

Main result

The study found that "a clear upward trend in the number of reviews over time is evident" with "a notable surge between 2019 and 2020." Analysis revealed that "reviews with an RQM greater than 0.8, an ARQ within the range of 0.7 to 0.9, and an SMP not exceeding 10 are more likely to achieve higher academic impact." Furthermore, "recent systems are able to generate coherent and reasonably organized reviews, sometimes enriched with figures or taxonomies, yet they still suffer from critical shortcomings" including "a tendency to over-rely on highly cited but outdated references, limited capacity to recognize and integrate very recent work, and insufficient incorporation of explanatory visuals or appraisal criteria."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist with mixed-methods (quantitative bibliometric analysis + qualitative narrative synthesis)

Author conclusions

The authors conclude that "together, these findings provide both a critical appraisal of existing review practices and a forward-looking perspective on how AI-generated reviews can evolve into trustworthy, customizable, and transformative complements to traditional human-authored surveys." They further note that "AI-generated reviews—when not intended for direct publication—can still provide substantial value. In particular, they may serve as an efficient aid for researchers seeking to keep pace with rapid developments in fast-evolving fields."

Risk of bias

Selection bias: Literature screening limited to papers with 'review' or 'survey' in title from specific journals/conferences; Data source bias: Reliance on arXiv API may exclude peer-reviewed reviews not on arXiv; Annotator bias: Manual validation performed by 6 annotators with potential inconsistency despite double-checking; Citation bias: Google Scholar citation counts subject to manipulation as noted by authors; Temporal bias: Database snapshot taken October 2024; historical data may have changed on API-based sources; Selection bias: Limited to arXiv-indexed papers; exclusion of non-preprint venues may skew representation; Manual validation bias: 6 annotators conducted double-checking with potential for inter-rater disagreement; Google Scholar bias: Authors acknowledge that "Google Scholar places a high weight on citation counts in its ranking algorithm and has therefore been criticized for exacerbating the Matthew effect"; Field-specific bias: Analysis restricted to PAMI domain; findings may not generalize to other fields; Temporal bias: Data collected October 2024; citation counts fluctuate over time; LLM-based extraction bias: Reliance on GPT-based filtering and information extraction introduces model-dependent errors; Selection bias: Automated filtering may miss relevant reviews not indexed in arXiv or using non-standard terminology; Indexing bias: Over-representation of arXiv papers compared to other publishers; underrepresentation of non-English reviews; Temporal bias: Citation counts measured at single time point (October 2024), subject to Matthew effect favoring highly-cited papers; Language bias: Search limited to English-language keywords and abstracts; Publisher bias: Focus on peer-reviewed venues may exclude relevant pre-prints or grey literature; Annotation bias: Manual validation conducted by 4 junior + 2 senior annotators; no inter-rater reliability reported

Limitations

  • The authors note that "this paper will focus only on reviews in the PAMI field" and acknowledge that "the snapshot is a mirror image of relevant papers before a specific time, its data remains unchanged over time compared to the API-based sources." Additionally, they state "Considering the potential legal risks of crawling to obtain academic data from Google Scholar, Semantic Scholar API was employed" rather than Google Scholar directly
  • The authors also emphasize that "no existing metric, including the proposed ones, can truly capture the intrinsic value of a review (particularly for systematic review)
  • The proposed impact or quality indicators should be regarded as supplementary instruments rather than definitive measures of quality."

Open questions raised

  • Limited scholarly attention to structural conventions and statistical patterns in PAMI reviews
  • Difficulty navigating expanding corpus of literature reviews to identify most relevant surveys
  • Need for improved bibliometric indicators beyond simple citation counts for review selection
  • Gaps in AI-generated review systems regarding reference retrieval, coverage of recent work, and incorporation of visual elements
  • Limited understanding of how AI-generated reviews can become reliable complements to human-authored surveys
  • The authors identify several gaps: (1) characteristics of literature reviews in PAMI remain underexplored; (2) researchers face difficulties navigating the expanding corpus of reviews; (3) reliability and limitations of AI-generated reviews compared to human-authored reviews remain unclear; (4) systematic methods for evaluating literature reviews are underdeveloped; (5) AI-generated review systems suffer from weaknesses in reference retrieval, coverage of recent work, and incorporation of visual elements.
Data: RiPAMI database; RiPAMI database with 3,099 review articles; authors state "All the data and code framework used in this paper are publicly available at https://sway.cloud.microsoft/2TXEuPuNlDKEmC9p"; RiPAMI database containing 3,099+ review articles with metadata (title, authors, publication date, citations, references, word count, pages, figures, tables). Authors state: "All the data and code framework used in this paper are publicly available at https://sway.cloud.microsoft/2TXEuPuNlDKEmC9p"Code: Code framework; https://sway.cloud.microsoft/2TXEuPuNlDKEmC9p (data and code framework); https://sway.cloud.microsoft/2TXEuPuNlDKEmC9p (stated to contain data and code framework)Extracted from: pdfAgreement 50%

Explore related topics

Related papers