12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Unraveling the Ai2 Asta Scholarly Research Assistant Citation System

Enrique Orduña‐Malea, Carlos Lopezosa · Revista Panamericana de Comunicación · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.21555/rpc.v7i2.3675

Methodology & findings

Study design

Exploratory empirical study using systematic query analysis.

Sample

N = 10, 4 groups

Primary method

Descriptive statistical analysis (mean, median, range calculations); Pearson correlation analysis (Rp = 0.86 reported); Venn diagram analysis for citation overlap visualization; frequency analysis of venues and publication years; proportional analysis of recent publications (2023 onwards); comparative tabulation across data collections

Main result

The study found that "the reports tend to include approximately between 20 and 40 cited references, which is considered a high number given the length of the reports (around 2.3k-2.4k words on average)" and that "formulating the same query at two different times has allowed us to verify the existence of significant variability in the references cited in the generated reports." Additionally, "the findings of this study indicate that Ai2 Asta's citation system displays a distinctive combination of high citation intensity, moderate bibliographic diversity, and considerable instability across repeated queries."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "Ai2 Asta's citation system displays a distinctive combination of high citation intensity, moderate bibliographic diversity, and considerable instability across repeated queries. Although the tool produces well-structured reports enriched with numerous citation-backed claims... the underlying set of cited publications varies substantially even when identical queries are issued at different times." They further state that "while Ai2 Asta offers high-quality and informative reports that can significantly support early-stage literature exploration, its limitations call for a cautious and informed use, as well as continued research into its technical behaviour and epistemic impact."

Risk of bias

Semantic Scholar database coverage bias: system emphasizes direct title matches and highly-cited papers with recent publication dates; arXiv favoritism bias: system positively favors researchers in fields with greater arXiv availability; Selection bias: only 10 queries analyzed, all focused on quantitative science studies; Database-level bias: 55.2% of cited publications published since 2020, skewing toward recent literature; Field of study assignment changes: four queries showed field of study changes between data collections that could alter retrieval; Inherent bias from Semantic Scholar database coverage and ranking algorithms; Selection bias: only 10 queries used, all from quantitative science studies field; Semantic Scholar emphasizes direct title matches and highly-cited papers with recent publication dates; Positive bias toward arXiv.org publications; Field of study assignment variability (changed in 4 cases between data collections); Limited time window for data collection (3 days only); Opaque selection mechanisms during report generation not fully transparent; Semantic Scholar database coverage bias - the tool inherits biases from the underlying database; Emphasis on direct title matches and highly-cited papers with recent publication dates in Semantic Scholar Academic Graph; Positive favoring of researchers in fields with greater arXiv publication availability; Selection bias from the system's internal ranking algorithms and neural re-ranking processes; Query formulation effects - different prompt precision may affect result coverage; Field of study assignment variability - in four cases, assigned field of study changed between data collections; Temporal instability - identical queries produce different citation sets when performed at different times

Limitations

  • The authors acknowledge several limitations: "Many aspects still remain to be examined, both in terms of in-depth document retrieval and the final selection of references during report generation, in order to obtain a more complete picture of its behavior." Additionally, they note "certain limitations in source identification are observed, which restrict the accuracy of raw venue-related statistics
  • For example, 23 cited references in the first data collection and 19 in the second do not provide source information." Furthermore, they identify normalization problems in venue naming, and acknowledge that "the tool's inherent biases" are inherited from Semantic Scholar, which "emphasizes direct title matches and highly-cited papers with recent publication dates."

Open questions raised

  • Need to expand the number and variety of queries across different scientific disciplines and research topics
  • In-depth examination of document retrieval mechanisms and final selection of references during report generation
  • Quality assessment dimensions: response length, structural coherence, argumentation strength, and presence of biases/inaccuracies
  • Detailed examination of cited reference suitability and reliability
  • Longitudinal studies analyzing Asta's behavior over time and impact of system updates
  • The authors identify multiple future research directions: (1) expanding the number and variety of queries to different scientific disciplines and research topics to determine if variability patterns persist or vary by field; (2) including new dimensions of analysis focused on overall report quality, including response length, structural coherence, argumentation strength, and presence of biases/inaccuracies; (3) detailed examination of cited reference suitability and reliability, assessing whether references are properly constructed, have clear authorship, scholarly rigor, and appropriate relevance; (4) developing longitudinal studies analyzing Asta's behavior over time to identify trends related to model updates and corpus expansions.
Data: Asta Summary Citation Counts: https://huggingface.co/datasets/allenai/asta-summary-citation-counts; Supplementary material: https://riunet.upv.es/handle/10251/230713; File 'sqa_citation_ranking_all_time.parquet' (downloaded October 28th); Asta Summary Citation Counts (https://huggingface.co/datasets/allenai/asta-summary-citation-counts) - open dataset compiled by Ai2 Lab team based on analysis of 113,000 queries collecting ~4 million citations from ~2 million publications, updated weekly; Supplementary material available at https://riunet.upv.es/handle/10251/230713; sqa_citation_ranking_all_time.parquet file downloaded October 28th; Asta Summary Citation Counts; Supplementary materialExtracted from: pdfAgreement 55%

Explore related topics

Related papers