vitaLITy 2: Reviewing Academic Literature Using Large Language Models
Hongye An, Arpit Narechania, Emily Wall, Kai Xu · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2408.13450
Methodology & findings
Study design
Design science methodology with artifact development, system architecture implementation, and exemplary usage scenarios.
Primary method
Design science methodology with iterative development building upon predecessor system (VITALITY 1)
Main result
VITALITY 2 extends prior literature search visualization approaches by incorporating "a novel Retrieval Augmented Generation (RAG) architecture" that enables users to search for relevant literature using multiple methods including paper seeds, abstracts, keyword searches, and natural language queries to an LLM-powered chat interface. The system "augments the dataset of 59,000 papers from VITALITY 1 to cover publications across 38 visualization venues over the past 3 years for a total of 66,692 papers" and introduces new capabilities including RAG-powered chat, paper summarization, and automated literature review generation.
Research paradigm
Design science / pragmatism
Author conclusions
The authors conclude that "VITALITY 2 takes advantage of text embeddings created by the latest LLMs, which represent a breakthrough for many NLP tasks, achieving performance levels close to humans." They state that "This capability holds immense potential for revolutionizing scholarly information access, enabling users to effectively harness the capabilities of RAG and pose a wide spectrum of inquiries." However, they recommend caution, suggesting "we suggest adding a user prompt in VITALITY 2 to explicitly warn users about the limitations of content generated by VITALITY 2, thereby reminding and cautioning users to use this tool with care."
Risk of bias
No formal user study conducted - only exemplary usage scenarios presented; Limited evaluation of LLM hallucination frequency and impact; Metadata-only approach may introduce bias in semantic understanding; ADA embeddings selection not rigorously validated against alternatives; Single embedding model preference not systematically justified
Open questions raised
- The authors identify several future research directions: optimizing the web crawler to retrieve full text of papers and segment texts into appropriately sized chunks; incorporating external knowledge bases from reliable sources such as Google Scholar search to further reduce hallucinations; and leveraging advances in LLM research, noting that "Recent benchmarks suggest that the incidence of AI hallucination is relatively small for GPT-4 by OpenAI."
- Future work directions include: (1) optimizing the web crawler to retrieve full text of papers and segment them into appropriately sized chunks with embeddings stored in the vector database for more comprehensive summaries; (2) adding user prompts to explicitly warn about LLM limitations; (3) further reducing hallucinations by incorporating external knowledge bases from reliable sources such as Google Scholar; (4) exploring newer LLM models like GPT-4 which show reduced hallucination rates.
- Need for full-text paper retrieval and chunking to improve summarization quality
- Reduction of LLM hallucinations through improved RAG and external knowledge bases
- Integration of external knowledge sources such as Google Scholar
- Evaluation of newer LLM versions (e.g., GPT-4) for reduced hallucination rates
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations