12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai

Living Meta-Analysis

AI in Research

LivingMeta.ai

A living meta-analysis platform that continuously monitors, classifies, and analyzes research on generative AI in academic research practice — including literature review automation, AI-assisted writing, peer review, research integrity, institutional governance, and human-AI collaboration in scholarly workflows.

What that means in practice: every new paper in ai-in-research is scored for relevance to this field, read independently by several AI extractor personas whose readings are reconciled into one cross-validated consensus, and tagged with methodology and evidence type as it enters the corpus. Evidence gaps and research priorities emerge from the literature itself, and update as new papers are pulled in each night. Researchers can browse and filter papers, weigh pooled effect sizes in the Meta-Analysis view, discover the datasets and instruments others have used, or open a thread in The Lab to work through a question alongside an AI research agent. (See the FAQ below for what Relevance, cross-validation, and FWCI mean.)

13,061
Papers
6,807
Research gaps
2,655
Sources

Your own field

Want this for your own field?

Build a free trial of your own literature — your journals, your keywords. ~400 papers, six-perspective extractions, a priority research agenda and a grounded Lab agent, live for a week.

Build a free trial →

Free · one €0.50 anti-bot card check · no subscription, no install

On the LivingMeta platform

What you'll find inside this instance

Five surfaces, one continuously updated pipeline. Each one is populated directly from the literature — no manual curation backlog.

Papers

Browse the Literature

Every paper in the field, classified by relevance and cross-validated by multiple independent AI perspectives. Filter by paradigm, methodology, evidence model, or full-text availability; agreement scores flag where the AI is confident, and disagreement shows where human judgment belongs.

Resources

Curated Datasets & Tools

Datasets, questionnaires, measurement instruments, and research tools surfaced across the corpus, each linked back to the papers that use them. Reusable building blocks for the next study.

Meta-Analysis

Effect Sizes & Forest Plots

Effect sizes extracted from every paper that reports one, converted to a common metric where the statistics allow. Where enough comparable studies exist within a single paradigm, they are pooled into a random-effects estimate with a forest plot and heterogeneity — never pooled across different phenomena.

Research Agenda

Priority Research Agenda

A ranked list of the field's highest-impact evidence gaps, generated from the literature itself and scored for frequency, depth, and feasibility. Each priority links straight to an investigation in The Lab — and to its living review where one exists. Updates as the field evolves.

The Lab

Collaborative Research Threads

Ask a question in plain language and a server-side AI agent investigates the corpus for you — pulling in the right specialist automatically: evidence mapping, gap hunting, skeptical review, resource scouting, or academic writing. Attach your own data, papers, or drafts; role-adaptive coaching scales the scaffolding from layperson to expert, all the way from a first question to a written-up study.

Priority Research Agenda

AI-identified research priorities ranked by frequency, impact, and feasibility. How does this work?

1

Cross-Domain Generalization of AI Research Assistance Tools

Current AI-assisted research tools are overwhelmingly validated on computer science and closely related domains, leaving the vast majority of scientific disciplines underserved. This domain narrowness fundamentally limits the field's ability to claim generalizable progress and restricts adoption across the broader scientific community. Addressing this gap is essential for establishing AI-assisted research as a universal scientific capability rather than a niche CS tool.

2

Hallucination Detection and Mitigation in AI-Generated Scientific Content

Hallucinated citations, fabricated findings, and factually incorrect statements represent the most critical reliability barrier for deploying LLMs in scientific workflows. Despite being the largest cluster of identified gaps, the field lacks systematic frameworks for measuring, categorizing, and mitigating hallucinations specifically in scientific contexts. Without solving this, AI-assisted research tools cannot be trusted for consequential scientific tasks.

3

Standardized Evaluation Frameworks for AI-Assisted Scientific Review Quality

The field currently lacks consensus on how to measure whether AI-assisted literature reviews, peer reviews, or research summaries are actually better, worse, or biased compared to human-produced equivalents. Without standardized quality metrics, results across studies are incomparable and the field cannot accumulate reliable knowledge about system performance. This gap affects every researcher building or evaluating AI review tools.

Frequently Asked Questions

Everything about how this living review is built and how to use it — grouped by topic. Click any question to expand.

Getting started

This is a living review of an entire research field — a continuously updated map of the literature, built on the LivingMeta platform. Every relevant paper is screened, classified, and read by AI, and the results are laid out across five surfaces: browsable Papers, curated Resources, a Meta-Analysis view, a Priority Research Agenda, and The Lab, where you work through a question alongside an AI research agent. You can explore the evidence, find datasets and instruments, see where the field's gaps are, and get grounded, cited help with your own research — without doing the screening and extraction by hand.
Researchers, PhD and master's students, research groups, and anyone who needs to get on top of a field's literature quickly — including policy analysts, journalists, and practitioners. Whether you are running a systematic review, writing a thesis chapter, scoping a grant proposal, or just orienting yourself in an unfamiliar area, the platform does the heavy screening and extraction so your time goes to judgment and writing. The Lab adapts its guidance to your level, from complete newcomer to expert supervisor.
Browsing the papers, resources, meta-analysis, and research agenda is open on public instances — no account needed to read. Writing in The Lab (starting threads, running the AI agent) requires signing in, and some instances are private, where the whole site is gated to invited accounts. Where the AI agent is available to you, its usage is covered within a fair-use allowance shown in your account; you never bring your own API key.
Those tools each cover one stage: Elicit does per-paper extraction, Consensus searches abstracts, Cochrane produces static reviews, and a general chatbot answers from memory (and can invent citations). This platform integrates continuous screening, multi-perspective extraction, meta-analytic effect sizes, gap analysis, resource discovery, and a coached AI agent into one workflow that keeps updating — and every claim the agent makes is tied to a specific paper in the corpus, by identifier, with the source quote. The platform itself is the methodology, with an audit trail throughout.

The papers

The corpus is built from open academic sources — primarily OpenAlex, plus arXiv preprints and international thesis repositories — using the field's key journals and a validated set of search terms, going back to 2000. Every candidate is scored by an AI classifier for how central it is to this field, and clearly off-topic papers are set aside so the browsable corpus stays focused. New papers are pulled in automatically as they publish, so the collection tracks the field over time rather than being a one-off snapshot.
Every paper is scored 0–10 by the AI classifier for how central it is to THIS field — higher means more on-topic. The score is field-specific: a paper can be highly relevant here and irrelevant to another instance. Papers below 4 are treated as off-topic and left out of the browsable corpus; papers scoring 6+ get abstract-level extraction, and the most central (9–10) get the full six-perspective treatment wherever the full text is available.
New papers are pulled from OpenAlex automatically every night (around 03:30 UTC), then classified and, where relevant, extracted. The public site refreshes when the instance is next redeployed. That is what 'living' means here: the review keeps tracking the field as it publishes, instead of ageing the moment it is finished — the usual fate of a traditional systematic review.
On the Papers page you can search titles, authors, topics, and AI summaries, and filter by domain, methodology, theme, evidence model, publication year, and whether full text is available. Filters combine, so you can narrow to, say, experimental studies in one theme that report an extractable effect size. Each result shows its relevance score, evidence model, and citation impact — plus, where extracted, its methodology and key findings inline.
Not every paper produces knowledge the same way, so each is tagged with the kind of evidence it offers: Empirical (experiments, surveys, trials), Theoretical (proofs, formal derivations), Design (prototypes, artifacts, creative works), Interpretive (textual or qualitative analysis), or Computational (simulations, machine-learning models). The extraction adapts to the model — an empirical paper is read for sample and effect size, a design paper for its artifact and evaluation — so the fields you see are the ones that actually matter for that type of work. The quality score is tuned per model too.
It marks papers whose full text was available (as an open-access PDF or XML) and could therefore be read in depth, rather than only from the abstract. Full-text papers get the richest six-perspective extraction; abstract-only papers get a lighter single-pass extraction. The badge doubles as a filter, so you can restrict the view to papers with full text when depth matters.
Yes — from the Papers page you can export the current (filtered) set to CSV for a spreadsheet or to BibTeX for your reference manager, so you can screen and shortlist here and take the results straight into your own tools. (On the open demo this action is disabled, but it is part of every full living review.)

AI extraction & quality

Each paper is read by several independent AI 'personas', each with a different focus, and they extract the same structured fields — methodology, sample, effect sizes, findings, gaps, and more. A consensus algorithm then reconciles their answers field by field. Crucially, every factual claim has to carry a verbatim quote from the source; extractions without textual evidence are automatically flagged, which is what keeps the AI from inventing results. Where the personas disagree, that is surfaced as an agreement score rather than hidden — disagreement marks where human judgment belongs.
Each full-text paper is read by up to six extractor personas with different priorities — a default reader, a rigorist, a synthesizer, a skeptic, a scout, and a meta-analyst. Four of them validate the same core data while two deliberately look for additional information (overlooked gaps, extra resources), and their readings are reconciled into a single consensus. 'Cross-validated' means a value survived that multi-perspective check rather than resting on a single pass — and the strength of their agreement becomes the paper's quality score.
LMQS — the LivingMeta Quality Score — is a per-paper measure of methodological transparency, shown as a small fraction like '3/4'. It is adaptive: for each evidence model it checks the signals that actually matter (for an empirical study, things like sample transparency and statistical rigor; for a computational one, whether code and data were shared and results are reproducible). Green means most of those signals are present, amber some, red few. It is about how transparently the work was reported — separate from how often it has been cited.
Field-Weighted Citation Impact, from OpenAlex, is a citation count normalized for the paper's field, year, and document type. 1.0 is the world average for comparable papers; above 1.0 means more-cited than average, below 1.0 less. It is a signal of reach and attention — not of quality or correctness — so read it alongside the LMQS transparency score, not instead of it.
On papers that have not had a full extraction yet, an 'Extract full text' button lets you supply the PDF; the platform then runs the same six-persona extraction plus consensus, and the results appear on the paper, usually within a few minutes. It is the main way an instance's owner deepens the review over time — the new extraction, and anything the scout persona discovers, flows into the shared data on the next refresh. (On the open demo the button is visible but disabled.)

Meta-analysis

For every paper that reports a numeric effect size, we extract it and convert it to a common metric (Pearson's r) where the statistics allow. When a single paradigm has enough comparable studies, they are pooled into a random-effects estimate and shown as a forest plot with a heterogeneity measure. Effects are only ever pooled within one paradigm — never across different phenomena — so the numbers stay interpretable, and papers without an extractable effect size are listed transparently rather than dropped.
Treat them as exploratory. They are pools of AI-extracted effect sizes from the corpus, not systematic reviews with a pre-registered search and formal inclusion/exclusion criteria. That makes them excellent for hypothesis generation and for seeing where the evidence clusters — but for a confirmatory meta-analysis you would still verify the extracted effects against the source papers and apply your own protocol.

Research gaps & the agenda

The agenda is built automatically from the evidence in this instance. The 'further research is needed' statements across all extracted papers are clustered into themes and scored on frequency, source quality, feasibility, recency, and how much of the field they touch. An AI then filters the top-ranked themes against the field's context, so only genuinely relevant priorities survive. The list updates as new papers arrive, and each priority has its own page with an Investigate button that opens a Lab thread — plus a living review where one has been generated.
A research gap is a specific place where the evidence is missing, thin, or contradictory. The platform classifies them using an established taxonomy (Miles 2017) into eight types — evidence, knowledge, practice, methodological, empirical, theoretical, population, and integration gaps — so a gap in long-term outcomes reads differently from a gap in how findings translate to practice. Gaps are drawn straight from what the papers themselves flag as unresolved, then aggregated so you can see which are field-wide versus one-off.
A living review is a synthesis that keeps itself current — as new papers publish and get extracted, the classifications, gaps, and priorities update rather than freezing on a publication date. For a priority on the agenda, the Investigate button opens a thread in The Lab pre-loaded with that gap, so you can go straight from 'here is an important gap' to working it through with the AI agent, grounded in the relevant papers.

Curated resources

Datasets, validated questionnaires, measurement instruments, tools, code repositories, and APIs that were discovered across the extracted papers, then curated and deduplicated into one registry. Each resource links back to the papers that use it, so you can see the evidence base — and, for instruments, details like what it measures — before deciding whether to adopt it. Link health is checked, and the registry grows as new papers are added.

The Lab & the AI agent

The Lab is where you and an AI research agent work on a question together, in plain-language chat. You ask something, and the agent searches this instance's papers, extractions, resources, and gaps to answer — automatically pulling in the right specialist for the job (evidence mapping, gap hunting, skeptical review, resource scouting, or academic writing). It runs on the server, so there is nothing to install and no API key to bring; you just type.
No API key, and no per-question charge to you. The operator covers the AI usage, and where a balance applies it is shown in your account as a token allowance so you always know where you stand; if it runs low the agent says so plainly rather than failing silently. (On the open demo the agent runs in a lighter single-pass mode; full instances use the deeper multi-pass agent.)
It can map the evidence on a question, compare papers side by side, surface and rank gaps, find datasets and instruments, pull the priority agenda, read documents you upload, and help you draft — always citing the specific papers it used. What it will not do is decide for you or answer from outside the corpus: it does not browse the open web, and it will not present a paper that is not in the database as evidence. The judgment, the conclusions, and the manuscript stay yours.
When you start a thread you pick a role — layperson, junior researcher, senior researcher, or supervisor/coach — and the agent adjusts how much it scaffolds. A layperson gets plain-language explanation and stops short of study design; a junior gets Socratic, step-by-step guidance; a senior gets fast, peer-level discourse; a coach gets meta-level help, including teaching materials. It is set per thread, so you can dial the support up or down depending on what you are doing.
A research project moves through eight phases — Explore, Understand, Evaluate, Synthesize, Design, Execute, Write, and Reflect — and The Lab scaffolds each one. So a thread can carry you from a first vague question, through weighing the evidence and designing a study, all the way to writing it up, with the agent's help shaped to the phase you are in. You are never forced through them in lockstep; the phases just keep the work oriented.
Yes. You can attach PDFs, Word documents, and text or Markdown files to a thread, and the agent reads their content — a dataset description, a questionnaire, your notes, or a draft — and works with it alongside the corpus. That lets you ground the conversation in your own material, not just the published literature.
Yes — there is a dedicated writing mode. The agent can co-write with you against a target journal's guidelines, keep every claim tied to a cited source, and it deliberately avoids the tell-tale 'AI polish' that reviewers notice. It is built to help you draft and structure honestly, with the evidence attached — not to hand you a finished paper to submit unread.
Every reference the agent gives points to a real paper in this instance's corpus, cited by its identifier with the source quote — it cannot cite a paper the database does not hold, and it is instructed not to reach for the open web. If relevant work is missing from the corpus, it says so ('potentially relevant, not yet in the database') instead of inventing a citation. It is also honest about coverage: if only a couple of papers bear on your question, it tells you the answer is thin rather than overclaiming.

Trust, data & your own field

Each instance is isolated, with its own corpus and access controls, and anything you upload to a thread is used to answer within that thread — not sold on or fed into a public chatbot. Threads and posts live in the platform's backend so you can return to them; private instances gate the whole site to invited accounts. If you ever want to leave, the underlying data is exportable.
Start a thread in The Lab describing what you found — a misclassification, a wrong extracted value, or a paper that should be in the corpus. The instance's owner reviews submissions and can trigger a correction or add the paper, which then flows through the same classification and extraction pipeline as everything else. Because the review is living, fixes and additions become part of it rather than errata on a frozen document.
Yes — it is one of the most common uses. The Lab's coaching guides you across the whole arc, from exploring the literature to writing the manuscript; pick the 'Junior Researcher' role and it scaffolds the process with Socratic questions and human checkpoints. For a grant, the Priority Research Agenda gives a defensible, evidence-based answer to 'where should we be working next?'. And you can export the papers you rely on to CSV or BibTeX at any point.
Yes. Every instance publishes machine-readable instructions (an agent.json describing its endpoints, taxonomy, and how to contribute) and follows the llms.txt convention, so an AI agent can discover and use it. There is also an MCP server for power users who want to query the corpus, extractions, gaps, and resources from their own tools. The same anti-hallucination rules apply: cite papers by identifier, and do not present outside work as evidence.
Yes — this instance is one field on the LivingMeta platform, and you can have your own. The fastest way is a free trial: describe your field and a couple of key papers, and the platform builds you a small private instance — a few hundred papers with extractions, resources, a priority agenda, and a Lab agent — live for a week. From there it can be grown into a full, maintained review of your field. Look for the 'build a free trial' link, or visit livingmeta.ai.

Built on the LivingMeta platform. Read the foundation paper for methodology and architecture details.