12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

From Static Repositories to Agentic Knowledge Webs: ResearchTwin and the S-Index for Federated Human-AI Research Discovery

Martin G. Frasch · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

System description with preliminary evaluation based on a deployed prototype.

Primary method

Design science research with system architecture design; Bimodal Glial-Neural Optimization (BGNO) inspired by biological neural tissue organization

Main result

The study found that "two researchers who appear nearly identical under citation-based metrics can have substantially different impact profiles when dataset contributions, code reuse, and collaboration breadth are incorporated." Specifically, Researcher A and Researcher B had comparable H-indexes (33 vs. 31) and Paper Impact scores (151.28 vs. 139.12), but their S-indexes differed by approximately 34% (1,049 vs. 782), demonstrating the discriminative power of the multi-modal S-index metric.

Research paradigm

Design science / artifact-centered research

Author conclusions

The authors conclude: "We believe that the transition from static repositories to agentic knowledge webs represents a meaningful shift in how scientific knowledge is organized, discovered, and reused. ResearchTwin represents a concrete step toward this vision, and we invite the research community to deploy, extend, and critique both the platform and the S-index formulation." They also emphasize that "the current v2 parameterization-while substantially simplified from the v1 formulation through the FAIR gate, field normalization, and parameter-free collaboration term-remains a principled but heuristic baseline; rigorous empirical calibration against expert assessments of research impact remains essential future work before the S-index can be recommended for evaluative purposes."

Risk of bias

Selection bias: only two researchers registered in preliminary evaluation; No ground truth for multi-modal impact ranking; Geographic bias in latency measurements (single location: Frankfurt); Cold-start vs. warm-cache performance differences not fully characterized; Citation source considerations: max-merge strategy may inflate counts relative to curated databases; Small sample size (n=2 researchers) leading to high selection bias; No ground truth for validating multi-modal impact rankings; Potential bias from max-merge strategy inflating citation counts relative to curated databases like Web of Science or Scopus; Heuristic parameterization of S-index not empirically calibrated; Dependence on unauthorized Google Scholar scraping may introduce data quality issues; No randomization or control conditions in case study comparison; Small sample size (n=2 researchers) introduces severe selection bias and limited generalizability; No ground truth validation against expert assessments; Google Scholar scraping dependency creates access fragility and potential sampling bias; S-index parameter choices (bonus weights, field medians, reuse event weightings) are heuristic rather than empirically calibrated; Geographic bias in latency measurement (single location: Frankfurt); Deduplication procedures may introduce hidden biases in artifact grouping

Limitations

  • The authors explicitly state that "This preliminary evaluation has several limitations that must be acknowledged: Sample size: Only two researchers are currently registered on the deployed instance
  • Any conclusions about the S-index's discriminative properties are necessarily anecdotal at this scale
  • No ground truth: There is no established ground truth for 'correct' multi-modal research impact ranking
  • Validating that higher S-index scores correspond to greater actual impact requires expert panel assessments, which are planned but not yet conducted." Additionally, "The bonus weights, field medians, and reuse event weightings (e.g., forks = 3× stars) are heuristic rather than empirically calibrated" and "Google Scholar does not provide an official API
  • The scholarly library relies on web scraping, which is subject to rate limiting, IP blocking, CAPTCHA challenges."

Open questions raised

  • Migration from Google Scholar (scraping-based) to official APIs (OpenAlex, Crossref)
  • Implementation of Tier 2 Hub layer for cross-institutional federation
  • Full-text semantic search with tiered permissions (local full-text, hosted metadata-only)
  • Researcher recommendation based on artifact similarity
  • Empirical calibration of S-index parameters through expert assessment panels (20-30 researchers across disciplines)
  • Expanded evaluation to 10-20 researchers across diverse fields
Data: Figshare datasets (accessed via Figshare Search API); Semantic Scholar Academic Graph (via API); Google Scholar profiles (via scholarly Python library)Code: https://github.com/martinfrasch/ResearchTwin (MIT license); https://github.com/martinfrasch/S-index (S-index specification); PyPI package: mcp-server-researchtwin; MCP Registry: io.github.martinfrasch/researchtwin; mcp-server-researchtwin (Python Package Index/PyPI, MCP Registry as io.github.martinfrasch/researchtwin); mcp-server-researchtwin (Python Package Index, MCP Registry as io.github.martinfrasch/researchtwin)Extracted from: pdfAgreement 60%

Explore related topics

Related papers