From Static Repositories to Agentic Knowledge Webs: ResearchTwin and the S-Index for Federated Human-AI Research Discovery
Martin G. Frasch · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
System description with preliminary evaluation based on a deployed prototype.
Primary method
Design science research with system architecture design; Bimodal Glial-Neural Optimization (BGNO) inspired by biological neural tissue organization
Main result
The study found that "two researchers who appear nearly identical under citation-based metrics can have substantially different impact profiles when dataset contributions, code reuse, and collaboration breadth are incorporated." Specifically, Researcher A and Researcher B had comparable H-indexes (33 vs. 31) and Paper Impact scores (151.28 vs. 139.12), but their S-indexes differed by approximately 34% (1,049 vs. 782), demonstrating the discriminative power of the multi-modal S-index metric.
Research paradigm
Design science / artifact-centered research
Author conclusions
The authors conclude: "We believe that the transition from static repositories to agentic knowledge webs represents a meaningful shift in how scientific knowledge is organized, discovered, and reused. ResearchTwin represents a concrete step toward this vision, and we invite the research community to deploy, extend, and critique both the platform and the S-index formulation." They also emphasize that "the current v2 parameterization-while substantially simplified from the v1 formulation through the FAIR gate, field normalization, and parameter-free collaboration term-remains a principled but heuristic baseline; rigorous empirical calibration against expert assessments of research impact remains essential future work before the S-index can be recommended for evaluative purposes."
Risk of bias
Selection bias: only two researchers registered in preliminary evaluation; No ground truth for multi-modal impact ranking; Geographic bias in latency measurements (single location: Frankfurt); Cold-start vs. warm-cache performance differences not fully characterized; Citation source considerations: max-merge strategy may inflate counts relative to curated databases; Small sample size (n=2 researchers) leading to high selection bias; No ground truth for validating multi-modal impact rankings; Potential bias from max-merge strategy inflating citation counts relative to curated databases like Web of Science or Scopus; Heuristic parameterization of S-index not empirically calibrated; Dependence on unauthorized Google Scholar scraping may introduce data quality issues; No randomization or control conditions in case study comparison; Small sample size (n=2 researchers) introduces severe selection bias and limited generalizability; No ground truth validation against expert assessments; Google Scholar scraping dependency creates access fragility and potential sampling bias; S-index parameter choices (bonus weights, field medians, reuse event weightings) are heuristic rather than empirically calibrated; Geographic bias in latency measurement (single location: Frankfurt); Deduplication procedures may introduce hidden biases in artifact grouping
Limitations
- The authors explicitly state that "This preliminary evaluation has several limitations that must be acknowledged: Sample size: Only two researchers are currently registered on the deployed instance
- Any conclusions about the S-index's discriminative properties are necessarily anecdotal at this scale
- No ground truth: There is no established ground truth for 'correct' multi-modal research impact ranking
- Validating that higher S-index scores correspond to greater actual impact requires expert panel assessments, which are planned but not yet conducted." Additionally, "The bonus weights, field medians, and reuse event weightings (e.g., forks = 3× stars) are heuristic rather than empirically calibrated" and "Google Scholar does not provide an official API
- The scholarly library relies on web scraping, which is subject to rate limiting, IP blocking, CAPTCHA challenges."
Open questions raised
- Migration from Google Scholar (scraping-based) to official APIs (OpenAlex, Crossref)
- Implementation of Tier 2 Hub layer for cross-institutional federation
- Full-text semantic search with tiered permissions (local full-text, hosted metadata-only)
- Researcher recommendation based on artifact similarity
- Empirical calibration of S-index parameters through expert assessment panels (20-30 researchers across disciplines)
- Expanded evaluation to 10-20 researchers across diverse fields
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Systematic review of research on artificial intelligence applications in higher education – where are the educators?Olaf Zawacki‐Richter · 2019 · 5,282 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- State of the art and practice in AI in educationW. Holmes · 2022 · 758 citations
- AI chatbots in programming education: Students’ use in a scientific computing course and consequences for learningS.E.A. Groothuijsen · 2024 · 65 citations