Multi-Agent AI Systems for Detecting Emerging Therapeutic Targets and Intervention Patterns in Neuroplasticity Research
Raza Hasan, Salman Mahmood · Journal of Informatics and Web Engineering · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.33093/jiwe.2026.5.2.20
Methodology & findings
Study design
Case study using multi-agent AI system with computational text analysis.
Sample
N = 533, 4 groups
Primary method
K-Means clustering (unsupervised; k optimized via sensitivity analysis); TF-IDF (Term Frequency-Inverse Document Frequency) transformation; Chi-squared test of independence (p<0.05 for validating cluster-defining terms); Silhouette Coefficient (cluster quality/separation metric); Within-Cluster Sum of Squares (WCSS; cluster compactness metric); Elbow Method (optimal k determination); Linear regression (temporal trend analysis; y=3.2x-6378.4, R²=0.86, p<0.001); Mann-Kendall test (monotonic trend test; τ=0.73, p<0.001); Welch's t-test (pre-2015 vs. post-2015 comparison; t=5.81, p<0.0001); Latent Dirichlet Allocation (LDA) for comparison/baseline; BERTopic (transformer-based topic modeling for comparison); Coherence score (CV; for topic model evaluation); Software: Python (algorithms implemented); spaCy, ScispaCy (NLP); Semantic Scholar API (data source)
Main result
The system successfully identified four statistically significant thematic clusters in neuroplasticity stroke rehabilitation literature: "The system performs the first end-to-end automated analysis on neuroplasticity stroke rehabilitation literature, with the system processing a corpus of 533 articles. This application has enabled the data-driven identification of four different research sub-domains with statistical validation (p<0.05)." The multi-agent system processed 533 articles in 12 minutes, identified 4,393 biomedical entities, and achieved "a Silhouette Score of 0.38, signifying a better separation distance among the formed clusters" with "94% of the top terms attaining statistical significance."
Reports effect sizes and confidence intervals.
Research paradigm
Computational empiricism with data-driven discovery
Author conclusions
The authors conclude: "Through our case study of stroke rehabilitation, we were able to derive valid or appropriate areas of focus for this landscape of scientific study, including its themes of Vagus Nerve Stimulation, molecular mechanisms, musically based therapies, and BCI-assisted trainings. This was all done in a fraction of the time that would normally be required manually." They also state that "The end goal of such a system, therefore, is clearly the enhancement of human knowledge, rather than the replacement of human knowledge by the system. This work, by virtue of automating the labour-intensive task of literature synthesis, allows scientists to better utilize their time on higher-level tasks such as experiment design, hypothesis formulation, and the application of these hypotheses to the clinical setting."
Risk of bias
Entity extraction errors: 10-15% miss rate; abbreviation ambiguity (CA as California vs. Cornu Ammonis); Co-occurrence bias: relationship inference assumes co-occurrence implies association; manual audit showed 30% false positive rate in extracted relationships; Cluster boundary uncertainty: 12% of documents had silhouette values <0.2, indicating ambiguous cluster membership; Subjective cluster interpretation despite statistical validation; Selection bias in literature source: Semantic Scholar API may not capture all relevant literature (e.g., gray literature, non-indexed sources); Fallback model precision loss: general spaCy model (en_core_web_sm) has lower biomedical specificity than primary model (en_core_sci_lg); API pagination strategy: automatic expansion of query results beyond requested limit (533 returned vs. 100 requested); No manual validation of all extracted entities; Entity extraction precision limited to 85-90%, introducing classification errors; Co-occurrence-based relationship inference may produce spurious associations (30% false positive rate reported); Binary keyword-based heuristic for quality assessment cannot distinguish between high-quality and low-quality studies; Cluster interpretation inherently subjective despite statistical validation; 12% of documents poorly classified (silhouette value <0.2) near cluster boundaries; Ambiguity in abbreviation extraction (e.g., 'CA' could refer to California or Cornu Ammonis); API pagination strategy and semantic expansion may introduce unintended documents into corpus; Entity extraction bias due to ambiguity (e.g., 'CA' as California vs. Cornu Ammonis); Fallback NLP model specificity loss reducing biomedical precision; Co-occurrence spurious associations (30% false positive rate in relationships); Subjectivity in cluster semantic interpretation despite statistical validation; API pagination and expansion potentially including off-topic documents; Publication bias inherent in literature corpus (only published studies included); Selection bias in corpus (Semantic Scholar API coverage limitations)
Limitations
- The authors acknowledge multiple limitations: "we acknowledge potential sources of error including entity extraction precision (not manually validated for all 4,393 entities), relationship inference based on co-occurrence (which may include spurious associations), and cluster interpretation (subjective)." Additional limitations include that "The current keyword-based heuristic used by the VA is a simplified heuristic" with "natural limitations in terms of the level of semantic analysis" and "contextual ignorance." The relationship extraction method is limited because "The current definition of relationships used by the CEA is based on co-occurrences of words in a sentence
- Although this is useful to map relationships in general, it doesn't capture the type of relationship (activates, inhibits, causes, etc.)." Entity extraction precision was estimated at 85-90%, with false positive rate in relationship extraction at roughly 30%, and 12% of documents had poor cluster assignment (silhouette value <0.2).
Open questions raised
- Semantic depth of relationships: Current system uses word co-occurrence; future work should employ fine-tuned transformer models (e.g., BioBERT) to capture relationship types (activates, inhibits, causes)
- Heuristic validation enhancement: Move beyond keyword-based heuristic to trained classifiers evaluating study quality
- Citation network analysis: Implement full citation network analysis including co-citation and coupling analysis
- Study design classification: Integrate study design classifiers to move beyond keyword matching
- PICO element extraction: Upgrade NLP module for more accurate relevance matching using PICO (Population, Intervention, Comparison, Outcome)
- Risk of bias integration: Incorporate structured risk of bias algorithms (GRADE, Cochrane Risk of Bias tool) and citation-weighted quality metrics
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations