SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
Farhana Keya, Gollam Rabby, Sören Auer, Sahar Vahdati, Prasenjit Mitra, Mohamad Yaser Jaradeh · Machine Learning · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s10994-026-07036-8
Methodology & findings
Study design
Hybrid empirical evaluation combining: (1) automated LLM-as-judge evaluation using GPT-4.1 across 60 LLM-prompt-embedding configurations evaluated on 100 researcher profiles (~6,000 runs with three evaluation seeds each); (2) domain expert evaluation with 15 PhD-level computer science researchers independently rating ideas on 1-10 scales across four dimensions (Novelty, Excitement, Feasibility, Effectiveness); (3) ablation studies isolating effects of embedding granularity (no embeddings vs.
Sample
N = 100, 6 groups
Primary method
Paired two-tailed t-tests for comparisons among embedding strategies; two-way ANOVA for assessing main effects of embedding and prompting strategies; pairwise t-tests for specific LLM-embedding combinations; weighted Cohen's kappa for inter-rater reliability; Pearson correlation coefficients for inter-rater agreement; independent t-test for ORCID-only vs. facet-based comparison; Fleiss' kappa for three-annotator keyphrase/publication verification (κ = 0.82 and κ = 0.79 respectively). Statistical significance threshold: p < 0.05. Software not explicitly stated.
Main result
Token-level embeddings deliver statistically significant feasibility gains over the no-embedding baseline for GPT-4o (∆= + 0.29, p<0.01) and DeepSeek-70B (∆= + 0.24, p<0.001), converting configurations into usable outputs, with expert scores confirming the same directional trend. Additionally, "full-section processing makes 100% context-window infeasibility the standard for most profiles and increases hallucination from 4% to 42% beyond 20 papers, whereas structured facet extraction resolves both issues."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/positivist with pragmatic mixed-methods (automated + human evaluation)
Author conclusions
"We presented SCI-IDEA, a two-stage framework for context-aware scientific idea generation that addresses 4 structural limitations of prior LLM-based ideation systems: absent structured contextual modelling, no quantitative novelty measurement, no explicit surprise signal, and poor scalability, by decomposing publications into compact structured facets, integrating token-level semantic novelty with a likelihood-based surprise metric, and embedding both signals into an iterative human-in-the-loop refinement loop. With different researcher profiles and LLM, prompts, and embedding configurations, SCI-IDEA with token-level embeddings achieves a mean quality score of 6.97 versus 6.77 for the LLM only baseline (∆ = +0.20, p < 0.05), driven by statistically significant feasibility gains (∆ = +0.21, p < 0.001), with GPT-4o reaching 100% usability and DeepSeek-70B peaking at 84% under this configuration." The authors emphasize that "these gains reflect relative architectural advantages under automated evaluation, a qualification supported by the substantial human and LLM score divergence."
Risk of bias
Self-evaluation bias: GPT-4.1 serves as primary evaluation LLM while GPT-4o and GPT-4.5 are among evaluated generators; Evaluator selection bias: 15 human experts from computer science only; limited geographic/institutional diversity; Positivity bias in LLM evaluation: LLM-based scores systematically 3-4 points higher than expert ratings; Scope bias: Dataset limited to 100 researcher profiles in computer science domain; may not generalize to other fields; Measurement bias: Novelty metric measures dissimilarity within candidate idea set, not against broader published literature; Threshold bias: Aha-moment thresholds (θ n = 0.7, θ s = 2.0) set heuristically without systematic evaluation; Inter-rater reliability concerns: Expert inter-rater correlations with LLM near zero (r = 0.02-0.17); Publication selection bias: Most participants had prior publications in top-tier venues (NeurIPS, ICML, ACL, SIGIR); Self-evaluation bias: GPT-4.1 (evaluator) shares architectural lineage with GPT-4o and GPT-4.5 (generators), potentially inflating scores for GPT-family models; Selection bias: Study limited to computer science researchers with prior publications in top-tier venues (NeurIPS, ICML, ACL, SIGIR), not representative of broader research populations; LLM positivity bias: LLM-based evaluator systematically 3-4 points higher than expert ratings on 10-point scale with near-zero inter-rater correlation (r = 0.02-0.17); Scope limitation: 100 researcher profiles and 15 human evaluators insufficient for cross-domain generalization; Heuristic threshold setting: Aha-moment thresholds (θ n = 0.7, θ s = 2.0) set without systematic sensitivity analysis; Self-evaluation bias: GPT-4.1 used as evaluator while GPT-4o and GPT-4.5 are among evaluated generators; Selection bias: Limited to computer science researchers only (100 profiles); Evaluator bias: Only 15 PhD-level domain experts; predominantly academic venue publication requirement; Measurement bias: LLM-based evaluation shows near-zero correlation with human expert evaluation (r = 0.02-0.17); LLM scores systematically 3-4 points higher than expert ratings; Architectural bias: Shared lineage between GPT-4.1 (evaluator) and GPT-4o/4.5 (generators) may introduce optimism bias; Scope limitation: Restricted to computer science domain; unclear generalizability
Limitations
- "Several limitations bound the scope of the present findings and should guide their interpretation
- (i) Evaluation validity
- Because the human and LLM score correlation is near zero, all quantitative rankings in this paper are conditional on automated evaluation and may not generalise to expert-verified quality
- Calibrated LLM judges with pre-registered scoring protocols and confidence intervals are a necessary next step before stronger claims can be made
- (ii) Novelty measurement
- The current novelty metric measures semantic dissimilarity within the candidate idea set, not against the broader published literature
Open questions raised
- Replacing candidate-set novelty with literature-anchored novelty over large corpora
- Calibrating automated evaluation against expert panels under pre-registered protocols
- Conducting forward-prediction tests to assess whether high-scoring Aha ideas anticipate subsequent publications
- Extending the framework to domains beyond computer science
- Evaluating threshold sensitivity for Aha-moment detection (θ n = 0.7, θ s = 2.0)
- Multi-domain longitudinal studies with larger human evaluator cohorts
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations