12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings

Farhana Keya, Gollam Rabby, Sören Auer, Sahar Vahdati, Prasenjit Mitra, Mohamad Yaser Jaradeh · Machine Learning · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s10994-026-07036-8

Methodology & findings

Study design

Hybrid empirical evaluation combining: (1) automated LLM-as-judge evaluation using GPT-4.1 across 60 LLM-prompt-embedding configurations evaluated on 100 researcher profiles (~6,000 runs with three evaluation seeds each); (2) domain expert evaluation with 15 PhD-level computer science researchers independently rating ideas on 1-10 scales across four dimensions (Novelty, Excitement, Feasibility, Effectiveness); (3) ablation studies isolating effects of embedding granularity (no embeddings vs.

Sample

N = 100, 6 groups

Primary method

Paired two-tailed t-tests for comparisons among embedding strategies; two-way ANOVA for assessing main effects of embedding and prompting strategies; pairwise t-tests for specific LLM-embedding combinations; weighted Cohen's kappa for inter-rater reliability; Pearson correlation coefficients for inter-rater agreement; independent t-test for ORCID-only vs. facet-based comparison; Fleiss' kappa for three-annotator keyphrase/publication verification (κ = 0.82 and κ = 0.79 respectively). Statistical significance threshold: p < 0.05. Software not explicitly stated.

Main result

Token-level embeddings deliver statistically significant feasibility gains over the no-embedding baseline for GPT-4o (∆= + 0.29, p<0.01) and DeepSeek-70B (∆= + 0.24, p<0.001), converting configurations into usable outputs, with expert scores confirming the same directional trend. Additionally, "full-section processing makes 100% context-window infeasibility the standard for most profiles and increases hallucination from 4% to 42% beyond 20 papers, whereas structured facet extraction resolves both issues."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/positivist with pragmatic mixed-methods (automated + human evaluation)

Author conclusions

"We presented SCI-IDEA, a two-stage framework for context-aware scientific idea generation that addresses 4 structural limitations of prior LLM-based ideation systems: absent structured contextual modelling, no quantitative novelty measurement, no explicit surprise signal, and poor scalability, by decomposing publications into compact structured facets, integrating token-level semantic novelty with a likelihood-based surprise metric, and embedding both signals into an iterative human-in-the-loop refinement loop. With different researcher profiles and LLM, prompts, and embedding configurations, SCI-IDEA with token-level embeddings achieves a mean quality score of 6.97 versus 6.77 for the LLM only baseline (∆ = +0.20, p < 0.05), driven by statistically significant feasibility gains (∆ = +0.21, p < 0.001), with GPT-4o reaching 100% usability and DeepSeek-70B peaking at 84% under this configuration." The authors emphasize that "these gains reflect relative architectural advantages under automated evaluation, a qualification supported by the substantial human and LLM score divergence."

Risk of bias

Self-evaluation bias: GPT-4.1 serves as primary evaluation LLM while GPT-4o and GPT-4.5 are among evaluated generators; Evaluator selection bias: 15 human experts from computer science only; limited geographic/institutional diversity; Positivity bias in LLM evaluation: LLM-based scores systematically 3-4 points higher than expert ratings; Scope bias: Dataset limited to 100 researcher profiles in computer science domain; may not generalize to other fields; Measurement bias: Novelty metric measures dissimilarity within candidate idea set, not against broader published literature; Threshold bias: Aha-moment thresholds (θ n = 0.7, θ s = 2.0) set heuristically without systematic evaluation; Inter-rater reliability concerns: Expert inter-rater correlations with LLM near zero (r = 0.02-0.17); Publication selection bias: Most participants had prior publications in top-tier venues (NeurIPS, ICML, ACL, SIGIR); Self-evaluation bias: GPT-4.1 (evaluator) shares architectural lineage with GPT-4o and GPT-4.5 (generators), potentially inflating scores for GPT-family models; Selection bias: Study limited to computer science researchers with prior publications in top-tier venues (NeurIPS, ICML, ACL, SIGIR), not representative of broader research populations; LLM positivity bias: LLM-based evaluator systematically 3-4 points higher than expert ratings on 10-point scale with near-zero inter-rater correlation (r = 0.02-0.17); Scope limitation: 100 researcher profiles and 15 human evaluators insufficient for cross-domain generalization; Heuristic threshold setting: Aha-moment thresholds (θ n = 0.7, θ s = 2.0) set without systematic sensitivity analysis; Self-evaluation bias: GPT-4.1 used as evaluator while GPT-4o and GPT-4.5 are among evaluated generators; Selection bias: Limited to computer science researchers only (100 profiles); Evaluator bias: Only 15 PhD-level domain experts; predominantly academic venue publication requirement; Measurement bias: LLM-based evaluation shows near-zero correlation with human expert evaluation (r = 0.02-0.17); LLM scores systematically 3-4 points higher than expert ratings; Architectural bias: Shared lineage between GPT-4.1 (evaluator) and GPT-4o/4.5 (generators) may introduce optimism bias; Scope limitation: Restricted to computer science domain; unclear generalizability

Limitations

  • "Several limitations bound the scope of the present findings and should guide their interpretation
  • (i) Evaluation validity
  • Because the human and LLM score correlation is near zero, all quantitative rankings in this paper are conditional on automated evaluation and may not generalise to expert-verified quality
  • Calibrated LLM judges with pre-registered scoring protocols and confidence intervals are a necessary next step before stronger claims can be made
  • (ii) Novelty measurement
  • The current novelty metric measures semantic dissimilarity within the candidate idea set, not against the broader published literature

Open questions raised

  • Replacing candidate-set novelty with literature-anchored novelty over large corpora
  • Calibrating automated evaluation against expert panels under pre-registered protocols
  • Conducting forward-prediction tests to assess whether high-scoring Aha ideas anticipate subsequent publications
  • Extending the framework to domains beyond computer science
  • Evaluating threshold sensitivity for Aha-moment detection (θ n = 0.7, θ s = 2.0)
  • Multi-domain longitudinal studies with larger human evaluator cohorts
Data: 100 researcher profiles from ORCID IDs with publications retrieved from CORE (https://core.ac.uk), arXiv (https://arxiv.org), and Semantic Scholar (https://semanticscholar.org); Topic distribution and research profile metrics reported in Table 2 (researchers with 2-2,320 publications; median: 52 papers); 100 researcher profiles with ORCID IDs and publications (not explicitly stated as publicly available); Publications retrieved from: CORE (https://core.ac.uk), arXiv (https://arxiv.org), Semantic Scholar (https://semanticscholar.org); 100 researcher profiles dataset curated from ORCID IDs and publications; related publications retrieved from CORE (https://core.ac.uk), arXiv (https://arxiv.org), and Semantic Scholar (https://semanticscholar.org). Dataset includes researchers with 2-2,320 publications (median: 52); topic distribution includes AI Alignment/RL/NLP (36), Knowledge Representation (12), Security/Privacy/Robustness in AI (10), Materials Science/Chemistry (17), Quantum Computing/Optimization (8), Others (17). Specific dataset URL not provided.Extracted from: pdfAgreement 58%

Explore related topics

Related papers