12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

TCMIIES: A Browser-Based LLM-Powered Intelligent Information Extraction System for Academic Literature

Hanqing Zhao · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Design science research with empirical evaluation.

Primary method

design science research with pragmatic implementation; user-centered design with interface patterns based on empirical HCI principles

Main result

The system achieved structured output compliance rates exceeding 94% and extraction accuracy averaging 81.6% across six extraction fields and three LLM providers. Specifically, "the overall compliance rate of 94.2% aligns with findings from Wang et al. [5], who reported 97% compliance using schema-guided prompting. The slight discrepancy may be attributed to the use of cost-efficient 'flash' and 'turbo' model variants rather than top-tier models, and to the Chinese-language content which presents additional parsing challenges."

Research paradigm

pragmatist/design science

Author conclusions

The authors conclude: "TCMIIES represents a practical bridge between cutting-edge LLM capabilities and the real-world needs of domain researchers. By eliminating barriers of programming expertise, infrastructure setup, and data privacy concerns, the system enables researchers to focus on their core scientific work while leveraging AI for efficient literature mining and knowledge synthesis." They state that "The system is actively deployed at the TCM Informatics Laboratory of Hebei University and is available as open-source software."

Risk of bias

Use of cost-efficient flash/turbo model variants rather than top-tier models may introduce systematic bias in accuracy assessments; Selection bias in corpus composition: 500 TCM papers from CNKI may not represent broader academic literature; Annotation bias: Only two expert annotators used for ground truth (κ=0.82), both potentially from same domain/institution; Language-specific bias: System tested primarily on Chinese-language content, generalizability to other languages not established; Provider selection bias: Only three LLM providers tested; results may not generalize to other providers; Model selection bias: Use of cost-efficient 'flash' and 'turbo' model variants rather than top-tier models may bias results toward lower performance; Language bias: Chinese-language content presents additional parsing challenges compared to English; Domain sampling bias: Evaluation limited to TCM-specific papers may not generalize to other biomedical domains; Annotator bias: Accuracy measured against only two domain experts rather than independent expert panels; Model selection bias: only tested three LLM providers (DeepSeek, Qwen, GLM-4); Cost-efficiency bias: used 'flash' and 'turbo' model variants which may not represent top-tier performance; Language-specific challenges: evaluation limited to Chinese-language academic content from specific databases (CNKI); Domain specificity: evaluation limited to TCM papers; generalization to other domains unclear; Sample size: only 100 papers (20 per sub-discipline) were annotated by domain experts for accuracy evaluation

Limitations

  • The paper states several limitations: "API Dependency: The system requires active API subscriptions, creating a dependency on commercial LLM providers." Additionally, "JSON Parsing Robustness: Despite the 94.2% compliance rate, approximately 5-8% of responses fail to parse as valid JSON." The authors note "Limited Context Handling: The system sends each paper as a single prompt, which may exceed context window limits for very long abstracts or full-text papers." They also acknowledge "Accuracy Ceiling: The zero-shot extraction accuracy of approximately 82% suggests that for high-stakes applications requiring near-perfect accuracy, human verification or domain-specific finetuning remains necessary." Finally, the system has "Lack of Full-Text Processing: Currently, the system primarily processes metadata (title, abstract, keywords)."

Open questions raised

  • Multi-Step Extraction Pipelines: Implementing staged extraction where the LLM first identifies relevant text passages, then extracts targeted information, and finally validates outputs through self-consistency checking
  • Integration with TCM Knowledge Bases: Linking extracted information to structured TCM knowledge bases such as TCMID and SymMap for entity normalization and validation
  • Active Learning for Prompt Optimization: Implementing automatic prompt optimization based on user corrections
  • PDF Parsing Support: Integrating browser-based PDF parsing to enable extraction from full-text papers rather than database export metadata
  • Cross-Database Normalization: Extending intelligent field mapping to support exports from additional databases such as PubMed, Web of Science, and Scopus
  • Cross-Database Normalization: Extending the intelligent field mapping to support exports from additional databases such as PubMed, Web of Science, and Scopus
Data: 500 TCM research papers from CNKI (China National Knowledge Infrastructure) spanning five sub-disciplines: herbal pharmacology (100 papers), acupuncture clinical trials (100 papers), formula composition studies (100 papers), TCM syndrome research (100 papers), and integrative medicine reviews (100 papers). No explicit statement of public availability provided.; 500 TCM research papers from CNKI (China National Knowledge Infrastructure) - dataset composition: 100 herbal pharmacology papers, 100 acupuncture clinical trials, 100 formula composition studies, 100 TCM syndrome research papers, 100 integrative medicine reviews; 500 TCM research papers from CNKI (China National Knowledge Infrastructure) spanning five sub-disciplines: herbal pharmacology (100 papers), acupuncture clinical trials (100 papers), formula composition studies (100 papers), TCM syndrome research (100 papers), and integrative medicine reviews (100 papers)Code: The paper states "The system is actively deployed at the TCM Informatics Laboratory of Hebei University and is available as open-source software" but does not provide a specific GitHub or repository URL.; System available as open-source software (specific repository URL not provided in paper); deployed at TCM Informatics Laboratory of Hebei University; System is available as open-source software (specific URL not provided in paper; mentioned as deployed at TCM Informatics Laboratory of Hebei University)Extracted from: pdfAgreement 63%

Explore related topics

Related papers