12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Towards an Integrative Approach for Automated Literature Reviews Using Machine Learning

Christoph Tauchert, Marco Bender, Neda Mesbah, Peter Buxmann · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences · 2020

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
19
Citations
0.89
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24251/hicss.2020.095

Methodology & findings

Study design

Design science research with iterative artifact development and three-stage evaluation: (1) functional evaluation of data acquisition, filtering, preprocessing and clustering; (2) human evaluation by researchers determining cluster correctness; (3) discussion by four IS researchers.

Primary method

Design science research with iterative development

Main result

The study found that "our extension is particularly suitable for capturing the topic of clusters without looking directly into each paper in detail" and that "the developed artifact delivers better results than known previous approaches and can be a helpful tool to support researchers in conducting literature reviews." The evaluation on 308 publications demonstrated that combining keyword extraction, topic modeling, and hierarchical clustering effectively organizes scientific literature into meaningful clusters with 17 clusters ranging from 9 to 29 documents each.

Research paradigm

Design science research (DSR) / pragmatic positivism

Author conclusions

"In summary, we have developed an artifact based on the word2vec algorithm, LDA topic modeling, rapid automatic keyword extraction, and agglomerative hierarchical clustering. This artifact is a first step towards simplifying the task of literature reviews within scientific research." The authors conclude that "the artifact provides a supporting mechanism to speed up and standardize the process of literature reviews and increases automation of an otherwise entirely manual process to ultimately improve the quality and reproducibility of this important aspect of research."

Risk of bias

Small evaluation dataset (308 papers) limits generalizability; Manual evaluation by researchers may introduce subjective bias in cluster assessment; Single domain tested (industrial manufacturing sensor data) - cross-domain applicability unclear; Potential evaluator bias as authors evaluated their own artifact; OCR conversion quality may vary affecting text extraction; Selection bias: Small, domain-specific dataset (308 papers on industrial manufacturing sensor data) may not generalize; Evaluator bias: Manual evaluation of clusters by researchers may introduce subjectivity despite attempting to standardize process; OCR bias: Tesseract OCR conversion of PDF to text may introduce errors, especially for scanned older manuscripts; Algorithm-human mismatch: Authors note that "the decision criteria might differ from a human interpretation since humans tend to interpret the meaning of topics and they do not solely rely on statistics and logical reasoning"; Limited ground truth validation: No comparison with expert-curated gold standard classifications; Small evaluation dataset (n=308) may not generalize; Manual cluster validation by only 4 IS researchers introduces subjective bias; Licensing issues led to use of existing dataset rather than crawler-generated dataset; Limited to sensor data in industrial manufacturing domain

Limitations

  • "The rather small evaluation data set of just 308 full-text papers which were manually checked if the proposed clustering of our model reflects the expectation of IS researchers, can only serve as a starting point for future research." Additionally, "Another limitation results from the conversion in plain text documents since all information stored in images and figures is not considered by the artifact." The authors also note that "a more rigorous evaluation of the artifact's utility for researchers during the creation of literature reviews in different contexts should be subject to future research."

Open questions raised

  • Need for evaluation with larger datasets beyond 308 papers
  • Comparative performance analysis of different topic modeling approaches (LDA vs. LSA vs. pLSA)
  • More rigorous evaluation of artifact utility in different research contexts and domains
  • Optimization of implemented algorithms
  • Investigation of how to handle information stored in images and figures during text extraction
  • Need for evaluation on larger and more diverse datasets across different research domains
Data: Dataset of 308 publications on sensor data application in industrial manufacturing used for evaluation. Dataset availability not explicitly stated as publicly available.; An existing set of 308 documents on the application and usage of sensor data in industrial manufacturing was used for evaluation (licensing issues prevented publication of this dataset)Code: Not mentionedExtracted from: pdfAgreement 62%

Explore related topics

Related papers