Towards an Integrative Approach for Automated Literature Reviews Using Machine Learning
Christoph Tauchert, Marco Bender, Neda Mesbah, Peter Buxmann · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences · 2020
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24251/hicss.2020.095
Methodology & findings
Study design
Design science research with iterative artifact development and three-stage evaluation: (1) functional evaluation of data acquisition, filtering, preprocessing and clustering; (2) human evaluation by researchers determining cluster correctness; (3) discussion by four IS researchers.
Primary method
Design science research with iterative development
Main result
The study found that "our extension is particularly suitable for capturing the topic of clusters without looking directly into each paper in detail" and that "the developed artifact delivers better results than known previous approaches and can be a helpful tool to support researchers in conducting literature reviews." The evaluation on 308 publications demonstrated that combining keyword extraction, topic modeling, and hierarchical clustering effectively organizes scientific literature into meaningful clusters with 17 clusters ranging from 9 to 29 documents each.
Research paradigm
Design science research (DSR) / pragmatic positivism
Author conclusions
"In summary, we have developed an artifact based on the word2vec algorithm, LDA topic modeling, rapid automatic keyword extraction, and agglomerative hierarchical clustering. This artifact is a first step towards simplifying the task of literature reviews within scientific research." The authors conclude that "the artifact provides a supporting mechanism to speed up and standardize the process of literature reviews and increases automation of an otherwise entirely manual process to ultimately improve the quality and reproducibility of this important aspect of research."
Risk of bias
Small evaluation dataset (308 papers) limits generalizability; Manual evaluation by researchers may introduce subjective bias in cluster assessment; Single domain tested (industrial manufacturing sensor data) - cross-domain applicability unclear; Potential evaluator bias as authors evaluated their own artifact; OCR conversion quality may vary affecting text extraction; Selection bias: Small, domain-specific dataset (308 papers on industrial manufacturing sensor data) may not generalize; Evaluator bias: Manual evaluation of clusters by researchers may introduce subjectivity despite attempting to standardize process; OCR bias: Tesseract OCR conversion of PDF to text may introduce errors, especially for scanned older manuscripts; Algorithm-human mismatch: Authors note that "the decision criteria might differ from a human interpretation since humans tend to interpret the meaning of topics and they do not solely rely on statistics and logical reasoning"; Limited ground truth validation: No comparison with expert-curated gold standard classifications; Small evaluation dataset (n=308) may not generalize; Manual cluster validation by only 4 IS researchers introduces subjective bias; Licensing issues led to use of existing dataset rather than crawler-generated dataset; Limited to sensor data in industrial manufacturing domain
Limitations
- "The rather small evaluation data set of just 308 full-text papers which were manually checked if the proposed clustering of our model reflects the expectation of IS researchers, can only serve as a starting point for future research." Additionally, "Another limitation results from the conversion in plain text documents since all information stored in images and figures is not considered by the artifact." The authors also note that "a more rigorous evaluation of the artifact's utility for researchers during the creation of literature reviews in different contexts should be subject to future research."
Open questions raised
- Need for evaluation with larger datasets beyond 308 papers
- Comparative performance analysis of different topic modeling approaches (LDA vs. LSA vs. pLSA)
- More rigorous evaluation of artifact utility in different research contexts and domains
- Optimization of implemented algorithms
- Investigation of how to handle information stored in images and figures during text extraction
- Need for evaluation on larger and more diverse datasets across different research domains
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations