A Critical Assessment of Large Language Models for Systematic Reviews: Utilizing ChatGPT for Complex Data Extraction
JMIR AI, volume 4, 2025[10.2196/68097]
12 Pages Posted: 19 Apr 2024 Last revised: 20 Oct 2025
Date Written: April 17, 2024
Abstract
Objectives: To assess the effectiveness of large language models (LLMs) for systematic reviews, focusing on measures from basic details to complex criteria requiring nuanced evaluations.
Methods: Screening the full text of 10 COVID-19 modeling studies, we analyzed three basic measures of study settings, e.g., analysis location, and three complex measures of behavioral components in models, e.g., risk perception. To extract data on these measures, we conducted 60 manual codings and compared them with 420 queries spanning seven iterations in ChatGPT.
Results: ChatGPT demonstrated 72% overall accuracy in extracting 60 data elements, performing better in extracting explicitly stated study settings (93%) than subjective behavioral components (50%).
Discussion: While ChatGPT’s accuracy improved as prompts were refined, varying accuracy across measures highlights its limitations.
Conclusion: We underscore LLMs’ utility in systematic reviews for basic, explicit data extraction but reveal significant limitations in handling nuanced, subjective criteria, emphasizing the current necessity for human oversight.
Suggested Citation: Suggested Citation