A Critical Assessment of Large Language Models for Systematic Reviews: Utilizing ChatGPT for Complex Data Extraction

JMIR AI, volume 4, 2025[10.2196/68097]

12 Pages Posted: 19 Apr 2024 Last revised: 20 Oct 2025

See all articles by Hesam Mahmoudi

Hesam Mahmoudi

Harvard University - Harvard Medical School

Doris Chang

Harvard University - Harvard Medical School

Hannah Lee

Harvard University - Harvard Medical School

Navid Ghaffarzadegan

Virginia Tech - Grado Department of Industrial and Systems Engineering (ISE)

Mohammad S. Jalali

Harvard University - Harvard Medical School; Massachusetts Institute of Technology (MIT)

Date Written: April 17, 2024

Abstract

Objectives: To assess the effectiveness of large language models (LLMs) for systematic reviews, focusing on measures from basic details to complex criteria requiring nuanced evaluations.

Methods: Screening the full text of 10 COVID-19 modeling studies, we analyzed three basic measures of study settings, e.g., analysis location, and three complex measures of behavioral components in models, e.g., risk perception. To extract data on these measures, we conducted 60 manual codings and compared them with 420 queries spanning seven iterations in ChatGPT.

Results: ChatGPT demonstrated 72% overall accuracy in extracting 60 data elements, performing better in extracting explicitly stated study settings (93%) than subjective behavioral components (50%).

Discussion: While ChatGPT’s accuracy improved as prompts were refined, varying accuracy across measures highlights its limitations.

Conclusion: We underscore LLMs’ utility in systematic reviews for basic, explicit data extraction but reveal significant limitations in handling nuanced, subjective criteria, emphasizing the current necessity for human oversight.

Suggested Citation

Mahmoudi, Hesam and Chang, Doris and Lee, Hannah and Ghaffarzadegan, Navid and Jalali, Mohammad S., A Critical Assessment of Large Language Models for Systematic Reviews: Utilizing ChatGPT for Complex Data Extraction (April 17, 2024). JMIR AI, volume 4, 2025[10.2196/68097], Available at SSRN: https://ssrn.com/abstract=4797024 or http://dx.doi.org/10.2196/68097

Hesam Mahmoudi

Harvard University - Harvard Medical School ( email )

125 Nashua St
Boston, MA 02114
United States

Doris Chang

Harvard University - Harvard Medical School ( email )

25 Shattuck St
Boston, MA 02115
United States

Hannah Lee

Harvard University - Harvard Medical School ( email )

25 Shattuck St
Boston, MA 02115
United States

Navid Ghaffarzadegan

Virginia Tech - Grado Department of Industrial and Systems Engineering (ISE) ( email )

Mohammad S. Jalali (Contact Author)

Harvard University - Harvard Medical School ( email )

101 Merrimac St
Suite 1010
Boston, MA 02114
United States

HOME PAGE: http://mj-lab.mgh.harvard.edu

Massachusetts Institute of Technology (MIT) ( email )

77 Massachusetts Avenue
50 Memorial Drive
Cambridge, MA 02139-4307
United States

Do you have a job opening that you would like to promote on SSRN?

Paper statistics

Downloads
388
Abstract Views
1,864
Rank
195,842
PlumX Metrics