Assessing the Accuracy of Generative Conversational Artificial Intelligence in Debunking Sleep Health Myths: Comparative Study with Expert Analysis

28 Pages Posted: 17 Jan 2024

Date Written: December 22, 2023

Abstract

Background: Adequate sleep is essential for maintaining both individual and public health, positively affecting cognition and well-being, and reducing chronic disease risks. It also plays a significant role in driving the economy, public safety, and managing healthcare costs. Digital tools, including websites, sleep trackers, and apps, are key in promoting sleep health education. Conversational Artificial Intelligence (AI) like ChatGPT offers accessible, personalized advice on sleep health but raises concerns about potential misinformation. This underscores the importance of ensuring AI-driven sleep health information is accurate, given its significant impact on individual and public health, and the spread of sleep-related myths.

Objective: The study aims to examine ChatGPT’s capability to debunk sleep-related disbeliefs.

Methods: ChatGPT was asked to categorize twenty sleep-related false myths identified by ten sleep experts and to rate them in terms of falseness and public health significance, on a 5-point Likert scale. Sensitivity, positive predictive value, and inter-rater agreement were also calculated.

Results: ChatGPT labeled a significant portion (85%, n=17) of the statements as "false" (45%, n=9) or "generally false" (40%, n=8), with varying accuracy across different domains. For instance, it correctly identified most myths about "sleep timing", "sleep duration", and "behaviors during sleep", while it had varying degrees of success with other categories like "pre-sleep behaviors" and "brain function and sleep". ChatGPT's assessment of the degree of falseness and public health significance, on the 5-point Likert scale, showed an average score of 3.45 (SD=0.85) and 3.15 (SD=0.96), respectively, indicating a good level of accuracy in identifying the falseness of statements and a good understanding of their impact on public health. The AI-based tool showed a sensitivity of 85% and a perfect positive predictive value of 100%. Overall, this indicates that when ChatGPT labels a statement as false, it is highly reliable, but it may miss identifying some false statements. When comparing with expert ratings, high intra-class correlation coefficients (ICCs) between ChatGPT’s appraisals and expert opinions could be found, suggesting that the AI’s ratings were generally aligned with expert views on falseness (ICC=.83, P<.0001) and public health significance (ICC=.79, P=.001) of sleep-related myths.

Conclusions: ChatGPT-4 can accurately address sleep-related queries and debunk sleep-related myths, with a performance comparable to sleep experts, even if, given its limitations, the AI cannot completely replace expert opinions, especially in nuanced and complex fields like sleep health, but can be a valuable complement in the dissemination of updated information and promotion of healthy behaviors.

Note:
Funding Information: No funding received.

Conflict of Interests: No competing interest to declare.

Keywords: sleep; generative conversational artificial intelligence; chatbot; ChatGPT; misinformation

Suggested Citation

Bragazzi, Nicola Luigi and Garbarino, Sergio, Assessing the Accuracy of Generative Conversational Artificial Intelligence in Debunking Sleep Health Myths: Comparative Study with Expert Analysis (December 22, 2023). Available at SSRN: https://ssrn.com/abstract=4673743 or http://dx.doi.org/10.2139/ssrn.4673743

Sergio Garbarino

University of Genoa

Genoa
Italy

Do you have a job opening that you would like to promote on SSRN?

Paper statistics

Downloads
79
Abstract Views
581
Rank
668,986
PlumX Metrics