<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="brief-report" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2024.1506676</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Perspective</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>AI-assisted human clinical reasoning in the ICU: beyond &#x201C;to err is human&#x201D;</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" equal-contrib="yes">
<name><surname>El Gharib</surname> <given-names>Khalil</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref rid="fn1001" ref-type="author-notes"><sup>&#x2020;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2858851/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author" equal-contrib="yes">
<name><surname>Jundi</surname> <given-names>Bakr</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref rid="fn1001" ref-type="author-notes"><sup>&#x2020;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2871933/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Furfaro</surname> <given-names>David</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Abdulnour</surname> <given-names>Raja-Elie E.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2904028/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Division of Pulmonary and Critical Care Medicine, Rutgers Robert Wood Johnson Medical School</institution>, <addr-line>New Brunswick, NJ</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Division of Pulmonary and Critical Care Medicine, Brigham and Women&#x2019;s Hospital and Harvard Medical School</institution>, <addr-line>Boston, MA</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Division of Pulmonary and Critical Care Medicine, Beth Israel Deaconess Medical Center and Harvard Medical School</institution>, <addr-line>Boston, MA</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001">
<p>Edited by: Tse-Yen Yang, Asia University, Taiwan</p>
</fn>
<fn fn-type="edited-by" id="fn0002">
<p>Reviewed by: Dinh Tuan Phan Le, New York City Health and Hospitals Corporation, United States</p>
<p>Maria Hallo, Escuela Polit&#x00E9;cnica Nacional, Ecuador</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Raja-Elie E. Abdulnour, <email>rabdulnour@bwh.harvard.edu</email></corresp>
<fn fn-type="equal" id="fn1001">
<p><sup>&#x2020;</sup>These authors have contributed equally to this work</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>12</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>7</volume>
<elocation-id>1506676</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>10</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>11</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2024 El Gharib, Jundi, Furfaro and Abdulnour.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>El Gharib, Jundi, Furfaro and Abdulnour</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Diagnostic errors pose a significant public health challenge, affecting nearly 800,000 Americans annually, with even higher rates globally. In the ICU, these errors are particularly prevalent, leading to substantial morbidity and mortality. The clinical reasoning process aims to reduce diagnostic uncertainty and establish a plausible differential diagnosis but is often hindered by cognitive load, patient complexity, and clinician burnout. These factors contribute to cognitive biases that compromise diagnostic accuracy. Emerging technologies like large language models (LLMs) offer potential solutions to enhance clinical reasoning and improve diagnostic precision. In this perspective article, we explore the roles of LLMs, such as GPT-4, in addressing diagnostic challenges in critical care settings through a case study of a critically ill patient managed with LLM assistance.</p>
</abstract>
<kwd-group>
<kwd>large language models</kwd>
<kwd>clinical reasoning</kwd>
<kwd>diagnostic errors</kwd>
<kwd>artificial intelligence</kwd>
<kwd>critical care</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="25"/>
<page-count count="5"/>
<word-count count="3665"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Medicine and Public Health</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<title>Introduction</title>
<p>Diagnostic error is a public health concern. It is estimated that nearly 800,000 Americans die or are permanently disabled by diagnostic error in various clinical settings each year (<xref ref-type="bibr" rid="ref13">Newman-Toker et al., 2024</xref>). Globally, the incidence of diagnostic error is likely even higher as access to basic diagnostic testing resources can be limited in low-resource contexts, resulting in diagnostic delays for life-threatening diseases (<xref ref-type="bibr" rid="ref13">Newman-Toker et al., 2024</xref>). Central goals of the initial clinical reasoning process are to reduce diagnostic uncertainty and communicate a plausible differential diagnosis for safe and effective patient care. However, the process frequently faces a variety of challenges including cognitive load, high patient complexity, and burnout leading to inefficiencies and diagnostic errors (<xref ref-type="bibr" rid="ref12">National Academies Press, 2015</xref>). All these factors predispose clinicians to cognitive biases. Burnout may lead to an &#x2018;availability bias&#x2019;, wherein a clinician defaults to a familiar diagnosis rather than considering a broader range of possibilities, as it requires less mental effort. Addressing these challenges is crucial to mitigate reliance on heuristic shortcuts and improve diagnostic accuracy.</p>
<p>Critically ill patients are particularly prone to the harms from diagnostic errors, and by some estimates the prevalence of diagnostic error in patients admitted from the emergency department (ED) to the ICU exceeds 40% (<xref ref-type="bibr" rid="ref1">Aaronson et al., 2020</xref>; <xref ref-type="bibr" rid="ref3">Bergl et al., 2018</xref>). In addition, a meta-analysis demonstrated that ICU patients are twice as likely to have major misdiagnoses when compared to the patients admitted to the medical wards (<xref ref-type="bibr" rid="ref18">Shojania et al., 2003</xref>). Moreover, it is estimated that up to 30% of patients with a diagnostic error in the ICU die secondary to this error (<xref ref-type="bibr" rid="ref2">Auerbach et al., 2024</xref>). There is a critical need to uncover new approaches in clinical diagnosis and reasoning to improve patient outcomes, especially in the ICU. In this perspective article, we delve into the role of large language models (LLM) to address this important unmet clinical need as a framework to enhance the paradigm of human clinical reasoning in the ICU.</p>
</sec>
<sec id="sec2">
<title>Methods to improve diagnosis by enhancing clinical reasoning</title>
<p>The National Academy of Medicine describes improving diagnosis in healthcare as a &#x201C;moral, professional, and public health imperative&#x201D; (<xref ref-type="bibr" rid="ref19">Singh and Graber, 2015</xref>). It articulated eight objectives to improve diagnosis, many of which target clinical reasoning, support clinical decision-making, and encourage cognitive forcing strategies and checklists (<xref ref-type="bibr" rid="ref12">National Academies Press, 2015</xref>). Existing cognitive reasoning tools, such as reflection strategies or checklists result in clinically important improvements in diagnostic accuracy; however the overall impact is limited (<xref ref-type="bibr" rid="ref21">Staal et al., 2022</xref>). In recent years, LLMs have been an emerging tool used in clinical settings with the goal of having a more meaningful effect (<xref ref-type="bibr" rid="ref8">Lee et al., 2023</xref>). Although Artificial Intelligence (AI) tools have been used in healthcare for many decades, most have been trained on narrow datasets and provide support in specific contexts. In contrast, LLMs are generative AI tools trained on a vast text corpus. Therefore, by &#x201C;hacking the operating system of human civilization&#x201D; (<xref ref-type="bibr" rid="ref25">Williams, n.d.</xref>), LLMs can provide support in many language-dependent domains, including medicine. In recent years, several LLMs have been developed, including BERT, XLNet, Pathways Language Model (PaLM), Open Pretrained Transformer (OPT), and the most globally used GPT (<xref ref-type="bibr" rid="ref15">OSF, n.d.</xref>). The sheer number of parameters and the size of the training data of modern LLMs have opened many opportunities to support human cognitive tasks in the workplace (<xref ref-type="bibr" rid="ref23">Thirunavukarasu et al., 2023</xref>). Currently, LLM applications are already being leveraged by clinicians as these new tools showed broad use cases, from drafting pre-authorization documents to transcribing and summarizing encounter notes (<xref ref-type="bibr" rid="ref8">Lee et al., 2023</xref>; <xref ref-type="bibr" rid="ref10">Locke et al., 2021</xref>). As such, LLMs have the potential to assist in many of the obstacles to diagnostic excellence in the ICU.</p>
</sec>
<sec id="sec3">
<title>Emerging diagnostic reasoning properties of LLM</title>
<p>Providing high-quality responses to medical questions requires an understanding of the medical context, recollection of pertinent knowledge, and human-like reasoning (<xref ref-type="bibr" rid="ref20">Singhal et al., 2023</xref>). To validate the reasoning capabilities of LLMs in this domain, investigators tested them with licensing examinations (<xref ref-type="bibr" rid="ref20">Singhal et al., 2023</xref>; <xref ref-type="bibr" rid="ref5">Jin et al., 2020</xref>; <xref ref-type="bibr" rid="ref22">Suchman et al., 2023</xref>). The results indicated that while their performance did not excel in certain assessments, they reached the requisite in others (<xref ref-type="bibr" rid="ref7">Kung et al., 2023</xref>; <xref ref-type="bibr" rid="ref14">Nori et al., 2023</xref>). Beyond answering test questions, LLMs have been evaluated for assessing patient scenarios and providing guidance about diagnosis and clinical reasoning (<xref ref-type="bibr" rid="ref9">Liu et al., 2024</xref>). A recent study assessed the ability of LLMs&#x2019; to answer questions within critical care by extracting and responding to clinical concepts from the MIMIC III dataset, which contains medical information on patients admitted to critical care units (<xref ref-type="bibr" rid="ref9">Liu et al., 2024</xref>). GPT-4 demonstrated superior performance compared to its predecessors, including GPT and LLaMA, providing answers that were relevant, clear, logical and more complete (<xref ref-type="bibr" rid="ref9">Liu et al., 2024</xref>).</p>
<p>Research on LLMs&#x2019; diagnostic processes is advancing. When compared to human diagnosis, mixed results were seen with simple medical cases (<xref ref-type="bibr" rid="ref16">Rao et al., 2023</xref>), complex cases (<xref ref-type="bibr" rid="ref6">Kanjee et al., 2023</xref>) and gerontic ones (<xref ref-type="bibr" rid="ref17">Shea et al., 2023</xref>), suggesting that with clinician-guided prompting and appropriate data input (<xref ref-type="bibr" rid="ref4">Cabral et al., 2024</xref>), models could become more reliable. More recently, clinicians presented with challenging medical cases from the New England Journal of Medicine were compared in terms of their responses with and without LLM assistance. The LLM-assisted responses were more comprehensive and appropriate than those generated solely with textbooks and internet searches, highlighting the potential of these models as assistive tools in clinical decision-making (<xref ref-type="bibr" rid="ref11">McDuff et al., 2023</xref>). When clinicians were provided with LLM support, their diagnostic accuracy improved, as these models could articulate clinical reasoning arguments assessed by validated rubrics. In situations where human clinicians and GPT-4 were given cases with unstructured data and asked to produce problem representations, clinical reasoning, and differential diagnoses, GPT-4 demonstrated better clinical reasoning with a similar level of diagnostic accuracy compared to humans (<xref ref-type="bibr" rid="ref4">Cabral et al., 2024</xref>).</p>
<p>However, literature on LLMs&#x2019; capabilities in resolving critical care cases is limited. Critically ill patients often present with serious and complex multi-organ involvement, and require simultaneous diagnostic investigation and therapy, which makes the application of LLMs to these real-world situations fraught with difficulty. Herein, we present a case study of the diagnostic process of LLMs in a patient in the ICU at the Brigham and Women&#x2019;s Hospital to assess their potential to enhance human diagnosis.</p>
</sec>
<sec id="sec4">
<title>Case study</title>
<p>A 61-year-old female with a history of hypertension, breast cancer status post bilateral mastectomy in 4&#x202F;years prior to presentation, and recurrent ovarian cancer complicated by chemotherapy-induced thrombocytopenia, presenting with a five-day history of nausea, vomiting, diarrhea with poor oral intake, and three days of headache, altered mental status, confusion, slurred speech, and gait instability. The patient initially presented to an out-of-state hospital and was diagnosed with a urinary tract infection and was treated with intravenous fluid and piperacillin-tazobactam and was then sent home with a prescription for ciprofloxacin. The patient presented to her primary oncology provider the day after discharge where she was referred to the ED for further evaluation. The physical examination at the ED was notable for tachycardia, facial myoclonus, gait instability, and lower back tenderness. Labs demonstrated leukocytosis with a white blood cell of 36&#x202F;K/uL, AST 148&#x202F;U/L / ALT 81&#x202F;U/L, ALP 328&#x202F;U/L. Urinalysis revealed pyuria. Computed tomography (CT) of the head demonstrated no acute findings. The patient was admitted to the hospital. The next day, the patient had a fever with maximum temperature of 101.5F with worsening tachycardia and new oxygen requirement of 2&#x202F;L nasal cannula to maintain oxygen saturation at &#x003E;90%. The patient then developed worsening respiratory failure requiring emergent intubation shortly after undergoing computed tomography pulmonary angiogram for concerns for a pulmonary embolus. The patient was transferred to the ICU after intubation for further management. On hospital day 3, anti-microbials were broadened to Vancomycin/Cefepime/Ampicillin/Acyclovir per neurology recommendations to empirically treat meningitis while awaiting lumbar puncture. The next day, brain magnetic resonance imaging was negative for acute abnormality. Electroencephalogram revealed moderate bilateral cerebral dysfunction consistent with encephalopathy. Lumbar puncture was performed on hospital day 4 and yielded colonies of <italic>Listeria monocytogenes</italic> and blood cultures from two days prior also grew <italic>Listeria monocytogenes</italic>. A summary of the timeline of events is presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>. More details of the events are presented in <xref rid="SM1" ref-type="supplementary-material">Supplementary material</xref>.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Timeline graphic comparing the clinical course of a patient as managed by clinicians (black) versus the recommendations made by GPT-4 (green).</p>
</caption>
<graphic xlink:href="frai-07-1506676-g001.tif"/>
</fig>
</sec>
<sec sec-type="discussion" id="sec5">
<title>Discussion</title>
<p>In <xref ref-type="fig" rid="fig1">Figure 1</xref>, we present the clinical course of the patient from the time they initially presented to the out-of-state hospital until the time of the final diagnosis. Using GPT-4, we provided a prompt (see <xref rid="SM1" ref-type="supplementary-material">Supplementary material</xref>) to instruct the model on providing a clinical summary with the top 10 differential diagnoses, additional diagnostic tests a physician should obtain, and initial management based on the history and physical written by the physician in the electronic medical record. The full response from GPT-4 is provided in <xref rid="SM1" ref-type="supplementary-material">Supplementary material</xref>. As indicated in red, the GPT-4 response recommended performing a lumbar puncture upon presentation at the ED and initiating broad-spectrum anti-microbials, including vancomycin, cefepime, ampicillin, and acyclovir, which was delayed by 48&#x202F;h in the real-life scenario. This suggests a clinical utility of using LLM models in the care of critically ill patients to enhance our clinical reasoning and diagnostic processes, ultimately aiming to improve patient care.</p>
<p>In <xref ref-type="fig" rid="fig2">Figure 2</xref>, we highlight potential targets where LLMs can assist in the care of the critically ill. During the initial patient evaluation, physicians spend a significant amount of time reviewing patients&#x2019; previous diagnostic work-ups, which can sometimes be overwhelming. GPT-4 can aid in reviewing a patient&#x2019;s history to identify potential diagnostic anchoring biases (<xref ref-type="bibr" rid="ref11">McDuff et al., 2023</xref>). Additionally, LLMs could assist in triaging and prioritizing patients who need immediate intervention based on the acuity of their presentations. In addition, GPT-4 can help organize and summarize patient data, including history, lab results, and imaging, to provide clinicians with a concise overview (<xref ref-type="bibr" rid="ref11">McDuff et al., 2023</xref>). GPT-4 can also provide differential diagnosis suggestions based on the presented symptoms and test results, helping to ensure that clinicians consider less common diagnoses they may have overlooked alongside the more common ones. In situations where complex management questions arise, GPT-4 can serve as an educational resource, providing quick access to relevant guidelines and literature (<xref ref-type="bibr" rid="ref11">McDuff et al., 2023</xref>). While it cannot replace human compassion, GPT-4 can offer support to healthcare workers under stress by providing a space to quickly debrief or reflect on difficult cases, which may help manage the emotional toll of healthcare work (<xref ref-type="bibr" rid="ref24">Tu et al., 2024</xref>).</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Enhancing ICU Patient Care with GPT-4: Addressing common hindrances to diagnosing critically ill patients (adapted with permission of the American Thoracic Society. Copyright &#x00A9; 2024 American Thoracic Society. All rights reserved (<xref ref-type="bibr" rid="ref3">Bergl et al., 2018</xref>). Annals of the American Thoracic Society is an official journal of the American Thoracic Society. Readers are encouraged to read the entire article for the correct context at <ext-link xlink:href="https://www.atsjournals.org/doi/10.1513/AnnalsATS.201801-068PS" ext-link-type="uri">https://www.atsjournals.org/doi/10.1513/AnnalsATS.201801-068PS</ext-link>. The authors, editors, and The American Thoracic Society are not responsible for errors or omissions in adaptations).</p>
</caption>
<graphic xlink:href="frai-07-1506676-g002.tif"/>
</fig>
<p>While the proof-of-concept case presented here demonstrates the potential benefits of LLMs, we acknowledge the limitations inherent in using a single case study to justify the broader application of LLMs in clinical reasoning. It is essential to approach the use of LLMs with caution, recognizing their limitations and potential biases. LLMs are trained on extensive datasets that may include biased information, leading to skewed responses. Additionally, the complexity of medical decision-making, characterized by nuanced and context-specific knowledge, can be challenging for LLMs, which rely heavily on pattern recognition rather than deep understanding. LLMs can also produce hallucinations, generating plausible-sounding but incorrect information, which can be dangerous in a clinical setting. Moreover, ethical and legal implications must be carefully considered, including potential malpractice issues and the necessity of informed consent for patients. Developing robust regulatory frameworks will be crucial to responsibly harness the potential of LLMs in clinical practice. Further research is needed to evaluate the effectiveness and safety of LLMs in supporting ICU clinical reasoning.</p>
</sec>
<sec sec-type="conclusions" id="sec6">
<title>Conclusion</title>
<p>Humans err, and errors are expensive and harmful in the healthcare setting. Diagnostic error remains a hidden epidemic in the ICU. It is time for the critical care community to acknowledge the gravity of the issue and recognize the potential for emerging technologies like LLMs to serve as pivotal allies in the ICU. The proof-of-concept case study presented here, along with the proposed integration of GPT-4 into the clinical workflow, has demonstrated the potential advantages and enhancements to diagnostic accuracy that LLMs can offer. As we find ourselves on the brink of a new era in medicine, it is becoming increasingly clear that the judicious use of AI, exemplified by LLMs, can usher in a paradigm shift toward more precise, efficient, and compassionate care in the ICU. To fully realize this potential, it is essential to educate physicians in the use of LLMs to augment their diagnostic and clinical reasoning skills. With ongoing research, refinement, and integration, LLMs could well become an indispensable component of critical care, mitigating the risk of diagnostic errors and elevating the standard of patient care to unprecedented heights.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="sec7">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/<xref rid="SM1" ref-type="supplementary-material">Supplementary material</xref>.</p>
</sec>
<sec sec-type="ethics-statement" id="sec8">
<title>Ethics statement</title>
<p>Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from the participants or the participants&#x2019; legal guardians/next of kin in accordance with the national legislation and the institutional requirements. Written informed consent was obtained from the individual(s) for the publication of any potentially identifiable images or data included in this article.</p>
</sec>
<sec sec-type="author-contributions" id="sec9">
<title>Author contributions</title>
<p>KE: Writing &#x2013; original draft. BJ: Writing &#x2013; original draft. DF: Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. R-EA: Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="funding-information" id="sec10">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<ack>
<p>We would like to thank Dr. Titi Olanipekun for providing us with the clinical case used in this article. ChatGPT 4o was used for grammatical check. The authors validated the output and take full responsibility for the content. Chat GPT4 generative AI model was used to compare its performance as opposed to physician performance in clinical reasoning regarding a clinical case. It was not used for any other purposes.</p>
</ack>
<sec sec-type="COI-statement" id="sec11">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="sec12">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="sec13">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2024.1506676/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/frai.2024.1506676/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.DOCX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aaronson</surname> <given-names>E.</given-names></name> <name><surname>Jansson</surname> <given-names>P.</given-names></name> <name><surname>Wittbold</surname> <given-names>K.</given-names></name> <name><surname>Flavin</surname> <given-names>S.</given-names></name> <name><surname>Borczuk</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Unscheduled return visits to the emergency department with ICU admission: a trigger tool for diagnostic error</article-title>. <source>Am. J. Emerg. Med.</source> <volume>38</volume>, <fpage>1584</fpage>&#x2013;<lpage>1587</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ajem.2019.158430</pub-id>, PMID: <pub-id pub-id-type="pmid">31699427</pub-id></citation>
</ref>
<ref id="ref2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Auerbach</surname> <given-names>A. D.</given-names></name> <name><surname>Lee</surname> <given-names>T. M.</given-names></name> <name><surname>Hubbard</surname> <given-names>C. C.</given-names></name> <name><surname>Ranji</surname> <given-names>S. R.</given-names></name> <name><surname>Raffel</surname> <given-names>K.</given-names></name> <name><surname>Valdes</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Diagnostic errors in hospitalized adults who died or were transferred to intensive care</article-title>. <source>JAMA Intern. Med.</source> <volume>184</volume>, <fpage>164</fpage>&#x2013;<lpage>173</lpage>. doi: <pub-id pub-id-type="doi">10.1001/jamainternmed.2023.7347</pub-id>, PMID: <pub-id pub-id-type="pmid">38190122</pub-id></citation>
</ref>
<ref id="ref3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergl</surname> <given-names>P. A.</given-names></name> <name><surname>Nanchal</surname> <given-names>R. S.</given-names></name> <name><surname>Singh</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>Diagnostic error in the critically III: defining the problem and exploring next steps to advance intensive care unit safety</article-title>. <source>Ann. Am. Thorac. Soc.</source> <volume>15</volume>, <fpage>903</fpage>&#x2013;<lpage>907</lpage>. doi: <pub-id pub-id-type="doi">10.1513/AnnalsATS.201801-068PS</pub-id>, PMID: <pub-id pub-id-type="pmid">29742359</pub-id></citation>
</ref>
<ref id="ref4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cabral</surname> <given-names>S.</given-names></name> <name><surname>Restrepo</surname> <given-names>D.</given-names></name> <name><surname>Kanjee</surname> <given-names>Z.</given-names></name> <name><surname>Wilson</surname> <given-names>P.</given-names></name> <name><surname>Crowe</surname> <given-names>B.</given-names></name> <name><surname>Abdulnour</surname> <given-names>R.-E.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Clinical reasoning of a generative artificial intelligence model compared with physicians</article-title>. <source>JAMA Intern. Med.</source> <volume>184</volume>:<fpage>581</fpage>. doi: <pub-id pub-id-type="doi">10.1001/jamainternmed.2024.0295</pub-id></citation>
</ref>
<ref id="ref5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jin</surname> <given-names>D.</given-names></name> <name><surname>Pan</surname> <given-names>E.</given-names></name> <name><surname>Oufattole</surname> <given-names>N.</given-names></name> <name><surname>Weng</surname> <given-names>W.-H.</given-names></name> <name><surname>Fang</surname> <given-names>H.</given-names></name> <name><surname>Szolovits</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>What disease does this patient have? A large-scale open domain question answering dataset from medical exams</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2009.13081</pub-id></citation>
</ref>
<ref id="ref6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kanjee</surname> <given-names>Z.</given-names></name> <name><surname>Crowe</surname> <given-names>B.</given-names></name> <name><surname>Rodman</surname> <given-names>A.</given-names></name></person-group> (<year>2023</year>). <article-title>Accuracy of a generative artificial intelligence model in a complex diagnostic challenge</article-title>. <source>JAMA</source> <volume>330</volume>, <fpage>78</fpage>&#x2013;<lpage>80</lpage>. doi: <pub-id pub-id-type="doi">10.1001/jama.2023.8288</pub-id>, PMID: <pub-id pub-id-type="pmid">37318797</pub-id></citation>
</ref>
<ref id="ref7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kung</surname> <given-names>T. H.</given-names></name> <name><surname>Cheatham</surname> <given-names>M.</given-names></name> <name><surname>Medenilla</surname> <given-names>A.</given-names></name> <name><surname>Sillos</surname> <given-names>C.</given-names></name> <name><surname>De Leon</surname> <given-names>L.</given-names></name> <name><surname>Elepa&#x00F1;o</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models</article-title>. <source>PLOS Digit Health</source> <volume>2</volume>:<fpage>e0000198</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pdig.0000198</pub-id>, PMID: <pub-id pub-id-type="pmid">36812645</pub-id></citation>
</ref>
<ref id="ref8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>P.</given-names></name> <name><surname>Bubeck</surname> <given-names>S.</given-names></name> <name><surname>Petro</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Benefits, limits, and risks of GPT-4 as an AI Chatbot for medicine</article-title>. <source>N. Engl. J. Med.</source> <volume>388</volume>, <fpage>1233</fpage>&#x2013;<lpage>1239</lpage>. doi: <pub-id pub-id-type="doi">10.1056/NEJMsr2214184</pub-id>, PMID: <pub-id pub-id-type="pmid">36988602</pub-id></citation>
</ref>
<ref id="ref9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>D.</given-names></name> <name><surname>Ding</surname> <given-names>C.</given-names></name> <name><surname>Bold</surname> <given-names>D.</given-names></name> <name><surname>Bouvier</surname> <given-names>M.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Shickel</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>[2401.13588] evaluation of general large language models in contextually assessing semantic concepts extracted from adult critical care electronic health record notes</article-title>. <source>arXiv</source>.</citation>
</ref>
<ref id="ref10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Locke</surname> <given-names>S.</given-names></name> <name><surname>Bashall</surname> <given-names>A.</given-names></name> <name><surname>Al-Adely</surname> <given-names>S.</given-names></name> <name><surname>Moore</surname> <given-names>J.</given-names></name> <name><surname>Wilson</surname> <given-names>A.</given-names></name> <name><surname>Kitchen</surname> <given-names>G. B.</given-names></name></person-group> (<year>2021</year>). <article-title>Natural language processing in medicine: a review</article-title>. <source>Trends Anaesthesia Crit. Care</source> <volume>38</volume>, <fpage>4</fpage>&#x2013;<lpage>9</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tacc.2021.02.007</pub-id></citation>
</ref>
<ref id="ref11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDuff</surname> <given-names>D.</given-names></name> <name><surname>Schaekermann</surname> <given-names>M.</given-names></name> <name><surname>Tu</surname> <given-names>T.</given-names></name> <name><surname>Palepu</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>A.</given-names></name> <name><surname>Garrison</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Towards Accurate Differential Diagnosis with Large Language Models</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2312.00164</pub-id></citation>
</ref>
<ref id="ref12">
<citation citation-type="book"><person-group person-group-type="author">
<collab id="coll1">National Academies Press</collab>
</person-group> (<year>2015</year>) in <source>Committee on diagnostic error in health care, board on health care services, Institute of Medicine, the National Academies of sciences, engineering, and medicine. Improving diagnosis in health care</source>. eds. <person-group person-group-type="editor"><name><surname>Balogh</surname> <given-names>E. P.</given-names></name> <name><surname>Miller</surname> <given-names>B. T.</given-names></name> <name><surname>Ball</surname> <given-names>J. R.</given-names></name></person-group> (<publisher-loc>Washington (DC)</publisher-loc>: <publisher-name>National Academies Press (US)</publisher-name>).</citation>
</ref>
<ref id="ref13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newman-Toker</surname> <given-names>D. E.</given-names></name> <name><surname>Nassery</surname> <given-names>N.</given-names></name> <name><surname>Schaffer</surname> <given-names>A. C.</given-names></name> <name><surname>Yu-Moe</surname> <given-names>C. W.</given-names></name> <name><surname>Clemens</surname> <given-names>G. D.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Burden of serious harms from diagnostic error in the USA</article-title>. <source>BMJ Qual. Saf.</source> <volume>33</volume>, <fpage>109</fpage>&#x2013;<lpage>120</lpage>. doi: <pub-id pub-id-type="doi">10.1136/bmjqs-2021-014130</pub-id>, PMID: <pub-id pub-id-type="pmid">37460118</pub-id></citation>
</ref>
<ref id="ref14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nori</surname> <given-names>H.</given-names></name> <name><surname>King</surname> <given-names>N.</given-names></name> <name><surname>McKinney</surname> <given-names>S. M.</given-names></name> <name><surname>Carignan</surname> <given-names>D.</given-names></name> <name><surname>Horvitz</surname> <given-names>E.</given-names></name></person-group> (<year>2023</year>). <article-title>Capabilities of GPT-4 on medical challenge problems</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2303.13375</pub-id></citation>
</ref>
<ref id="ref15">
<citation citation-type="other"><person-group person-group-type="author">
<collab id="coll2">OSF</collab>
</person-group>. (Accessed 2023 Dec 29). Available at: <ext-link xlink:href="https://osf.io/preprints/edarxiv/5er8f" ext-link-type="uri">https://osf.io/preprints/edarxiv/5er8f</ext-link></citation>
</ref>
<ref id="ref16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>A.</given-names></name> <name><surname>Pang</surname> <given-names>M.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Kamineni</surname> <given-names>M.</given-names></name> <name><surname>Lie</surname> <given-names>W.</given-names></name> <name><surname>Prasad</surname> <given-names>A. K.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Assessing the utility of chatgpt throughout the entire clinical workflow</article-title>. <source>medRxiv</source>. doi: <pub-id pub-id-type="doi">10.2196/48659</pub-id></citation>
</ref>
<ref id="ref17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shea</surname> <given-names>Y.-F.</given-names></name> <name><surname>Lee</surname> <given-names>C. M. Y.</given-names></name> <name><surname>Ip</surname> <given-names>W. C. T.</given-names></name> <name><surname>Luk</surname> <given-names>D. W. A.</given-names></name> <name><surname>Wong</surname> <given-names>S. S. W.</given-names></name></person-group> (<year>2023</year>). <article-title>Use of GPT-4 to analyze medical Records of Patients with Extensive Investigations and Delayed Diagnosis</article-title>. <source>JAMA Netw. Open</source> <volume>6</volume>:<fpage>e2325000</fpage>. doi: <pub-id pub-id-type="doi">10.1001/jamanetworkopen.2023.25000</pub-id>, PMID: <pub-id pub-id-type="pmid">37578798</pub-id></citation>
</ref>
<ref id="ref18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shojania</surname> <given-names>K. G.</given-names></name> <name><surname>Burton</surname> <given-names>E. C.</given-names></name> <name><surname>McDonald</surname> <given-names>K. M.</given-names></name> <name><surname>Goldman</surname> <given-names>L.</given-names></name></person-group> (<year>2003</year>). <article-title>Changes in rates of autopsy-detected diagnostic errors over time: a systematic review</article-title>. <source>JAMA</source> <volume>289</volume>, <fpage>2849</fpage>&#x2013;<lpage>2856</lpage>. doi: <pub-id pub-id-type="doi">10.1001/jama.289.21.2849</pub-id>, PMID: <pub-id pub-id-type="pmid">12783916</pub-id></citation>
</ref>
<ref id="ref19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>H.</given-names></name> <name><surname>Graber</surname> <given-names>M. L.</given-names></name></person-group> (<year>2015</year>). <article-title>Improving diagnosis in health care--the next imperative for patient safety</article-title>. <source>N. Engl. J. Med.</source> <volume>373</volume>, <fpage>2493</fpage>&#x2013;<lpage>2495</lpage>. doi: <pub-id pub-id-type="doi">10.1056/NEJMp1512241</pub-id>, PMID: <pub-id pub-id-type="pmid">26559457</pub-id></citation>
</ref>
<ref id="ref20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singhal</surname> <given-names>K.</given-names></name> <name><surname>Azizi</surname> <given-names>S.</given-names></name> <name><surname>Tu</surname> <given-names>T.</given-names></name> <name><surname>Mahdavi</surname> <given-names>S. S.</given-names></name> <name><surname>Wei</surname> <given-names>J.</given-names></name> <name><surname>Chung</surname> <given-names>H. W.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Large language models encode clinical knowledge</article-title>. <source>Nature</source> <volume>620</volume>, <fpage>172</fpage>&#x2013;<lpage>180</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41586-023-06291-2</pub-id>, PMID: <pub-id pub-id-type="pmid">37438534</pub-id></citation>
</ref>
<ref id="ref21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staal</surname> <given-names>J.</given-names></name> <name><surname>Hooftman</surname> <given-names>J.</given-names></name> <name><surname>Gunput</surname> <given-names>S. T. G.</given-names></name> <name><surname>Mamede</surname> <given-names>S.</given-names></name> <name><surname>Frens</surname> <given-names>M. A.</given-names></name> <name><surname>Van den Broek</surname> <given-names>W. W.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Effect on diagnostic accuracy of cognitive reasoning tools for the workplace setting: systematic review and meta-analysis</article-title>. <source>BMJ Qual. Saf.</source> <volume>31</volume>, <fpage>899</fpage>&#x2013;<lpage>910</lpage>. doi: <pub-id pub-id-type="doi">10.1136/bmjqs-2022-014865</pub-id>, PMID: <pub-id pub-id-type="pmid">36396150</pub-id></citation>
</ref>
<ref id="ref22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suchman</surname> <given-names>K.</given-names></name> <name><surname>Garg</surname> <given-names>S.</given-names></name> <name><surname>Trindade</surname> <given-names>A. J.</given-names></name></person-group> (<year>2023</year>). <article-title>Chat generative Pretrained transformer fails the multiple-choice American College of Gastroenterology self-assessment test</article-title>. <source>Am. J. Gastroenterol.</source> <volume>118</volume>, <fpage>2280</fpage>&#x2013;<lpage>2282</lpage>. doi: <pub-id pub-id-type="doi">10.14309/ajg.0000000000002320</pub-id>, PMID: <pub-id pub-id-type="pmid">37212584</pub-id></citation>
</ref>
<ref id="ref23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thirunavukarasu</surname> <given-names>A. J.</given-names></name> <name><surname>Ting</surname> <given-names>D. S. J.</given-names></name> <name><surname>Elangovan</surname> <given-names>K.</given-names></name> <name><surname>Gutierrez</surname> <given-names>L.</given-names></name> <name><surname>Tan</surname> <given-names>T. F.</given-names></name> <name><surname>Ting</surname> <given-names>D. S. W.</given-names></name></person-group> (<year>2023</year>). <article-title>Large language models in medicine</article-title>. <source>Nat. Med.</source> <volume>29</volume>, <fpage>1930</fpage>&#x2013;<lpage>1940</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41591-023-02448-8</pub-id>, PMID: <pub-id pub-id-type="pmid">37460753</pub-id></citation>
</ref>
<ref id="ref24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tu</surname> <given-names>T.</given-names></name> <name><surname>Palepu</surname> <given-names>A.</given-names></name> <name><surname>Schaekermann</surname> <given-names>M.</given-names></name> <name><surname>Saab</surname> <given-names>K.</given-names></name> <name><surname>Freyberg</surname> <given-names>J.</given-names></name> <name><surname>Tanno</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Towards conversational diagnostic AI</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2401.05654</pub-id></citation>
</ref>
<ref id="ref25">
<citation citation-type="other"><person-group person-group-type="author">
<name><surname>Williams</surname> <given-names>D.</given-names></name>
</person-group> Yuval Noah Harari argues that AI has hacked the operating system of human civilisation. [cited 2024 Feb 13]. Available at: <ext-link xlink:href="https://www.economist.com/by-invitation/2023/04/28/yuval-noah-harari-argues-that-ai-has-hacked-the-operating-system-of-human-civilisation?utm_content=section_content&#x0026;gad_source=1&#x0026;gclid=CjwKCAiAq4KuBhA6EiwArMAw1LjfQ_0TN5W2zjVDoD5AZEZRPkqo-4JmWTZyVikACDI4GJJ9wa0WUBoCbOkQAvD_BwE&#x0026;gclsrc=aw.ds" ext-link-type="uri">https://www.economist.com/by-invitation/2023/04/28/yuval-noah-harari-argues-that-ai-has-hacked-the-operating-system-of-human-civilisation?utm_content=section_content&#x0026;gad_source=1&#x0026;gclid=CjwKCAiAq4KuBhA6EiwArMAw1LjfQ_0TN5W2zjVDoD5AZEZRPkqo-4JmWTZyVikACDI4GJJ9wa0WUBoCbOkQAvD_BwE&#x0026;gclsrc=aw.ds</ext-link></citation>
</ref>
</ref-list>
</back>
</article>