<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Med.</journal-id>
<journal-title>Frontiers in Medicine</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Med.</abbrev-journal-title>
<issn pub-type="epub">2296-858X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmed.2021.785711</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Medicine</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Rule-Based Models for Risk Estimation and Analysis of In-hospital Mortality in Emergency and Critical Care</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Haas</surname> <given-names>Oliver</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1233930/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Maier</surname> <given-names>Andreas</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1103292/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Rothgang</surname> <given-names>Eva</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1426235/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Industrial Engineering and Health, Institute of Medical Engineering, Technical University Amberg-Weiden</institution>, <addr-line>Weiden</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Pattern Recognition Lab, Department of Computer Science, Technical Faculty, Friedrich-Alexander University</institution>, <addr-line>Erlangen</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Zhongheng Zhang, Sir Run Run Shaw Hospital, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Cledy Eliana Santos, Community Health Service of Hospital Concei&#x000E7;&#x000E3;o Group, Brazil; Rui Guo, First Affiliated Hospital of Chongqing Medical University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Oliver Haas <email>o.haas&#x00040;oth-aw.de</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Intensive Care Medicine and Anesthesiology, a section of the journal Frontiers in Medicine</p></fn></author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>785711</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Haas, Maier and Rothgang.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Haas, Maier and Rothgang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license></permissions>
<abstract><p>We propose a novel method that uses associative classification and odds ratios to predict in-hospital mortality in emergency and critical care. Manual mortality risk scores have previously been used to assess the care needed for each patient and their need for palliative measures. Automated approaches allow providers to get a quick and objective estimation based on electronic health records. We use association rule mining to find relevant patterns in the dataset. The odds ratio is used instead of classical association rule mining metrics as a quality measure to analyze association instead of frequency. The resulting measures are used to estimate the in-hospital mortality risk. We compare two prediction models: one minimal model with socio-demographic factors that are available at the time of admission and can be provided by the patients themselves, namely gender, ethnicity, type of insurance, language, and marital status, and a full model that additionally includes clinical information like diagnoses, medication, and procedures. The method was tested and validated on MIMIC-IV, a publicly available clinical dataset. The minimal prediction model achieved an area under the receiver operating characteristic curve value of 0.69, while the full prediction model achieved a value of 0.98. The models serve different purposes. The minimal model can be used as a first risk assessment based on patient-reported information. The full model expands on this and provides an updated risk assessment each time a new variable occurs in the clinical case. In addition, the rules in the models allow us to analyze the dataset based on data-backed rules. We provide several examples of interesting rules, including rules that hint at errors in the underlying data, rules that correspond to existing epidemiological research, and rules that were previously unknown and can serve as starting points for future studies.</p></abstract>
<kwd-group>
<kwd>in-hospital mortality</kwd>
<kwd>critical care</kwd>
<kwd>odds ratio</kwd>
<kwd>associative classification</kwd>
<kwd>machine learning</kwd>
<kwd>artificial intelligence</kwd>
</kwd-group>
<contract-sponsor id="cn001">BayWISS Doctoral Consortium Health Research</contract-sponsor>
<contract-sponsor id="cn002">bidt Fellowship</contract-sponsor>
<contract-sponsor id="cn003">Bayerisches Staatsministerium f&#x000FC;r Wissenschaft, Forschung und Kunst<named-content content-type="fundref-id">10.13039/501100005341</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="3"/>
<equation-count count="2"/>
<ref-count count="42"/>
<page-count count="13"/>
<word-count count="7887"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The term in-hospital mortality defines the death of a patient during their stay at the hospital. Especially in emergency and critical care, many patients die in the hospital. While not all of these deaths can be prevented, early knowledge of a patient&#x00027;s in-hospital mortality risk can be used to assess the patient&#x00027;s status and necessary adjustments to this patient&#x00027;s care, reducing missed care and decreasing mortality rates (<xref ref-type="bibr" rid="B1">1</xref>). Apart from individual changes like the start of palliative care, organizational changes like a different allocation of nurse time or other resources can be informed by such risk scores. In-hospital mortality rates have also been used in the assessment of hospital care quality, as is the case in the United Kingdom (<xref ref-type="bibr" rid="B2">2</xref>).</p>
<p>In-hospital mortality risks are commonly determined manually (<xref ref-type="bibr" rid="B3">3</xref>). Manual scoring systems are built upon expert knowledge and have gone through significant development time. Another approach is Machine Learning (ML), which uses data and statistical methods to build a predictive model (<xref ref-type="bibr" rid="B4">4</xref>). Several methods for ML-based methods have been developed in recent years (<xref ref-type="bibr" rid="B5">5</xref>&#x02013;<xref ref-type="bibr" rid="B7">7</xref>). While these approaches offer data-based evidence that is independent of expert knowledge, they face two challenges.</p>
<p>First, they often lack interpretability. As critical care is a life-and-death situation, providers need to be able to understand why a patient&#x00027;s status is assessed the way it is. Interpretability in Machine Learning is a broad field. Two distinctions to be made are local vs. global interpretability, i.e., whether the interpretation concerns one observation or the whole population, and model-specific vs. model-agnostic interpretability, i.e., whether the interpretation comes from within a specific model or is built on top of an existing model (<xref ref-type="bibr" rid="B8">8</xref>). We consider global model-specific interpretability. This offers two decisive advantages. First, global interpretability allows us to not only make predictions based on data, but also to explore the complex and heterogeneous data underlying our prediction model. Second, model-specific interpretability allows us to explain the reasoning behind our model and its inner workings to providers. Interpretability in in-hospital mortality risk estimation has previously been discussed (<xref ref-type="bibr" rid="B6">6</xref>, <xref ref-type="bibr" rid="B9">9</xref>) and a trade-off between a prediction model&#x00027;s interpretability and predictive performance has been identified. Among the commonly used algorithms, Decision Trees (<xref ref-type="bibr" rid="B10">10</xref>) often offer high predictive performance while being highly interpretable (<xref ref-type="bibr" rid="B6">6</xref>).</p>
<p>The second challenge arises from the high number of possible variables in in-hospital mortality risk estimation. As critical care is complex, many variables can potentially play a role in estimating a patient&#x00027;s in-hospital mortality risk. This renders many common ML algorithms challenging to use and increases models&#x00027; complexity, further hindering interpretability. Manual methods use expert knowledge from years of scientific research to identify which variables to include. Variables that were previously not considered but are readily available could play a role in in-hospital mortality risk estimation because they are highly correlated.</p>
<p>Association rule mining (ARM) is often used to detect patterns in high-dimensional data (<xref ref-type="bibr" rid="B11">11</xref>). This data mining method analyzes a given dataset for rules of the form &#x0201C;<italic>A</italic>&#x021D2;<italic>B</italic>,&#x0201D; where <italic>A</italic> and <italic>B</italic> are sets of <italic>items</italic>, which in our case describe variable-value pairs. Such a rule denotes that in an observation in the dataset in which the items in <italic>A</italic> occur, the items in <italic>B</italic> will also likely occur. This algorithm produces a set of rules that fulfill pre-configured quality constraints. In the neighboring field of associative classification (AC), this class of algorithms is used to mine rules that help in the classification task at hand (<xref ref-type="bibr" rid="B12">12</xref>). AC algorithms first mine rules in which the right-hand side is the outcome of interest and then build a classifier based on these rules, e.g., by using the best rule that applies to an observation or by aggregating all applying rules (<xref ref-type="bibr" rid="B12">12</xref>). This leads to predictive models that are easy to interpret and offer a human-readable, model-specific, and global interpretation. Additionally, the rules in the model can be used to analyze the dataset itself due to their statistical nature. This helps detect interesting patterns in the data that correspond to correlations between clinical variables and the outcome of interest.</p>
<p>AC methods have previously been used in various fields of healthcare, including the prediction of outcomes and adverse events (<xref ref-type="bibr" rid="B13">13</xref>&#x02013;<xref ref-type="bibr" rid="B16">16</xref>), the prediction of diseases and wellness (<xref ref-type="bibr" rid="B17">17</xref>&#x02013;<xref ref-type="bibr" rid="B20">20</xref>), as well as biochemistry and genetics (<xref ref-type="bibr" rid="B21">21</xref>&#x02013;<xref ref-type="bibr" rid="B25">25</xref>). AC methods have also been used in the field of in-hospital mortality risk estimation using the results of 12 lab tests (<xref ref-type="bibr" rid="B26">26</xref>). This shows the feasibility of AC methods in in-hospital mortality risk estimation.</p>
<p>We aim to improve automated, ML-based in-hospital mortality risk estimation methods by including heterogeneous variables as well as more variables in general. ARM methods, and thus also AC models, can incorporate large numbers of variables of different types, which is why we analyze the feasibility of AC models for in-hospital mortality risk estimation. We propose to use AC models to estimate the risk for in-hospital mortality risk estimation in critical and intensive care. The goal of the present study is 1) to analyze whether this approach is feasible and 2) which kind of rules are found to by such methods and which variables play a role in the prediction. This expands our previous work (<xref ref-type="bibr" rid="B27">27</xref>) in major ways. First, we present two models to analyze the temporal evolution of the prediction. Second, we expand our analysis of the resulting rules. Third, we compare our models to Decision Tree models. Lastly, we add more experiments to gain more insight into the model size.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<p>The method presented in this section was developed in the programming language C&#x00023;. The overall concept of the proposed method, as well as a comparison to Decision Trees, can be seen in <xref ref-type="fig" rid="F1">Figure 1</xref>. Based on a publicly available clinical dataset, we first mine association rules from the data. The resulting rules are then combined into a prediction model that can be applied to previously unseen cases. Finally, we test and validate the proposed method using experiments.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The overall concept of the method in comparison to decision trees.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0001.tif"/>
</fig>
<sec>
<title>2.1. Dataset</title>
<p>We used data from the MIMIC-IV project, version 0.4 (<xref ref-type="bibr" rid="B28">28</xref>). MIMIC-IV is a collection of around 525,000 emergency department and intensive care unit cases, collected between 2008 and 2019 at the Beth Israel Deaconess Medical Center in Boston, Massachusetts, United States of America. Recorded variables include diagnoses, procedures, drug prescriptions, socio-demographic factors like gender, insurance type, and marital status, and organizational information like diagnosis-related groups (DRGs), wards, and services.</p>
<p><xref ref-type="table" rid="T1">Table 1</xref> lists all types of variables used in this study. We included all categorical information that can be assumed to be available in most clinical contexts. This excludes rapidly changing information like vital parameters as well as textual notes. Not all of this information is available at the beginning of a clinical case. Some information, like diagnoses, is recorded later in the case. We thus divided the variable types into two classes. In the first class, only variables that are available at the beginning of the cases are considered. These will be used to build a minimal model to estimate the in-hospital mortality risk right at the beginning of the case. The second class contains all variable types and is used to build a full model based on all available information.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>A list of all variables from MIMIC-IV that were used in this study.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Variable type</bold></th>
<th valign="top" align="center"><bold>Minimal model?</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="center"><bold>&#x00023; variables</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Diagnosis</td>
<td/>
<td valign="top" align="left">Diagnoses (coded as ICD)</td>
<td valign="top" align="center">86,751</td>
</tr>
<tr>
<td valign="top" align="left">Ethnicity</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="left">White, Asian, Black/African American etc.</td>
<td valign="top" align="center">8</td>
</tr>
<tr>
<td valign="top" align="left">Gender</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="left">Binary: male/female</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Insurance type</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="left">Medicare, Medicaid, or Other</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Language</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="left">Binary: English/other</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Marital status</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="left">Single, married, divorced, widowed, or missing</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Prescription</td>
<td/>
<td valign="top" align="left">Drugs described to a patient</td>
<td valign="top" align="center">10,259</td>
</tr>
<tr>
<td valign="top" align="left">Procedure</td>
<td/>
<td valign="top" align="left">Procedures (coded as ICD)</td>
<td valign="top" align="center">82,763</td>
</tr>
<tr>
<td valign="top" align="left">Service</td>
<td/>
<td valign="top" align="left">Clinical services: neonatal, psychological etc.</td>
<td valign="top" align="center">21</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Ward</td>
<td/>
<td valign="top" align="left">Clinical wards: intensive care unit, surgery etc.</td>
<td valign="top" align="center">43</td>
</tr> <tr>
<td valign="top" align="left">Overall</td>
<td valign="top" align="center">5 of 10</td>
<td/>
<td valign="top" align="center">179,857</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The table includes the name of the variable type, whether it is part of the minimal model, a description, as well as the number of variables. ICD International Classification of Diseases</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>All cases have been transformed into a set of items to enable ARM. This was done by adding all variables that occurred during a case to the case&#x00027;s itemset. We did not exclude any case in order to get a general in-hospital mortality risk estimation model.</p>
<p>This results in 179,857 clinical variables and 524,520 cases, in 9,369 (1.79%) of which the patient died during their stay in the hospital. This in-hospital mortality rate is comparable to similar populations like England (<xref ref-type="bibr" rid="B2">2</xref>).</p>
</sec>
<sec>
<title>2.2. Rule Mining</title>
<p>Due to the presence of rare diseases or smaller patient subgroups, infrequent rules can be of interest in healthcare. This is why we do not use the classical ARM metrics like support and confidence (<xref ref-type="bibr" rid="B11">11</xref>) that focus on frequency, but instead epidemiological metrics that are widely used to measure associations between clinical variables and outcomes. We use the odds ratio (OR) as the primary metric to measure this association. Odds ratios have previously been used in ARM in healthcare (<xref ref-type="bibr" rid="B13">13</xref>). Starting from a contingency table like the one in <xref ref-type="table" rid="T2">Table 2</xref>, the OR can be calculated as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">OR</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>100</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mn>400</mml:mn></mml:mrow><mml:mrow><mml:mn>200</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mn>300</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mover accent="true"><mml:mrow><mml:mn>6</mml:mn></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>indicating that there is a negative association between the items in <italic>A</italic> and in-hospital mortality. The set <italic>A</italic> can contain one or more items of one or more different types. This flexibility makes ARM techniques easy to use with large, heterogeneous data. ORs range from 0 to &#x0002B;&#x0221E;. An OR of 1.0 denotes no classification, while an OR &#x02264; 1.0 denotes a positive or negative association, respectively. The higher (in the case OR &#x0003E; 1.0) or lower (OR &#x0003C; 1.0), the stronger the association.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>A fictional contingency table to show how the OR is calculated.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>Died</bold></th>
<th valign="top" align="center"><bold>Survived</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Items in <italic>A</italic> occurred</td>
<td valign="top" align="center"><italic>a</italic> &#x0003D; 100</td>
<td valign="top" align="center"><italic>b</italic> &#x0003D; 200</td>
</tr>
<tr>
<td valign="top" align="left">Items in <italic>A</italic> did not occur</td>
<td valign="top" align="center"><italic>c</italic> &#x0003D; 300</td>
<td valign="top" align="center"><italic>d</italic> &#x0003D; 400</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The mining process consists of two steps. In the first step, a contingency table like the one in <xref ref-type="table" rid="T2">Table 2</xref> is constructed for each variable by counting the cases with and without the variable as well as with and without in-hospital mortality. From this table, the odds ratio is calculated according to Equation (1). Note that this also includes ORs lower than 1.0, which denotes a negative association. ORs of 0 or &#x0221E; are discarded. This corresponds to one of the cell entries <italic>a, b, c</italic> and <italic>d</italic> being 0.</p>
<p>In the second step, the model size is reduced by applying a filter to the rules constructed in the first step. As an OR of 1.0 denotes no association, we use a statistical hypothesis test to ensure that the calculated OR differs significantly from 1.0. The normal approximation of the log odds ratio (<xref ref-type="bibr" rid="B29">29</xref>) is used. The null hypothesis is &#x0201C;OR &#x0003D; 1.0,&#x0201D; and the test returns a two-sided <italic>p</italic>-value. Only ORs with a <italic>p</italic>-value below a configurable value <italic>p</italic><sub>max</sub> are kept. As this method results in many tests on the same dataset, Bonferroni correction (<xref ref-type="bibr" rid="B30">30</xref>) can be used. This procedure divides <italic>p</italic><sub>max</sub> by the number of tests to be executed and uses this quotient as the threshold value instead of <italic>p</italic><sub>max</sub>. All rules of the form &#x0201C;variable&#x021D2;in-hospital mortality,&#x0201D; together with their OR, that are left after this filtering step then form the prediction model.</p>
</sec>
<sec>
<title>2.3. Prediction</title>
<p>Given a new observation, the prediction model can be used to estimate the corresponding patient&#x00027;s in-hospital mortality risk using the model&#x00027;s rules. First, all the rules that apply to the model are determined. The remaining rules do not play a role in this observation&#x00027;s prediction. The ORs of these applying rules are then aggregated by calculating their average value. A decision boundary &#x003B4; is used to decide which average OR leads to the prediction of high in-hospital mortality risk. If OR<sub><italic>p</italic></sub> is the average OR of all rules that apply to observation <italic>p</italic>, then the prediction model&#x00027;s decision function <italic>f</italic>(<italic>p</italic>) is</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mtext>high&#x000A0;in-hospital&#x000A0;mortality&#x000A0;risk,</mml:mtext></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mrow><mml:mtext>if&#x000A0;OR</mml:mtext></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mtext>low&#x000A0;in-hospital&#x000A0;mortality&#x000A0;risk,</mml:mtext></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mrow><mml:mtext>if&#x000A0;OR</mml:mtext></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
</sec>
<sec>
<title>2.4. Experiments</title>
<p>The following hyperparameters were used. The <italic>p</italic>-value threshold <italic>p</italic><sub>max</sub> decides which rules are kept in the filtering step. We used the thresholds 10<sup>&#x02212;<italic>n</italic></sup> for <italic>n</italic> in {0, 1, &#x02026;, 10} as well as the commonly used value 0.05. All of these thresholds were used with and without Bonferroni correction. Additionally, a value of 0 was used to only keep ORs with a <italic>p</italic>-value of exactly 0. This was possible due to the high number of observations. As computers use finite representations of floating-point numbers, this should be understood to be exactly zero after internal rounding. This resulted in 13&#x000B7;2 &#x0003D; 26 experiments. Each experiment was repeated ten times. Each time, the dataset was randomly split into 90% training set and 10% test set. The primary metric of interest was the area under the receiver operating characteristic curve (AUC) (<xref ref-type="bibr" rid="B31">31</xref>). The AUC values were calculated using the Accord Framework, version 3.8.0 (<xref ref-type="bibr" rid="B32">32</xref>). All metrics have been calculated both on the training and the test set to analyze possible overfitting. The experiments were executed for both the full model and the minimal model.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<p>The AUC values achieved by the proposed method can be seen in <xref ref-type="fig" rid="F2">Figure 2</xref>. For the full model, the mean AUC values on the test set range from 0.977 to 0.980, which shows that the filtering step has a minor impact on the predictive performance. With a range of 0.687 to 0.690, this is also true for the minimal model. In both cases, the difference to the training set is low, indicating that the method does not suffer from overfitting. The largest deviation in mean AUC values was 0.001 for the full model, <italic>p</italic><sub>max</sub> &#x0003D; 1 and no Bonferroni correction.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The AUC values of the full model (top) and the minimal model (bottom). Error bars show one standard deviation, the color indicates whether the test or the training set was used. The left panel is without, the right panel with Bonferroni correction. The x-axis shows different <italic>p</italic>-values used in the filtering step. The leftmost dot denotes the <italic>p</italic>-value zero. The full model clearly outperforms the minimal model. Both models are very stable with respect to both the <italic>p</italic>-value threshold and Bonferroni correction.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0002.tif"/>
</fig>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> shows the number of rules in both models. In the minimal model, the number of rules is almost constant. This can be explained by the small number of variables, as only 20 variables are considered. In the full model, on the other hand, almost 180,000 variables can be included in the model. The number of rules grows exponentially, with Bonferroni correction slowing the growth down considerably. Still, the number of rules reaches into the thousands. As such a large number of rules is hard to handle, a strict <italic>p</italic>-value filter is advisable if interpretability is of interest. As the performance stays almost the same with lower <italic>p</italic><sub>max</sub> values, but the number of rules is considerably lower, we can deduce that the statistical significance test is an effective filter that greatly reduces the size of the prediction model without compromising the predictive performance. Still, around 1,220 rules remain even for the full model with <italic>p</italic><sub>max</sub> &#x0003D; 0. As can be seen in <xref ref-type="fig" rid="F4">Figure 4</xref>, this is due to the complexity of the problem. In-hospital mortality can be caused and influenced by many factors. Per observation, however, only around 22 to 36 rules apply when using the full model, depending on <italic>p</italic><sub>max</sub> and the use of Bonferroni correction. On average, around 4 to 5 rules apply when using minimal model.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The number of rules the full model (left) and the minimal model (right). Error bars (which are very small) show one standard deviation, the color indicates whether Bonferroni correction was used. The x-axis shows different <italic>p</italic>-values used in the filtering step. The leftmost dot denotes the <italic>p</italic>-value zero. The number of rules increases exponentially in the full model, while this effect is much smaller in the minimal model. The minimal model has fewer rules than the full model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>The number of rules that apply to an average case in the full model (left) and the minimal model (right). Error bars (which are very small) show one standard deviation, the color indicates whether Bonferroni correction was used. The x-axis shows different <italic>p</italic>-values used in the filtering step. The leftmost dot denotes the <italic>p</italic>-value zero. Only a small fraction of all rules apply to an individual case, making it easy to interpret the model and its decisions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0004.tif"/>
</fig>
<p>In the remainder of this study, we further analyze small prediction models, as they are easier to manage and interpret and have a comparable predictive performance. For both the full model and the minimal model, the lowest threshold <italic>p</italic><sub>max</sub> &#x0003D; 0 was used. In this case, Bonferroni correction makes no difference. This results in a full model with 1,217 rules and an AUC of 0.98 and a minimal model with 13 rules and an AUC of 0.69.</p>
<p>The full model clearly outperforms the minimal model. This is to be expected, as the minimal model only contains very limited data. The variable types ethnicity, gender, insurance type, language, and marital status contain no information that could help determine the cause of the clinical stay or the patient&#x00027;s health status. However, it is noteworthy that the minimal model still achieves an AUC of 0.69 with this limited information. As all this information is readily available and patient-reported, the model can be used as a first assessment before any provider encounter. The corresponding receiver operating characteristic (ROC) curve can be seen in <xref ref-type="fig" rid="F5">Figure 5</xref>. To choose a decision boundary, we searched for the highest Youden index (<xref ref-type="bibr" rid="B33">33</xref>), which is equivalent to optimizing the sum of sensitivity and specificity. The corresponding minimal model uses a decision boundary of 1.015, which results in an accuracy of 63%, a sensitivity of 66%, and a specificity of 63%. Other decision boundaries are possible, depending on the context in which they are used. Varying the decision boundary will affect both sensitivity and specificity of the model.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The receiver operating characteristic of the minimal model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0005.tif"/>
</fig>
<p>This additional information in the full model improves the predictive performance considerably. With an AUC of 0.98, the model can almost perfectly predict the in-hospital mortality risk. The corresponding ROC curve can be seen in <xref ref-type="fig" rid="F6">Figure 6</xref>. With an optimal decision boundary of 5.306, this results in an accuracy of 93%, a sensitivity of 95%, and a specificity of 93%.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The receiver operating characteristic of the full model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0006.tif"/>
</fig>
<p>All analyses were done in the R programming language (<xref ref-type="bibr" rid="B34">34</xref>), using the tidyverse packages (<xref ref-type="bibr" rid="B35">35</xref>). The rules of both models can be found in the <xref ref-type="sec" rid="s10">Supplementary Materials</xref>.</p>
<sec>
<title>3.1. Analysis of Rules</title>
<p>Apart from their predictive qualities, the rules in the models also allow us to analyze the algorithm&#x00027;s reasoning. The 13 rules in the minimal model span all five included variable types. These rules indicate that male patients die more often than female patients (ORs 1.25 vs. 0.80), that English speakers are more likely to survive than others (0.73 vs. 1.36), and that Medicare patients are more likely to die than Medicaid and other patients (2.40 vs. 0.57 vs. 0.50). The ethnicities &#x02018;Hispanic/Latino&#x02019; and &#x02018;Black/African American&#x02019; are negatively associated with in-hospital mortality (0.47 and 0.58), while &#x02018;unknown&#x02019; and &#x02018;unable to obtain&#x02019; have higher ORs (4.59 and 2.25). The marital status &#x02018;widowed&#x02019; is positively associated with in-hospital mortality (2.05), while being single is negatively associated with in-hospital mortality (0.50).</p>
<p>From the five variable types in the minimal model, the same 13 rules are in the full model. The remaining 1,204 rules span over the five other variable types with 602 diagnosis, 445 prescription, 131 procedure, nine service, and 17 ward rules.</p>
<p>Only five diagnosis rules are negative associations, namely three diagnoses on single liveborn infants (0.08-0.16), encounters for immunizations (0.06), and unspecified chest pain (0.11). The remaining 597 rules can be categorized into various groups of rules that describe variants of the same pattern. Examples include alcohol abuse and its consequences (2.34&#x02013;11.69), anemia (1.67&#x02013;4.56), various forms of hemorrhage (3.32&#x02013;582.13), neoplasms (3.50&#x02013;13.66), pneumonia (6.27&#x02013;18.49), pressure ulcers (5.29&#x02013;17.96), sepsis and septicemia (4.15&#x02013;33.18), and diabetes (1.56&#x02013;13.20). Some doubling occurs due to two versions of the International Classification of Diseases being used in the dataset.</p>
<p>This also concerns the 131 procedure rules. Here, examples of groups of rules are catheterizations (2.59&#x02013;21.11), drainages (5.53&#x02013;22.06), ventilation (2.44&#x02013;26.02), and transfusions (3.66&#x02013;13.51). No procedure rule has an OR below 1.0.</p>
<p>Only 17 out of 445 prescription rules are negative. As some drugs are recorded with slightly different names, different doses or different mode of administration, some doubling occurs. One extreme example is sodium chloride with 20 rules. The rules &#x0201C;0.45% Sodium Chloride&#x0201D; and &#x0201C;0.45 % Sodium Chloride&#x0201D; (note the space) have ORs of 2.11 and 94.94, respectively. The difference in ORs can be explained by the unequal distribution across units. While 18 of the 23 (78%) cases with &#x0201C;0.45 % Sodium Chloride&#x0201D; in MIMIC-IV were in at least one Intensive Care Unit, only 3375 of the 15,180 (22%) cases with &#x0201C;0.45% Sodium Chloride&#x0201D; were in at least one Intensive Care Unit. This indicates that there are some unexpected inconsistencies in the data. With almost 180,000 variables, it would be impossible to check all variables manually for inconsistencies due to variations in practice. Other prescription rule groups with high variation include Heparin (2.10&#x02013;68.58), vaccines (0.01&#x02013;2.82), and lidocaine (2.10&#x02013;15.55).</p>
<p>Three service rules are negative, namely obstetrics (0.01), newborn (0.17), and orthopedic (0.19). The positive rules include four medical services (1.64&#x02013;2.62), trauma (2.39), and neurologic surgical (2.73).</p>
<p>Finally, five out of 17 ward rules are negative, with three being at least partially about childbirth (0.02&#x02013;0.48). The six rules with the highest ORs (7.55&#x02013;13.25) are all the intensive care units in the dataset. Both the service and ward rules reflect different patient populations with different reasons for the clinical stay. The in-hospital mortality risk is vastly different between pregnant women and traffic accident victims that need intensive care, for example.</p>
<p>Overall, only 37 of 1,217 rules are negative. The five rules with the lowest ORs are about childbirth (0.00&#x02013;0.04) and a Hepatitis B vaccine (0.01), the five highest ORs are medications used in palliative care (Morphine and Angiotensin II, 856.26 and 332.77), subdural hemorrhage (582.14), and two rules on brain death (413.38 and 668.41). These last two rules are unexpected. Brain death should always co-occur with in-hospital mortality, which would exclude this OR of &#x0221E;. Their occurrence hints at an error in the data. Braindead patients who are organ donors are recorded as having died, but then another case is opened for them with the diagnosis of brain death, but it is recorded that they survived this second case. This leads to more inconsistent data contained in MIMIC-IV, which can now be seen in the prediction model.</p>
</sec>
<sec>
<title>3.2. Interpretability</title>
<p>To study the proposed method&#x00027;s interpretability, we give a fictional example of how the prediction model can be used. <xref ref-type="table" rid="T3">Table 3</xref> shows the rules that apply to a fictional patient in the emergency department. This list of rules is how the prediction model&#x00027;s decision could be shown to providers. Apart from the decision, the rules also explain what the case is about. The items in the minimal model give us a first assessment. The patient is female, Black/African American, single, insured under Medicaid, and speaks English. This results in an average OR of 0.64, which is below the decision boundary of 1.015. This tells us that, based on the minimal model, our patient is low-risk.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>A list of rules that apply to a fictional patient.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Variable type</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="center"><bold>OR</bold></th>
<th valign="top" align="center"><bold>Minimal model?</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Ethnicity</td>
<td valign="top" align="left">Black/African American</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">Gender</td>
<td valign="top" align="left">Female</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">Insurance type</td>
<td valign="top" align="left">Medicaid</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">Language</td>
<td valign="top" align="left">English</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">&#x02713;</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Marital status</td>
<td valign="top" align="left">Single</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">&#x02713;</td>
</tr> <tr>
<td valign="top" align="left">Diagnosis</td>
<td valign="top" align="left">Acidosis</td>
<td valign="top" align="center">10.96</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Diagnosis</td>
<td valign="top" align="left">Anuria and oliguria</td>
<td valign="top" align="center">17.27</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Prescription</td>
<td valign="top" align="left">Sodium Bicarbonate</td>
<td valign="top" align="center">11.54</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Prescription</td>
<td valign="top" align="left">Furosemide</td>
<td valign="top" align="center">5.00</td>
<td/>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>This is also how the applying rules could be shown to providers</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>In the full model, more information becomes available. For example, suppose that the following information has become available after 20 min in the emergency department. We now know that acidosis and anuria or oliguria was diagnosed. The patient was given Sodium Bicarbonate and Furosemide and is currently in the emergency department. While this is not much information due to the short length of stay, it allows us to update our risk estimation. With an average OR of 5.328, which is greater than the full model&#x00027;s decision boundary of 5.306, we conclude that the patient currently has a high in-hospital mortality risk, but she is close to the decision boundary. We see that not only are the models interpretable, but they can also be used to get a quick overview of the case at hand.</p>
</sec>
<sec>
<title>3.3. Comparison to Decision Trees</title>
<p>We compared our models to Decision Trees models, which are the state of the art in interpretable machine learning for in-hospital mortality as they show the best trade-off between performance and interpretability (<xref ref-type="bibr" rid="B6">6</xref>).</p>
<p>We trained Decision Tree classifiers using scikit-learn, version 0.24.2 (<xref ref-type="bibr" rid="B36">36</xref>), with varying maximal depths. As the dataset is highly imbalanced, with over 98% of patients surviving, we used the balanced class weight option to give more weight to the rarer class. These experiments were executed for both the full and the minimal model.</p>
<p>The results are visualized in <xref ref-type="fig" rid="F7">Figure 7</xref>. For the minimal model, the results are very similar to our proposed method. Above a maximal tree depth of five, the metrics are stable, with only small deviations in the sensitivities of the training and test set. <xref ref-type="fig" rid="F8">Figure 8</xref> shows the number of leaves in the trees, which can be understood as the number of possible paths in the tree and thus as the number of rules that can be extracted from such a tree. The number of leaves in the minimal Decision Tree model grows quickly to around 433. The growth stops at a maximal tree depth of around 15, and the model&#x00027;s size stays relatively stable. A maximal depth of at least six is needed to achieve a stable model with low standard deviations and satisfactory metrics, at which point the model already contains around 60 leaves. Our proposed method&#x00027;s minimal model only needs 13 rules for comparable performance. We thus see that while the predictive performance is similar to our proposed minimal model, the complexity of the model is higher.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Various metrics on the Decision Tree models. The left panels are the full model, while the right panels are the minimal model. The color denotes whether the training or the test set was used to produce these metrics. On the x-axis, various maximal tree depths were analyzed. For the full model, there is major overfitting. The test set sensitivity quickly drop below the training set sensitivity for higher maximal tree depths. For the minimal model, the sensitivity also drops below the training set&#x00027;s sensitivity.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0007.tif"/>
</fig>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>The number of leaves in the Decision Tree models. The full model increases exponentially. The minimal model&#x00027;s size is constant above a depth of around 15.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmed-08-785711-g0008.tif"/>
</fig>
<p>The full Decision Tree model suffers from major overfitting, as can be seen in the sensitivity values of the test set, which drop below the training set sensitivity for maximal tree depth values above five. With the highest sensitivity of 90%, Decision Trees cannot outperform our proposed method&#x00027;s full model in terms of detecting high-risk cases. On the other hand, decision Trees achieve a slightly higher specificity of 97% and thus detect low-risk cases more reliably. Note that the choice of a decision boundary impacts both sensitivity and specificity of the proposed method, and another decision boundary for the full model (namely 6.081) leads to almost the same metrics as the full Decision Tree model (namely a sensitivity of 90% and a specificity of 96%). In this way, the higher specificity at the cost of a lower sensitivity can also be achieved with the full model.</p>
<p>As can be seen in <xref ref-type="fig" rid="F8">Figure 8</xref>, the number of leaves grows exponentially in the full model, just as the number of rules does. In comparison to <xref ref-type="fig" rid="F3">Figure 3</xref>, there are more rules in our proposed method then leaves in the Decision Tree methods. With Decision Trees, however, the paths are much more complex, as their length can go up to the maximal tree depth. We provide one example full Decision Tree model with a maximal tree depth of seven in the <xref ref-type="sec" rid="s10">Supplementary Materials</xref>. At this depth, the Decision Tree model does not yet suffer from major overfitting, resulting in a test accuracy of 93%, sensitivity of 90%, and specificity of 94%. The tree contains 103 leaves and 204 decision nodes.</p>
<p>Following the first chain in the tree, the model decides that the patient is low-risk due to the variables &#x0201C;5% Dextrose,&#x0201D; &#x0201C;Morphine Sulfate,&#x0201D; &#x0201C;Encounter for palliative care,&#x0201D; &#x0201C;Insertion of endotracheal tube,&#x0201D; &#x0201C;Encounter for palliative care,&#x0201D; and &#x0201C;Vasopressin&#x0201D; being absent in the case. If the diagnosis &#x0201C;Less than 24 completed weeks of gestation&#x0201D; were present, the model would decide that the patient is high-risk. This means that the variables&#x00027; influence is combined into a chain of yes-no choices, and one change in the decision chain can alter the model&#x00027;s decision. This hides each variable&#x00027;s influence on the decision as it could occur multiple times in the tree, but each patient only activates one chain. Our proposed method separates all the variables into one-to-one rules, making it easier to identify each variable&#x00027;s importance individually. The combination of the variables into one decision takes place during the prediction, where the whole context of the case is taken into account and we get an overview of the case.</p>
<p>Another challenge in the interpretation of decision tree models is the combination of variables in each chain. The variable &#x0201C;Encounter for palliative care&#x0201D; occurs twice in the same chain due to two International Classification of Diseases versions being used. Additionally, the variable &#x0201C;Less than 24 completed weeks of gestation&#x0201D; does not match the palliative care diagnoses. The tree structure does not allow us to get an overview of the case at hand as it also highlights variables that did not occur during the case, obfuscating what actually happened. Our full model shows which rules apply to the case and thus give us an explanation of what happened during the case. Building the decision on top of this knowledge helps providers understand the model&#x00027;s decision.</p>
<p>In summary, Decision Tree models do not offer performance improvements and show limited interpretability compared to our proposed model. This is due to the fact that more complexity in the Decision Tree Models is needed to achieve a comparable predictive performance and because Decision Trees combine various variables in chains of yes-no-choices. Additionally, Decision Tree models introduce order to the variables, as the tree-like structure divides the dataset sequentially. On the other hand, our model treats every variable individually, independent of the other variables, making it easier to analyze each variable&#x00027;s effect.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>Associative classification based on ORs is a feasible method for in-hospital mortality risk estimation. Apart from the high predictive quality of the full model, the proposed method also allowed us to analyze the dataset based on the rules contained in the prediction model. The proposed method is not prone to overfitting and generalizes well.</p>
<p>While the AUCs of the minimal model are much lower, its predictive qualities are relatively high. Only socio-demographic information that can be provided by the patient is added, there is no information on the case at hand. This hints at underlying structural differences. The information what its effects are on individual patients can be valuable information for providers.</p>
<p>The fictional patient described above is an example of how the prediction models can be easily explained to providers. This interpretability also helps us understand the clinical data and discover patterns and inconsistencies in the data. As shown in the examples of brain death and sodium chloride, inconsistencies in the data show up as unexpected rules or rules with similar variables but vastly different ORs.</p>
<p>Other rules can serve as starting points for future studies. One example are the language rules. The evidence in support of an effect of primary language on in-hospital mortality is limited (<xref ref-type="bibr" rid="B37">37</xref>, <xref ref-type="bibr" rid="B38">38</xref>). Nevertheless, our rules suggest some underlying effect exists with ORs of 0.73 for English speakers and 1.36 for others. Further research is needed to analyze the differences in these groups and whether this effect is due to documentation practice.</p>
<p>The majority of rules reproduce known associations or associations that are evident and explainable with expert knowledge. Examples for the latter include the higher mortality in intensive care units due to the higher number of critical cases, the low mortality in childbirth rules, the high mortality in Medicare patients due to confounding by age, and various diagnoses that are either indicative of palliative care (like Morphine or Angiotensin II) or of low-risk cases like immunizations.</p>
<p>Some of the rules for which previous research exists include the following. Diabetes insipidus (13.20) has previously been identified as a potential cause of missed care and increased mortality (<xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>). Alkalosis (5.39&#x02013;8.10) is associated with increased mortality (<xref ref-type="bibr" rid="B41">41</xref>). Alcohol dependence (2.34-11.69) can lead to various health problems of cognitive, cardiovascular, and gastrointestinal nature, among others (<xref ref-type="bibr" rid="B42">42</xref>). Each of these, in turn, can contribute to increased mortality and emergencies that lead to increased in-hospital mortality.</p>
<sec>
<title>4.1. Limitations and Outlook</title>
<p>The present study has major limitations. First, only categorical data that changes slowly is considered. This excludes other interesting data like vital parameters and many biomarkers. Due to the rule-based approach, quantitative data have to be grouped into bins. Quickly changing data could easily be added to the method, but this would require analysis on a higher temporal resolution. Timestamps for all of the variables in the underlying dataset would help get more information from each case.</p>
<p>A second limitation is the lack of causal explanations. All the rules in both models are correlations between the variable and in-hospital mortality. They do not explain why this correlation exists. Future work is needed to introduce causal inference mechanisms into the presented approach.</p>
<p>There are several further possibilities for future work.</p>
<p>While the inconsistencies encountered in the dataset do not harm the proposed method&#x00027;s predictive performance, it might be helpful to remove them. However, it requires efforts to fix potential inconsistencies in thousands of clinical variables. Future research is needed to assess the consequences of resolving the inconsistencies as well as the potential for automated solutions to do so.</p>
<p>Longer rules, i.e., rules with more than one item on the left-hand side, can be created using ORs. This has not been analyzed in this study for two reasons. First, the predictive qualities are very good as-is, so more complex rules are not expected to bring much improvement. Second, more complex rules hinder interpretability. In the present form, the rules separate all variables and make them analyzable in isolation. This sets the proposed approach apart from Decision Trees, where many variables are merged into more complex rules, lowering the method&#x00027;s interpretability.</p>
<p>The proposed method was tested and validated using in-hospital mortality as an example. Other variables in MIMIC-IV could be used as the outcome of interest without changes to the model. This, of course, results in other sets of rules and different predictive performance, but the interpretable nature of the model remains the same.</p>
<p>One open question is the usability of the proposed method in clinics and hospitals. A usability study could be used to assess whether the proposed method is helpful to providers in realistic scenarios and whether it can replace or complement existing manual scoring methods. This is expected to highlight potential challenges in and benefits from the implementation of the proposed method.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>We proposed a novel associative classification method to estimate a patient&#x00027;s in-hospital mortality risk. With a minimal model that uses patient-reported information that is quickly available and a full model that uses more information that is available at a later point in time, providers can objectively track a patient&#x00027;s in-hospital mortality risk. Apart from the high predictive performance, the rule-based nature of the proposed method allowed us to analyze which among around 180,000 variables play a role in the estimation of in-hospital mortality risk, resulting in a model with around 1,200 variables.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://physionet.org/content/mimiciv/0.4/">https://physionet.org/content/mimiciv/0.4/</ext-link>.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>OH: concept, programming, analysis of the results, and writing of the manuscript. AM and ER: substantial revision of the work. All authors have approved the submitted version and take responsibility for the scientific integrity of the work.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This project was funded by the Bavarian State Ministry of Science and the Arts, coordinated by the Bavarian Research Institute for Digital Transformation (bidt), and supported by the Bavarian Academic Forum (BayWISS)&#x02013;Doctoral Consortium Health Research.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec></body>
<back>
<ack><p>The authors wish to thank Felix Meister for the valuable discussions that have greatly improved the present study.</p>
</ack>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmed.2021.785711/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmed.2021.785711/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.XLSX" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Table 1</label>
<caption><p>The rules in the full model.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Table_2.XLSX" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Table 2</label>
<caption><p>The rules in the minimal model.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Table_3.DOCX" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Table 3</label>
<caption><p>One full Decision Tree model.</p></caption>
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wieczorek-Wojcik</surname> <given-names>B</given-names></name> <name><surname>Gaworska-Krzemi&#x00144;ska</surname> <given-names>A</given-names></name> <name><surname>Owczarek</surname> <given-names>AJ</given-names></name> <name><surname>Kila&#x00144;ska</surname> <given-names>D</given-names></name></person-group>. <article-title>In-hospital mortality as the side effect of missed care</article-title>. <source>J Nurs Manag</source>. (<year>2020</year>) <volume>28</volume>:<fpage>2240</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1111/jonm.12965</pub-id><pub-id pub-id-type="pmid">32239793</pub-id></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stewart</surname> <given-names>K</given-names></name> <name><surname>Choudry</surname> <given-names>MI</given-names></name> <name><surname>Buckingham</surname> <given-names>R</given-names></name></person-group>. <article-title>Learning from hospital mortality</article-title>. <source>Clin Med</source>. (<year>2016</year>) <volume>16</volume>:<fpage>530</fpage>. <pub-id pub-id-type="doi">10.7861/clinmedicine.16-6-530</pub-id><pub-id pub-id-type="pmid">27927816</pub-id></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salluh</surname> <given-names>JI</given-names></name> <name><surname>Soares</surname> <given-names>M</given-names></name></person-group>. <article-title>ICU severity of illness scores: APACHE, SAPS and MPM</article-title>. <source>Curr Opin Crit Care</source>. (<year>2014</year>) <volume>20</volume>:<fpage>557</fpage>&#x02013;<lpage>565</lpage>. <pub-id pub-id-type="doi">10.1097/MCC.0000000000000135</pub-id><pub-id pub-id-type="pmid">25137401</pub-id></citation></ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bishop</surname> <given-names>CM</given-names></name></person-group>. <source>Pattern Recognition and Machine Learning (Information Science and Statistics)</source>. <publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name> (<year>2006</year>).</citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>AEW</given-names></name> <name><surname>Pollard</surname> <given-names>TJ</given-names></name> <name><surname>Mark</surname> <given-names>RG</given-names></name></person-group>. <article-title>Reproducibility in critical care: a mortality prediction case study</article-title>. In: <person-group person-group-type="editor"><name><surname>Doshi-Velez</surname> <given-names>F</given-names></name> <name><surname>Fackler</surname> <given-names>J</given-names></name> <name><surname>Kale</surname> <given-names>D</given-names></name> <name><surname>Ranganath</surname> <given-names>R</given-names></name> <name><surname>Wallace</surname> <given-names>B</given-names></name> <name><surname>Wiens</surname> <given-names>J</given-names></name></person-group>, editors. <source>Proceedings of the 2nd Machine Learning for Healthcare Conference, Vol. 68 of Proceedings of Machine Learning Research</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>PMLR</publisher-name> (<year>2017</year>). p. <fpage>361</fpage>&#x02013;<lpage>76</lpage>.</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>J</given-names></name> <name><surname>Su</surname> <given-names>B</given-names></name> <name><surname>Li</surname> <given-names>C</given-names></name> <name><surname>Lin</surname> <given-names>K</given-names></name> <name><surname>Li</surname> <given-names>H</given-names></name> <name><surname>Hu</surname> <given-names>Y</given-names></name> <etal/></person-group>. <article-title>A review of modeling methods for predicting in-hospital mortality of patients in intensive care unit</article-title>. <source>J Emerg Crit Care Med</source>. (<year>2017</year>) <volume>1</volume>:<fpage>e18</fpage>. <pub-id pub-id-type="doi">10.21037/jeccm.2017.08.03</pub-id></citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keuning</surname> <given-names>BE</given-names></name> <name><surname>Kaufmann</surname> <given-names>T</given-names></name> <name><surname>Wiersema</surname> <given-names>R</given-names></name> <name><surname>Granholm</surname> <given-names>A</given-names></name> <name><surname>Pettil&#x000E4;</surname> <given-names>V</given-names></name> <name><surname>M&#x000F8;ller</surname> <given-names>MH</given-names></name> <etal/></person-group>. <article-title>Mortality prediction models in the adult critically ill: a scoping review</article-title>. <source>Acta Anaesthesiol Scand</source>. (<year>2020</year>) <volume>64</volume>:<fpage>424</fpage>&#x02013;<lpage>442</lpage>. <pub-id pub-id-type="doi">10.1111/aas.13527</pub-id><pub-id pub-id-type="pmid">31828760</pub-id></citation></ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stiglic</surname> <given-names>G</given-names></name> <name><surname>Kocbek</surname> <given-names>P</given-names></name> <name><surname>Fijacko</surname> <given-names>N</given-names></name> <name><surname>Zitnik</surname> <given-names>M</given-names></name> <name><surname>Verbert</surname> <given-names>K</given-names></name> <name><surname>Cilar</surname> <given-names>L</given-names></name></person-group>. <article-title>Interpretability of machine learning-based prediction models in healthcare</article-title>. <source>WIREs Data Min Knowl Discov</source>. (<year>2020</year>) <volume>10</volume>:<fpage>e1379</fpage>. <pub-id pub-id-type="doi">10.1002/widm.1379</pub-id></citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>LH</given-names></name> <name><surname>Schwartz</surname> <given-names>J</given-names></name> <name><surname>Moy</surname> <given-names>A</given-names></name> <name><surname>Knaplund</surname> <given-names>C</given-names></name> <name><surname>Kang</surname> <given-names>MJ</given-names></name> <name><surname>Schnock</surname> <given-names>KO</given-names></name> <etal/></person-group>. <article-title>Development and validation of early warning score system: a systematic literature review</article-title>. <source>J Biomed Inform</source>. (<year>2020</year>) <volume>105</volume>:<fpage>e103410</fpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2020.103410</pub-id><pub-id pub-id-type="pmid">32278089</pub-id></citation></ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L</given-names></name> <name><surname>Friedman</surname> <given-names>JH</given-names></name> <name><surname>Olshen</surname> <given-names>RA</given-names></name> <name><surname>Stone</surname> <given-names>CJ</given-names></name></person-group>. <source>Classification and Regression Trees</source>. <publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>Wadsworth International Group</publisher-name> (<year>1984</year>).</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Agrawal</surname> <given-names>R</given-names></name> <name><surname>Imieli&#x00144;ski</surname> <given-names>T</given-names></name> <name><surname>Swami</surname> <given-names>A</given-names></name></person-group>. <article-title>Mining association rules between sets of items in large databases</article-title>. In: <source>Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data. SIGMOD &#x00027;93</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name> (<year>1993</year>). p. <fpage>207</fpage>&#x02013;<lpage>16</lpage>.</citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thabtah</surname> <given-names>F</given-names></name></person-group>. <article-title>A review of associative classification mining</article-title>. <source>Knowl Eng Rev</source>. (<year>2007</year>) 03;<volume>22</volume>:<fpage>37</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1017/S0269888907001026</pub-id></citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>WY</given-names></name> <name><surname>Li</surname> <given-names>HY</given-names></name> <name><surname>Du</surname> <given-names>JW</given-names></name> <name><surname>Feng</surname> <given-names>WY</given-names></name> <name><surname>Lo</surname> <given-names>CF</given-names></name> <name><surname>Soo</surname> <given-names>VW</given-names></name></person-group>. <article-title>iADRs: towards online adverse drug reaction analysis</article-title>. <source>Springerplus</source>. (<year>2012</year>) <volume>1</volume>:<fpage>e72</fpage>. <pub-id pub-id-type="doi">10.1186/2193-1801-1-72</pub-id><pub-id pub-id-type="pmid">23420567</pub-id></citation></ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>El Houby</surname> <given-names>EM</given-names></name></person-group>. <article-title>A framework for prediction of response to HCV therapy using different data mining techniques</article-title>. <source>Adv Bioinformatics</source>. (<year>2014</year>) <volume>2014</volume>:<fpage>e181056</fpage>. <pub-id pub-id-type="doi">10.1155/2014/181056</pub-id><pub-id pub-id-type="pmid">25580118</pub-id></citation></ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Uriarte-Arcia</surname> <given-names>AV</given-names></name> <name><surname>L&#x000F3;pez-Y&#x000E1;&#x000F1;ez</surname> <given-names>I</given-names></name> <name><surname>Y&#x000E1;&#x000F1;ez-M&#x000E1;rquez</surname> <given-names>C</given-names></name></person-group>. <article-title>One-hot vector hybrid associative classifier for medical data classification</article-title>. <source>PLoS ONE</source>. (<year>2014</year>) <volume>9</volume>:<fpage>e95715</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0095715</pub-id><pub-id pub-id-type="pmid">24752287</pub-id></citation></ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kadkhoda</surname> <given-names>M</given-names></name> <name><surname>Akbarzadeh-T</surname> <given-names>MR</given-names></name> <name><surname>Sabahi</surname> <given-names>F</given-names></name></person-group>. <article-title>FLeAC: a human-centered associative classifier using the validity concept</article-title>. <source>IEEE Trans Cybern</source>. (<year>2020</year>). <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2020.3025479</pub-id><pub-id pub-id-type="pmid">33104515</pub-id></citation></ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dua</surname> <given-names>S</given-names></name> <name><surname>Singh</surname> <given-names>H</given-names></name> <name><surname>Thompson</surname> <given-names>HW</given-names></name></person-group>. <article-title>Associative classification of mammograms using weighted rules</article-title>. <source>Expert Syst Appl</source>. (<year>2009</year>) <volume>36</volume>:<fpage>9250</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2008.12.050</pub-id><pub-id pub-id-type="pmid">20160889</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rea</surname> <given-names>S</given-names></name> <name><surname>Huff</surname> <given-names>S</given-names></name></person-group>. <article-title>Cohort amplification: an associative classification framework for identification of disease cohorts in the electronic health record</article-title>. <source>AMIA Ann Symp Proc</source>. (<year>2010</year>) <volume>2010</volume>:<fpage>862</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="pmid">21347101</pub-id></citation></ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ujager</surname> <given-names>FS</given-names></name> <name><surname>Mahmood</surname> <given-names>A</given-names></name></person-group>. <article-title>A context-aware accurate wellness determination (CAAWD) model for elderly people using lazy associative classification</article-title>. <source>Sensors (Basel)</source>. (<year>2019</year>) <volume>19</volume>:<fpage>e1613</fpage>. <pub-id pub-id-type="doi">10.3390/s19071613</pub-id><pub-id pub-id-type="pmid">30987246</pub-id></citation></ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meena</surname> <given-names>K</given-names></name> <name><surname>Tayal</surname> <given-names>DK</given-names></name> <name><surname>Gupta</surname> <given-names>V</given-names></name> <name><surname>Fatima</surname> <given-names>A</given-names></name></person-group>. <article-title>Using classification techniques for statistical analysis of Anemia</article-title>. <source>Artif Intell Med</source>. (<year>2019</year>) <volume>94</volume>:<fpage>138</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2019.02.005</pub-id><pub-id pub-id-type="pmid">30871679</pub-id></citation></ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kianmehr</surname> <given-names>K</given-names></name> <name><surname>Alhajj</surname> <given-names>R</given-names></name></person-group>. <article-title>CARSVM: a class association rule-based classification framework and its application to gene expression data</article-title>. <source>Artif Intell Med</source>. (<year>2008</year>) <volume>44</volume>:<fpage>7</fpage>&#x02013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2008.05.002</pub-id><pub-id pub-id-type="pmid">18586476</pub-id></citation></ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>Y</given-names></name> <name><surname>Hui</surname> <given-names>SC</given-names></name></person-group>. <article-title>Exploring ant-based algorithms for gene expression data analysis</article-title>. <source>Artif Intell Med</source>. (<year>2009</year>) <volume>47</volume>:<fpage>105</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2009.03.004</pub-id><pub-id pub-id-type="pmid">19376690</pub-id></citation></ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>P</given-names></name> <name><surname>Wild</surname> <given-names>DJ</given-names></name></person-group>. <article-title>Fast rule-based bioactivity prediction using associative classification mining</article-title>. <source>J Cheminform</source>. (<year>2012</year>) <volume>4</volume>:<fpage>e29</fpage>. <pub-id pub-id-type="doi">10.1186/1758-2946-4-29</pub-id><pub-id pub-id-type="pmid">23176548</pub-id></citation></ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>P</given-names></name> <name><surname>Wild</surname> <given-names>DJ</given-names></name></person-group>. <article-title>Discovering associations in biomedical datasets by link-based associative classifier (LAC)</article-title>. <source>PLoS ONE</source>. (<year>2012</year>) <volume>7</volume>:<fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0051018</pub-id><pub-id pub-id-type="pmid">23227228</pub-id></citation></ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>ElHefnawi</surname> <given-names>M</given-names></name> <name><surname>Sherif</surname> <given-names>FF</given-names></name></person-group>. <article-title>Accurate classification and hemagglutinin amino acid signatures for influenza a virus host-origin association and subtyping</article-title>. <source>Virology</source>. (<year>2014</year>) <volume>449</volume>:<fpage>328</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1016/j.virol.2013.11.010</pub-id><pub-id pub-id-type="pmid">24418567</pub-id></citation></ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>CW</given-names></name> <name><surname>Wang</surname> <given-names>MD</given-names></name></person-group>. <article-title>Improving personalized clinical risk prediction based on causality-based association rules</article-title>. <source>ACM BCB</source>. (<year>2015</year>) <volume>2015</volume>:<fpage>386</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1145/2808719.2808759</pub-id><pub-id pub-id-type="pmid">27532063</pub-id></citation></ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haas</surname> <given-names>O</given-names></name> <name><surname>Maier</surname> <given-names>A</given-names></name> <name><surname>Rothgang</surname> <given-names>E</given-names></name></person-group>. <article-title>Using associative classification and odds ratios for in-hospital mortality risk estimation</article-title>. In: <source>Workshop on Interpretable ML in Healthcare at International Conference on Machine Learning (ICML)</source>. (<year>2021</year>).</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>A</given-names></name> <name><surname>Bulgarelli</surname> <given-names>L</given-names></name> <name><surname>Pollard</surname> <given-names>T</given-names></name> <name><surname>Horng</surname> <given-names>S</given-names></name> <name><surname>Celi</surname> <given-names>LA</given-names></name> <name><surname>Mark</surname> <given-names>R</given-names></name></person-group>. <article-title>MIMIC-IV (version 0.4)</article-title>. <source>PhysioNet</source>. (<year>2020</year>). <pub-id pub-id-type="doi">10.13026/a3wn-hq05</pub-id></citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morris</surname> <given-names>JA</given-names></name> <name><surname>Gardner</surname> <given-names>MJ</given-names></name></person-group>. <article-title>Calculating confidence intervals for relative risks (odds ratios) and standardised ratios and rates</article-title>. <source>Br Med J (Clin Res Ed)</source>. (<year>1988</year>) <volume>296</volume>:<fpage>1313</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1136/bmj.296.6632.1313</pub-id><pub-id pub-id-type="pmid">3133061</pub-id></citation></ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shaffer</surname> <given-names>JP</given-names></name></person-group>. <article-title>Multiple hypothesis testing</article-title>. <source>Annu Rev Psychol</source>. (<year>1995</year>) <volume>46</volume>:<fpage>561</fpage>&#x02013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.ps.46.020195.003021</pub-id></citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fawcett</surname> <given-names>T</given-names></name></person-group>. <article-title>An introduction to ROC analysis</article-title>. <source>Pattern Recognit Lett</source>. (<year>2006</year>) <volume>27</volume>:<fpage>861</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2005.10.010</pub-id></citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Souza</surname> <given-names>C</given-names></name> <name><surname>Kirillov</surname> <given-names>A</given-names></name> <name><surname>Catalano</surname> <given-names>MD</given-names></name> <name><surname>Contributors</surname> <given-names>AN</given-names></name></person-group>. <source>The Accord</source>.NET Framework (<year>2014</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="http://accord-framework.net">http://accord-framework.net</ext-link>.</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Youden</surname> <given-names>WJ</given-names></name></person-group>. <article-title>Index for rating diagnostic tests</article-title>. <source>Cancer</source>. (<year>1950</year>) <volume>3</volume>:<fpage>32</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1002/1097-014219503:1&#x0003C;32::AID-CNCR2820030106&#x0003E;3.0.CO;2-3</pub-id></citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="web"><person-group person-group-type="author"><collab>R Core Team</collab></person-group>. <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna</publisher-loc> (<year>2021</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.R-project.org/">https://www.R-project.org/</ext-link>.</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wickham</surname> <given-names>H</given-names></name> <name><surname>Averick</surname> <given-names>M</given-names></name> <name><surname>Bryan</surname> <given-names>J</given-names></name> <name><surname>Chang</surname> <given-names>W</given-names></name> <name><surname>McGowan</surname> <given-names>LD</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>R</given-names></name> <etal/></person-group>. <article-title>Welcome to the tidyverse</article-title>. <source>J Open Source Softw</source>. (<year>2019</year>) <volume>4</volume>:<fpage>1686</fpage>. <pub-id pub-id-type="doi">10.21105/joss.01686</pub-id></citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F</given-names></name> <name><surname>Varoquaux</surname> <given-names>G</given-names></name> <name><surname>Gramfort</surname> <given-names>A</given-names></name> <name><surname>Michel</surname> <given-names>V</given-names></name> <name><surname>Thirion</surname> <given-names>B</given-names></name> <name><surname>Grisel</surname> <given-names>O</given-names></name> <etal/></person-group>. <article-title>Scikit-learn: machine learning in python</article-title>. <source>J Mach Learn Res</source>. (<year>2011</year>) <volume>12</volume>:<fpage>2825</fpage>&#x02013;<lpage>30</lpage>.</citation>
</ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>John-Baptiste</surname> <given-names>A</given-names></name> <name><surname>Naglie</surname> <given-names>G</given-names></name> <name><surname>Tomlinson</surname> <given-names>G</given-names></name> <name><surname>Alibhai</surname> <given-names>S</given-names></name> <name><surname>Etchells</surname> <given-names>E</given-names></name> <name><surname>Cheung</surname> <given-names>A</given-names></name> <etal/></person-group>. <article-title>The effect of english language proficiency on length of stay and in-hospital mortality</article-title>. <source>J Gen Intern Med</source>. (<year>2004</year>) <volume>19</volume>:<fpage>221</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1111/j.1525-1497.2004.21205.x</pub-id><pub-id pub-id-type="pmid">15009776</pub-id></citation></ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cano-Ib&#x000E1;&#x000F1;ez</surname> <given-names>N</given-names></name> <name><surname>Zolfaghari</surname> <given-names>Y</given-names></name> <name><surname>Amezcua-Prieto</surname> <given-names>C</given-names></name> <name><surname>Khan</surname> <given-names>K</given-names></name></person-group>. <article-title>Physician&#x02013;patient language discordance and poor health outcomes: a systematic scoping review</article-title>. <source>Front Public Health</source>. (<year>2021</year>) <volume>9</volume>:<fpage>629041</fpage>. <pub-id pub-id-type="doi">10.3389/fpubh.2021.629041</pub-id><pub-id pub-id-type="pmid">33816420</pub-id></citation></ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gleeson</surname> <given-names>H</given-names></name> <name><surname>Bonfield</surname> <given-names>A</given-names></name> <name><surname>Hackett</surname> <given-names>E</given-names></name> <name><surname>Crasto</surname> <given-names>W</given-names></name></person-group>. <article-title>Concerns about patient safety in patients with diabetes insipidus admitted as inpatients</article-title>. <source>Clin Endocrinol</source>. (<year>2016</year>) <volume>84</volume>:<fpage>950</fpage>&#x02013;<lpage>1</lpage>. <pub-id pub-id-type="doi">10.1111/cen.13028</pub-id><pub-id pub-id-type="pmid">26824191</pub-id></citation></ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ebrahimi</surname> <given-names>F</given-names></name> <name><surname>Kutz</surname> <given-names>A</given-names></name> <name><surname>Wagner</surname> <given-names>U</given-names></name> <name><surname>Illigens</surname> <given-names>B</given-names></name> <name><surname>Siepmann</surname> <given-names>T</given-names></name> <name><surname>Schuetz</surname> <given-names>P</given-names></name> <etal/></person-group>. <article-title>Excess mortality among hospitalized patients with hypopituitarism-a population based matched cohort study</article-title>. <source>J Clin Endocrinol Metab</source>. (<year>2020</year>) <volume>105</volume>:<fpage>dgaa517</fpage>. <pub-id pub-id-type="doi">10.1210/clinem/dgaa517</pub-id><pub-id pub-id-type="pmid">32785679</pub-id></citation></ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>L</given-names></name> <name><surname>Henrich</surname> <given-names>W</given-names></name></person-group>. <article-title>Alkalemia-associated morbidity and mortality in medical and surgical patients</article-title>. <source>South Med J</source>. (<year>1987</year>) <volume>80</volume>:<fpage>729</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1097/00007611-198706000-00016</pub-id><pub-id pub-id-type="pmid">3589765</pub-id></citation></ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schuckit</surname> <given-names>MA</given-names></name></person-group>. <article-title>Alcohol-use disorders</article-title>. <source>Lancet</source>. (<year>2009</year>) <volume>373</volume>:<fpage>492</fpage>&#x02013;<lpage>501</lpage>. <pub-id pub-id-type="doi">10.1016/S0140-6736(09)60009-X</pub-id><pub-id pub-id-type="pmid">19168210</pub-id></citation></ref>
</ref-list>
</back>
</article>