<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Digit. Health</journal-id>
<journal-title>Frontiers in Digital Health</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Digit. Health</abbrev-journal-title>
<issn pub-type="epub">2673-253X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdgth.2024.1198904</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Digital Health</subject>
<subj-group>
<subject>Technology and Code</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Promoting appropriate medication use by leveraging medical big data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" equal-contrib="yes"><name><surname>Hong</surname><given-names>Linghong</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="an1"><sup>&#x2020;</sup></xref></contrib>
<contrib contrib-type="author" equal-contrib="yes"><name><surname>Huang</surname><given-names>Shiwang</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="an1"><sup>&#x2020;</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2267890/overview"/></contrib>
<contrib contrib-type="author"><name><surname>Cai</surname><given-names>Xiaohai</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lin</surname><given-names>Zhiming</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Shao</surname><given-names>Yunting</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Chen</surname><given-names>Longbiao</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/932226/overview" /></contrib>
<contrib contrib-type="author" corresp="yes"><name><surname>Zhao</surname><given-names>Min</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/1800633/overview" /></contrib>
<contrib contrib-type="author" corresp="yes"><name><surname>Yang</surname><given-names>Chenhui</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref></contrib>
</contrib-group>
<aff id="aff1"><label><sup>1</sup></label><institution>Department of Drug Clinical Trial Institution, Xiang&#x0027;an Hospital of Xiamen University, School of Medicine, Xiamen University</institution>, <addr-line>Xiamen</addr-line>, <country>China</country></aff>
<aff id="aff2"><label><sup>2</sup></label><institution>Fujian Key Laboratory of Sensing and Computing for Smart Cities, School of Informatics, Xiamen University</institution>, <addr-line>Xiamen</addr-line>, <country>China</country></aff>
<aff id="aff3"><label><sup>3</sup></label><institution>Big Data Center, The First Affiliated Hospital of Xiamen University, School of Medicine, Xiamen University</institution>, <addr-line>Xiamen</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> Sanjat Kanjilal, Harvard Pilgrim Health Care and Harvard Medical School, United States</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> Arinjita Bhattacharyya, University of Louisville, United States</p>
<p>Raphael Zozimus Sangeda, Muhimbili University of Health and Allied Sciences, Tanzania</p>
<p>Bodhayan Prasad, University of Glasgow, United Kingdom</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Min Zhao <email>xmzmdyyy@xmu.edu.cn</email> Chenhui Yang <email>chenhuiyang@xmu.edu.cn</email></corresp>
<fn fn-type="equal" id="an1"><label><sup>&#x2020;</sup></label><p>These authors have contributed equally to this work and share first authorship</p></fn>
</author-notes>
<pub-date pub-type="epub"><day>07</day><month>11</month><year>2024</year></pub-date>
<pub-date pub-type="collection"><year>2024</year></pub-date>
<volume>6</volume><elocation-id>1198904</elocation-id>
<history>
<date date-type="received"><day>02</day><month>04</month><year>2023</year></date>
<date date-type="accepted"><day>11</day><month>09</month><year>2024</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Hong, Huang, Cai, Lin, Shao, Chen, Zhao and Yang.</copyright-statement>
<copyright-year>2024</copyright-year><copyright-holder>Hong, Huang, Cai, Lin, Shao, Chen, Zhao and Yang</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>According to World Health Organization statistics, inappropriate medication has become an important factor affecting the safety of rational medication. In the gray area of medical insurance supervision, such as designated drugstores and medical institutions, there are lots of inappropriate medication phenomena regarding &#x201C;big prescription for minor ailments.&#x201D; A traditional clinical decision support system is mostly based on established rules to regulate inappropriate prescriptions, which are not suitable for clinical environments and require intelligent review. In this study, we model the complex relationships between patients, diseases, and drugs based on medical big data to promote appropriate medication use. More specifically, we first construct the medication knowledge graph based on the historical prescription big data of tertiary hospitals and medical text data. Second, based on the medication knowledge graph, we employ a Gaussian mixture model to group patient population representation as physiological features. For diagnostic features, we employ pre-training word vector Bidirectional Encoder Representations from Transformers to enhance the semantic representation between diagnoses. In addition, to reduce adverse drug interactions caused by drug combinations, we employ a graph convolution network to transform drug interaction information into drug interaction features. Finally, we employ the sequence generation model to learn the complex relationships between patients, diseases, and drugs and provide an appropriate medication evaluation for doctor prescriptions in small hospitals from two aspects: drug list and medication course of treatment. In this study, we utilize the MIMIC III dataset alongside data from a tertiary hospital in Fujian Province to validate our model. The results show that our method is more effective than other baseline methods in the accuracy of the medication regimen prediction of rational medication. In addition, it achieved high accuracy in the appropriate medication detection of prescription in small hospitals.</p>
</abstract>
<kwd-group>
<kwd>rational use of drugs</kwd>
<kwd>appropriate medication</kwd>
<kwd>NLP</kwd>
<kwd>knowledge graph</kwd>
<kwd>transformer</kwd>
</kwd-group><counts>
<fig-count count="10"/>
<table-count count="10"/><equation-count count="108"/><ref-count count="33"/><page-count count="16"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Health Informatics</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body><sec id="s1" sec-type="intro"><label>1</label><title>Introduction</title>
<p>The rational use of medicines is safe, effective, affordable, and appropriate for treating or curing the patient (<xref ref-type="bibr" rid="B1">1</xref>). The inappropriate use of medicines is a major problem worldwide. The World Health Organization (WHO) estimates that more than half of all medicines are prescribed, dispensed, or sold inappropriately and that half of all patients fail to take them correctly (<xref ref-type="bibr" rid="B2">2</xref>). In addition, in the gray area of medical insurance supervision, such as designated pharmacies and medical institutions, there may be &#x201C;big prescription for minor ailments&#x201D; healthcare fraud (<xref ref-type="bibr" rid="B3">3</xref>, <xref ref-type="bibr" rid="B4">4</xref>). Inappropriate drug use behaviors such as the overuse, underuse, or misuse of medicines not only waste medical resources but also lead to significant patient harm in terms of medication errors (MEs) and adverse drug events (ADEs) (<xref ref-type="bibr" rid="B1">1</xref>). The WHO is committed to promoting the <italic>rational use of medicines</italic> for clinical physicians and pharmacists to ensure that &#x201C;patients receive the appropriate medicines, in doses that meet their own individual requirements, for an adequate period of time&#x201D; (<xref ref-type="bibr" rid="B1">1</xref>).</p>
<p>One of the key challenges in the rational use of medicines is appropriate medication use. Compared with the safety, effectiveness, and economics of rational drug use, the evaluation of the appropriate use of medicines is more complicated, involving hyper-medication, under-medication, and inappropriate medication.</p>
<p>To address these issues, experienced investigators are assigned to hospitals to manage Medicare fraud detection. However, this method becomes time-consuming and inefficient due to the large amount of data collection. With the advent of the big data era, healthcare big data analysis can offer predictive modeling, clinical decision support, disease or safety monitoring, and other capabilities for public healthcare (<xref ref-type="bibr" rid="B5">5</xref>). Improvements in data mining and deep learning tools have turned attention to automated systems for fraud detection. Several deep learning-based clinical decision support systems (CDSSs) have been developed and deployed in hospitals to reduce the incidence of improper drug use.</p>
<p>Leveraging the application of knowledge graph construction and sequence model generation makes medication decision-making in the field of pharmacy more scientifically rational (<xref ref-type="bibr" rid="B6">6</xref>). Healthcare practitioners can gain a comprehensive understanding of the interrelationships between medications, and sequence generation can optimize medication plans based on patients&#x2019; medical histories, symptoms, and physiological data. For example, in the safety of rational medicines, Shao et al. (<xref ref-type="bibr" rid="B7">7</xref>) construct a probabilistic probability model of massive prescription data based on a knowledge graph to evaluate the risk of a drug combination by a graph search algorithm. In the rational use of medicines, Shang et al. (<xref ref-type="bibr" rid="B8">8</xref>) jointly model the longitudinal patient records as an electronic health record (EHR) graph and the drug knowledge base as a drug&#x2013;drug interaction (DDI) graph through the generation of sequence models that train end-to-end to provide effective and safe medication recommendations. Based on the experimental results on real-world EHR, GAMENet outperformed all baselines in DDI rate reduction (<xref ref-type="bibr" rid="B8">8</xref>). After analyzing a large number of medical records, the diagnosis-related groups (DRGs) payment system (<xref ref-type="bibr" rid="B9">9</xref>) based on disease type has been launched by The National Medical Insurance Administration to specify uniform drug delivery rules and prevent excessive medical treatment. However, the single and rigid pharmaceutical rules cannot achieve more accurate personal medication, which also poses a major challenge to the promotion of DRGs (<xref ref-type="bibr" rid="B10">10</xref>). To address this issue, we need a more flexible and intelligent method for the evaluation of appropriate medication.</p>
<p>Fortunately, with the emergence of medical consortia and the sinking of medical resources, the professional prescription experience of tertiary hospitals can be accessed, providing us with new perspectives to address the problems existing in designated pharmacies and medical institutions. Therefore, we integrate the clinical medication experience of tertiary hospitals and medical knowledge and transfer the learned knowledge to small hospitals and clinics so that their prescriptions are more in line with professional standards. To achieve these goals, we need to address the following issues.</p>
<p>Owing to the large individual differences in patients, such as being children, adults, or older, and differences in their liver and kidney functions, nervous system, and other physiological characteristics, the same diagnosis may lead to different treatment regimens. The majority of drugs are administered based on the patient&#x2019;s age or weight (mg/kg) (<xref ref-type="bibr" rid="B11">11</xref>). Therefore, to remedy the case with greater precision, we need to consider the individualized use of medicines.</p>
<p>Since the relationship between disease and symptoms is not a simple one-to-one relationship, the occurrence of a single disease may cause the simultaneous occurrence of multiple symptoms (<xref ref-type="bibr" rid="B12">12</xref>); therefore, doctors must treat patients through the combination of multiple drugs. Multimorbidity (<xref ref-type="bibr" rid="B13">13</xref>) is becoming more common and is a growing global challenge. Therefore, it is a challenge for us to address the complex relationship between disease and drug use.</p>
<p>The increase in drug species shows that the compatibility relationships between drugs are more complicated. In addition, there would be more drug overuse and abuse in the case of &#x201C;big prescription for minor ailments,&#x201D; and polypharmacy may increase drug side effects and even more adverse drug&#x2013;drug interactions (<xref ref-type="bibr" rid="B14">14</xref>&#x2013;<xref ref-type="bibr" rid="B16">16</xref>). Therefore, DDIs should be taken into account when evaluating the appropriateness of rational drug use to reduce adverse reactions associated with combined drug prescriptions.</p>
<p>In the preceding discussion, we delved extensively into the interconnections among patient characteristics, diagnosis, and prescription medications. However, in real-world scenarios, the relationships among these three components are even more intricately intertwined. A patient&#x2019;s individual attributes, such as gender, age, and medical history, exert a significant influence on the susceptibility to diseases, progression of the illness, and response to treatment (<xref ref-type="bibr" rid="B11">11</xref>). Diverse patient characteristics may give rise to distinct pathophysiological processes, thereby impacting the selection of diagnostic and therapeutic strategies for the ailment. This, in turn, substantially affects the physician&#x2019;s ability to accurately diagnose the condition and formulate an effective treatment regimen (<xref ref-type="bibr" rid="B17">17</xref>). The precision of the diagnosis is pivotal in devising a successful therapeutic plan. Simultaneously, the choice of medications must take patient-specific features into account, including age, gender, baseline health status, and potential interactions with other medications (<xref ref-type="bibr" rid="B18">18</xref>). Furthermore, patient attributes can also influence the individual&#x2019;s response to and tolerance of pharmaceuticals; for instance, certain medications may be metabolized at a slower rate in older patients, necessitating dose adjustments to avert adverse reactions. To sum up, these three elements intricately intertwine, paving the way for patients to access optimal treatment pathways and furnish a robust foundation for scientifically sound medication recommendations.</p>
<p>To address the above issues, in this study, we propose a regulatory framework of rational drug use based on medical consortia and big data through mining the clinical experience of prescription big data and medical knowledge of drug instructions. To be specific, we first extract information from the big data of prescription of tertiary hospitals and medical text data, and establish the medication knowledge graph based on the extracted information. Second, based on the medication knowledge graph, we extract physiological, diagnostic, and drug interaction features through feature enhancement. Finally, we construct the sequence generation model to solve the complex relationship between patients, diseases, and drugs and then evaluate the appropriate medication prescribed by doctors in small hospitals using the model learned from a tertiary hospital.</p>
<p>In conclusion, the contribution of this study is as follows:
<list list-type="simple">
<list-item><label>1.</label>
<p>To the best of our knowledge, this is the first study on the data-driven evaluation of appropriate medication use. By utilizing extensive prescription data from tertiary hospitals and integrating medical text information, we provide a practical tool for assessing and improving prescription practices in small hospitals, focusing on drug selection and treatment courses.</p></list-item>
<list-item><label>2.</label>
<p>We propose a data-driven experience extraction of clinical rational drug use and an appropriate medication evaluation framework based on advanced deep learning techniques. This approach facilitates the transfer of rational drug use practices from tertiary hospitals to primary care settings, thereby ensuring safer and more effective medication management in these environments.</p></list-item>
<list-item><label>3.</label>
<p>We evaluate the proposed framework with two medical record datasets: Medical Information Mart for Intensive Care III (MIMIC&#x005F;III) and real-world prescription big data collected from tertiary hospitals. Results show that our method has more accurate medication regimen prediction ability and consistently outperforms other baselines. In addition, it has achieved high accuracy in the appropriate medication detection of prescription in small hospitals.</p></list-item>
<list-item><label>4.</label>
<p>Our research utilizes medical big data to improve medication use practices by addressing important public health challenges, such as MEs and ADEs. Through the analysis of data on prescriptions and patient outcomes, our study aims to support the development of drug safety monitoring and medication management practices.</p></list-item>
<list-item><label>5.</label>
<p>Furthermore, the methodologies and findings of our study have profound implications for clinical trials. Our data-driven approach allows for a better understanding of drug efficacy and safety across diverse patient demographics, aiding in the design and evaluation of clinical trials. This is particularly crucial in trials that aim to tailor medical treatments to individual patient needs, a cornerstone of personalized medicine.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. We first elaborate on the proposed framework in <xref ref-type="sec" rid="s2">Section 2</xref>, and then present our experiments in <xref ref-type="sec" rid="s3">Section 3</xref>. Finally, a comprehensive summary of our work is encapsulated in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
</sec>
<sec id="s2" sec-type="methods"><label>2</label><title>Methods</title>
<p>We propose a framework for the experience extraction of clinical rational drug use and appropriate index evaluation, as illustrated in <xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>. In the <italic>medication knowledge graph construction</italic> stage, we first extract drug triads from historical prescriptions and medical text data, then establish a patient&#x2013;disease&#x2013;drug knowledge graph. In the <italic>modeling</italic> phase, we first employ a Gaussian mixture model (GMM) (<xref ref-type="bibr" rid="B19">19</xref>) to group patient population representation as physiological features, based on the four physiological variables of gender, age, height, and weight. Second, we transform patients&#x2019; diagnostic information into word vectors as diagnostic features through pre-training word vector Bidirectional Encoder Representations from Transformers (BERT) (<xref ref-type="bibr" rid="B20">20</xref>) to enhance the semantic representation between diagnoses. Third, to reduce adverse drug interactions caused by drug combinations, we employ a graph convolution network (GCN) (<xref ref-type="bibr" rid="B21">21</xref>) to transform drug interaction information into drug interaction features. Finally, we exploit the medication regimen from historical prescription data to train a sequence generation model. In the <italic>analysis</italic> stage, given a new prescription from small hospitals or clinics, we use the trained model to predict the rational medication regimen for the prescription, and provide an evaluation of appropriate drug use in terms of the drug list and medication course of treatment to physicians and pharmacists in small hospitals. We elaborate the details of the key components in the following.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>An overview of the proposed framework.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g001.tif"/>
</fig>
<sec id="s2a"><label>2.1</label><title>Medication knowledge graph construction</title>
<p>In this section, our objective is to construct a medication knowledge graph to model medication rules for co-prescription in big data. However, relying only on historical prescription data is not enough to simulate the comprehensive medication rules, because adverse drug reactions (ADRs) may not be reflected in clinical practice. Therefore, we also incorporate the drug interaction information extracted from the drug instructions as a supplement. First, owing to the large amount of non-(semi-structured) data in historical prescription big data and drug instructions, we need to transform these data into structured triplet data. Second, we build the clinical experience edges and medical knowledge edges based on the structured data of historical prescriptions and drug instruction. Finally, we construct our medication knowledge graph according to two kinds of edges. We elaborate the details as follows.</p>
<sec id="s2a1"><label>2.1.1</label><title>Information extraction</title>
<p>In this step, to extract drug entities and relationships, we modeled the problem as an information extraction task in natural language processing (NLP) and solved it using information extraction technology. First, for historical prescription data, we transformed semi-structured disease&#x2013;drug&#x2013;diagnosis information into structured clinical triples to achieve a complete delineation of clinical experience. Then, for auxiliary medical text data, we extracted medical knowledge triples to supplement the medication knowledge graph. The information extract details are elaborated as follows.</p>
<p><italic>Clinical experience extraction.</italic> In this step, we extracted clinical experience based on the collected prescription big data. The historical prescription data mainly include the prescription number, patient&#x2019;s age, height, weight, and other personal signs, the diagnosis of disease, drugs, and their course of treatment, and other information. To better show the clinical medication experience, we established explicit attributes of entities and implicit triple relationships between entities according to medication knowledge.</p>
<p>Specifically, we first extract different entities in the prescription, including patients, diseases, and drugs. Then, we regard physiological characteristics such as gender, age, height, and weight as the attributes of patient entity. In addition, if there is a diagnosis on the prescription that is associated with a pregnant woman, such as at 14&#x2009;weeks of gestation, the patient will be given the role of pregnant woman. Furthermore, we construct implicit relationships between different entities based on prescriptions, such as the relationship between the patient and the drug, the relationship between the patient and the disease, and the relationship between the diagnosis and the disease. Finally, we iterate over each prescription and use a triple to represent all the entities in the prescription and their relationships, such as &#x201C;Influenza&#x2013;Prescribe&#x2013;Ribavirin Spray,&#x201D; etc.</p>
<p><italic>Medical knowledge extraction.</italic> In this step, we extract medical knowledge based on the collected dataset of drug instructions. As there is an implicit regional structure in each of the drug instructions, as shown on the left in <xref ref-type="fig" rid="F2">Figure&#x00A0;2</xref>, we first divide a part of the collected drug instruction data into blocks to extract the required structured information, such as drug name, main ingredients, indications, contraindications, adverse reactions, precautions, and drug interactions. Second, we manually label the pre-processed dataset based on the open source labeling tool YEDDA, as shown in <xref ref-type="fig" rid="F2">Figure&#x00A0;2</xref>. In addition, the labeled entity includes but is not limited to the drug names, diseases, ingredients, indications, adverse reactions, and contraindications. Third, we also marked another dataset in the format of (text, entity, relationship, entity) on the module data of annotated notes and drug interaction to extract drug interactions. According to the harm degree of drug interaction to the human body, the relationship fields of drug interaction are divided into four categories, which are beneficial, no effect, unknown, and harmful. Finally, we model the medical knowledge extraction problem as named entity recognition and relation extraction tasks in NLP to extract medical triplet information, as shown in <xref ref-type="fig" rid="F3">Figure&#x00A0;3</xref>.</p>
<fig id="F2" position="float"><label>Figure 2</label>
<caption><p>We use the YEDDA tool to label the entities in the drug instructions, and the labeled entities include the drug names, diseases, ingredients, indications, adverse reactions, and contraindications.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g002.tif"/>
</fig>
<fig id="F3" position="float"><label>Figure 3</label>
<caption><p>The medical knowledge extraction framework.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g003.tif"/>
</fig>
<p>Specifically, we first train the Bert-BilSTM-CRF model to recognize medical entities, including drug, ingredient, disease, indication, and contraindication. The Bert-Bi-LSTM-CRF model was proven to outperform all other models in the NLP of Chinese electronic health documents (<xref ref-type="bibr" rid="B22">22</xref>). Second, to extract the relationships between entities, such as drug interactions, we construct the relation extraction model (RE model): BertModel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM1"><mml:mo>+</mml:mo></mml:math></inline-formula> Dropout <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM2"><mml:mo>+</mml:mo></mml:math></inline-formula> Linear. Finally, we employ the trained entity recognition model to extract medical entities. In addition, for drug interaction data, we identify the relationships based on the extracted entities. There are two approaches to form medical triplet data. The first method is to take the drug name and other entities as the first entity and the second entity, respectively, and label as the relation, such as &#x201C;Ribavirin spray&#x2013;Ingredient&#x2013;Ribavirin.&#x201D; The second method is to use the triplet data extracted from the RE model, such as &#x201C;Cefoperazone sodium for injection&#x2013;Contraindication&#x2013;Amikacin.&#x201D;</p>
</sec>
<sec id="s2a2"><label>2.1.2</label><title>Graph node and edge construction</title>
<p>To construct the medication knowledge graph <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM3"><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>E</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, we define entities (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM4"><mml:mi>E</mml:mi></mml:math></inline-formula>) and the relationships (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM5"><mml:mi>R</mml:mi></mml:math></inline-formula>) between them, as illustrated in <xref ref-type="fig" rid="F4">Figure&#x00A0;4</xref>. In graph theory, the triple <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM6"><mml:mi>Q</mml:mi></mml:math></inline-formula> is defined as the set of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM7"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM8"><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM9"><mml:msub><mml:mi>e</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> denote two different entities, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM10"><mml:mi>r</mml:mi></mml:math></inline-formula> denotes the relationship between node <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM11"><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and node <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM12"><mml:msub><mml:mi>e</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>. As shown in <xref ref-type="fig" rid="F4">Figure&#x00A0;4</xref>, the black edge sets represent the clinical experience edges <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM13"><mml:msub><mml:mi>R</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula>, and the red edge sets represent the medical knowledge edges <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM14"><mml:msub><mml:mi>R</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula>. The detailed construction information can be found in the Appendix.</p>
<fig id="F4" position="float"><label>Figure 4</label>
<caption><p>The structure of the medication knowledge graph.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g004.tif"/>
</fig>
</sec>
</sec>
<sec id="s2b"><label>2.2</label><title>Drug recommendation model based on knowledge graph</title>
<p>In this section, our objective is to model the complex relationships between patients, diagnoses, and drugs based on the medication knowledge graph constructed in the previous phase. As there are many prescription features in the medication knowledge graph, we first extract the features of patients, diagnoses, and drugs. Then, we employ the sequential generation model to model the sequential decision-making process of the drug regimen. We specify the specific work as follows.</p>
<sec id="s2b1"><label>2.2.1</label><title>Feature extraction from graph</title>
<p>In clinical practice, most pediatric medicines are dosed according to the patient&#x2019;s age (<xref ref-type="bibr" rid="B23">23</xref>), body height, or body weight (mg/kg) (<xref ref-type="bibr" rid="B11">11</xref>). Moreover, treatments also vary according to the patient&#x2019;s symptom and indication; therefore, diagnostic information is helpful when developing medication regimens (<xref ref-type="bibr" rid="B24">24</xref>). As combination drugs are more common in complex prescriptions, they are more likely to cause ADRs. Therefore, drug interactions should also be considered in the rational and appropriate use of drugs. Based on this previous knowledge, we extract the corresponding physiological diagnostic features and drug interaction feature from the medical knowledge graph constructed in the previous phase. Detailed information is provided in the Appendix.</p>
</sec>
<sec id="s2b2"><label>2.2.2</label><title>Sequence generation model</title>
<p>In this step, our objective is to predict the rational medication regimen based on the extracted features. One of the intuitive methods is to concatenate the physiology and indication features into a vector and build a regression or classification model to predict the rational medication regimen. However, owing to the considerable variety of the two categories of features, such a direct concatenation of the two heterogeneous features does not perform well, especially when some features play a dominate role in specific medication conditions (<xref ref-type="bibr" rid="B25">25</xref>). To address these challenges, we use the sequence generation model to transform the problem into a sequence decision process of the drug regimen, including the medication list and treatment of drug use. Detailed information is provided in the Appendix.</p>
</sec>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Experiments</title>
<p>In this section, we evaluate our method with a medical record dataset collected from the MIMIC&#x005F;III dataset and real-world anonymized prescription big data collected from tertiary hospitals. We first introduce the experiment settings and then present the evaluation results. Finally, we display our analysis results on the visualization platform.</p>
<sec id="s3a"><label>3.1</label><title>Experiment settings</title>
<sec id="s3a1"><label>3.1.1</label><title>Dataset</title>
<p>After data cleansing, we obtain a dataset containing 1,084,594 prescriptions with 23,225 patients, 2,393 medicines, and 5,591 diagnoses from the MIMIC&#x005F;III database, and another dataset containing 230,390 prescriptions with 19,146 patients, 3,782 diagnoses, and 1,198 medicines from tertiary hospitals. The summary of the dataset is shown in <xref ref-type="table" rid="T1">Table&#x00A0;1</xref>.</p>
<table-wrap id="T1" position="float"><label>Table 1</label>
<caption><p>Summary of datasets.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">MIMIC&#x005F;III</th>
<th valign="top" align="center">FUJIAN</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Data collection period</td>
<td valign="top" align="center">2001&#x2013;2008</td>
<td valign="top" align="center">January 2015&#x2013;January 2017</td>
</tr>
<tr>
<td valign="top" align="left">&#x0023; Prescriptions</td>
<td valign="top" align="center">1,084,594</td>
<td valign="top" align="center">230,390</td>
</tr>
<tr>
<td valign="top" align="left">&#x0023; Medicines</td>
<td valign="top" align="center">2,393</td>
<td valign="top" align="center">1,198</td>
</tr>
<tr>
<td valign="top" align="left">&#x0023; Diagnoses</td>
<td valign="top" align="center">5,591</td>
<td valign="top" align="center">3,782</td>
</tr>
<tr>
<td valign="top" align="left">&#x0023; Patients</td>
<td valign="top" align="center">23,225</td>
<td valign="top" align="center">19,146</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><italic>MIMIC&#x005F;III dateset</italic>: In this study, we first perform the pre-processing operation of removing invalid patient prescriptions with medical devices, no weight field, and incorrect age statistics. As shown in <xref ref-type="table" rid="T2">Table&#x00A0;2</xref>, after pre-processing, the dataset contains 11 attributes: (1) patient ID; (2) case number; (3) sex; (4) age, calculated from the patient&#x2019;s date of birth and admission date, measured in years; (5) Weight, measured in kilograms; (6) diagnosis name; (7) drug ID, National Drug Code for medications; (8) drug name; (9) dosage; (10) dosage unit; and (11) days of administration (the duration of medication usage prescribed by the doctor, calculated from the start and end dates of medication usage, measured in days). The dataset&#x2019;s characteristics include the presence of multiple hospital admissions for some patients, as evidenced by records 4&#x2013;5 in <xref ref-type="table" rid="T2">Table&#x00A0;2</xref>, in which patient &#x201C;109&#x201D; has two case numbers: &#x201C;173633&#x201D; and &#x201C;172335&#x201D;. In addition, the dataset includes instances in which multiple diagnoses were assigned to a patient during a single prescription, with multiple medications prescribed for treatment, as illustrated by records 1&#x2013;3 in <xref ref-type="table" rid="T2">Table&#x00A0;2</xref>. For example, for patient &#x201C;23,&#x201D; with case number &#x201C;124321,&#x201D; the physician assigned two diagnoses, &#x201C;2252&#x201D; and &#x201C;V4581,&#x201D; and prescribed three medications for treatment: &#x201C;vancomycin,&#x201D; &#x201C;levofloxacin,&#x201D; and &#x201C;dexamethasone.&#x201D;</p>
<table-wrap id="T2" position="float"><label>Table 2</label>
<caption><p>Example of the MIMIC&#x005F;III medical record dataset.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="left"/>
<col align="center"/>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Patient ID</th>
<th valign="top" align="center">Case Number</th>
<th valign="top" align="center">Sex</th>
<th valign="top" align="center">Age</th>
<th valign="top" align="center">Weight</th>
<th valign="top" align="center">Diagnosis name</th>
<th valign="top" align="center">Drug ID</th>
<th valign="top" align="center">Drug name</th>
<th valign="top" align="center">Dosage</th>
<th valign="top" align="center">Dosage unit</th>
<th valign="top" align="center">Days</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">23</td>
<td valign="top">124321</td>
<td valign="top">M</td>
<td valign="top">75</td>
<td valign="top">66.8</td>
<td valign="top">Meningitis</td>
<td valign="top">00338355248</td>
<td valign="top">Vancomycin</td>
<td valign="top">1,000</td>
<td valign="top">mg</td>
<td valign="top">4</td>
</tr>
<tr>
<td valign="top" align="left">23</td>
<td valign="top">124321</td>
<td valign="top">M</td>
<td valign="top">75</td>
<td valign="top">66.8</td>
<td valign="top">Meningitis</td>
<td valign="top">00045006601</td>
<td valign="top">Levofloxacin</td>
<td valign="top">750</td>
<td valign="top">mg</td>
<td valign="top">2</td>
</tr>
<tr>
<td valign="top" align="left">23</td>
<td valign="top">124321</td>
<td valign="top">M</td>
<td valign="top">75</td>
<td valign="top">66.8</td>
<td valign="top">Meningitis</td>
<td valign="top">00054817525</td>
<td valign="top">Dexamethasone</td>
<td valign="top">4</td>
<td valign="top">mg</td>
<td valign="top">2</td>
</tr>
<tr>
<td valign="top" align="left">109</td>
<td valign="top">173633</td>
<td valign="top">F</td>
<td valign="top">24</td>
<td valign="top">44.9</td>
<td valign="top">Hypertensive chronic kidney disease</td>
<td valign="top">00172438210</td>
<td valign="top">Gabapentin</td>
<td valign="top">300</td>
<td valign="top">mg</td>
<td valign="top">6</td>
</tr>
<tr>
<td valign="top" align="left">109</td>
<td valign="top">172335</td>
<td valign="top">F</td>
<td valign="top">24</td>
<td valign="top">66.8</td>
<td valign="top">Other primary cardiomyopathies</td>
<td valign="top">00182055589</td>
<td valign="top">Hydralazine</td>
<td valign="top">50</td>
<td valign="top">mg</td>
<td valign="top">2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><italic>Fujian dataset</italic>: The second historical medical record dataset used in this study is derived from the clinical outpatient data of a tertiary hospital in Fujian Province, China. After pre-processing, this dataset contains a total of 11 attributes, as shown in <xref ref-type="table" rid="T3">Table&#x00A0;3</xref>. The difference between this dataset and the MIMIC&#x005F;III dataset lies in the inclusion of prescription numbers and patient heights, while excluding case numbers and drug numbers. As illustrated by the examples in <xref ref-type="table" rid="T3">Table&#x00A0;3</xref>, it can be observed that the characteristics of this dataset are consistent with those of the MIMIC&#x005F;III medical record dataset. These characteristics include multiple hospital admissions for patients and instances in which multiple diagnosis information and multiple medications are prescribed for a single hospitalization.</p>
<table-wrap id="T3" position="float"><label>Table 3</label>
<caption><p>Example of medical record dataset from Fujian Province.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="left"/>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Patient ID</th>
<th valign="top" align="center">Prescription number</th>
<th valign="top" align="center">Sex</th>
<th valign="top" align="center">Age</th>
<th valign="top" align="center">Height</th>
<th valign="top" align="center">Weight</th>
<th valign="top" align="center">Diagnosis name</th>
<th valign="top" align="center">Drug name</th>
<th valign="top" align="center">Dosage</th>
<th valign="top" align="center">Dosage unit</th>
<th valign="top" align="center">Days</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">12&#x002A;&#x002A;26</td>
<td valign="top">d1&#x002A;&#x002A;2c</td>
<td valign="top">M</td>
<td valign="top">57</td>
<td valign="top">164</td>
<td valign="top">58.5</td>
<td valign="top">Septicemia</td>
<td valign="top">Alfacalcidol</td>
<td valign="top">0.25</td>
<td valign="top"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM15"><mml:mrow><mml:mtext fontfamily="times">&#x3BC;</mml:mtext></mml:mrow></mml:math></inline-formula>g</td>
<td valign="top">10</td>
</tr>
<tr>
<td valign="top" align="left">12&#x002A;&#x002A;26</td>
<td valign="top">d1&#x002A;&#x002A;2c</td>
<td valign="top">M</td>
<td valign="top">57</td>
<td valign="top">164</td>
<td valign="top">58.5</td>
<td valign="top">Septicemia</td>
<td valign="top">Rebamipide Tablets</td>
<td valign="top">0.1</td>
<td valign="top">g</td>
<td valign="top">14</td>
</tr>
<tr>
<td valign="top" align="left">12&#x002A;&#x002A;26</td>
<td valign="top">c0&#x002A;&#x002A;28</td>
<td valign="top">M</td>
<td valign="top">57</td>
<td valign="top">164</td>
<td valign="top">58.5</td>
<td valign="top">Nausea and vomiting</td>
<td valign="top">Trivitamins ferrous chewable tablets</td>
<td valign="top">20</td>
<td valign="top">tablet</td>
<td valign="top">14</td>
</tr>
<tr>
<td valign="top" align="left">10&#x002A;&#x002A;26</td>
<td valign="top">ae&#x002A;&#x002A;02</td>
<td valign="top">F</td>
<td valign="top">79</td>
<td valign="top">166</td>
<td valign="top">70</td>
<td valign="top">Hypertension</td>
<td valign="top">Nifedipine controlled release tablets</td>
<td valign="top">30.0</td>
<td valign="top">mg</td>
<td valign="top">14</td>
</tr>
<tr>
<td valign="top" align="left">10&#x002A;&#x002A;26</td>
<td valign="top">ae&#x002A;&#x002A;02</td>
<td valign="top">F</td>
<td valign="top">79</td>
<td valign="top">166</td>
<td valign="top">70</td>
<td valign="top">Hyperlipemia</td>
<td valign="top">Pitavastatin calcium tablets</td>
<td valign="top">2.0</td>
<td valign="top">mg</td>
<td valign="top">7</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3a2"><label>3.1.2</label><title>Experiment plan</title>
<p>First, we randomly select 80&#x0025; of the prescriptions collected from the constructed medication knowledge graph for training, and the left 20&#x0025; for evaluation. Then, we collected 100 problematic prescriptions with inappropriate drug use from small hospitals as a test dataset to evaluate our trained model. Specifically, for each prescription, we use our model to classify whether these prescriptions are an inappropriate use of drugs. We evaluated our model by measuring the proportion of correctly classified prescriptions in terms of the medication sequence list and medication treatment, and using the rational medication regimen to represent the predicted results of both.</p>
</sec>
<sec id="s3a3"><label>3.1.3</label><title>Evaluation metrics</title>
<p>To measure the accuracy of the proposed model, we used the Jaccard Similarity Score (Jaccard, as defined in <xref ref-type="disp-formula" rid="disp-formula1">Equation 1</xref>), precision (as defined in <xref ref-type="disp-formula" rid="disp-formula2">Equation 2</xref>), recall (as defined in <xref ref-type="disp-formula" rid="disp-formula3">Equation 3</xref>), and average F1 (as defined in <xref ref-type="disp-formula" rid="disp-formula4">Equation 4</xref>), as shown in <xref ref-type="table" rid="T4">Table&#x00A0;4</xref>. Jaccard is defined as the size of the intersection divided by the size of the union of the ground truth medication regimen <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM16"><mml:msubsup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and predicted medication regimen <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM17"><mml:msubsup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>:<disp-formula id="disp-formula1"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM1"><mml:mtext>Jaccard</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>&#x22C2;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>&#x222A;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM18"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of patients in test set.</p>
<table-wrap id="T4" position="float"><label>Table 4</label>
<caption><p>Evaluation metrics overview.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Metric</th>
<th valign="top" align="left">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Jaccard</td>
<td valign="top" align="left">Measures the similarity between predicted drug prescriptions and actual drug prescriptions.</td>
</tr>
<tr>
<td valign="top" align="left">Precision</td>
<td valign="top" align="left">Measures the proportion of predicted drug prescriptions correctly identified by the model.</td>
</tr>
<tr>
<td valign="top" align="left">Recall</td>
<td valign="top" align="left">Measures the percentage of successful identifications by the model in actual drug prescriptions.</td>
</tr>
<tr>
<td valign="top" align="left">F1</td>
<td valign="top" align="left">Combining the accuracy and recall of the model is a comprehensive metric for evaluating the performance of the model.</td>
</tr>
<tr>
<td valign="top" align="left">DDI Rate</td>
<td valign="top" align="left">Measures the probability of a drug interaction in a predicted drug sequence.</td>
</tr>
</tbody>
</table>
</table-wrap><disp-formula id="disp-formula2"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM2"><mml:mtext>Precision</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>&#x22C2;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula> <disp-formula id="disp-formula3"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM3"><mml:mtext>Recall</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>&#x22C2;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula><disp-formula id="disp-formula4"><label>(4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM4"><mml:mtext>F1</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mtext>\,precision</mml:mtext><mml:mo>&#x00D7;</mml:mo><mml:mtext>recall</mml:mtext></mml:mrow><mml:mrow><mml:mtext>\,precision</mml:mtext><mml:mo>+</mml:mo><mml:mtext>recall</mml:mtext></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>When considering the accuracy of drug prediction, we also need to measure the safety of drug prediction; therefore, to measure medication safety, we define the DDI rate (as defined in <xref ref-type="disp-formula" rid="disp-formula5">Equation 5</xref>) to judge the probability of drug interactions in the predicted drug sequence:<disp-formula id="disp-formula5"><label>(5)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM5"><mml:mtext>DDI Rate</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>where the set will count each medication pair <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM19"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in the recommendation set if the pair belongs to the drug interaction adjacency matrix constructed. Here, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM20"><mml:mi>N</mml:mi></mml:math></inline-formula> is the size of the test dataset.</p>
</sec>
<sec id="s3a4"><label>3.1.4</label><title>Baseline methods</title>
<p>We compared our method with several baseline methods with regard to medication regimen prediction and medication regimen adequate evaluation. For medication regimen prediction, we compared our model with several baseline methods as follows.
<list list-type="simple">
<list-item><label>1.</label>
<p><bold>Bi-LSTM</bold>: this baseline is a sequence-sequence model. At the encoding end, BI-LSTM is used to learn the diagnostic information at the input end, and at the decoding end, ordinary LSTM is used to predict drugs.</p></list-item>
<list-item><label>2.</label>
<p><bold>GAMENet</bold>: this baseline is a memory-enhancing neural network model that inherits a drug interaction knowledge graph through the graph convolutional network storage module to provide safe and personalized drug combination recommendations.</p></list-item>
</list></p>
<p>For appropriate medication evaluation, we compare our model with the following baselines.
<list list-type="simple">
<list-item><label>1.</label>
<p><bold>Empirical</bold>: This method is based only on the medical experience of professional doctors in tertiary hospitals, without considering the drug contraindication information from existing drug instructions.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s3b"><label>3.2</label><title>Evaluation results</title>
<sec id="s3b1"><label>3.2.1</label><title>Medication regimen prediction evaluation results</title>
<p><xref ref-type="table" rid="T5">Table&#x00A0;5</xref> shows the rational drug use prediction results from the MIMIC&#x005F;III dataset using our proposed method as well as the baselines. Results show our proposed method has the highest score among all baselines with respect to Jaccard, precision, recall, and F1. The model we used benefited from the advantages of its structure, which could obtain the relationship between patients&#x2019; multiple diagnoses, making it closer to the real doctor&#x2019;s prescription when making drug predictions.</p>
<table-wrap id="T5" position="float"><label>Table 5</label>
<caption><p>The rational medication regimen prediction results of the MIMIC&#x005F;III dataset.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Methods</th>
<th valign="top" align="center">Jaccard</th>
<th valign="top" align="center">Precision</th>
<th valign="top" align="center">Recall</th>
<th valign="top" align="center">F1</th>
<th valign="top" align="center">DDI rate</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Bi-LSTM</td>
<td valign="top" align="center">0.5115</td>
<td valign="top" align="center">0.6705</td>
<td valign="top" align="center">0.5697</td>
<td valign="top" align="center">0.6160</td>
<td valign="top" align="center">0.1624</td>
</tr>
<tr>
<td valign="top" align="left">GAMENet</td>
<td valign="top" align="center">0.6517</td>
<td valign="top" align="center">0.7501</td>
<td valign="top" align="center">0.6535</td>
<td valign="top" align="center">0.6985</td>
<td valign="top" align="center">0.1324</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Ours</bold></td>
<td valign="top" align="center"><bold>0.8685</bold></td>
<td valign="top" align="center"><bold>0.9555</bold></td>
<td valign="top" align="center"><bold>0.8927</bold></td>
<td valign="top" align="center"><bold>0.9173</bold></td>
<td valign="top" align="center"><bold>0.0867</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn111"><p>Ours indicate the proposed method in this paper and bold values indicate the best performance in the corresponding metric.</p></fn>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T6">Table&#x00A0;6</xref> shows the rational drug use regimen prediction results for the outpatient medical record dataset using our proposed method as well as the baselines. Results show our proposed method achieves the best performance compared with other baseline methods. As this dataset is different from MIMIC&#x005F;III, no authoritative drug classification has been performed, and drugs with therapeutic equivalence exist in this dataset as multiple drugs. Therefore, we chose to provide three alternative elements for each element generated by the sequence model. It is deemed to be the correct prediction when the actual use of the drug appears in the three alternatives. The calculation formula of its evaluation index is shown in <xref ref-type="disp-formula" rid="disp-formula6">Equation 6</xref>.<disp-formula id="disp-formula6"><label>(6)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM6"><mml:mtext>Precision</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mi>M</mml:mi></mml:mfrac></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM21"><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo>,</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>.</p>
<table-wrap id="T6" position="float"><label>Table 6</label>
<caption><p>The prediction results of the outpatient medical record dataset.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Methods</th>
<th valign="top" align="center">Precision</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Bi-LSTM</td>
<td valign="top" align="center">0.4355</td>
</tr>
<tr>
<td valign="top" align="left">GAMENet</td>
<td valign="top" align="center">0.6355</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Ours</bold></td>
<td valign="top" align="center"><bold>0.7769</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn114"><p>Bold values indicate the best performance in the corresponding metric, and Ours indicate the proposed method in this paper.</p></fn>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T7">Table&#x00A0;7</xref> clearly shows the predicted results of rational drug use of the model in the Fujian province dataset. The actual prescriptions shown on the left of the table below are two drugs taken by a boy for allergic rhinitis, mycoplasma infection, and bronchitis. On the right are the recommendations for drug therapy provided by our model. It can be found that the model in this paper can accurately cover the real prescription after providing three alternatives, and most of the other alternatives provided are also drugs for the treatment of respiratory diseases such as rhinitis. It indicates that the model in this study can obtain the drug recommendations of actual doctors according to patient diagnosis and other characteristics.</p>
<table-wrap id="T7" position="float"><label>Table 7</label>
<caption><p>Drug recommendation cases in the Fujian Province dataset.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Real prescription</th>
<th valign="top" align="left">Predict prescription</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Ganan mixture</td>
<td valign="top" align="left">Ganan mixture,</td>
</tr>
<tr>
<td valign="top" align="left"/>
<td valign="top" align="left">Josamycinpropionate granules,</td>
</tr>
<tr>
<td valign="top" align="left"/>
<td valign="top" align="left">Mometasone furoate aque</td>
</tr>
<tr>
<td valign="top" align="left">Loratadine tablets</td>
<td valign="top" align="left">Montelukast sodium oral granules,</td>
</tr>
<tr>
<td valign="top" align="left"/>
<td valign="top" align="left">Loratadine tablets,</td>
</tr>
<tr>
<td valign="top" align="left"/>
<td valign="top" align="left">Ambroterol oral solution</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, we conducted two comparative experiments to determine whether the addition of patients&#x2019; physiological and DDI features could improve the effect of our model. <xref ref-type="table" rid="T8">Table&#x00A0;8</xref> shows the accuracy rate of drug recommendation is improved after the introduction of physiological characteristics in our model. <xref ref-type="table" rid="T9">Table&#x00A0;9</xref> shows that after the introduction of DDI features in our model, although the accuracy of model prediction is slightly sacrificed, the probability of adverse drug interactions caused by recommended drugs is reduced. Therefore, we can adjust the weight proportion of DDI characteristics to meet the application requirements.</p>
<table-wrap id="T8" position="float"><label>Table 8</label>
<caption><p>The influence of physiological features.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">Jaccard</th>
<th valign="top" align="center">Precision</th>
<th valign="top" align="center">Recall</th>
<th valign="top" align="center">F1</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">No physical features</td>
<td valign="top" align="center">0.8556</td>
<td valign="top" align="center">0.9546</td>
<td valign="top" align="center">0.8727</td>
<td valign="top" align="center">0.9069</td>
</tr>
<tr>
<td valign="top" align="left">Have physical features</td>
<td valign="top" align="center"><bold>0.8627</bold></td>
<td valign="top" align="center"><bold>0.9564</bold></td>
<td valign="top" align="center"><bold>0.8855</bold></td>
<td valign="top" align="center"><bold>0.9131</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn115"><p>Bold values indicate the best performance in the corresponding metric.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="T9" position="float"><label>Table 9</label>
<caption><p>The influence of DDI features.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">DDI rate</th>
<th valign="top" align="center">Jaccard</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">No DDI features</td>
<td valign="top" align="center">0.0906</td>
<td valign="top" align="center">0.8239</td>
</tr>
<tr>
<td valign="top" align="left">Have DDI features</td>
<td valign="top" align="center"><bold>0.0837</bold></td>
<td valign="top" align="center"><bold>0.8167</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn112"><p>Bold values indicate the best performance in the corresponding metric.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3b2"><label>3.2.2</label><title>Medication regimen appropriateness evaluation results</title>
<p>In this study, we use the trained models as classifiers to judge the appropriateness of prescriptions in small hospitals. <xref ref-type="table" rid="T10">Table&#x00A0;10</xref> shows the average accuracy scores of medicine use appropriate evaluation using the proposed method and the baselines. We can see that the proposed method achieves the best performance with regard to evaluation accuracy scores. Specifically, the baseline method <italic>Empirical</italic> attempts to evaluate the appropriate use of drugs based only on the experience of drug use of professional doctors in tertiary hospitals, which results in some combination drugs with adverse reactions being misjudged. In summary, the proposed method integrates the two heterogeneous information to model sequential patterns and therefore improves the accuracy of evaluation. The calculation formula of its evaluation index is shown in <xref ref-type="disp-formula" rid="disp-formula7">Equation 7</xref>.<disp-formula id="disp-formula7"><label>(7)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM7"><mml:mtext>Accuracy</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mo fence="false" stretchy="false">&#x230A;</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mi>M</mml:mi></mml:mfrac></mml:mrow><mml:mo fence="false" stretchy="false">&#x230B;</mml:mo><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM22"><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo>,</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM23"><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> represents the rationality marker of prescriptions in small hospital.</p>
<table-wrap id="T10" position="float"><label>Table 10</label>
<caption><p>The appropriate evaluation of prescription in small hospitals.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Methods</th>
<th valign="top" align="center">Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Empirical</td>
<td valign="top" align="center">83&#x0025;</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>89&#x0025;</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn113"><p>Bold values indicate the best performance in the corresponding metric.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s3c"><label>3.3</label><title>Clinical appropriate medication evaluation system</title>
<p>To demonstrate the work in this paper more clearly, we have built a platform for the appropriateness of rational drug use and applied it in a small hospital to evaluate the appropriateness of doctors&#x2019; prescriptions. As shown in <xref ref-type="fig" rid="F5">Figure&#x00A0;5</xref>, this platform is mainly divided into two parts, among which the left view is divided into three subgraphs (patient&#x2013;disease&#x2013;drug), and the right view shows the evaluation results. In the left frame, you can first fill in patient information, diagnostic information, and drug information. Then, click the button to evaluate for the appropriateness of prescribing. Finally, the predicted medication results are displayed on the right side of the frame, with a tabulated comparison of the doctor&#x2019;s medication regimen and the predicted medication regimen. At the same time, the entity relationship involved with the disease in the medication graph would be visualized below the evaluation results, which would be convenient for doctors and pharmacists to further review and modify prescriptions.</p>
<fig id="F5" position="float"><label>Figure 5</label>
<caption><p>Clinical prescription appropriateness evaluation system.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g005.tif"/>
</fig>
</sec>
<sec id="s3d"><label>3.4</label><title>Case study</title>
<p>We conduct a case study of one prescription randomly selected from 100 problematic prescriptions of inappropriate medications in small hospitals. As shown in <xref ref-type="fig" rid="F5">Figure&#x00A0;5</xref>, in the input prescription, the patient is a male, has a height of 112&#x2009;cm, has a weight of 19 kg, and is 5 years old. The patient&#x2019;s diagnosed symptoms were acute bronchitis, bronchitis, and an acute upper respiratory tract infection. The prescribed medicines were amoxicillin and clavulanate potassium for oral suspension, Combivent, calcium gluconate oral solution, and the corresponding treatment courses are 3, 1, and 1&#x2009;day(s). Based on our proposed framework, the predicted medication regimen was amoxicillin and clavulanate potassium for oral suspension for 3 days and Combivent for 1&#x2009;day. After evaluation, the system would provide default color labels and red labels to represent the consistent medication regimen and inconsistent medication regimen, respectively. The red label in the picture indicated whether calcium gluconate oral solution are unnecessary drugs.</p>
</sec>
</sec>
<sec id="s4" sec-type="conclusions"><label>4</label><title>Conclusion</title>
<p>In this study, we investigate one of the key problems in rational medication, i.e., the evaluation of appropriate medication use. We propose a framework of appropriate drug use based on medical association and big data to accurately predict the medication regimen by leveraging prescription big data and medical text data. Specifically, a medication knowledge graph is first constructed based on historical prescription big data and medical text data from tertiary hospitals. Then, we employ a GMM for physiological features, BERT for diagnostics, and graph convolutions for drug interactions, yielding accurate medication regimens. Our approach surpasses baselines in predicting regimens and detecting appropriate medications, and was validated on MIMIC&#x005F;III and real-world prescription data from tertiary hospitals.</p>
<p>One of the limitations of this study is the <italic>feature selection</italic>. There might be other indication or physiology features that could be associated with medication regimens and used as predictive features for example. For example, for adolescents, the probability of developing corresponding diseases during adolescence can also be considered in the prediction model to improve the prediction accuracy in teenagers. We are currently working with hospitals to retrieve richer information related to prescription datasets, such as picture archiving and communication systems and inspection results from laboratory information systems, which we believe will provide useful and important features for drug regimen prediction.</p>
<p>In the future, we plan to extend our work in the following directions. First, we plan to involve more data sources from other hospital information systems, especially data from clinical laboratories, to investigate more relevant factors of doctor medication. Second, we plan to investigate the reasons for the wrong medication sequence list, including overtreatment or undertreatment by model overfitting, and then leverage the knowledge to improve our predictive models. Third, we plan to integrate our method with the existing clinical decision support systems to provide dosing recommendations for doctors and pharmacists in small clinics.</p>
</sec>
</body>
<back>
<sec id="s6" sec-type="data-availability"><title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s111" sec-type="author-contributions"><title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p>
</sec>
<sec id="s112" sec-type="funding-information"><title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This research is supported by the Youth Science Foundation of Xiang&#x0027;an Hospital of Xiamen University (NO. XM03030003).</p>
</sec>
<ack><title>Acknowledgments</title>
<p>We would like to thank the reviewers for their constructive suggestions.</p>
</ack>
<sec id="s7" sec-type="COI-statement"><title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s8" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="book"><collab>WHO</collab>, <source>Promoting Rational Use of Medicines</source>. <publisher-loc>Geneva</publisher-loc>: <publisher-name>WHO Regional Office for South-East Asia</publisher-name> (<year>2011</year>).</citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bogert</surname><given-names>C</given-names></name><name><surname>Mestrinaro</surname><given-names>M</given-names></name><name><surname>Weerasuriya</surname><given-names>K</given-names></name></person-group>, <source>The Pursuit of Responsible Use of Medicines: Sharing and Learning from Country Experiences</source>. <publisher-loc>Geneva</publisher-loc>: <publisher-name>World Health Organization</publisher-name> (<year>2012</year>).</citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname><given-names>ZX</given-names></name><name><surname>Hohmann</surname><given-names>L</given-names></name><name><surname>Banjara</surname><given-names>B</given-names></name><name><surname>Zhao</surname><given-names>Y</given-names></name><name><surname>Diggs</surname><given-names>K</given-names></name><name><surname>Westrick</surname><given-names>SC</given-names></name></person-group>. <article-title>Recommendations to protect patients and health care practices from medicare and medicaid fraud</article-title>. <source>J Am Pharm Assoc (2003)</source>. (<year>2020</year>) <volume>60</volume>:<fpage>e60</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1016/j.japh.2020.05.011</pub-id><pub-id pub-id-type="pmid">32616445</pub-id></citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fei</surname><given-names>Y</given-names></name><name><surname>Fu</surname><given-names>Y</given-names></name><name><surname>Yang</surname><given-names>D</given-names></name><name><surname>Hu</surname><given-names>C.</given-names></name></person-group>. <article-title>Research on the formation mechanism of health insurance fraud in China: from the perspective of the tripartite evolutionary game</article-title>. <source>Front Public Health</source>. (<year>2022</year>) <volume>10</volume>:<fpage>930120</fpage>. <pub-id pub-id-type="doi">10.3389/fpubh.2022.930120</pub-id><pub-id pub-id-type="pmid">35812495</pub-id></citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khoury</surname><given-names>MJ</given-names></name><name><surname>Ioannidis</surname><given-names>JP</given-names></name></person-group>. <article-title>Big data meets public health</article-title>. <source>Science</source>. (<year>2014</year>) <volume>346</volume>:<fpage>1054</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaa2709</pub-id><pub-id pub-id-type="pmid">25430753</pub-id></citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Najafabadi</surname><given-names>MM</given-names></name><name><surname>Villanustre</surname><given-names>F</given-names></name><name><surname>Khoshgoftaar</surname><given-names>TM</given-names></name><name><surname>Seliya</surname><given-names>N</given-names></name><name><surname>Wald</surname><given-names>R</given-names></name><name><surname>Muharemagic</surname><given-names>E</given-names></name></person-group>. <article-title>Deep learning applications and challenges in big data analytics</article-title>. <source>J Big Data</source>. (<year>2015</year>) <volume>2</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.1186/s40537-014-0007-7</pub-id></citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shao</surname><given-names>Y</given-names></name><name><surname>Hong</surname><given-names>L</given-names></name><name><surname>Chen</surname><given-names>J</given-names></name><name><surname>Chen</surname><given-names>L</given-names></name><name><surname>Fan</surname><given-names>X</given-names></name><name><surname>Xu</surname><given-names>Z</given-names></name><etal/></person-group>. <article-title>Medicine concomitant modeling and risk evaluation based on knowledge graph</article-title>. <source>China Digit Med</source>. (<year>2018</year>) <volume>13</volume>:<fpage>39</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.3969/j.issn.1673-7571.2018.10.013</pub-id></citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Shang</surname><given-names>J</given-names></name><name><surname>Xiao</surname><given-names>C</given-names></name><name><surname>Ma</surname><given-names>T</given-names></name><name><surname>Li</surname><given-names>H</given-names></name><name><surname>Sun</surname><given-names>J</given-names></name></person-group>. <article-title>Gamenet: graph augmented memory networks for recommending medication combination</article-title>. In: <source>>Proceedings of the AAAI Conference on Artificial Intelligence</source>. Vol. <volume>33</volume> (<year>2019</year>). p. <fpage>1126</fpage>&#x2013;<lpage>33</lpage>.</citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yao</surname><given-names>Y</given-names></name><name><surname>Weng</surname><given-names>Y</given-names></name><name><surname>Deng</surname><given-names>J</given-names></name><name><surname>Zou</surname><given-names>L</given-names></name></person-group>. <article-title>Review on domestic and foreign development, application of the diagnosis related groups and the research of payment standard</article-title>. <source>Chin Health Econ</source>. (<year>2018</year>) <volume>37</volume>(<issue>1</issue>):<fpage>24</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.7664/CHE20180106</pub-id></citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kahn</surname><given-names>K</given-names></name><name><surname>Rogers</surname><given-names>W</given-names></name><name><surname>Rubenstein</surname><given-names>L</given-names></name><name><surname>Sherwood</surname><given-names>M</given-names></name><name><surname>Reinisch</surname><given-names>E</given-names></name><name><surname>Keeler</surname><given-names>E</given-names></name><etal/></person-group>. <article-title>Measuring quality of care with explicit process criteria before and after implementation of the DRG-based prospective payment system</article-title>. <source>JAMA J Am Med Assoc</source>. (<year>1990</year>) <volume>264</volume>:<fpage>1969</fpage>&#x2013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1001/jama.1990.03450150069033</pub-id></citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frey</surname><given-names>AMM</given-names></name></person-group>. <article-title>Pediatric dosage calculations</article-title>. <source>J Infus Nurs</source>. (<year>1985</year>) <volume>8</volume>:<fpage>373</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname><given-names>X</given-names></name><name><surname>Menche</surname><given-names>J</given-names></name><name><surname>Barab&#x00E1;si</surname><given-names>AL</given-names></name><name><surname>Sharma</surname><given-names>A</given-names></name></person-group>. <article-title>Human symptoms&#x2013;disease network</article-title>. <source>Nat Commun</source>. (<year>2014</year>) <volume>5</volume>:<fpage>4212</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms5212</pub-id><pub-id pub-id-type="pmid">24967666</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skou</surname><given-names>ST</given-names></name><name><surname>Mair</surname><given-names>FS</given-names></name><name><surname>Fortin</surname><given-names>M</given-names></name><name><surname>Guthrie</surname><given-names>B</given-names></name><name><surname>Nunes</surname><given-names>BP</given-names></name><name><surname>Miranda</surname><given-names>JJ</given-names></name><etal/></person-group>. <article-title>Multimorbidity</article-title>. <source>Nat Rev Dis Primers</source>. (<year>2022</year>) <volume>8</volume>:<fpage>48</fpage>. <pub-id pub-id-type="doi">10.1038/s41572-022-00376-4</pub-id><pub-id pub-id-type="pmid">35835758</pub-id></citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wallace</surname><given-names>J</given-names></name><name><surname>Paauw</surname><given-names>D</given-names></name></person-group>. <article-title>Appropriate prescribing and important drug interactions in older adults</article-title>. <source>Med Clin North Am</source>. (<year>2015</year>) <volume>99</volume>:<fpage>295</fpage>&#x2013;<lpage>310</lpage>. <pub-id pub-id-type="doi">10.1016/j.mcna.2014.11.005</pub-id><pub-id pub-id-type="pmid">25700585</pub-id></citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname><given-names>J</given-names></name><name><surname>Parish</surname><given-names>A</given-names></name></person-group>. <article-title>Polypharmacy and medication management in older adults</article-title>. <source>Nurs Clin North Am</source>. (<year>2017</year>) <volume>52</volume>:<fpage>457</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1016/j.cnur.2017.04.007</pub-id><pub-id pub-id-type="pmid">28779826</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kozyra</surname><given-names>E</given-names></name><name><surname>Lau</surname><given-names>T</given-names></name></person-group>. <article-title>Medication strategies: switching, tapering, cross-over, overmedication, drug&#x2013;drug interactions, and discontinuation syndromes</article-title>. In: <source>Inpatient Geriatric Psychiatry</source>. Springer (<year>2019</year>). p. <fpage>325</fpage>&#x2013;<lpage>38</lpage>.</citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wurcel</surname><given-names>V</given-names></name><name><surname>Cicchetti</surname><given-names>A</given-names></name><name><surname>Garrison</surname><given-names>L</given-names></name><name><surname>Kip</surname><given-names>M</given-names></name><name><surname>Koffijberg</surname><given-names>H</given-names></name><name><surname>Kolbe</surname><given-names>A</given-names></name><etal/></person-group>. <article-title>The value of diagnostic information in personalised healthcare: a comprehensive concept to facilitate bringing this technology into healthcare systems</article-title>. <source>Public Health Genomics</source>. (<year>2019</year>) <volume>22</volume>:<fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1159/000501832</pub-id><pub-id pub-id-type="pmid">31390644</pub-id></citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joyner</surname><given-names>M</given-names></name><name><surname>Paneth</surname><given-names>N</given-names></name></person-group>. <article-title>Seven questions for personalized medicine</article-title>. <source>JAMA</source>. (<year>2015</year>) <volume>314</volume>(<issue>10</issue>):<fpage>999</fpage>&#x2013;<lpage>1000</lpage>. <pub-id pub-id-type="doi">10.1001/jama.2015.7725</pub-id><pub-id pub-id-type="pmid">26098474</pub-id></citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Vermet</surname><given-names>F</given-names></name></person-group>, <source>Statistical Learning Methods</source>. <publisher-name>Big Data for Insurance Companies</publisher-name> (<year>2018</year>).</citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Devlin</surname><given-names>J</given-names></name><name><surname>Chang</surname><given-names>MW</given-names></name><name><surname>Lee</surname><given-names>K</given-names></name><name><surname>Toutanova</surname><given-names>K</given-names></name></person-group>. <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:1810.04805</italic></comment> (<year>2018</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1810.04805">https://arxiv.org/abs/1810.04805</ext-link>.</citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kipf</surname><given-names>T</given-names></name><name><surname>Welling</surname><given-names>M</given-names></name></person-group>. <article-title>Semi-supervised classification with graph convolutional networks.</article-title> <comment><italic>arXiv</italic>. <italic>arXiv:1609.02907</italic></comment> (<year>2016</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1609.02907">https://arxiv.org/abs/1609.02907</ext-link>.</citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Dai</surname><given-names>Z</given-names></name><name><surname>Wang</surname><given-names>X</given-names></name><name><surname>Ni</surname><given-names>P</given-names></name><name><surname>Li</surname><given-names>Y</given-names></name><name><surname>Li</surname><given-names>G</given-names></name><name><surname>Bai</surname><given-names>X</given-names></name></person-group>. <article-title>Named entity recognition using BERT BiLSTM CRF for Chinese electronic health records</article-title>. In: <source>2019 12th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI)</source> (<year>2019</year>). p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.<pub-id pub-id-type="doi">10.1109/CISP-BMEI48845.2019.8965823</pub-id></citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Kelly</surname><given-names>AB</given-names></name><name><surname>Weier</surname><given-names>M</given-names></name><name><surname>Hall</surname><given-names>WD</given-names></name></person-group>, <source>The Age of Onset of Substance Use Disorders</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name> (<year>2019</year>). p. <fpage>149</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-72619-9-8</pub-id></citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname><given-names>M</given-names></name><name><surname>Hong</surname><given-names>L</given-names></name><name><surname>Zhao</surname><given-names>Y</given-names></name><name><surname>Chen</surname><given-names>L</given-names></name><name><surname>Wang</surname><given-names>J</given-names></name></person-group>. <article-title>Dosage prediction in pediatric medication leveraging prescription big data</article-title>. <source>IEEE Access</source>. (<year>2019</year>) <volume>7</volume>:<fpage>94285</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2928457</pub-id></citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname><given-names>L</given-names></name><name><surname>Fan</surname><given-names>X</given-names></name><name><surname>Wang</surname><given-names>L</given-names></name><name><surname>Zhang</surname><given-names>D</given-names></name><name><surname>Yu</surname><given-names>Z</given-names></name><name><surname>Li</surname><given-names>J</given-names></name><etal/></person-group>. <article-title>RADAR: road obstacle identification for disaster response leveraging cross-domain urban data</article-title>. <source>Proc ACM Interact Mob Wear Ubiquitous Technol</source>. (<year>2018</year>) <volume>1</volume>:<fpage>130:1</fpage>&#x2013;<lpage>130:23</lpage>. <pub-id pub-id-type="doi">10.1145/3161159</pub-id></citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Akaike</surname><given-names>H</given-names></name></person-group>, <source>Akaike&#x2019;s Information Criterion</source>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2011</year>).</citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neath</surname><given-names>AA</given-names></name><name><surname>Cavanaugh</surname><given-names>JE</given-names></name></person-group>. <article-title>The Bayesian information criterion: background, derivation, and applications</article-title>. <source>Wiley Interdiscip Rev Comput Stat</source>. (<year>2012</year>) <volume>4</volume>:<fpage>199</fpage>&#x2013;<lpage>203</lpage>. <pub-id pub-id-type="doi">10.1002/wics.199</pub-id></citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watanabe</surname><given-names>S</given-names></name></person-group>. <article-title>A widely applicable Bayesian information criterion</article-title>. <source>J Mach Learn Res</source>. (<year>2013</year>) <volume>14</volume>:<fpage>867</fpage>&#x2013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.5555/2567709.2502609</pub-id></citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dempster</surname><given-names>AP</given-names></name><name><surname>Laird</surname><given-names>NM</given-names></name><name><surname>Rubin</surname><given-names>DB</given-names></name></person-group>. <article-title>Maximum likelihood from incomplete data via the EM algorithm</article-title>. <source>J R Stat Soc Ser B (Methodol)</source>. (<year>2018</year>) <volume>39</volume>:<fpage>1</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1111/j.2517-6161.1977.tb01600.x</pub-id></citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wu</surname><given-names>Y</given-names></name><name><surname>Schuster</surname><given-names>M</given-names></name><name><surname>Chen</surname><given-names>Z</given-names></name><name><surname>Le</surname><given-names>QV</given-names></name><name><surname>Norouzi</surname><given-names>M</given-names></name><name><surname>Macherey</surname><given-names>W</given-names></name><etal/></person-group>. <article-title>Google&#x2019;s neural machine translation system: bridging the gap between human and machine translation</article-title>. <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:1609.08144</italic></comment> (<year>2016</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1609.08144">https://arxiv.org/abs/1609.08144</ext-link>.</citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaswani</surname><given-names>A</given-names></name><name><surname>Shazeer</surname><given-names>N</given-names></name><name><surname>Parmar</surname><given-names>N</given-names></name><name><surname>Uszkoreit</surname><given-names>J</given-names></name><name><surname>Jones</surname><given-names>L</given-names></name><name><surname>Gomez</surname><given-names>AN</given-names></name><etal/></person-group> <article-title>Attention is all you need</article-title>. <source>Adv Neural Inf Process Syst</source>. (<year>2017</year>):<fpage>6000</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.5555/3295222.3295349</pub-id></citation></ref>
<ref id="B32"><label>32.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Luong</surname><given-names>MT</given-names></name><name><surname>Pham</surname><given-names>H</given-names></name><name><surname>Manning</surname><given-names>CD</given-names></name></person-group>. <article-title>Effective approaches to attention-based neural machine translation</article-title>. <comment><italic>arXiv</italic> [Preprint]. <italic>arXiv:1508.04025</italic></comment> (<year>2015</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1508.04025">https://arxiv.org/abs/1508.04025</ext-link>.</citation></ref>
<ref id="B33"><label>33.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cho</surname><given-names>K</given-names></name><name><surname>Van Merrienboer</surname><given-names>B</given-names></name><name><surname>Gulcehre</surname><given-names>C</given-names></name><name><surname>Bahdanau</surname><given-names>D</given-names></name><name><surname>Bougares</surname><given-names>F</given-names></name><name><surname>Schwenk</surname><given-names>H</given-names></name><etal/></person-group> <article-title>Learning phrase representations using RNN encoder&#x2013;decoder for statistical machine translation</article-title>. In: <source>Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); Doha, Qatar</source>. <publisher-name>Association for Computational Linguistics</publisher-name> (<year>2014</year>). p. <fpage>1724</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.3115/v1/D14-1179</pub-id></citation></ref></ref-list>
<app-group><app id="app1"><title>Appendix</title>
<sec id="s9"><title>1 Graph node and edge construction</title>
<p><italic>Clinical experience edges.</italic> We first construct the clinical experience edge set <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM24"><mml:msub><mml:mi>R</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> based on the triples extracted from prescription big data. Specifically, the edge set <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM25"><mml:msub><mml:mi>R</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> of graph <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM26"><mml:mi>G</mml:mi></mml:math></inline-formula> are defined as follows: for the triples extracted from the prescription record, we set up two nodes and assign a directed edge between the corresponding nodes in graph <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM27"><mml:mi>G</mml:mi></mml:math></inline-formula>. The triples include &#x201C;Patient-Have-Prescription,&#x201D; &#x201C;Prescription-Diagnose-Disease,&#x201D; &#x201C;Prescription-Prescribe-Drug,&#x201D; &#x201C;Patient-Have-Disease,&#x201D; and &#x201C;Patient-Use-Drug.&#x201D;</p>
<p><italic>Medical knowledge edges.</italic> The second step is to construct medical knowledge edges <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM28"><mml:msub><mml:mi>R</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> based on the triples extracted from the drug instructions. Specifically, the edge set <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM29"><mml:msub><mml:mi>R</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> of graph <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM30"><mml:mi>G</mml:mi></mml:math></inline-formula> are defined as follows: for the triples extracted from drug instructions, we set up two nodes and assign a directed edge between the corresponding nodes in graph <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM31"><mml:mi>G</mml:mi></mml:math></inline-formula>. The triples include &#x201C;Drug-Ingredient-Drug,&#x201D; &#x201C;Drug-Indication-Disease,&#x201D; &#x201C;Drug-Contraindication-Patient,&#x201D; &#x201C;Drug-Contraindication-Disease,&#x201D; &#x201C;Drug-Interaction-Drug,&#x201D; and &#x201C;Drug-Contraindication-Drug.&#x201D;</p>
</sec>
<sec id="s10"><title>2 Feature extraction from graph</title>
<p><italic>Physiological feature extraction.</italic> In this step, our objective is to extract patients&#x2019; physiological information from historical prescription big data as the features of rational drug use. Physiology metrics of patients, such as sex, age group, body weight, and body height, are usually the most important considerations in clinical medication calculation. Although the combination of these factors provides the greatest accuracy in calculating medication regiments, simple digital groupings of different age groups, heights, or weights could lead to excessive discretion in the population sample. Therefore, we group patient populations as a physiological feature by modeling physiological information.</p>
<p>Owing to differences in gender, height, weight, and other physiological characteristics, the distribution of prescription data may be composed of N Gauss. For example, patients with chronic gastritis. By plotting the distribution of patients of different ages in male and female genders, as shown in <xref ref-type="fig" rid="F6">Figure A1</xref>, we can observe that there are approximately two component kernels in the distribution. The component kernals were contributed by male patients aged approximately 50 and female patients aged approximately 60. Therefore, we can use a mixed Gaussian distribution to fit all patient prescription data. We use the BMI to represent height and weight for each prescription data. First, we estimate the optimal number of component cores n for the patient&#x2019;s prescription data using the Akaike information criterion (AIC) (<xref ref-type="bibr" rid="B26">26</xref>) and Bayesian information criterion (BIC) (<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B28">28</xref>). The number of component cores is optimal when the AIC and BIC are as small as possible. Then, we calculate the probability distribution of the patient in each component kernel. Finally, we take the group with the highest probability as physiological characteristics, as shown in the <xref ref-type="fig" rid="F7">Figure A2</xref>. The probability density function of the GMM is given by <xref ref-type="disp-formula" rid="disp-formula8">Equation (A1)</xref>:<disp-formula id="disp-formula8"><label>(A1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM8"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="script">N</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM32"><mml:mi>X</mml:mi></mml:math></inline-formula> is the age distribution of prescriptions, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM33"><mml:mi>K</mml:mi></mml:math></inline-formula> is the number of sub-Gaussian models in GMMs, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM34"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the mixture coefficient, which is the probability that each observation data belongs to the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM35"><mml:mi>k</mml:mi></mml:math></inline-formula>th submodel. The <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM36"><mml:mrow><mml:mi mathvariant="script">N</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the distribution function, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM37"><mml:msub><mml:mi>u</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the expectation, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM38"><mml:msub><mml:mrow><mml:mi mathvariant="bold">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the covariance of the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM39"><mml:mi>k</mml:mi></mml:math></inline-formula>th component in the mixed model. The above variables satisfy <xref ref-type="disp-formula" rid="disp-formula9">Equations (A2)</xref> and <xref ref-type="disp-formula" rid="disp-formula10">(A3)</xref>:<disp-formula id="disp-formula9"><label>(A2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM9"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:math></disp-formula><disp-formula id="disp-formula10"><label>(A3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM10"><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>1</mml:mn></mml:math></disp-formula></p>
<fig id="F6" position="float"><label>Figure A1</label>
<caption><p>Distribution of patients with chronic gastritis.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g006.tif"/>
</fig>
<fig id="F7" position="float"><label>Figure A2</label>
<caption><p>The framework of physiological feature extraction.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g007.tif"/>
</fig>
<p>With the EM algorithm (expectation-maximization algorithm) (<xref ref-type="bibr" rid="B29">29</xref>), we can iteratively calculate the parameters in the GMM: (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM40"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">&#x03A3;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> ). In short, the EM algorithm has two steps. The first step is E (expectation), which updates the implicit variable. The second step is M (maximization), which is used to update the parameters of each Gaussian distribution in the GMM. Then, the above two steps are repeated until the iteration termination condition is reached.</p>
<p><italic>Diagnostic feature extraction.</italic> In this step, our objective is to extract diagnostic sequence information as diagnostic features of the patients. A simpler approach to represent diagnostic information is to digitize the diagnosis. However, the diagnostic text representation is rich in the semantic information of words, such as text similarity. Therefore, considering that word vectors are rich in more semantic features, we choose to represent diagnostic features by transforming text into word vectors through a BERT pre-trained word model (<xref ref-type="bibr" rid="B20">20</xref>), as shown in <xref ref-type="fig" rid="F8">Figure A3</xref>. We elaborate the details as follows.</p>
<fig id="F8" position="float"><label>Figure A3</label>
<caption><p>Framework of diagnostic feature extraction.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g008.tif"/>
</fig>
<p>When sending a diagnostic text into BERT, it will encode the text into an input vector, the length of which is always 512. For an input vector, it is composed of three embedding features: (1) WordPiece (<xref ref-type="bibr" rid="B30">30</xref>), (2) position embedding, and (3) segment embedding. Furthermore, as shown in <xref ref-type="fig" rid="F8">Figure A3</xref>, its framework consists of the multi-layer transformers proposed by Vaswani et al. (<xref ref-type="bibr" rid="B31">31</xref>). Transformers realize a series of encoding and decoding to transform an input text into a possible predicted result. Finally, the diagnostic text is converted to tokens by BERT, and a word vector at the corresponding position of each token is printed. We take out the results of the penultimate hidden layer and use the results of all vector mean pooling as diagnostic features.</p>
<p><italic>Drug interaction feature extraction.</italic> In this step, our objective is to establish drug interaction relationships from the medication knowledge graph as another feature. It can be found that the main form of drug interaction relation stored in the medication knowledge graph is pair, and it is more effective to use the graph structure to represent the drug interaction relation. Therefore, we used such a pair relationship to generate the drug interaction matrix <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM41"><mml:mi>A</mml:mi></mml:math></inline-formula> to represent the drug interaction graph. However, the interaction diagram exists as a two-dimensional matrix, whereas physiological and diagnostic features exist as one-dimensional vectors. Therefore, we need to transform the characteristics of drug interaction so that it can be spliced with the other two characteristics as the input of a reasonable drug recommendation model. We elaborate on the details as follows.</p>
<p>The size of the drug interaction matrix is <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM42"><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM43"><mml:mi>N</mml:mi></mml:math></inline-formula> is the size of the drug set. For the MIMIC&#x005F;III dataset, we first select two drugs in the drug set randomly, assuming that the coded values are <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM44"><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:math></inline-formula>. Then, we map its ATC4 code to the CID classification. Finally, we use the CID code as the keyword to search the drug interaction database for adverse drug interaction risks in these two categories. We iterate through all drug pairs in the database to generate an adjacency matrix of the final drug interaction diagram, which reflects the currently known drug combination contraindications and can reduce the number of treatment options that produce ADRs when a drug is recommended. For the medical record data of Fujian Province, we also randomly select two drugs in the drug set, and then search in the medication knowledge graph with the keyword of drug composition for whether these two drugs have an adverse drug interaction risk. If there is, the corresponding element of the marker matrix is 1, otherwise it is 0. The simplest way to convert a two-dimensional matrix to a one-dimensional vector is to compress the matrix. However, such compression destroys the connection between nodes, making it impossible for the model to learn the complete drug interaction relationship. Therefore, we employ the graph neural network method to embed the two-dimensional drug interaction characteristics into the one-dimensional space. The specific transformation process is described as follows.</p>
<p>Like the convolutional neural network in image vision, the GCN (<xref ref-type="bibr" rid="B21">21</xref>) is used for feature extraction. Therefore, we use the basic graph convolution network to construct a simple graph neural network, which maps the graph node representation to the low-dimensional vector space while preserving the topology and node information of the graph. A graph neural network with two GCN layers is established in this study, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM45"><mml:mi>A</mml:mi></mml:math></inline-formula> is the graph structure and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM46"><mml:mi>X</mml:mi></mml:math></inline-formula> is the matrix representation of the graph. The GCN layer compresses the hidden representation of each node by aggregating the feature information from the node neighborhood, and after the feature aggregation, non-linear permutation, such as ReLU, is applied to the generated output. Through the stacking of multiple layers of the GCN, the final hidden representation of each node in the diagram obtains information from subsequent neighborhoods. Finally, we connect it to a fully connected network to obtain a one-dimensional vector output. Through the transformation of the graph neural network, we obtain the one-dimensional representation of the characteristics of drug interactions.</p>
</sec>
<sec id="s11"><title>3 Sequence generation model</title>
<p>First, the symbols are defined, with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM47"><mml:mi>X</mml:mi></mml:math></inline-formula> representing the diagnostic space and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM48"><mml:mi>Y</mml:mi></mml:math></inline-formula> representing the drug treatment space. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM49"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> is a set of prescription records, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM50"><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2286;</mml:mo><mml:mi>X</mml:mi></mml:math></inline-formula> is a diagnostic sequence, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM51"><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mn>2</mml:mn><mml:mi>k</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM52"><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2286;</mml:mo><mml:mi>Y</mml:mi></mml:math></inline-formula> is a sequence of medication regimen, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM53"><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mn>2</mml:mn><mml:mi>k</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM54"><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM55"><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:math></inline-formula> are the sequence lengths of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM56"><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM57"><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>, respectively. There is no explicit mapping of the corresponding elements between diagnostic sequence <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM58"><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and drug sequence <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM59"><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. To avoid confusion, if there is no ambiguity, we omit <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM60"><mml:mi>k</mml:mi></mml:math></inline-formula> in the symbol. The purpose of drug prediction is to select the best sequence of medication regimen <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM61"><mml:msub><mml:mi>Y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> among all drugs <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM62"><mml:mi>Y</mml:mi></mml:math></inline-formula> based on the diagnostic sequence <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM63"><mml:mi>x</mml:mi></mml:math></inline-formula>. Therefore, the model in this study should have the ability to learn to map any diagnostic sequence to a corresponding medication regimen sequence, which requires the model in this study to learn not only the relationship between drugs and diagnosis but also the relationship between drugs and drugs. Therefore, we used the popular transformer approach in NLP to generate drug sequences. We will take a brief look at one of the key mechanisms in the transformer model and the overall architecture.</p>
<p><italic>Attention mechanism.</italic> The transformer is based on the attention mechanism; therefore, before introducing the overall framework of the transformer model, we first introduce the role of the attention mechanism. The sequence generation model transformer is composed of an encoder module and a decoding module. As shown in <xref ref-type="fig" rid="F9">Figure A4(a)</xref>, the encoder compresses the information expressed by the input sequence into a fixed-length semantic vector, and then the decoder decodes the information based on this semantic vector and generates the target sequence one by one. It means that elements at any point in the input sequence are equally important to the current target element. This method of learning input sequence cannot express the position information of the sequence, and if the sequence is too long, the decoding effect of the fixed semantic vector will be poor, because it is easy to lose the information contained in the sequence. To solve the above problems, Luong et al. (<xref ref-type="bibr" rid="B32">32</xref>) proposed the attention mechanism in 2015. As shown in <xref ref-type="fig" rid="F9">Figure A4B</xref>, the mechanism generates an independent semantic vector for each element in the output sequence, which could express the different importance of each input sequence to the decoded target element. The semantic vector <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM64"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the weighted sum of the elements in the input sequence, as described in <xref ref-type="disp-formula" rid="disp-formula11">Equations (A4)</xref> and <xref ref-type="disp-formula" rid="disp-formula12">(A5)</xref>:<disp-formula id="disp-formula11"><label>(A4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM11"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></disp-formula> <disp-formula id="disp-formula12"><label>(A5)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM12"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula> <disp-formula id="disp-formula13"><label>(A6)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM13"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM65"><mml:msub><mml:mi>L</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> is the length of the input sequence, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM66"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is the distribution of attention, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM67"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the importance of the element in the input sequence to the element that determines the output sequence, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM68"><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> represents the implicit state of the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM69"><mml:mi>j</mml:mi></mml:math></inline-formula>th element in the input sequence. In fact, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM70"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is a similarity measure, which is calculated according to the correlation between the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM71"><mml:mi>j</mml:mi></mml:math></inline-formula>th element in the input sequence and the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM72"><mml:mi>i</mml:mi></mml:math></inline-formula>th element in the output sequence. As shown in <xref ref-type="disp-formula" rid="disp-formula13">A6</xref>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM73"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is obtained from the output <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM74"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> of the hidden layer at the time of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM75"><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> in the decoder and the correlation degree of the hidden state <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM76"><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> corresponding to each element in the encoder. There are many methods to calculate the similarity between <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM77"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM78"><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>. In this study, we adopt the dot product method as shown in <xref ref-type="disp-formula" rid="disp-formula14">A7</xref> to calculate the similarity.</p>
<fig id="F9" position="float"><label>Figure A4</label>
<caption><p>The framework of diagnostic feature extraction.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g009.tif"/>
</fig><disp-formula id="disp-formula14"><label>(A7)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM14"><mml:mi>a</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover><mml:mi>h</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mi>s</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mover><mml:mi>h</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mi>s</mml:mi></mml:msub></mml:math></disp-formula>
<p>The above attention mechanism focuses on the relationship between the input sequence and output sequence, which can help us obtain more information when we model the relationship between drug and prescription. However, the relationship between elements in the input and output sequences is not taken into account. To solve this problem, another new attention mechanism, self-attention (<xref ref-type="bibr" rid="B33">33</xref>), has been proposed and widely used. Compared with the use of various cyclic neural networks that require longer information accumulation, the model with the introduction of the self-attention mechanism cannot only obtain the dependency relationship between two elements that are far apart, but also obtain the dependency relationship between the internal elements of the sequence more easily. In addition, the self-attention mechanism also improves the parallel computing capability of the model, greatly reducing the training time of the model. The transformer model used in this study is also based on the self-attention mechanism.</p>
<p>In the self-attention mechanism, the input sequence will be represented in the form of key value pairs, and then the input sequence with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM79"><mml:mi>N</mml:mi></mml:math></inline-formula> elements will be represented as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM80"><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where the key value is used to compute the attention distribution <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM81"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and the value is used to compute the semantic vector. The output sequence will be represented as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM82"><mml:mi>N</mml:mi></mml:math></inline-formula> queries; therefore, the semantic vector computation problem can be considered as an addressing operation. We use query to find the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM83"><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>q</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>y</mml:mi></mml:math></inline-formula> element in the input sequence, and the value obtained is the semantic vector or called attention. In particular, the self-attention mechanism can be thought of as soft addressing. Instead of looking only for elements with key values that are equal to the query value, a weighted sum is applied to all values to calculate the final attention. The weight of each value is determined by calculating the similarity between the query and each key. Therefore, the formula for calculating attention is shown in <xref ref-type="disp-formula" rid="disp-formula15">Equation (A8)</xref>:<disp-formula id="disp-formula15"><label>(A8)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM15"><mml:mtext>Attention</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mtext>softmax</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>Q</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>V</mml:mi></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM84"><mml:mi>Q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> , <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM85"><mml:mi>K</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM86"><mml:mi>V</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>; therefore, attention is a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM87"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> matrix. Different from the general attention mechanism, the self-attention mechanism also performs a division operation when calculating the attention distribution coefficient to avoid the similarity value calculated by the inner product method being too large, which results in 0 or 1 being generated when the softmax function is used for normalization, losing the meaning of soft addressing.</p>
<p><italic>Overall framework of transformer.</italic> The transformer model is a machine translation model proposed by Google in 2017 Vaswani et al. (<xref ref-type="bibr" rid="B31">31</xref>). As shown in <xref ref-type="fig" rid="F10">Figure A5</xref>, the transformer model is mainly composed of an encoder module and a decoder module. Each module is composed of several encoder and decoder layers, and the number of stacked layers can be adjusted according to the difficulty of tasks. This model abandons the traditional RNN or CNN architecture and adopts the structure of a full self-attention stack, which achieves excellent performance in NLP tasks. In this study, we build a two-layer transformer model to carry out the task of rational drug recommendation.</p>
<fig id="F10" position="float"><label>Figure A5</label>
<caption><p>The transformer model framework.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-06-1198904-g010.tif"/>
</fig>
<p>There are three kinds of attention in transformer, namely, self-attention in the encoder, self-attention in the decoder, and attention between the encoder and decoder. To capture all the spatial information in the input sequence, the attention calculation method in transformer is improved on the basis of the previous introduction and the concept of multi-head is introduced. It is mainly used to project query, key, and value to different spaces for H times through linear transformation, and then h self-attention matrix is obtained through calculation. As the feed-forward layer can only receive one matrix, we finally splice the h matrices and multiply by a weight matrix to generate the final attention matrix. When self-attention between the input and output sequences is calculated, set <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM88"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mi>V</mml:mi></mml:math></inline-formula>, and the specific calculation process is shown in <xref ref-type="disp-formula" rid="disp-formula15">Equation (A8)</xref> and <xref ref-type="disp-formula" rid="disp-formula16">(A9)</xref>:<disp-formula id="disp-formula16"><label>(A9)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM16"><mml:mtext>MultiHead</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mtext>Concat</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mtext>head</mml:mtext><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mtext>head</mml:mtext><mml:mi>h</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msup></mml:math></disp-formula><disp-formula id="disp-formula17"><label>(A10)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM17"><mml:msub><mml:mtext>head</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mtext>Attention</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mi>i</mml:mi><mml:mi>Q</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mi>i</mml:mi><mml:mi>K</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mi>i</mml:mi><mml:mi>V</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>As we can see from <xref ref-type="fig" rid="F10">Figure A5</xref>, the decoder structure is similar to that of the encoder, with the difference that the decoder has two attention layers. The decoder&#x2019;s goal in decoding is to give a probability distribution for the first output element. Therefore, we need to calculate the attention value of the last decoder layer in the decoder module. The input to this attention layer consists of two parts: the output of the last encoder module and the output of the first decoder module in the decoder module. Therefore, in calculating the attention of the decoder layer, the value <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM89"><mml:mi>K</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM90"><mml:mi>V</mml:mi></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="disp-formula16">A9</xref> comes from the encoder and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM91"><mml:mi>Q</mml:mi></mml:math></inline-formula> comes from the decoder. Decoder decoding is different from the encoder in that it can compute in parallel because it needs to use the output of the previous decoder layer as a query; therefore, the decoder is used one by one to generate the elements of the output sequence. In the model training process, the decoder layer uses real values; therefore, the mask method should be used to calculate the self-attention between output sequences to ensure that the current model cannot obtain more information than the current location.</p>
<p>Based on the transformer model, as shown in <xref ref-type="fig" rid="F10">Figure A5</xref>, we join the features together as the input sequence. After a multistep process of encoding and decoding, we select the element with the highest probability in the probability distribution as the prediction result for the current position, until we either reach the end of the generated identifier or the maximum sequence length. The resulting sequence is used as the model&#x2019;s recommendation for the current patient. Next, we will verify the effectiveness of the transformer model in the prediction of rational drug use through several comparative experiments.</p>
</sec></app>
</app-group>
</back>
</article>