<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Antibiot.</journal-id>
<journal-title>Frontiers in Antibiotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Antibiot.</abbrev-journal-title>
<issn pub-type="epub">2813-2467</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frabi.2023.1126468</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Antibiotics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A comparison of various feature extraction and machine learning methods for antimicrobial resistance prediction in <italic>streptococcus pneumoniae</italic>
</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Kaya</surname>
<given-names>Deniz Ece</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1994948"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>&#xdc;lgen</surname>
<given-names>Ege</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/615604"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Kocag&#xf6;z</surname>
<given-names>Ay&#x15f;e Sesin</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sezerman</surname>
<given-names>Osman U&#x11f;ur</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/28066"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Biostatistics and Medical Informatics, School of Medicine, Acibadem Mehmet Ali Aydinlar University</institution>, <addr-line>Istanbul</addr-line>, <country>T&#xfc;rkiye</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Infectious Diseases, School of Medicine, Acibadem Mehmet Ali Aydinlar University</institution>, <addr-line>Istanbul</addr-line>, <country>T&#xfc;rkiye</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Akhilesh K. Chaurasia, Sungkyunkwan University, Republic of Korea</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Amen Shamim, Sungkyunkwan University, Republic of Korea; Anubrata Das, University of Birmingham, United Kingdom</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Deniz Ece Kaya, <email xlink:href="mailto:denizecek@gmail.com">denizecek@gmail.com</email>
</p>
</fn>
<fn fn-type="other" id="fn002">
<p>This article was submitted to Antibiotic Resistance, a section of the journal Frontiers in Antibiotics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>24</day>
<month>03</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>2</volume>
<elocation-id>1126468</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>03</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Kaya, &#xdc;lgen, Kocag&#xf6;z and Sezerman</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Kaya, &#xdc;lgen, Kocag&#xf6;z and Sezerman</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Streptococcus pneumoniae is one of the major concerns of clinicians and one of the global public health problems. This pathogen is associated with high morbidity and mortality rates and antimicrobial resistance (AMR). In the last few years, reduced genome sequencing costs have made it possible to explore more of the drug resistance of S. pneumoniae, and machine learning (ML) has become a popular tool for understanding, diagnosing, treating, and predicting these phenotypes. Nucleotide k-mers, amino acid k-mers, single nucleotide polymorphisms (SNPs), and combinations of these features have rich genetic information in whole-genome sequencing. This study compares different ML models for predicting AMR phenotype for S. pneumoniae. We compared nucleotide k-mers, amino acid k-mers, SNPs, and their combinations to predict AMR in S. pneumoniae for three antibiotics: Penicillin, Erythromycin, and Tetracycline. 980 pneumococcal strains were downloaded from the European Nucleotide Archive (ENA). Furthermore, we used and compared several machine learning methods to train the models, including random forests, support vector machines, stochastic gradient boosting, and extreme gradient boosting. In this study, we found that key features of the AMR prediction model setup and the choice of machine learning method affected the results. The approach can be applied here to further studies to improve AMR prediction accuracy and efficiency.</p>
</abstract>
<kwd-group>
<kwd>AMR</kwd>
<kwd>machine learning</kwd>
<kwd>streptococcus pneumonaie</kwd>
<kwd>SNP</kwd>
<kwd>kmer</kwd>
<kwd>whole genome sequencing (WGS)</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="5"/>
<equation-count count="0"/>
<ref-count count="58"/>
<page-count count="12"/>
<word-count count="6496"/>
</counts>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Antimicrobial resistance (AMR) has caused a significant increase in morbidity and mortality rate in infectious diseases all over the world. According to World Health Organization (WHO), AMR is one of the top 10 global public health threats humanity faces. The global death rate from infectious diseases is projected to rise to 10 million per year by 2050 (<xref ref-type="bibr" rid="B52">World Health Organization, 2019</xref>; <xref ref-type="bibr" rid="B53">World Health Organization, 2022a</xref>). Many commonly used antibiotics have become ineffective due to rapidly increasing antimicrobial resistance in pathogens (<xref ref-type="bibr" rid="B4">Blair et&#xa0;al., 2015</xref>). In recent years, the development of new antimicrobial compounds has not been as rapid as the spread of resistance (<xref ref-type="bibr" rid="B23">Henriques-Normark and Tuomanen, 2013</xref>; <xref ref-type="bibr" rid="B36">Michael et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B7">Christaki et&#xa0;al., 2019</xref>), and this raises global health concerns. Rapid antibiotic susceptibility tests (AST) can guide the use of antibiotics and reduce drug-resistant strains. Currently, classical phenotypic AST methods, based on culturing target pathogens, are the gold standard. However, these methods take a few days to result and delay urgent treatment decisions. This delay also contributes to the spread of drug resistance (<xref ref-type="bibr" rid="B1">AMR Review, 2015</xref>). Molecular approaches have significantly improved over the years and play a critical role in the fight against antimicrobial resistance (<xref ref-type="bibr" rid="B25">Inouye et&#xa0;al., 2014</xref>). Due to the rapid development of sequencing technology and the decreasing cost, whole genome sequencing (WGS) or direct metagenomic sequencing of clinical materials has been proposed as the next-generation genotypic AST (<xref ref-type="bibr" rid="B18">Dunne et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B58">Zhang et&#xa0;al., 2019</xref>). In the face of growing AMR threats, it is increasingly vital to develop methods for interpreting minimum inhibitor concentrations (MICs) tests (<xref ref-type="bibr" rid="B37">Michael et&#xa0;al., 2020</xref>). Epidemiological cutoff values are set by the European Committee on Antimicrobial Susceptibility Testing (EUCAST) (<xref ref-type="bibr" rid="B20">ESCMID - European Society of Clinical Microbiology and Infectious Diseases, 2008</xref>) and by the Clinical and Laboratory Standards Institute (CLSI) for its epidemiological cutoff values (<xref ref-type="bibr" rid="B8">CLSI guidelines, 2022</xref>). Clinical breakpoints are another popular method of categorization. As a result of this process, MIC values are categorized according to different clinical outcomes (<xref ref-type="bibr" rid="B37">Michael et&#xa0;al., 2020</xref>). According to CLSI, these classes are &#x201c;resistant&#x201d; (R), &#x201c;susceptible&#x201d; (S), and &#x201c;intermediate&#x201d; (I) (<xref ref-type="bibr" rid="B8">CLSI guidelines, 2022</xref>).</p>
<p>As discussed above, due to reduced genome sequencing costs, detecting AMR phenotypes directly from sequence data has become a preferred method. In the last few years, the use of machine learning (ML) for understanding, diagnosing, treating, and predicting AMR phenotypes has aroused interest in the literature, and it has been shown in publications (<xref ref-type="bibr" rid="B55">Yang et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B41">Nguyen et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B15">Deelder et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B42">Nguyen et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B27">Khaledi et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B50">Wang et&#xa0;al., 2022</xref>) that for many bacterial species. Antimicrobial resistance can be predicted quite accurately based on the genome sequence. ML techniques applied to WGS can accurately predict MIC results. However, some MIC data are only shared as classes, while the remaining are shared as concentration, which may cause discrepancies while training ML models.</p>
<p>AMR has been extensively studied <italic>via</italic> ML in various microorganisms, including <italic>Mycobacterium tuberculosis</italic> (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B17">Drouin et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B55">Yang et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B15">Deelder et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B2">Aytan-Aktug et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B50">Wang et&#xa0;al., 2022</xref>), <italic>Escherichia col</italic>i (<xref ref-type="bibr" rid="B39">Moradigaravand et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B43">Pataki et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B2">Aytan-Aktug et&#xa0;al., 2020</xref>), <italic>Salmonella enterica</italic> (<xref ref-type="bibr" rid="B2">Aytan-Aktug et&#xa0;al., 2020</xref>), nontyphoidal Salmonella (<xref ref-type="bibr" rid="B42">Nguyen et&#xa0;al., 2019</xref>), <italic>Staphylococcus aureus</italic> (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B2">Aytan-Aktug et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al., 2020</xref>), <italic>Acinetobacter baumannii</italic> (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>), <italic>Streptococcus pneumoniae</italic> (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B17">Drouin et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B34">Li et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B33">Li et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B58">Zhang et&#xa0;al., 2019</xref>), <italic>Clostridium difficile</italic> (<xref ref-type="bibr" rid="B17">Drouin et&#xa0;al., 2016</xref>), <italic>Pseudomonas aeruginosa</italic> (<xref ref-type="bibr" rid="B17">Drouin et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B27">Khaledi et&#xa0;al., 2020</xref>), <italic>Actinobacillus pleuropneumoniae</italic> (<xref ref-type="bibr" rid="B35">Liu et&#xa0;al., 2020</xref>), <italic>Elizabethkingia</italic> (<xref ref-type="bibr" rid="B40">Naidenov et&#xa0;al., 2019</xref>), <italic>Klebsiella pneumoniae</italic> (<xref ref-type="bibr" rid="B41">Nguyen et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al., 2020</xref>), <italic>Campylobacter jejuni</italic> (<xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al., 2020</xref>), and <italic>Neisseria gonorrhoeae</italic> (<xref ref-type="bibr" rid="B21">Eyre et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al., 2020</xref>). While setting up ML models, k-mer counts of various lengths (8-mers to 11-mers (<xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al., 2020</xref>), 10-mers (<xref ref-type="bibr" rid="B41">Nguyen et&#xa0;al., 2018</xref>), 15-mers (<xref ref-type="bibr" rid="B42">Nguyen et&#xa0;al., 2019</xref>), 31-mers (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B17">Drouin et&#xa0;al., 2016</xref>)), AMR genes (<xref ref-type="bibr" rid="B24">Her and Wu, 2018</xref>), SNPs (<xref ref-type="bibr" rid="B55">Yang et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B15">Deelder et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B47">Shi et&#xa0;al., 2019</xref>), or a combination of these (<xref ref-type="bibr" rid="B39">Moradigaravand et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B40">Naidenov et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B27">Khaledi et&#xa0;al., 2020</xref>) have been successfully used as features. In a recently published study by <xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al. (2020)</xref>, amino acid k-mers were also utilized as features and yielded successful results.</p>
<p>
<italic>Streptococcus pneumoniae</italic> is known to be one of the bacteria with the most common AMR problem (<xref ref-type="bibr" rid="B49">van der Poll et&#xa0;al., 2009</xref>). S. <italic>pneumoniae</italic> is a gram-positive human pathogen that is the primary cause of respiratory tract infection and diseases such as pneumonia and meningitis. This bacterium is also found in the nasopharyngeal flora in childhood and often causes invasive infectious diseases such as acute otitis media and sinusitis (<xref ref-type="bibr" rid="B23">Henriques-Normark and Tuomanen, 2013</xref>). According to the WHO, diseases caused by Streptococcus pneumoniae are an important public health problem worldwide. It is estimated that about one million children die yearly from pneumococcal disease (<xref ref-type="bibr" rid="B54">World Health Organization, 2022b</xref>-2).</p>
<p>Many pneumococcal isolates are resistant to common antibacterial drugs like fluoroquinolones, macrolides, and &#x3b2;- lactams (<xref ref-type="bibr" rid="B46">Sader et&#xa0;al., 2019</xref>). The main targets of penicillin are penicillin-binding proteins (PBPs). For many years, penicillin has been the primary choice for treating S.pneumoniae-associated infections (<xref ref-type="bibr" rid="B57">Zapun et&#xa0;al., 2008</xref>). &#x3b2;-lactams bind to enzymes essential for bacterial cell wall synthesis and reducing peptidoglycan synthesis. (<xref ref-type="bibr" rid="B57">Zapun et&#xa0;al., 2008</xref>). The main resistance mechanism to resist &#x3b2;-lactams is mutating PBPs to reduce their affinity to antibiotics (<xref ref-type="bibr" rid="B44">Poole, 2004</xref>). During the same time as penicillin resistance spread, macrolide-resistant pneumococci also increased. Moreover, the removal of the antimicrobial from the cell by the acquisition of <italic>mef</italic> and <italic>erm</italic> genes and modification of the target site are the two main mechanisms of macrolide-like erythromycin resistance in S. pneumoniae (<xref ref-type="bibr" rid="B9">Cornick and Bentley, 2012</xref>). Tetracyclines inhibit the growth of bacteria by binding to the 30S subunit of the bacterial ribosome. Pneumococcal resistance to tetracycline occurs <italic>via</italic> ribosomal protection (tet(O) and tet(M) genes) (<xref ref-type="bibr" rid="B38">Montanari et&#xa0;al., 2003</xref>).</p>
<p>This study compares different ML models for predicting AMR phenotype for S. pneumoniae. We compared nucleotide k-mers, amino acid k-mers, and SNPs to predict AMR in S. pneumoniae for three antibiotics: Penicillin, Erythromycin, and Tetracycline. Further, we attempted to use and compare various ML methods: random forest (RF), support vector machine (SVM), stochastic gradient boosting (GBM), and extreme gradient boosting (XGBoost) to train the models. We discuss the strengths and limitations of feature and ML model selection for MIC prediction. We observed and concluded that the choice of features and the selection of the ML model affects the performance of prediction (as measured by F1 score and accuracy) differently for each antibiotic.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Overview</title>
<p>The overview of the study is presented in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. Following feature generation, feature selection, and various ML models were trained to predict the MIC class for each antibiotic. The performances were evaluated using several metrics, and the results were compared. The details of data, feature generation, feature selection, model training, and evaluation are described in the following subsections. All analyses were performed in R version 4.0.2 (<ext-link ext-link-type="uri" xlink:href="http://www.R-project.org">http://www.R-project.org</ext-link>). The R scripts utilized in this study are available on GitHub at <ext-link ext-link-type="uri" xlink:href="https://github.com/denizecek/AMRprediction">https://github.com/denizecek/AMRprediction</ext-link>.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Overall pipeline for all feature extraction and classification approaches.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frabi-02-1126468-g001.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Datasets and pre-processing</title>
<p>S. pneumoniae metagenomic sequences and the related MIC class information of the antibiotics penicillin, erythromycin, and tetracycline were included in our study. We used four publicly available datasets: 980 pneumococcal strains were downloaded from the European Nucleotide Archive (ENA) (<ext-link ext-link-type="uri" xlink:href="http://www.ebi.ac.uk/ena/">http://www.ebi.ac.uk/ena/</ext-link>) with the project accession code in PRJEB2632 (<xref ref-type="bibr" rid="B11">Croucher et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B10">Croucher et&#xa0;al., 2015</xref>), PRJNA34791 (<xref ref-type="bibr" rid="B16">Demczuk et&#xa0;al., 2017</xref>), PRJEB3084 (<xref ref-type="bibr" rid="B22">Gladstone et&#xa0;al., 2019</xref>), PRJEB2255 (<xref ref-type="bibr" rid="B12">Croucher et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B10">Croucher et&#xa0;al., 2015</xref>). The MIC class information was downloaded from the PATRIC database (<xref ref-type="bibr" rid="B14">Davis et&#xa0;al., 2020</xref>) and PubMLST (<xref ref-type="bibr" rid="B26">Jolley et&#xa0;al., 2018</xref>), and the &#x201c;Resistant&#x201d; and &#x201c;Susceptible&#x201d; classes were matched to the genome data. For each antibiotic, we discarded the &#x201c;Intermediate&#x201d; class because these were underrepresented. <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref> presents the sample number of the four datasets, and <xref ref-type="supplementary-material" rid="SM1">
<bold>Table S1</bold>
</xref> (<xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary 1</bold>
</xref>) contains detailed sample information.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Datasets and the corresponding numbers of samples per MIC class for the three antibiotics.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" rowspan="2" align="left">Antibiotic</th>
<th valign="top" colspan="2" align="left">DS 1 - PRJEB2632</th>
<th valign="top" colspan="2" align="left">DS 2 &#x2013; PRJEB3084</th>
<th valign="top" colspan="2" align="left">DS 3 &#x2013; PRJNA347910</th>
<th valign="top" colspan="2" align="left">DS 4 - PRJEB2255</th>
<th valign="top" rowspan="2" align="left">Total</th>
</tr>
<tr>
<th valign="top" align="left">S</th>
<th valign="top" align="left">R</th>
<th valign="top" align="left">S</th>
<th valign="top" align="left">R</th>
<th valign="top" align="left">S</th>
<th valign="top" align="left">R</th>
<th valign="top" align="left">S</th>
<th valign="top" align="left">R</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Penicillin</td>
<td valign="top" align="left">453</td>
<td valign="top" align="left">68</td>
<td valign="top" align="left">11</td>
<td valign="top" align="left">49</td>
<td valign="top" align="left">2</td>
<td valign="top" align="left">109</td>
<td valign="top" align="left">128</td>
<td valign="top" align="left">49</td>
<td valign="top" align="left">869</td>
</tr>
<tr>
<td valign="top" align="left">Erythromycin</td>
<td valign="top" align="left">470</td>
<td valign="top" align="left">78</td>
<td valign="top" align="left">18</td>
<td valign="top" align="left">41</td>
<td valign="top" align="left">25</td>
<td valign="top" align="left">132</td>
<td valign="top" align="left">17</td>
<td valign="top" align="left">114</td>
<td valign="top" align="left">895</td>
</tr>
<tr>
<td valign="top" align="left">Tetracycline</td>
<td valign="top" align="left">287</td>
<td valign="top" align="left">36</td>
<td valign="top" align="left">35</td>
<td valign="top" align="left">19</td>
<td valign="top" align="left">3</td>
<td valign="top" align="left">131</td>
<td valign="top" align="left">11</td>
<td valign="top" align="left">121</td>
<td valign="top" align="left">643</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Feature generation and selection</title>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>Nucleotide K-mers</title>
<p>SPAdes (<xref ref-type="bibr" rid="B3">Bankevich et&#xa0;al., 2012</xref>) in the PATRIC assembly service (<xref ref-type="bibr" rid="B14">Davis et&#xa0;al., 2020</xref>) was used for genome assembly. Contigs with less than 5-fold coverage and lengths less than 500 bp were removed. The contigs were divided into 10-mers, and the frequencies of these 10-mers were obtained using the R &#x201c;kmer&#x201d; package (<xref ref-type="bibr" rid="B51">Wilkinson, 2018</xref>). For the AMR classification task, the k-mer counts were used as one set of features, and antibiotic MIC classes were used as labels.</p>
<p>In this work, we chose to use a 10-mers instead of a longer k-mer length to reduce the size of the resulting k-mer matrix. Longer k-mers were not selected because of memory limitations, and we did not utilize shorter k-mers due to lower initial accuracy. Next, k-mer counts were converted to depict the presence &#x201c;1&#x201d; or absence &#x201c;0&#x201d; of each k-mer in each genome.</p>
<p>The dataset was very large, and fitting ML models using this data might have caused significant challenges, including high computational cost and processing time. Moreover, it is known that ML models trained on large datasets (i.e., large sets of features) tend to have poorer performance compared to using an optimal set of features (<xref ref-type="bibr" rid="B56">Yu and Liu, 2004</xref>; <xref ref-type="bibr" rid="B45">Pudjihartono et&#xa0;al., 2022</xref>). Since the total number of 10-kmers is 1,048,578 in our dataset, the absolute mean difference between resistant and susceptible samples of each feature was used as the first feature selection step. For every feature, we calculated the mean of the resistance and susceptible samples, features with an absolute mean difference of at least 0.3 were selected for penicillin and erythromycin and 0.4 were selected for tetracycline. The main reason we choose 0.3 as the threshold is to reduce the number of features below 10000. When this cutoff value was set to 0.2, 11149 features remained for penicillin, and when it was set to 0.3, 2099 features remained. When 0.4 was selected, we had 141 features remaining. Final number of features before the next feature selection step for penicillin, erythromycin and tetracycline are 1591, 2099 and 3376.</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Single nucleotide polymorphisms</title>
<p>We used single nucleotide polymorphisms (SNPs) as another set of features for ML model training. Our reference genome for SNP calling was <italic>S. pneumoniae</italic> TIGR4. For variant calling, BWA-mem (<xref ref-type="bibr" rid="B31">Li, 2013</xref>) and SAMtools (<xref ref-type="bibr" rid="B32">Li et&#xa0;al., 2009</xref>) were used <italic>via</italic> the PATRIC variant calling service (<xref ref-type="bibr" rid="B14">Davis et&#xa0;al., 2020</xref>). Bcftools (<xref ref-type="bibr" rid="B30">Li, 2011</xref>) was used for filtering variants with DP &gt; 20 and qual &gt; 50 parameters. A total of 221,304 SNPs were obtained. SNP positions (compared to the reference genome) were the columns of the resulting matrix, and the samples were rows. A sample with an SNP at a given site was shown as 1, and those without any SNPs were shown as 0. Compared to the 10-mer features, the number of SNP features was much lower (221,304); hence the absolute mean difference cutoff value was also decreased. The absolute mean difference between resistant and susceptible samples was calculated for each SNP and filtered at least 0.2 for all three antibiotics. The features with an absolute mean difference lower than this cutoff were removed. With this first step of feature selection, for penicillin, 4,954 features remained; for erythromycin, 8,844 features remained and for tetracycline, 6,695 features remained.</p>
</sec>
<sec id="s2_3_3">
<label>2.3.3</label>
<title>Amino acid K-mers</title>
<p>An amino acid k-mer model for predicting MIC classes for the three antibiotics was built following the method previously described by <xref ref-type="bibr" rid="B48">ValizadehAslani et&#xa0;al. (2020)</xref>. To provide annotation of genomic features, the Genome Annotation Service in PATRIC (<xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>), which uses the RAST toolkit (RASTtk) (<xref ref-type="bibr" rid="B5">Brettin et&#xa0;al., 2015</xref>), was utilized. Protein FASTA sequences were downloaded from the PATRIC database. For counting the amino acid k-mers, we used the &#x201c;kmer&#x201d; R package (<xref ref-type="bibr" rid="B51">Wilkinson, 2018</xref>). The Dayhoff-6 alphabet (<xref ref-type="bibr" rid="B19">Edgar, 2004</xref>) was used to minimize computation time when counting longer k-mers. 5-mers of the amino acid were counted for the genome of each strain. We did not use shorter amino acid k-mers due to lower initial accuracy, and longer k-mers were not chosen because of memory limitations. Since the total number of features was less than 10,000, the pre-elimination used in other feature extraction methods (absolute mean difference between MIC classes) was not used here.</p>
</sec>
<sec id="s2_3_4">
<label>2.3.4</label>
<title>Combinations of features</title>
<p>10-mer nucleotides, 5-mer amino acid content, and SNP features were combined as binary combinations and tested as another feature extraction method. Sections 2.3.1, 2.3.2, and 2.3.3 were used for feature selection, and the optimal features were combined.</p>
</sec>
<sec id="s2_3_5">
<label>2.3.5</label>
<title>Boruta</title>
<p>Feature selection algorithm Boruta, implemented as an R package, was used for the second and final feature selection step. Boruta is an ML algorithm used in feature selection (<xref ref-type="bibr" rid="B28">Kursa and Rudnicki, 2010</xref>). It is a wrapper feature selection method built around the Random Forest classification algorithm. The algorithm adds randomness to the data set by creating a shuffled copy of all features. These features are called &#x201c;Shadow Features&#x201d;. The shadow features and original features are then merged, and the algorithm builds a random forest classifier, which determines each feature&#x2019;s importance using Z-score and mean decreased accuracy. Boruta then checks whether an original feature has higher importance than the shadow features. At each iteration, significant features are kept, and unimportant ones are constantly removed. This iteration repeats until all features are confirmed or rejected (<xref ref-type="bibr" rid="B28">Kursa and Rudnicki, 2010</xref>). In this study, the important features (<xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>) were chosen using the default settings of Boruta.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Boruta Results and number of Important (Imp), Unimportant (Unimp), and Tentative (Ten) features.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" rowspan="2" align="left">Antibiotic</th>
<th valign="top" colspan="3" align="left">Nucleotide 10-mer Boruta Results</th>
<th valign="top" colspan="3" align="left">Amino Acid 5-mer Boruta Results</th>
<th valign="top" colspan="3" align="left">SNP Boruta Results</th>
</tr>
<tr>
<th valign="top" align="left">Imp</th>
<th valign="top" align="left">Unimp</th>
<th valign="top" align="left">Ten</th>
<th valign="top" align="left">Imp</th>
<th valign="top" align="left">Unimp</th>
<th valign="top" align="left">Ten</th>
<th valign="top" align="left">Imp</th>
<th valign="top" align="left">Unimp</th>
<th valign="top" align="left">Ten</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Penicillin</td>
<td valign="top" align="left">138</td>
<td valign="top" align="left">4690</td>
<td valign="top" align="left">126</td>
<td valign="top" align="left">162</td>
<td valign="top" align="left">8394</td>
<td valign="top" align="left">288</td>
<td valign="top" align="left">25</td>
<td valign="top" align="left">6628</td>
<td valign="top" align="left">42</td>
</tr>
<tr>
<td valign="top" align="left">Erythromycin</td>
<td valign="top" align="left">153</td>
<td valign="top" align="left">4633</td>
<td valign="top" align="left">168</td>
<td valign="top" align="left">45</td>
<td valign="top" align="left">8660</td>
<td valign="top" align="left">139</td>
<td valign="top" align="left">124</td>
<td valign="top" align="left">6394</td>
<td valign="top" align="left">177</td>
</tr>
<tr>
<td valign="top" align="left">Tetracycline</td>
<td valign="top" align="left">108</td>
<td valign="top" align="left">4642</td>
<td valign="top" align="left">204</td>
<td valign="top" align="left">18</td>
<td valign="top" align="left">8736</td>
<td valign="top" align="left">90</td>
<td valign="top" align="left">106</td>
<td valign="top" align="left">6432</td>
<td valign="top" align="left">157</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_3_6">
<label>2.3.6</label>
<title>Statistical Analysis</title>
<p>To understand the importance of nucleotide 10-mers, top ten features were determined for each model. The union of these most successful features consisting of nucleotide 10-mers has been prepared. Hypergeometric distribution test was used to determine whether these kmers were over-represented in AMR genes. For penicillin samples, pbp2B, pbp2x, pbp1a genes were downloaded from NCBI, the contigs were divided into 10-mers, and hypergeometric distribution was calculated for each antibiotic.</p>
<p>Ermb, mefE, mefA were used for erythromysin nucleotide 10-mers and tetS and tetM were for tetracycline.</p>
</sec>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Machine learning models for AMR classification</title>
<p>As samples to generate the ML models, as discussed above, since the small number of intermediate samples created an imbalance, we only included isolates categorized as either &#x201c;resistant&#x201d; or &#x201c;susceptible&#x201d; for each antibiotic. As features, we separately utilized the presence/absence of SNPs, nucleotide 10-mers, and amino acid 5-mers. Pairs of combinations of these features were analyzed in the same manner. An experiment of tenfold cross validation was used to evaluate the model&#x2019;s stability and accuracy and was built according to the methodology previously described by <xref ref-type="bibr" rid="B14">Davis et&#xa0;al. (2020)</xref>. For each drug, we randomly assigned isolates to a training set comprising 80% of the resistant and susceptible isolates, respectively. The remaining 20% were divided equally into a test set and a validation set. Parameters of ML models were optimized on the validation set, and their accuracy was assessed in cross-validation, while the test set was used to obtain another independent performance estimate.</p>
<p>The accuracy and sensitivity of the ML models generated by this study were evaluated by 10-fold cross-validation. The data were divided into training and test sets as 8:2. The matrix is divided into ten equal parts by cross-validation, with an equal number of antibiotic-MIC combinations in each part. One part is used for testing, one for validation, and eight for training. Each model used the validation set to avoid overfitting. 10-fold cross-validation was performed in the hyperparameter tuning stage. Optimal combinations of hyper-parameters were selected for each fold based on the mean squared error of validation. Ten sets of hyper-parameters were generated from the tests, one for each fold. Different ML algorithms were compared based on accuracy, F1 score, and Cohen&#x2019;s (unweighted) Kappa statistic averaged across the resampling results.</p>
<p>To detect penicillin, erythromycin, and tetracycline resistance <italic>Streptococcus pneumonia</italic>, we trained random forest (RF), support vector machine (SVM), stochastic gradient boosting (GBM), and extreme gradient boosting (XGBoost) classifiers. Three different models were tested for RF. The model with the default for each parameter, random search, and grid search was performed. For the SVM classifier, SVM with linear kernel, polynomial and radial kernel functions were tested. For the GBM we tried the tuning parameters. For XGBoost models, we used a grid search to tune our important hyperparameters. The models with the highest F1 score among all created models were compared. The optimal parameters for each ML approach are presented in <xref ref-type="supplementary-material" rid="SM1">
<bold>Table S2</bold>
</xref>.</p>
<p>Notably, the relative contribution of the different information sources to the susceptibility and resistance sensitivity strongly depended on the antibiotic. To assess the effect of the classification technique, we compared the performance of different classifiers. The 980-genome model contained data from all antibiotics and MICs, making feature extraction challenging to determine which k-mers contribute to the MIC predictions for each antibiotic. To address this limitation, we modified the protocol by building separate models for each antibiotic. Another reason why we set up a separate model for each antibiotic is that not all samples have MIC information for all three antibiotics. As you can see in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref> we have 869 penicillin samples, however we have 643 tetracycline samples with MIC information. We did not want to reduce the number of samples to train our models with a single large integrated model. We also worried about the computational problems like memory, RAM and training times to performing best classifier for a single large integrated models for all antibiotics. </p>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results</title>
<p>As detailed in Methods, we trained several ML classification methods on features individually and in combination for predicting antibiotic susceptibility or resistance of isolates and evaluated the classifier performances. We calculated the accuracy, sensitivity, specificity, and the F1-score, as an overall performance measure based on a classifier trained on a specific combination and shown in <xref ref-type="table" rid="T3">
<bold>Tables&#xa0;3</bold>
</xref>&#x2013;<xref ref-type="table" rid="T5">
<bold>5</bold>
</xref>. Training and validation sets accuracy and kappa results are presented in <xref ref-type="supplementary-material" rid="SM1">
<bold>Tables S3</bold>
</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">
<bold>S5</bold>
</xref>.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Penicillin models performances.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="bottom" align="left">Algorithm</th>
<th valign="bottom" align="left">Input</th>
<th valign="bottom" align="left">F1-score</th>
<th valign="bottom" align="left">Kappa</th>
<th valign="bottom" align="left">Accuracy</th>
<th valign="bottom" align="left">Sensitivity</th>
<th valign="bottom" align="left">Specificity</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">k-mer</td>
<td valign="top" align="left">0.848</td>
<td valign="top" align="left">0.814</td>
<td valign="top" align="left">0.943</td>
<td valign="top" align="left">0.849</td>
<td valign="top" align="left">0.965</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">k-mer</td>
<td valign="top" align="left">0.865</td>
<td valign="top" align="left">0.834</td>
<td valign="top" align="left">0.950</td>
<td valign="top" align="left">0.852</td>
<td valign="top" align="left">0.972</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">k-mer</td>
<td valign="top" align="left">0.879</td>
<td valign="top" align="left">0.851</td>
<td valign="top" align="left">0.955</td>
<td valign="top" align="left">0.879</td>
<td valign="top" align="left">0.972</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">k-mer</td>
<td valign="top" align="left">0.866</td>
<td valign="top" align="left">0.843</td>
<td valign="top" align="left">0.950</td>
<td valign="top" align="left">0.853</td>
<td valign="top" align="left">0.972</td>
</tr>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">AA k-mer</td>
<td valign="top" align="left">0.900</td>
<td valign="top" align="left">0.871</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.882</td>
<td valign="top" align="left">0.980</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">AA k-mer</td>
<td valign="top" align="left">0.899</td>
<td valign="top" align="left">0.874</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.861</td>
<td valign="top" align="left">0.986</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">AA k-mer</td>
<td valign="top" align="left">0.882</td>
<td valign="top" align="left">0.854</td>
<td valign="top" align="left">0.955</td>
<td valign="top" align="left">0.857</td>
<td valign="top" align="left">0.979</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">AA k-mer</td>
<td valign="top" align="left">0.886</td>
<td valign="top" align="left">0.858</td>
<td valign="top" align="left">0.955</td>
<td valign="top" align="left">0.838</td>
<td valign="top" align="left">0.986</td>
</tr>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">SNP</td>
<td valign="top" align="left">0.786</td>
<td valign="top" align="left">0.747</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.957</td>
<td valign="top" align="left">0.929</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">SNP</td>
<td valign="top" align="left">0.772</td>
<td valign="top" align="left">0.730</td>
<td valign="top" align="left">0.927</td>
<td valign="top" align="left">0.917</td>
<td valign="top" align="left">0.928</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">SNP</td>
<td valign="top" align="left">0.759</td>
<td valign="top" align="left">0.712</td>
<td valign="top" align="left">0.921</td>
<td valign="top" align="left">0.880</td>
<td valign="top" align="left">0.928</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">SNP</td>
<td valign="top" align="left">0.786</td>
<td valign="top" align="left">0.747</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.957</td>
<td valign="top" align="left">0.929</td>
</tr>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">SNP/AA k-mer</td>
<td valign="bottom" align="left">0.889</td>
<td valign="bottom" align="left">0.865</td>
<td valign="bottom" align="left">0.961</td>
<td valign="bottom" align="left">0.933</td>
<td valign="bottom" align="left">0.966</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">SNP/AA k-mer</td>
<td valign="bottom" align="left">0.889</td>
<td valign="bottom" align="left">0.865</td>
<td valign="bottom" align="left">0.961</td>
<td valign="bottom" align="left">0.933</td>
<td valign="bottom" align="left">0.966</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">SNP/AA k-mer</td>
<td valign="bottom" align="left">0.857</td>
<td valign="bottom" align="left">0.826</td>
<td valign="bottom" align="left">0.950</td>
<td valign="bottom" align="left">0.900</td>
<td valign="bottom" align="left">0.959</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">SNP/AA k-mer</td>
<td valign="bottom" align="left">0.871</td>
<td valign="bottom" align="left">0.844</td>
<td valign="bottom" align="left">0.955</td>
<td valign="bottom" align="left">0.931</td>
<td valign="bottom" align="left">0.960</td>
</tr>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">SNP/k-mer</td>
<td valign="bottom" align="left">0.831</td>
<td valign="bottom" align="left">0.793</td>
<td valign="bottom" align="left">0.938</td>
<td valign="bottom" align="left">0.844</td>
<td valign="bottom" align="left">0.959</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">SNP/k-mer</td>
<td valign="bottom" align="left">0.820</td>
<td valign="bottom" align="left">0.782</td>
<td valign="bottom" align="left">0.938</td>
<td valign="bottom" align="left">0.893</td>
<td valign="bottom" align="left">0.946</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">SNP/k-mer</td>
<td valign="bottom" align="left">0.852</td>
<td valign="bottom" align="left">0.822</td>
<td valign="bottom" align="left">0.949</td>
<td valign="bottom" align="left">0.929</td>
<td valign="bottom" align="left">0.963</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">SNP/k-mer</td>
<td valign="bottom" align="left">0.867</td>
<td valign="bottom" align="left">0.840</td>
<td valign="bottom" align="left">0.955</td>
<td valign="bottom" align="left">0.953</td>
<td valign="bottom" align="left">0.953</td>
</tr>
<tr>
<td valign="bottom" align="left">Random Forest</td>
<td valign="bottom" align="left">AA k-mer/k-mer</td>
<td valign="bottom" align="left">0.867</td>
<td valign="bottom" align="left">0.840</td>
<td valign="bottom" align="left">0.955</td>
<td valign="bottom" align="left">0.963</td>
<td valign="bottom" align="left">0.953</td>
</tr>
<tr>
<td valign="bottom" align="left">Support Vector Machine</td>
<td valign="bottom" align="left">AA k-mer/k-mer</td>
<td valign="bottom" align="left">0.852</td>
<td valign="bottom" align="left">0.822</td>
<td valign="bottom" align="left">0.950</td>
<td valign="bottom" align="left">0.929</td>
<td valign="bottom" align="left">0.953</td>
</tr>
<tr>
<td valign="bottom" align="left">Stochastic Gradient Boosting</td>
<td valign="bottom" align="left">AA k-mer/k-mer</td>
<td valign="bottom" align="left">0.847</td>
<td valign="bottom" align="left">0.818</td>
<td valign="bottom" align="left">0.950</td>
<td valign="bottom" align="left">0.962</td>
<td valign="bottom" align="left">0.947</td>
</tr>
<tr>
<td valign="bottom" align="left">Extreme Gradient Boosting</td>
<td valign="bottom" align="left">AA k-mer/k-mer</td>
<td valign="bottom" align="left">0.847</td>
<td valign="bottom" align="left">0.818</td>
<td valign="bottom" align="left">0.950</td>
<td valign="bottom" align="left">0.962</td>
<td valign="bottom" align="left">0.947</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Erythromycin models performances.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Algorithm</th>
<th valign="top" align="left">input</th>
<th valign="top" align="left">F1-score</th>
<th valign="top" align="left">Kappa</th>
<th valign="top" align="left">Accuracy</th>
<th valign="top" align="left">Sensitivity</th>
<th valign="top" align="left">Specificity</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.977</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.984</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.940</td>
<td valign="top" align="left">0.916</td>
<td valign="top" align="left">0.965</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.968</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.951</td>
<td valign="top" align="left">0.931</td>
<td valign="top" align="left">0.971</td>
<td valign="top" align="left">0.942</td>
<td valign="top" align="left">0.984</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.971</td>
<td valign="top" align="left">0.959</td>
<td valign="top" align="left">0.983</td>
<td valign="top" align="left">0.962</td>
<td valign="top" align="left">0.991</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.971</td>
<td valign="top" align="left">0.926</td>
<td valign="top" align="left">0.992</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.971</td>
<td valign="top" align="left">0.926</td>
<td valign="top" align="left">0.992</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.977</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.983</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.977</td>
<td valign="top" align="left">0.961</td>
<td valign="top" align="left">0.983</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.813</td>
<td valign="top" align="left">0.748</td>
<td valign="top" align="left">0.902</td>
<td valign="top" align="left">0.925</td>
<td valign="top" align="left">0.895</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.821</td>
<td valign="top" align="left">0.754</td>
<td valign="top" align="left">0.902</td>
<td valign="top" align="left">0.886</td>
<td valign="top" align="left">0.907</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.816</td>
<td valign="top" align="left">0.744</td>
<td valign="top" align="left">0.896</td>
<td valign="top" align="left">0.851</td>
<td valign="top" align="left">0.913</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.792</td>
<td valign="top" align="left">0.712</td>
<td valign="top" align="left">0.884</td>
<td valign="top" align="left">0.844</td>
<td valign="top" align="left">0.898</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.876</td>
<td valign="top" align="left">0.822</td>
<td valign="top" align="left">0.925</td>
<td valign="top" align="left">0.851</td>
<td valign="top" align="left">0.958</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.884</td>
<td valign="top" align="left">0.835</td>
<td valign="top" align="left">0.930</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.958</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.862</td>
<td valign="top" align="left">0.805</td>
<td valign="top" align="left">0.919</td>
<td valign="top" align="left">0.862</td>
<td valign="top" align="left">0.942</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.873</td>
<td valign="top" align="left">0.820</td>
<td valign="top" align="left">0.924</td>
<td valign="top" align="left">0.865</td>
<td valign="top" align="left">0.950</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.919</td>
<td valign="top" align="left">0.965</td>
<td valign="top" align="left">0.894</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.927</td>
<td valign="top" align="left">0.893</td>
<td valign="top" align="left">0.953</td>
<td valign="top" align="left">0.864</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.914</td>
<td valign="top" align="left">0.877</td>
<td valign="top" align="left">0.948</td>
<td valign="top" align="left">0.889</td>
<td valign="top" align="left">0.974</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.884</td>
<td valign="top" align="left">0.835</td>
<td valign="top" align="left">0.930</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.958</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.903</td>
<td valign="top" align="left">0.959</td>
<td valign="top" align="left">0.923</td>
<td valign="top" align="left">0.975</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.942</td>
<td valign="top" align="left">0.917</td>
<td valign="top" align="left">0.965</td>
<td valign="top" align="left">0.924</td>
<td valign="top" align="left">0.983</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.932</td>
<td valign="top" align="left">0.903</td>
<td valign="top" align="left">0.959</td>
<td valign="top" align="left">0.923</td>
<td valign="top" align="left">0.975</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.900</td>
<td valign="top" align="left">0.859</td>
<td valign="top" align="left">0.942</td>
<td valign="top" align="left">0.918</td>
<td valign="top" align="left">0.951</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Tetracycline models performances.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Algorithm</th>
<th valign="top" align="left">input</th>
<th valign="top" align="left">F1-score</th>
<th valign="top" align="left">Kappa</th>
<th valign="top" align="left">Accuracy</th>
<th valign="top" align="left">Sensitivity</th>
<th valign="top" align="left">Specificity</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.906</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.850</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.921</td>
<td valign="top" align="left">0.887</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.853</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.891</td>
<td valign="top" align="left">0.847</td>
<td valign="top" align="left">0.937</td>
<td valign="top" align="left">0.846</td>
<td valign="top" align="left">0.977</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">k-mer</td>
<td valign="top" align="left">0.891</td>
<td valign="top" align="left">0.847</td>
<td valign="top" align="left">0.937</td>
<td valign="top" align="left">0.846</td>
<td valign="top" align="left">0.977</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.933</td>
<td valign="top" align="left">0.905</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.875</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.931</td>
<td valign="top" align="left">0.903</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.894</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.933</td>
<td valign="top" align="left">0.905</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.875</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">AA k-mer</td>
<td valign="top" align="left">0.917</td>
<td valign="top" align="left">0.883</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.891</td>
<td valign="top" align="left">0.977</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.869</td>
<td valign="top" align="left">0.820</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.882</td>
<td valign="top" align="left">0.946</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.849</td>
<td valign="top" align="left">0.788</td>
<td valign="top" align="left">0.913</td>
<td valign="top" align="left">0.815</td>
<td valign="top" align="left">0.955</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.869</td>
<td valign="top" align="left">0.820</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.882</td>
<td valign="top" align="left">0.946</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP</td>
<td valign="top" align="left">0.873</td>
<td valign="top" align="left">0.824</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.861</td>
<td valign="top" align="left">0.956</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.902</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.916</td>
<td valign="top" align="left">0.978</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.901</td>
<td valign="top" align="left">0.863</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.888</td>
<td valign="top" align="left">0.967</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.916</td>
<td valign="top" align="left">0.883</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.891</td>
<td valign="top" align="left">0.997</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP/AA k-mer</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.902</td>
<td valign="top" align="left">0.960</td>
<td valign="top" align="left">0.916</td>
<td valign="top" align="left">0.978</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.906</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.850</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.906</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.850</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.876</td>
<td valign="top" align="left">0.827</td>
<td valign="top" align="left">0.929</td>
<td valign="top" align="left">0.842</td>
<td valign="top" align="left">0.966</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">SNP/k-mer</td>
<td valign="top" align="left">0.891</td>
<td valign="top" align="left">0.847</td>
<td valign="top" align="left">0.937</td>
<td valign="top" align="left">0.846</td>
<td valign="top" align="left">0.933</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.906</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.850</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Support Vector Machine</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.921</td>
<td valign="top" align="left">0.887</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.853</td>
<td valign="top" align="left">1.000</td>
</tr>
<tr>
<td valign="top" align="left">Stochastic Gradient Boosting</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.906</td>
<td valign="top" align="left">0.867</td>
<td valign="top" align="left">0.944</td>
<td valign="top" align="left">0.850</td>
<td valign="top" align="left">0.988</td>
</tr>
<tr>
<td valign="top" align="left">Extreme Gradient Boosting</td>
<td valign="top" align="left">AA k-mer/k-mer</td>
<td valign="top" align="left">0.918</td>
<td valign="top" align="left">0.885</td>
<td valign="top" align="left">0.952</td>
<td valign="top" align="left">0.871</td>
<td valign="top" align="left">0.988</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>F1 Scores of all penicillin, erythromycin, and tetracycline resistance models using six feature types and 4 ML approaches are presented in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. While the overall performances were adequate, different feature types and ML approaches yielded varying performances per each antibiotic.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>F1 Scores of penicillin, erythromycin, and tetracycline resistance classification models using six different input types and four different ML approaches.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frabi-02-1126468-g002.tif"/>
</fig>
<p>Of interest when the distribution of SNPs between resistant and susceptible samples were compared for each antibiotic, it was observed that the number of SNPs in the susceptible samples was significantly higher for all three antibiotics (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S1</bold>
</xref>).</p>
<p>For penicillin, parameters were optimized <italic>via</italic> cross-validation, and performance estimates averaged over five repeats of this setup on 869 samples. For the prediction of penicillin susceptibility and resistance, the machine learning classifiers performed almost equally well with the five feature types (k-mer, AA k-mer, k-mer/AA k-mer, SNP/k-mer, and SNP/AA k-mer) except SNP alone itself. (all F1 score&gt; 0.76). Comparisons of the models are shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref> and <xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2</bold>
</xref>&#x2013;<xref ref-type="fig" rid="f4">
<bold>4</bold>
</xref>. These figures shows the F1 score, accuracy and kappa results of penicillin, erythromycin, and tetracycline resistance classification models using six different input types and four different ML approaches.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Accuracy results of penicillin, erythromycin, and tetracycline resistance classification models using six different input types and four different ML approaches.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frabi-02-1126468-g003.tif"/>
</fig>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Kappa results of penicillin, erythromycin, and tetracycline resistance classification models using six different input types and four different ML approaches.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frabi-02-1126468-g004.tif"/>
</fig>
<p>When different features for penicillin were compared, it was observed that nucleotide 10-mer and amino acid 5-mer F1 scores were higher than other features; amino acid 5-mer combinations were other inputs that yielded relatively higher F1 scores. SNP, SNP/AA k-mer, and SNP/k-mer combinations yielded lower accuracy and F1 scores than other penicillin models. The highest F1 score was observed in the Random Forest model at 0.9. The accuracy of the same model was found to be 0.96. The second-best option was SVM, and the third-best option was XGBoost, which performed almost as well as RF in F1-Score. The RF model utilizing 5-mer AA k-mer features to classify penicillin resistance yielded a sensitivity of 0.882 and a specificity of 0.99. Similarly, the SVM linear model resulted in high predictive sensitivity and specificity values of 0.86 and 0.99. Moreover, the XGBoost resulted in a sensitivity of 0.96 and a specificity of 0.93.</p>
<p>For erythromycin, a total of 895 samples were used in our models. As measured by the accuracy and F1 score, the best performance was achieved by nucleotide k-mer model with XGBoost (F1-Score 0.97, Accuracy 0.98). For the erythromycin AMR prediction, all classifiers performed almost equally well with all feature types except for SNP features. The first and second highest F1 score measured by XGBoost and random forest, which performed close to GBM in F1-Score and accuracy. When different features in erythromycin models were compared, it was seen for combinations of inputs, including SNP/AA k-mer, SNP/k-mer, and AA k-mer/k-mer, F1-scores were higher than SNP. Amino acid 5-mer and SNP combinations yielded the highest F1 scores. Performances of the erythromycin models are presented in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>. The second-best option was Random Forest in terms of F1-Score. With nucleotide k-mer feature, the erythromycin resistance classification using the XGBoost model correctly predicted resistance with a sensitivity of 0.87 and a specificity of 0.88.</p>
<p>Tetracycline parameters were optimized <italic>via</italic> cross-validation, and performance estimates were averaged over five repeats of this setup using 643 samples. Performances of all tetracycline models are presented in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>. For the prediction of tetracycline resistance, the ML classifiers performed almost equally well with the six input data types (k-mer, AA k-mer, SNPs, k-mer/AA k-mer, SNP/k-mer, and SNP/AA k-mer) (F1 score &gt; 0.85, <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>). When different feature inputs in tetracycline were compared, it was observed that the AA k-mer/k-mer combination yielded a higher F1 score than other inputs. The highest F1 score was observed for the random forest model, with 0.93. The accuracy of the same model was found to be 0.96. The second-best option was GBM, and the third option was SVM, which performed close to RF in terms of F1-Score and accuracy. With the AA k-mer/k-mer feature tetracycline resistance RF model, tetracycline resistance could be predicted with a sensitivity of 0.85 and a specificity of 0.98.</p>
<p>As described in Methods, we used binary features (i.e., the presence or absence of k-mers) rather than k-mer counts to simplify the analyses. When a model is used to predict the MIC class for a new genome, the k-mers with the highest importance values are expected to be the most informative. Thus, by analyzing the feature importance values of each k-mer, we can use the models generated in this study to understand the genomic regions that differentiate MIC classes. Hence, to understand the relationship between known AMR genes and the important k-mers chosen by each model, we searched for k-mers with high-importance values within AMR genes or near an AMR gene.</p>
<p>In most cases, the top k-mers corresponded to known AMR genes. The top 10 10-mers with the highest feature importance values were checked against <italic>S. pneumoniae</italic>-related known AMR genes, including penicillin-binding proteins (PBPs), which have a major role in the cell wall synthesis (<italic>PBP2b</italic>, <italic>PBP2x</italic>, and <italic>PBP1a</italic>) and are most often associated with penicillin resistance. For macrolide resistance mechanisms in S. pneumoniae, <italic>ermB</italic> and <italic>mefE</italic> genes stand out, encoding an active efflux pump. Also, the most common resistance mechanism to tetracycline in S. pneumoniae is the acquisition of one of the three genes, <italic>tetM</italic>, <italic>tetO</italic>, and <italic>tetK</italic>. In the case of penicillin, for the top 10 features in <italic>S. pneumoniae</italic>, we used the hypergeometric test to calculate the probability of top 10 10-mers appearing in the resistance genes. As a result, p value (0.036) was found to be statistically significant when compared with pbp2b, pbp2x and pbp1a resistance genes. When we looked for tetracycline with tetS and tetM genes, the p value was 0.015. When we evaluated our 10-mers for Erythromycin, the rate of occurrence of nucleotide sequences in these genes for four genes (Ermb, mefE, msrD, mefA) was found to be 0.12.</p>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>AMR has caused a significant increase in morbidity and mortality rate in infectious diseases worldwide, raising global health concerns (<xref ref-type="bibr" rid="B52">World Health Organization, 2019</xref>; <xref ref-type="bibr" rid="B53">World Health Organization, 2022a</xref>). It is crucial to quickly detect AMR in bacterial genomes as the number of effective antibiotics decreases. Molecular approaches have significantly improved over the years and play a critical role in the fight against antimicrobial resistance. (<xref ref-type="bibr" rid="B29">Leski et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B25">Inouye et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B13">Davis et&#xa0;al., 2016</xref>) Building classifiers with a balanced number of susceptible and resistant genomes is also important for building accurate classifiers but is currently a major limitation. In most cases, the number of available genomes with AMR data is resistant because these are of clinical importance to hospitals and epidemiologists.</p>
<p>Given the current data sets available on PATRIC, we built RF, SVM, GBM, and XGBoost classifiers for penicillin, erythromycin, and tetracycline resistance <italic>Streptococcus pneumoniae</italic>. The classifiers were highly accurate and performed classifications based on nucleotide k-mers.</p>
<p>The feature extraction methods that we present here have different pre-processing steps. In the case of SNP features, alignment to the reference genome is required because each SNP must have a unique position in the reference genome. In order to work with SNP locations, variant calling is required. Although it is not a very long process, it is a pre-process that should be evaluated. Apart from the location of the SNPs, we looked to see if there was a difference in the number of SNPs between susceptible and resistant samples, it was observed that the number of SNPs in the susceptible samples was significantly higher for all three antibiotics (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S1</bold>
</xref>). The number of SNPs was significantly lower in the resistance samples regardless of antibiotic. This shows that simply assessing the number of SNPs in a sample might be a useful initial step when predicting MIC class.</p>
<p>When we compared the machine learning models, we could not find any obvious difference that could distinguish one from the other. When evaluating the results of Erythromycin, when we ran XGBoost, which gave F1 scores of 0.97 and 0.96, with the SNP feature, we saw that it gave the weakest result among the tested models (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2</bold>
</xref>&#x2013;<xref ref-type="fig" rid="f4">
<bold>4</bold>
</xref>). XGBoost is a popular machine learning algorithm it&#x2019;s because of high predictive accuracy. XGBoost is fast and ideal for big datasets, when we compare to other models like random forest.</p>
<p>For amino acid k-mers methods, by contrast, the input to the feature extraction method is the amino acid sequence of the genes. This means that just aligning the short reads to contigs is not sufficient. This adds an extra pre-processing step to these methods. However, predicting AMR as fast as possible and as cheaply as possible is the top priority. Thus, amino acid k-mers are the better option because of the smaller feature size and better interpretability of AA features. Overall, aa kmer can be a useful tool for prediction, This method, which has just started to be used for MIC prediction, is seen to give high results when compared to other feature inputs.</p>
<p>Our comparisons showed that different feature inputs yielded the optimal results for each antibiotic. Amino acid 5-mers resulted in the best performance for penicillin. In contrast, the SNP and amino acid 5-mers combination were the best for tetracycline, and the combination of nucleotide 10-mers and amino acid 5-mers yielded the best performance for tetracycline.</p>
<p>In machine learning, an excessive number of features can increase the required memory and lead to over-fitting. Using long k-mers is hard because the number of features increases; however, we have shown that for amino acid k-mers, this increase in feature size is less severe than for nucleotide k-mers. One advantage of amino acid k-mers over nucleotide k-mers is that they are more compact representations of biological information. Each codon consists of three nucleotides and translates into one amino acid. Moreover, amino acid k-mers and their combinations achieved better performance in terms of accuracy.</p>
<p>In this study, the k-mers relating to penicillin resistance in <italic>S. pneumoniae</italic> that were identified by RF corresponded with the <italic>pbp2x</italic> gene that was also identified in previous genome-wide association studies (<xref ref-type="bibr" rid="B6">Chewapreecha et&#xa0;al., 2014</xref>). In that study, <xref ref-type="bibr" rid="B6">Chewapreecha and colleagues (2014)</xref> also found significant variations relating to resistance in the <italic>pbp1a</italic> and <italic>pbp2a</italic> penicillin-binding proteins, which were also identified in this study using the RF model.</p>
<p>In this study, we compared feature sets and ML models for predicting AMR phenotype for S. pneumoniae. We compared nucleotide k-mers, amino acid k-mers, and SNPs to predict AMR for three antibiotics: Penicillin, Erythromycin, and Tetracycline. Further, we attempted to use and compare various ML methods: random forest, support vector machine, stochastic gradient boosting, and extreme gradient boosting to train the classification models. We attempted to discuss the strengths and limitations of feature and ML model selection for MIC prediction. As a result of our work, we have observed that the features used in the model setup and the choice of ML method affect the result. Especially the feature combinations giving high accuracy and F1 score for some antibiotics showed that these feature inputs should be evaluated in the future. We hope that the approach undertaken by this study can be used in further studies to improve AMR prediction performance and accuracy and help alleviate the burden of AMR in the clinical setting.</p>
</sec>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="s9">
<bold>Supplementary Material</bold>
</xref>. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>First authorship: DK Senior authorship: EU, AK. Last authorship: OS. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s7" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s8" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s9" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frabi.2023.1126468/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frabi.2023.1126468/full#supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document"/>
<supplementary-material xlink:href="DataSheet_2.xlsx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>AMR Review</collab>
</person-group> (<year>2015</year>)<article-title>Review on antimicrobial resistance</article-title>. In: <source>Rapid diagnostics: Stopping unnecessary use of antibiotics</source>. Available at: <uri xlink:href="https://amr-review.org/Publications.html">https://amr-review.org/Publications.html</uri> (Accessed <access-date>May 12, 2022</access-date>).</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aytan-Aktug</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Clausen</surname> <given-names>P. T.</given-names>
</name>
<name>
<surname>Bortolaia</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Aarestrup</surname> <given-names>F. M.</given-names>
</name>
<name>
<surname>Lund</surname> <given-names>O.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Prediction of acquired antimicrobial resistance for multiple bacterial species using neural networks</article-title>. <source>MSystems</source> <volume>5</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1128/msystems.00774-19</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bankevich</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Nurk</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Antipov</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Gurevich</surname> <given-names>A. A.</given-names>
</name>
<name>
<surname>Dvorkin</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Kulikov</surname> <given-names>A. S.</given-names>
</name>
<etal/>
</person-group>. (<year>2012</year>). <article-title>Spades: A new genome assembly algorithm and its applications to single-cell sequencing</article-title>. <source>J. Comput. Biol.</source> <volume>19</volume> (<issue>5</issue>), <fpage>455</fpage>&#x2013;<lpage>477</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1089/cmb.2012.0021</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blair</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Webber</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Baylay</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Ogbolu</surname> <given-names>D. O.</given-names>
</name>
<name>
<surname>Piddock</surname> <given-names>L. J.</given-names>
</name>
</person-group>. (<year>2015</year>). <article-title>Molecular mechanisms of antibiotic resistance</article-title>. <source>Nat. Rev. Microbiol.</source> <volume>13</volume>, <fpage>42</fpage>&#x2013;<lpage>51</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/nrmicro3380</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brettin</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Davis</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Disz</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Edwards</surname> <given-names>R. A.</given-names>
</name>
<name>
<surname>Gerdes</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Olsen</surname> <given-names>G. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>RASTtk: a modular and extensible implementation of the RAST algorithm for building custom annotation pipelines and annotating batches of genomes</article-title>. <source>Sci. Rep.</source> <volume>5</volume>, <fpage>8365</fpage>. doi: <pub-id pub-id-type="doi">10.1038/srep08365</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chewapreecha</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Marttinen</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Croucher</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>Salter</surname> <given-names>S. J.</given-names>
</name>
<name>
<surname>Harris</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>Mather</surname> <given-names>A. E.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Comprehensive identification of single nucleotide polymorphisms associated with beta-lactam resistance within pneumococcal mosaic genes</article-title>. <source>PLoS Genet.</source> <volume>10</volume> (<issue>8</issue>), <fpage>e1004547</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pgen.1004547</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Christaki</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Marcou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Tofarides</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Antimicrobial resistance in bacteria: Mechanisms, evolution, and persistence</article-title>. <source>J. Mol. Evol.</source> <volume>88</volume> (<issue>1</issue>), <fpage>26</fpage>&#x2013;<lpage>40</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00239-019-09914-3</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>CLSI guidelines</collab>
</person-group> (<year>2022</year>) <source>Clinical &amp; laboratory standards institute</source>. Available at: <uri xlink:href="https://clsi.org/">https://clsi.org/</uri>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cornick</surname> <given-names>J. E.</given-names>
</name>
<name>
<surname>Bentley</surname> <given-names>S. D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Streptococcus pneumoniae: The evolution of antimicrobial resistance to beta-lactams, fluoroquinolones, and macrolides</article-title>. <source>Microbes Infection</source> <volume>14</volume> (<issue>7-8</issue>), <fpage>573</fpage>&#x2013;<lpage>583</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.micinf.2012.01.012</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Croucher</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>Hanage</surname> <given-names>W. P.</given-names>
</name>
<name>
<surname>Harris</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>McGee</surname> <given-names>L.</given-names>
</name>
<name>
<surname>van der Linden</surname> <given-names>M.</given-names>
</name>
<name>
<surname>de Lencastre</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>Population genomic datasets describing the post-vaccine evolutionary epidemiology of</article-title>. <source>Streptococcus pneumoniae Sci. Data</source> <volume>2</volume>, <fpage>150058</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/sdata.2015.58</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Croucher</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>Finkelstein</surname> <given-names>J. A.</given-names>
</name>
<name>
<surname>Pelton</surname> <given-names>S. I.</given-names>
</name>
<name>
<surname>Mitchell</surname> <given-names>P. K.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>G. M.</given-names>
</name>
<name>
<surname>Parkhill</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Population genomics of post-vaccine changes in pneumococcal epidemiology</article-title>. <source>Nat. Genet.</source> <volume>45</volume> (<issue>6</issue>), <fpage>656</fpage>&#x2013;<lpage>663</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng.2625</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Croucher</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>Hanage</surname> <given-names>W. P.</given-names>
</name>
<name>
<surname>Harris</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>McGee</surname> <given-names>L.</given-names>
</name>
<name>
<surname>van der Linden</surname> <given-names>M.</given-names>
</name>
<name>
<surname>de Lencastre</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Variable recombination dynamics during the emergence, transmission, and &#x2018;disarming&#x2019; of a multidrug-resistant pneumococcal clone</article-title>. <source>BMC Biol.</source> <volume>12</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1741-7007-12-49</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davis</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Boisvert</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Brettin</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Kenyon</surname> <given-names>R. W.</given-names>
</name>
<name>
<surname>Mao</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Olson</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>Antimicrobial resistance prediction in patric and rast</article-title>. <source>Sci. Rep.</source> <volume>6</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1038/srep27930</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davis</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Wattam</surname> <given-names>A. R.</given-names>
</name>
<name>
<surname>Aziz</surname> <given-names>R. K.</given-names>
</name>
<name>
<surname>Brettin</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Butler</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Butler</surname> <given-names>R. M.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>The PATRIC bioinformatics resource center: expanding data and analysis capabilities</article-title>. <source>Nucleic Acids Res.</source> <volume>48</volume> (<issue>D1</issue>), <fpage>D606</fpage>&#x2013;<lpage>D612</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkz943</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deelder</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Christakoudi</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Phelan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Benavente</surname> <given-names>E. D.</given-names>
</name>
<name>
<surname>Campino</surname> <given-names>S.</given-names>
</name>
<name>
<surname>McNerney</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Machine learning predicts accurately mycobacterium tuberculosis drug resistance from whole genome sequencing data</article-title>. <source>Front. Genet.</source> <volume>10</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2019.00922</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Demczuk</surname> <given-names>W. H.</given-names>
</name>
<name>
<surname>Martin</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Hoang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Van Caeseele</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Lefebvre</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Horsman</surname> <given-names>G.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Phylogenetic analysis of emergent <italic>streptococcus pneumonia</italic>e serotype 22F causing invasive pneumococcal disease using whole genome sequencing</article-title>. <source>PLoS One.</source> <volume>12</volume> (<issue>5</issue>), <fpage>e0178040</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0178040</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Drouin</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Gigu&#xe8;re</surname> <given-names>S.</given-names>
</name>
<name>
<surname>D&#xe9;raspe</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Marchand</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Tyers</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Loo</surname> <given-names>V. G.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>Predictive computational phenotyping and biomarker discovery using reference-free genome comparisons</article-title>. <source>BMC Genomics</source> <volume>17</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12864-016-2889-6</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dunne</surname> <given-names>J. W.M.</given-names>
</name>
<name>
<surname>Jaillard</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Rochas</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Van Belkum</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Microbial genomics and antimicrobial susceptibility testing</article-title>. <source>Expert Rev. Mol. Diagnostics</source> <volume>17</volume> (<issue>3</issue>), <fpage>257</fpage>&#x2013;<lpage>269</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1080/14737159.2017.1283220</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Edgar</surname> <given-names>R. C.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Local homology recognition and distance measures in linear time using compressed amino acid alphabets</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>380</fpage>&#x2013;<lpage>385</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkh180</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>ESCMID - European Society of Clinical Microbiology and Infectious Diseases</collab>
</person-group> (<year>2008</year>). <source>Mic and zone diameter distributions and ecoffs</source> (<publisher-name>EUCAST</publisher-name>). Available at: <uri xlink:href="https://www.eucast.org/mic_and_zone_distributions_and_ecoffs">https://www.eucast.org/mic_and_zone_distributions_and_ecoffs</uri>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eyre</surname> <given-names>D. W.</given-names>
</name>
<name>
<surname>De Silva</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Cole</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Peters</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Cole</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Grad</surname> <given-names>Y. H.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>WGS to predict antibiotic mics for neisseria gonorrhoeae</article-title>. <source>J. Antimicrobial Chemotherapy</source> <volume>72</volume> (<issue>7</issue>), <fpage>1937</fpage>&#x2013;<lpage>1947</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/jac/dkx067</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gladstone</surname> <given-names>R. A.</given-names>
</name>
<name>
<surname>Lo</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Lees</surname> <given-names>J. A.</given-names>
</name>
<name>
<surname>Croucher</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>van Tonder</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Corander</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>International genomic definition of pneumococcal lineages, to contextualise disease, antibiotic resistance and vaccine impact</article-title>. <source>EBioMedicine</source> <volume>43</volume>, <fpage>338</fpage>&#x2013;<lpage>346</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.ebiom.2019.04.021</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Henriques-Normark</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Tuomanen</surname> <given-names>E. I.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>The pneumococcus: Epidemiology, microbiology, and pathogenesis</article-title>. <source>Cold Spring Harbor Perspect. Med.</source> <volume>3</volume> (<issue>7</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1101/cshperspect.a010215</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Her</surname> <given-names>H.-L.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y.-W.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A pan-genome-based machine learning approach for predicting antimicrobial resistance activities of the escherichia coli strains</article-title>. <source>Bioinformatics</source> <volume>34</volume> (<issue>13</issue>), <fpage>i89</fpage>&#x2013;<lpage>i95</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/bty276</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Inouye</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Dashnow</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Raven</surname> <given-names>L.-A.</given-names>
</name>
<name>
<surname>Schultz</surname> <given-names>M. B.</given-names>
</name>
<name>
<surname>Pope</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Tomita</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>SRST2: Rapid genomic surveillance for public health and hospital microbiology labs</article-title>. <source>Genome Med.</source> <volume>6</volume> (<issue>11</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s13073-014-0090-6</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jolley</surname> <given-names>K. A.</given-names>
</name>
<name>
<surname>Bray</surname> <given-names>J. E.</given-names>
</name>
<name>
<surname>Maiden</surname> <given-names>M. C. J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications</article-title>. <source>Wellcome Open Res.</source> <volume>3</volume>, <fpage>124</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.12688/wellcomeopenres.14826.1</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khaledi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Weimann</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Schniederjans</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Asgari</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Kuo</surname> <given-names>T. H.</given-names>
</name>
<name>
<surname>Oliver</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Predicting antimicrobial resistance in pseudomonas aeruginosa with machine learning-enabled molecular diagnostics</article-title>. <source>EMBO Mol. Med.</source> <volume>12</volume> (<issue>3</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.15252/emmm.201910264</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kursa</surname> <given-names>M. B.</given-names>
</name>
<name>
<surname>Rudnicki</surname> <given-names>W. R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Feature selection with theborutapackage</article-title>. <source>J. Stat. Software</source> <volume>36</volume> (<issue>11</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.18637/jss.v036.i11</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Leski</surname> <given-names>T. A.</given-names>
</name>
<name>
<surname>Vora</surname> <given-names>G. J.</given-names>
</name>
<name>
<surname>Barrows</surname> <given-names>B. R.</given-names>
</name>
<name>
<surname>Pimentel</surname> <given-names>G.</given-names>
</name>
<name>
<surname>House</surname> <given-names>B. L.</given-names>
</name>
<name>
<surname>Nicklasson</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Molecular characterization of multidrug-resistant hospital isolates using the antimicrobial resistance determinant microarray</article-title>. <source>PloS One</source> <volume>8</volume> (<issue>7</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0069507</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A statistical framework for SNP calling mutation discovery, association mapping, and population genetic parameter estimation from sequencing data</article-title>. <source>Bioinformatics</source> <volume>27</volume> (<issue>21</issue>), <fpage>2987</fpage>&#x2013;<lpage>2993</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btr509</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Aligning sequence reads, clone sequences, and assembly contigs with BWA-MEM</article-title>. <source>ArXiv</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/ARXIV.1303.3997</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Handsaker</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wysoker</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Fennell</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Ruan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Homer</surname> <given-names>N.</given-names>
</name>
<etal/>
</person-group>. (<year>2009</year>). <article-title>And 1000 genome project data processing subgroup the sequence alignment/map (SAM) format and SAMtools</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2078</fpage>&#x2013;<lpage>2079</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btp352</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Metcalf</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Chochua</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Gertz</surname> <given-names>R. E.</given-names>
</name>
<name>
<surname>Walker</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Validation of &#x3b2;-lactam minimum inhibitory concentration predictions for pneumococcal isolates with newly encountered penicillin-binding protein (PBP) sequences</article-title>. <source>BMC Genomics</source> <volume>18</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12864-017-4017-7</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Metcalf</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Chochua</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Gertz</surname> <given-names>R. E.</given-names>
</name>
<name>
<surname>Walker</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>Penicillin-binding protein transpeptidase signatures for tracking and predicting &#x3b2;-lactam resistance levels in <italic>streptococcus pneumonia</italic>e</article-title>. <source>MBio</source> <volume>7</volume> (<issue>3</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1128/mbio.00756-16</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lv</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Evaluation of machine learning models for predicting antimicrobial resistance of actinobacillus pleuropneumoniae from whole genome sequences</article-title>. <source>Front. Microbiol.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmicb.2020.00048</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Michael</surname> <given-names>C. A.</given-names>
</name>
<name>
<surname>Dominey-Howes</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Labbate</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>The antimicrobial resistance crisis: Causes, consequences, and management</article-title>. <source>Front. Public Health</source> <volume>2</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpubh.2014.00145</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Michael</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Kelman</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Pitesky</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Overview of quantitative methodologies to understand antimicrobial resistance <italic>via</italic> minimum inhibitory concentration</article-title>. <source>Animals</source> <volume>10</volume> (<issue>8</issue>), <elocation-id>1405</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/ani10081405</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Montanari</surname> <given-names>M. P.</given-names>
</name>
<name>
<surname>Cochetti</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Mingoia</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Varaldo</surname> <given-names>P. E.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Phenotypic and molecular characterization of tetracycline- and erythromycin-resistant strains of <italic>streptococcus pneumoniae</italic>
</article-title>. <source>Antimicrobial Agents Chemotherapy</source> <volume>47</volume> (<issue>7</issue>), <fpage>2236</fpage>&#x2013;<lpage>2241</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1128/aac.47.7.2236-2241.2003</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moradigaravand</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Palm</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Farewell</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Mustonen</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Warringer</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Parts</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Prediction of antibiotic resistance in escherichia coli from large-scale pan-genome data</article-title>. <source>PloS Comput. Biol.</source> <volume>14</volume> (<issue>12</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pcbi.1006258</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Naidenov</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lim</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Willyerd</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Torres</surname> <given-names>N. J.</given-names>
</name>
<name>
<surname>Johnson</surname> <given-names>W. L.</given-names>
</name>
<name>
<surname>Hwang</surname> <given-names>H. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Pan-genomic and polymorphic driven prediction of antibiotic resistance in elizabethkingia</article-title>. <source>Front. Microbiol.</source> <volume>10</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmicb.2019.01446</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nguyen</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Brettin</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Long</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Musser</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Olsen</surname> <given-names>R. J.</given-names>
</name>
<name>
<surname>Olson</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>Developing an in silico minimum inhibitory concentration panel test for klebsiella pneumoniae</article-title>. <source>Sci. Rep.</source> <volume>8</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41598-017-18972-w</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nguyen</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Long</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>McDermott</surname> <given-names>P. F.</given-names>
</name>
<name>
<surname>Olsen</surname> <given-names>R. J.</given-names>
</name>
<name>
<surname>Olson</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Stevens</surname> <given-names>R. L.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Using machine learning to predict antimicrobial mics and associated genomic features for nontyphoidal <italic>salmonella</italic>
</article-title>. <source>J. Clin. Microbiol.</source> <volume>57</volume> (<issue>2</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1128/jcm.01260-18</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pataki</surname> <given-names>B.&#xc1;.</given-names>
</name>
<name>
<surname>Matamoros</surname> <given-names>S.</given-names>
</name>
<name>
<surname>van der Putten</surname> <given-names>B. C. L.</given-names>
</name>
<name>
<surname>Remondini</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Giampieri</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Aytan-Aktug</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Understanding and predicting ciprofloxacin minimum inhibitory concentration in <italic>escherichia coli</italic> with machine learning</article-title>. <source>Sci Rep.</source> <volume>10</volume> (<issue>1</issue>), <fpage>15026</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1101/806760</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poole</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Resistance to b-lactam antibiotics</article-title>. <source>Cell. Mol. Life Sci.</source> <volume>61</volume> (<issue>17</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00018-004-4060-9</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pudjihartono</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Fadason</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Kempa-Liehr</surname> <given-names>A. W.</given-names>
</name>
<name>
<surname>O'Sullivan</surname> <given-names>J. M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A review of feature selection methods for machine learning-based disease risk prediction</article-title>. <source>Front. Bioinform.</source> <volume>2</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fbinf.2022.927312</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sader</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Mendes</surname> <given-names>R. E.</given-names>
</name>
<name>
<surname>Le</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Denys</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Flamm</surname> <given-names>R. K.</given-names>
</name>
<name>
<surname>Jones</surname> <given-names>R. N.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Antimicrobial susceptibility of streptococcus pneumoniae from north America, Europe, Latin America, and the Asia-pacific region: Results from 20 years of the sentry antimicrobial surveillance program, (1997&#x2013;2016)</article-title>. <source>Open Forum Infect. Dis.</source> <volume>6</volume> (<supplement>Supplement_1</supplement>). doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ofid/ofy263</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Links</surname> <given-names>M. G.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Dillon</surname> <given-names>J.-A. R.</given-names>
</name>
<name>
<surname>Horsch</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Antimicrobial resistance genetic factor identification from whole-genome sequence data using deep feature selection</article-title>. <source>BMC Bioinf.</source> <volume>20</volume> (<issue>S15</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12859-019-3054-4</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>ValizadehAslani</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Sokhansanj</surname> <given-names>B. A.</given-names>
</name>
<name>
<surname>Rosen</surname> <given-names>G. L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Amino acid K-mer feature extraction for quantitative antimicrobial resistance (AMR) prediction by machine learning and model interpretation for biological insights</article-title>. <source>Biology</source> <volume>9</volume> (<issue>11</issue>), <elocation-id>365</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/biology9110365</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>van der Poll</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Opal</surname> <given-names>S. M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Pathogenesis, treatment, and prevention of pneumococcal pneumonia</article-title>. <source>Lancet</source> <volume>374</volume> (<issue>9700</issue>), <fpage>1543</fpage>&#x2013;<lpage>1556</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/s0140-6736(09)61114-4</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Xiong</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Whole-genome sequencing of <italic>mycobacterium tuberculosis</italic> for prediction of drug resistance</article-title>. <source>Epidemiol. Infection</source> <volume>150</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.1017/s095026882100279x</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wilkinson</surname> <given-names>S. P.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Kmer an r package for fast alignment-free clustering of biological sequences</article-title>,&#x201d; in <source>R package version 1.0.0</source>. Available at: <uri xlink:href="https://cran.r-project.org/package=kmer">https://cran.r-project.org/package=kmer</uri>.</citation>
</ref>
<ref id="B52">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>World Health Organization</collab>
</person-group> (<year>2019</year>) <source>A new report calls for urgent action to avert the antimicrobial resistance crisis</source>. Available at: <uri xlink:href="https://www.who.int/news/item/29-04-2019-new-report-calls-for-urgent-action-to-avert-antimicrobial-resistance-crisis">https://www.who.int/news/item/29-04-2019-new-report-calls-for-urgent-action-to-avert-antimicrobial-resistance-crisis</uri> (Accessed <access-date>May 12, 2022</access-date>).</citation>
</ref>
<ref id="B53">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>World Health Organization</collab>
</person-group> (<year>2022</year>a) <source>Antimicrobial resistance</source>. Available at: <uri xlink:href="https://www.who.int/news-room/fact-sheets/detail/antimicrobial-resistance">https://www.who.int/news-room/fact-sheets/detail/antimicrobial-resistance</uri> (Accessed <access-date>May 12, 2022</access-date>).</citation>
</ref>
<ref id="B54">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>World Health Organization</collab>
</person-group> (<year>2022</year>b) <source>Pneumococcal disease. world health organization</source>. Available at: <uri xlink:href="https://www.who.int/teams/health-product-policy-and-standards/standards-and-specifications/vaccine-standardization/pneumococcal-disease">https://www.who.int/teams/health-product-policy-and-standards/standards-and-specifications/vaccine-standardization/pneumococcal-disease</uri> (Accessed <access-date>May 15, 2022</access-date>).</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Niehaus</surname> <given-names>K. E.</given-names>
</name>
<name>
<surname>Walker</surname> <given-names>T. M.</given-names>
</name>
<name>
<surname>Iqbal</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Walker</surname> <given-names>A. S.</given-names>
</name>
<name>
<surname>Wilson</surname> <given-names>D. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Machine learning for classifying tuberculosis drug resistance from DNA sequencing data</article-title>. <source>Bioinformatics</source> <volume>34</volume> (<issue>10</issue>), <fpage>1666</fpage>&#x2013;<lpage>1671</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btx801</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Efficient feature selection via analysis of relevance and redundancy</article-title>. <source>J. Mach. Learn. Res.</source> <volume>5</volume>, <fpage>1205</fpage>&#x2013;<lpage>1224</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.5555/1005332.1044700</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zapun</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Contreras-Martel</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Vernet</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Penicillin-binding proteins and &#x3b2;-lactam resistance</article-title>. <source>FEMS Microbiol. Rev.</source> <volume>32</volume> (<issue>2</issue>), <fpage>361</fpage>&#x2013;<lpage>385</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/j.1574-6976.2007.00095.x</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Ju</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Systematic analysis of supervised machine learning as an effective approach to predicate &#x3b2;-lactam resistance phenotype in <italic>Streptococcus pneumonia</italic>e</article-title>. <source>Briefings Bioinf.</source> <volume>21</volume> (<issue>4</issue>), <fpage>1347</fpage>&#x2013;<lpage>1355</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbz056</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>