<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Oncol.</journal-id>
<journal-title>Frontiers in Oncology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Oncol.</abbrev-journal-title>
<issn pub-type="epub">2234-943X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fonc.2022.888556</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Oncology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Highly Effective System for Predicting MHC-II Epitopes With Immunogenicity</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Xu</surname><given-names>Shi</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1821483"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname><given-names>Xiaohua</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1819448"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Fei</surname><given-names>Caiyi</given-names>
</name>
<xref ref-type="author-notes" rid="fn001"><sup>*</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1674738"/>
</contrib>
</contrib-group>    <aff id="aff1"><institution>Department of AI and Bioinformatics, Nanjing Chengshi BioTech (TheraRNA) Co., Ltd.</institution>, <addr-line>Nanjing</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Zexian Liu, Sun Yat-sen University Cancer Center (SYSUCC), China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Shaofeng Lin, Huazhong University of Science and Technology, China; Tianshun Gao, Sun Yat-sen University, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Caiyi Fei, <email xlink:href="mailto:feicaiyi@therarna.cn">feicaiyi@therarna.cn</email>
</p>
</fn>
<fn fn-type="other" id="fn002">
<p>This article was submitted to Molecular and Cellular Oncology, a section of the journal Frontiers in Oncology</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>16</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>12</volume>
<elocation-id>888556</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>03</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Xu, Wang and Fei</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Xu, Wang and Fei</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>In the past decade, the substantial achievements of therapeutic cancer vaccines have shed a new light on cancer immunotherapy. The major challenge for designing potent therapeutic cancer vaccines is to identify neoantigens capable of inducing sufficient immune responses, especially involving major histocompatibility complex (MHC)-II epitopes. However, most previous studies on T-cell epitopes were focused on either ligand binding or antigen presentation by MHC rather than the immunogenicity of T-cell epitopes. In order to better facilitate a therapeutic vaccine design, in this study, we propose a revolutionary new tool: a convolutional neural network model named FIONA (Flexible Immunogenicity Optimization Neural-network Architecture) trained on IEDB datasets. FIONA could accurately predict the epitopes presented by the given specific MHC-II subtypes, as well as their immunogenicity. By leveraging the human leukocyte antigen allele hierarchical encoding model together with peptide dense embedding fusion encoding, FIONA (with AUC = 0.94) outperforms several other tools in predicting epitopes presented by MHC-II subtypes in head-to-head comparison; moreover, FIONA has unprecedentedly incorporated the capacity to predict the immunogenicity of epitopes with MHC-II subtype specificity. Therefore, we developed a reliable pipeline to effectively predict CD4+ T-cell immune responses against cancer and infectious diseases.</p>
</abstract>
<kwd-group>
<kwd>neoantigen</kwd>
<kwd>cancer vaccine</kwd>
<kwd>deep learning</kwd>
<kwd>IEDB</kwd>
<kwd>CD4<sup>+</sup> T cell</kwd>
<kwd>MHC-II</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="2"/>
<equation-count count="9"/>
<ref-count count="59"/>
<page-count count="12"/>
<word-count count="5956"/>
</counts>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Therapeutic cancer vaccines (<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B3">3</xref>) are regarded as the most promising cancer immunotherapies (<xref ref-type="bibr" rid="B4">4</xref>&#x2013;<xref ref-type="bibr" rid="B8">8</xref>). The primary therapeutic mechanism of cancer vaccines is to &#x201c;educate&#x201d; the immune system to recognize and eliminate tumor cells as foreign substances. From the rationale above, the key of a vaccine design is to identify valuable antigens that can distinguish tumor cells from normal cells. Therefore, previous studies on therapeutic cancer vaccines have involved tumor-associated antigens (TAAs) (<xref ref-type="bibr" rid="B9">9</xref>&#x2013;<xref ref-type="bibr" rid="B12">12</xref>) and tumor-specific antigens (TSAs), that is, neoantigens (<xref ref-type="bibr" rid="B13">13</xref>&#x2013;<xref ref-type="bibr" rid="B16">16</xref>).</p>
<p>In the past decade, therapeutic cancer vaccines have achieved excellent clinical study results (<xref ref-type="bibr" rid="B17">17</xref>&#x2013;<xref ref-type="bibr" rid="B24">24</xref>). For instance, in patients with anti-PD1-refractory/relapsed unresectable Stage III or IV melanoma, BioNTech&#x2019;s therapeutic cancer vaccine candidate BNT111 combined with cemiplimab elicited durable objective responses (<xref ref-type="bibr" rid="B23">23</xref>), which received Food and Drug Administration (FDA) Fast Track Designation in 2021. In another study, treatment with dendritic cell vaccine primed with WT1 mRNA could prevent or delay relapse in 43% of patients with AML in remission after chemotherapy in a Phase II trial (<xref ref-type="bibr" rid="B25">25</xref>). These two clinical studies utilized therapeutic vaccines based on TAA.</p>
<p>Therapeutic vaccines based on personalized neoantigens also made remarkable progress. In a clinical trial on melanoma patients conducted by Otto et al., a synthetic long peptide vaccine consisting of multiple epitopes established tumor-specific T-cell responses and demonstrated effectiveness over five years (<xref ref-type="bibr" rid="B20">20</xref>, <xref ref-type="bibr" rid="B22">22</xref>, <xref ref-type="bibr" rid="B26">26</xref>). In end-stage colorectal cancer (CRC) patients (3rd line or more advanced), Gritstone&#x2019;s GRANITE personalized immunotherapy showed a 44% molecular response rate (4/9) by circulating tumor DNA analysis that can be considered as a surrogate endpoint [NCT03639714].</p>
<p>Theoretically, neoantigens are superior to TAA as targets for therapeutic cancer vaccine: although TAAs have relatively higher expression levels in tumor cells, they may still be present in particular types of normal cells at low levels; Her2 and survivin would be good examples (<xref ref-type="bibr" rid="B27">27</xref>&#x2013;<xref ref-type="bibr" rid="B30">30</xref>). In contrast, neoantigens originate from mutations and aberrant translations of tumor RNA transcriptome. Consequently, they are &#x201c;absolutely&#x201d; specific to tumor cells as normal cells do not have such mutations and aberrant translations. Such &#x201c;absolute&#x201d; specificity means that T-cell responses against neoantigens are unlikely to elicit an off-target effect on normal cells. Thus, the safety concern of neoantigen-based personalized vaccines would be minimal.</p>
<p>Despite its theoretical superiority on safety, neoantigen-based personalized vaccines still need to face a technical bottleneck: how to identify T-cell epitopes with sufficient immunogenicity from neoantigens for vaccine design. Especially, MHC-II epitopes are believed to be more necessary than MHC-I epitopes for preventing the immune escape of tumor cells (<xref ref-type="bibr" rid="B31">31</xref>, <xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B34">34</xref>). The insufficient capability of predicting MHC-II epitopes has obviously limited the development of neoantigen-based personalized vaccines, resulting in scarcely reported clinical studies involving MHC-II epitopes. Therefore, our research aims to break this bottleneck and provide a powerful tool for developing a neoantigen-based personalized vaccine.</p>
<p>Generally speaking, the conversion of aberrant peptides generated by genomic variations in tumor cells into epitopes eliciting <italic>in vivo</italic> T-cell immune responses is a complex process involving multiple hierarchical levels. Therefore, the prediction and identification of T-cell epitopes should preferably involve multiple levels to reflect complex biological processes. Such a &#x201c;funnel-like&#x201d; procedure (<xref ref-type="bibr" rid="B35">35</xref>&#x2013;<xref ref-type="bibr" rid="B37">37</xref>) that would eliminate most T-cell epitope candidates would necessarily involve several major steps:</p>
<list list-type="order">
<list-item>
<p>Mutation identification</p>
</list-item>
<list-item>
<p>Peptide&#x2013;MHC binding prediction</p>
</list-item>
<list-item>
<p>Peptide&#x2013;MHC presentation prediction</p>
</list-item>
<list-item>
<p>Peptide&#x2013;MHC immunogenicity prediction</p>
</list-item>
</list>
<p>Plenty of previous work has been accomplished by various research groups in the relevant field and thereafter generated several well-known software implements:</p>
<list list-type="bullet">
<list-item>
<p>The latest version of NetMHCIIpan uses binding and elution datasets deconvoluted by NNalign_MA (<xref ref-type="bibr" rid="B38">38</xref>) to predict peptide ligands that can be presented by MHC-I and MHC-II on the cell surface (<xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>).</p>
</list-item>
<list-item>
<p>MHCflurry improves the pan-allele prediction of MHC-I-presented peptide ligands by incorporating antigen processing and MHC ligandome elution (<xref ref-type="bibr" rid="B41">41</xref>).</p>
</list-item>
<list-item>
<p>ForestMHC applied the deconvolution of polyallelic datasets trained by MixMHCpred based on position weight matrices (PWMs) and MHC-I-presented peptide ligands (<xref ref-type="bibr" rid="B42">42</xref>).</p>
</list-item>
<list-item>
<p>MARIA adopts a multimodal recurrent neural network that summarizes <italic>in vitro</italic> binding measurements, mRNA abundance, and protease cleavage signatures to predict MHC-II-presented peptide ligands (<xref ref-type="bibr" rid="B43">43</xref>).</p>
</list-item>
</list>
<p>However, the well-known tools listed above never touched the 4th step of the funnel: immunogenicity. Considering the negative selection of T cells during thymus development (<xref ref-type="bibr" rid="B44">44</xref>, <xref ref-type="bibr" rid="B45">45</xref>), the vast majority of self-derived peptides will not trigger a downstream immune response even if presented by APC such as DC (<xref ref-type="bibr" rid="B46">46</xref>, <xref ref-type="bibr" rid="B47">47</xref>), and such peptides account for 90% of all presented peptides. Obviously, the current antigen presentation prediction tools are NOT the ultimate solutions for the design of neoantigen-based personalized vaccines because even the peptide ligands presented by MHC-I or MHC-II may not be immunogenic at all.</p>
<p>Recently, several emerging studies have taken MHC-I immunogenicity prediction into consideration. For example, deepHLApan incorporated both peptide&#x2013;MHC complex binding affinity and immunogenicity to predict the T-cell epitope (<xref ref-type="bibr" rid="B48">48</xref>). DeepNetBim extracted the attributes of the network as new features from peptide&#x2013;MHC binding and immunogenic models as a pan-specific MHC-I epitope prediction tool (<xref ref-type="bibr" rid="B49">49</xref>).</p>
<p>Nevertheless, there remains an unfilled gap in identifying MHC-II epitopes with sufficient immunogenicity, as neoantigen-driven B-cell and CD4+ T-helper cell collaboration promotes anti-tumor CD8 T-cell responses (<xref ref-type="bibr" rid="B50">50</xref>). In this work, we developed an overarching framework to predict MHC-II epitopes: our convolutional neural network (CNN) model predicts the probability of a peptide to be presented to the cell surface by a designated MHC subtype, as well as its immunogenicity to activate immune T cells. The overall research consists of the following parts:</p>
<list list-type="simple">
<list-item>
<p>(1) The datasets of peptide presentation and immunogenicity are obtained from an open database (IEDB) (<xref ref-type="bibr" rid="B51">51</xref>) and then processed with rigorous organization and cleaning.</p>
</list-item>
<list-item>
<p>(2) We constructed a semiotic-based human leukocyte antigen (HLA)-encoding method with three levels to associate the information of the HLA allele nomenclature, which better represents the characteristics of different MHC subtypes that are not entirely independent or discrete.</p>
</list-item>
<list-item>
<p>(3) The encoded MHC subtypes and peptides are integrated into the deep learning model based on a specially designed CNN.</p>
</list-item>
<list-item>
<p>(4) Independent validation datasets are used to evaluate the model&#x2019;s prediction performance.</p>
</list-item>
</list>
</sec>
<sec id="s2" sec-type="materials|methods">
<title>2 Materials and Methods</title>
<sec id="s2_1">
<title>2.1 Eluted Ligandome and Immunogenicity Data</title>
<p>The Eluted Ligandome date corresponding to various MHC-II subtypes is downloaded from the IEDB database; T-cell assay data reflecting the immunogenicity of peptides are extracted from the IEDB database. Python scripts are used to resolve raw XML data filtered with the following criteria:</p>
<list list-type="simple">
<list-item>
<p>(1) MHC-II alleles include HLA-DP, DQ, and DR &#x3b2; chains, whereas &#x3b1; chains are reasonably omitted as they contribute little to ligand specificity.</p>
</list-item>
<list-item>
<p>(2) Only MHC-II subtypes with explicit 2 fields in the HLA nomenclature such as HLA-DPB1*01:03 are retained.</p>
</list-item>
<list-item>
<p>(3) The peptide length is in the range of 9~25 amino acids, representing 98% of total peptides</p>
</list-item>
<list-item>
<p>(4) Peptide&#x2013;MHC pairs with controversial assay results are excluded.</p>
</list-item>
<list-item>
<p>(5) MHC-II subtypes with fewer than 10 corresponding peptides are excluded as the data size is too small to train our model, which leads to 65 available MHC-II subtypes.</p>
</list-item>
<list-item>
<p>(6) T-cell assay data are based on wet-lab assays rather than predictions in original dataset&#x2019;s column named Assay Type.</p>
</list-item>
</list>
</sec>
<sec id="s2_2">
<title>2.2 Negative Elution Training Data Generation</title>
<p>We generate the negative datasets corresponding to elution data treated as positive data from the global maximum dissimilarity scoring matrix based on sequence dissimilarity with an additional NetMHCIIpan binding filter:</p>
<list list-type="simple">
<list-item>
<p>(1) Full protein length F is extracted according to its accession ID (GenBank ID) given an eluted sequence P.</p>
</list-item>
<list-item>
<p>(2) We use a window with the same length of P to slide on the full-length sequence F to get a list of candidates from which 10 negative sequences with the lowest sequence similarity compared to the entire positive dataset and the lowest possibility to be eluted sequences calculated by NetMHCIIpan 4.0 as a filter.</p>
</list-item>
</list>
<p>In total, we obtained 273,102 non-redundant eluted ligands (as positive data) and corresponding to 61 MHC-II subtypes (<xref ref-type="supplementary-material" rid="ST1"><bold>Supplementary Table&#xa0;1</bold></xref>), amino acids frequency of most prevalence length of top 5 most corresponding restricted peptides of MHC-II subtypes is shown in <xref ref-type="supplementary-material" rid="SM1"><bold>Supplementary pdf</bold></xref>; 16,384 (10,131 positive and 6,253 negative) non-redundant T-cell assay data corresponding to 53 MHC-II subtypes (<xref ref-type="supplementary-material" rid="ST2"><bold>Supplementary Table&#xa0;2</bold></xref>).</p>
</sec>
<sec id="s2_3">
<title>2.3 MHC-II Subtype Encoding Based on Hierarchical Relationship</title>
<p>Antigen presentation and immunogenicity are both closely associated with MHC-II subtypes because peptide ligands are finally presented on the cell surface by MHC-II to T-cell receptors. In order to develop useful tools to predict MHC-II epitopes, we need to &#x201c;teach&#x201d; computer programs how to distinguish various MHC-II subtypes. Therefore, setting a reasonable coding method for MHC-II subtypes is an inevitable question. In quite a number of earlier studies, MHC-II subtypes are converted into orthogonal vectors using one-hot encoding. Although a one-hot coding approach is feasible and straightforward, it apparently does not fully reflect biological mechanisms. One-hot coding treats each MHC-II subtype as a unique dimension: for example, in the perspective of one-hot coding, HLA-DRB1*01:01 and HLA-DRB1*01:09 are assumed to have no relation at all, neither are their corresponding ligandomes. However, such an assumption conflicts with real-world biological mechanisms: the evolution of various MHC subtypes can be reflected in phylogenetic trees, and some MHC-II supertypes consisting of multiple subtypes have been characterized by a partially shared ligandome in previous studies.</p>
<p>As an imperfect approach, one-hot coding for MHC-II subtypes may waste lots of training data as it does not recognize the overlapping ligands of closely related MHC-II subtypes. Moreover, one-hot coding would cause the MHC-II subtypes without abundant training data (e.g., fewer than 10 corresponding peptides) to be neglected, as the segregated data amount may not be sufficient for training the model. In order to develop more powerful tools for predicting MHC-II epitopes, we propose a novel coding system that could quantitatively reflect the relation among various MHC-II subtypes. Our goal is to use training data in a more scientific way with maximal utilization and also enable epitope prediction for the MHC-II subtypes without many available data.</p>
<p>The nomenclature rationale of each HLA allele is like a leaf node based on a tree, which enriches the hierarchical information and truly reflects the categories and associations of different HLA alleles. We creatively propose a new HLA coding method named hierarchical relationship&#x2013;based HLA encoding, as shown in <xref ref-type="fig" rid="f1"><bold>Figure&#xa0;1</bold></xref>. In this model, we regard the HLA gene (HLA-DRB1, DPB1 and DQB1) as layer 0, the first field number (e.g., &#x2018;01&#x2019; of DRB1*01) as layer 1, and the second field number (e.g., &#x2018;02&#x2019; of DRB1*01:02) as layer 2. We encode each single layer according to an [99&#xd7;128] embedding table to get an <italic>E &#x2208; R</italic><sup>1&#xd7;128</sup> vector that represents each layer so that on the single layer, the same symbols have the same biological means while different symbols are discrete and orthogonal to each other mathematically. Afterwards, a transition matrix is adopted to transform the concatenated three-layer encoding matrix [3&#xd7;128] into a one-dimensional vector [1&#xd7;128] for later model training.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Schema of the MHC-II subtype hierarchical relationship encoding. Each layer representing one field of HLA name is converted into a [1&#xd7;128] vector and concatenated into a [3&#xd7;128] matrix, which is normalized and convoluted to get a [1&#xd7;128] vector for later calculation.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g001.tif"/>
</fig>
<disp-formula>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 1.</bold> <italic>embedding</italic><sub>0</sub>, <italic>embedding</italic><sub>1</sub>, <italic>embedding</italic><sub>2</sub>, represents the coding information of each layer, respectively; in each layer, the coding container size is [99&#xd7;128].</p>
<disp-formula>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mn>3</mml:mn>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 2.</bold> Convolution of feature extraction, <italic>W<sub>i</sub>
</italic> is the transfer matrix representing the weight of each layer. <italic>E<sub>i</sub>
</italic> is the information coding of each layer; <bold><italic>e</italic>
</bold> is the integrated HLA embedding value.</p>
</sec>
<sec id="s2_4">
<title>2.4 Normalization of HLA Embedding Value</title>
<p>The obtained HLA embedding value needs to be normalized before feeding to the deep learning model. Batch normalization (BN), a commonly used method, is used to normalize the whole batch of the dataset to a standard Gaussian distribution (<xref ref-type="bibr" rid="B52">52</xref>) so that differences in distinct data distribution from different samples can be normalized according to Equation 3:</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3c5;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mroot>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mroot>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 3.</bold> Batch normalization. <italic>v</italic> and <italic>&#x3c3;</italic><sup>2</sup> are the per-dimension mean and variance, respectively. Arbitrarily, the constant <italic>u</italic> is added in the denominator for numerical stability.</p>
<p>On the contrary, layer normalization (LN) normalizes all features of each sample in the sample scale (<xref ref-type="bibr" rid="B53">53</xref>) according to Equation4:</p>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mi>&#x3c5;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mroot>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mroot>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 4.</bold> Layer normalization <italic>v</italic> and <italic>&#x3c3;</italic><sup>2</sup> are the per-dimension mean and variance, respectively. Arbitrarily, constant <italic>u</italic> is added in the denominator for numerical stability for each single layer <italic>l</italic>.</p>
<p>Both batch normalization and layer normalization could be used to avoid gradient disappearance or gradient explosion caused by excessive fluctuation of the input value, so as to simplify subsequent model training. However, they still have substantial differences: batch normalization depends more on the statistical parameters between different samples; thus, feature extraction and normalization calculation within a single sample are insufficient, whereas layer normalization eliminates the characteristic relationship between different samples in a batch and only normalizes different eigenvalues in the same sample. Because of the reasons described above, both methods are not very suitable for current HLA embedding normalization. Because we need to consider not only the characteristics of the same layer but also the impact of differences at different layers, we developed a new method of HLA normalization (HLAN):</p>
<disp-formula>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>H</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>H</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mi>l</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>H</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mn>3</mml:mn>
</mml:munderover>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>H</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msubsup>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mroot>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mroot>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 5.</bold> HLA normalization equation <bold><italic>&#xb5;</italic>
</bold> is the mean value based on different levels, <italic>&#x3c3;</italic> is the level variance, <italic>x</italic> is the input value <italic>C</italic> is the layer according to the HLA-named system, and <italic>H</italic> is the length of the input value.</p>
<p>After normalization, the features are integrated through a convolution layer, and the final output results that are used as the input of the subsequent deep learning model are as follows:</p>
<disp-formula>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 6.</bold> Convolution layer to integrate an HLAN result.</p>
</sec>
<sec id="s2_5">
<title>2.5 HLA-Encoding Fusion Layer</title>
<p>We tested two different coding fusion layer schemas to fuse hierarchical representations from different layers representing an MHC-II subtype nomenclature in <xref ref-type="fig" rid="f2"><bold>Figures&#xa0;2A, B</bold></xref>. Considering that the numbers representing MHC-II subtypes are sparse in some datasets, inadequate training may occur during model training, we merged the embedding table at each level as shown in <xref ref-type="fig" rid="f2"><bold>Figure&#xa0;2C</bold></xref>; shared parameters are calculated as the same embedding table called HLA_Norm.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Schema of HLA-encoding fusion layer. <bold>(A)</bold> integrates sequence information by means of direct addition; <bold>(B)</bold> integrates the sequence information by means of concatenation; <bold>(C)</bold> shows that the sequence information is integrated by concatenation after the weight is processed by using the shared index embedding table.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g002.tif"/>
</fig>
</sec>
<sec id="s2_6">
<title>2.6 Variable Length Peptide Encoding</title>
<p>We first built a 21-character vocab, which uses J as the initial letter for the completion of the lengths of peptides, which are less than 25, plus 20 single-letter symbols of amino acids:</p>
<p>vocab = [&#x201c;J&#x201d;,&#x201c;A&#x201d;,&#x201c;C&#x201d;,&#x201c;D&#x201d;,&#x201c;E&#x201d;,&#x201c;F&#x201d;,&#x201c;G&#x201d;,&#x201c;H&#x201d;,&#x201c;I&#x201d;,&#x201c;K&#x201d;,&#x201c;L&#x201d;,&#x201c;M&#x201d;,&#x201c;N&#x201d;,&#x201c;P&#x201d;,&#x201c;Q&#x201d;,&#x201c;R&#x201d;,&#x201c;S&#x201d;,&#x201c;T&#x201d;,&#x201c;V&#x201d;,&#x201c;W&#x201d;,&#x201c;Y&#x201d;]</p>
<p>A [21&#xd7;128] size embedding table based on random normal distribution is developed according to the vocab shown in <xref ref-type="supplementary-material" rid="ST5"><bold>Supplementary Table&#xa0;5</bold></xref>.</p>
<p>For each input peptide sequence, we completed its length to 25 with the letter &#x201c;J&#x201d; and then converted each letter into a [1&#xd7;128] vector according to its position in vocab to tokenize the whole sequence and finally get a [25&#xd7;128] matrix presenting the input peptide.</p>
</sec>
<sec id="s2_7">
<title>2.7 HLA Subtype and Peptide Sequence Fusion Encoding</title>
<p>The MHC-II subtype and peptide sequence are paired, concatenated (<xref ref-type="fig" rid="f3"><bold>Figure&#xa0;3</bold></xref>), and sent to our model for further training and testing.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Schema of HLA&#x2013;peptide fusion encoding. The orange HLA_head is the result generated in <xref ref-type="fig" rid="f2"><bold>Figure&#xa0;2C</bold></xref>. HLA_head is loaded on top of peptide sequence in concatenation mode to form a new constituent sequence containing HLA_head and peptide.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g003.tif"/>
</fig>
<p>To characterize an HLA-peptide sequence after pairing, we need a unified model that can extract the paired information. Compared with the recurrent neural network, a full CNN can better model the information of adjacent positions. A one-dimensional CNN can be expressed as follows:</p>
<disp-formula>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>f</mml:mi>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>X</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msup>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
</mml:msup>
<mml:mo>&#x2217;</mml:mo>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mi>s</mml:mi>
</mml:msup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<p><bold>Equation 7.</bold> <italic>f</italic> is the convolution kernel, * is the convolution operator, and <italic>X</italic> is the input value. <italic>f<sup>c</sup>
</italic> is a one-dimensional convolution kernel of <bold><italic>c</italic>
</bold> dimension, and <italic>x<sup>s</sup>
</italic> is the input value decomposed according to its own dimension.</p>
</sec>
<sec id="s2_8">
<title>2.8 10-Fold Cross-Validation</title>
<p>Ten-fold cross-validation is applied to evaluate model robustness. Before training, the dataset is randomly partitioned into 10 non-overlapping subsets. The cross-validation process is repeated 10 times, with each subset used as a validation set while the remaining subsets are utilized as the training set. The results of the cross-validation sets are averaged to obtain the final result. One hundred epochs are executed, and the model is saved if the validation accuracy is better than previous epochs.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<sec id="s3_1">
<title>3.1 The Architecture of FIONA</title>
<p>We used the matrix <italic>p(x)</italic> obtained by matrix transformation in Equation 6 that converts a one-dimensional vector sequence of the MHC-II subtype and peptide into a [26&#xd7;128] matrix as input for the model to predict whether a peptide will be presented to the cell surface (FIONA-P) or trigger immunogenicity (FIONA-I) given a specific MHC subtype. In order to implement the above 2 predictive functions, we constructed two models with different training datasets (presentation and immunogenicity) explained in Section 2.1 and Section 2.2 with the same architecture shown in <xref ref-type="fig" rid="f4"><bold>Figure&#xa0;4</bold></xref>. FIONA includes a CNN layer for prediction, which focuses on integrating and extracting overall features from MHC-II subtype&#x2013;peptide pairs. In this process, HLA embedding and peptide embedding are integrated to play a synergistic role in improving the prediction performance. Additionally, in order to improve the prediction ability of our model, we added multiple pooling layers in the convolution layer to extract and integrate features.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Architecture of FIONA. The dataset is downloaded from the IEDB database according to Section 2.1. The first dotted box contains an HLA_head and peptide, which are encoded separately and then integrated for feature transformation. The middle part contains an HLA_head, and peptide_Embedding represents sequence feature information. The right dotted box is our CNN model for training and prediction.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g004.tif"/>
</fig>
<p>Our model accepts fused HLA_Peptide embedding as input. Referring to Resnet&#x2019;s design pattern shown in the left of <xref ref-type="fig" rid="f4"><bold>Figure&#xa0;4</bold></xref>, we created several aggregation modules in the form of blocks for stacked layers, connected layers, and convolution layers successively named BlockConv. The final convolution layer aggregates the internal characteristics of each embedding into a vector.</p>
</sec>
<sec id="s3_2">
<title>3.2 Ablation Experiment</title>
<p>We conducted ablation experiments to validate our HLA-encoding schema and its impact on the overall results by eliminating the HLA_Norm layer or replacing the normalization layer with batch normalization and layer normalization individually. We divided the comparison into two parts: the first part is based on HLA_ Embedding using different encoding and normalization methods, and the second part partially modifies the architecture of our model to find out the impacts of these modifications on the performance of our model.</p>
<p>As shown in <xref ref-type="table" rid="T1"><bold>Table&#xa0;1</bold></xref>, the ablation test shows that our MHC subtype hierarchical relationship&#x2013;encoding method greatly outperforms the traditional one-hot method regardless of subsequent normalization methods on both presentation data and immunogenicity data. In addition, the HLA_Norm method has the best performance on both presentation and immunogenicity datasets compared to the Batch Norm and Layer Norm. Meanwhile, the final architecture consisting of con_ANA and BlockConv has the best performance among all tests including eliminating the HLA coding content, which leads to a dramatic decrease of ROC and PR values.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Results of the ablation experiment. MSE (mean-squared error), AUC (area under the curve), and PR (precision rate) are evaluation indicators.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Method</th>
<th valign="top" colspan="3" align="center">MHC-II presentation</th>
<th valign="top" colspan="3" align="center">MHC-II immunogenicity</th>
</tr>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center">MSE (test)</th>
<th valign="top" align="center">AUC (test)</th>
<th valign="top" align="center">PR (test)</th>
<th valign="top" align="center">MSE (test)</th>
<th valign="top" align="center">AUC (test)</th>
<th valign="top" align="center">PR (test)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">PE+HLA_Norm+con_ANA+BLOCKConv</td>
<td valign="top" align="center"><bold>0.0421</bold>
</td>
<td valign="top" align="center"><bold>0.9391</bold>
</td>
<td valign="top" align="center"><bold>0.9513</bold>
</td>
<td valign="top" align="center"><bold>0.1819</bold>
</td>
<td valign="top" align="center"><bold>0.8876</bold>
</td>
<td valign="top" align="center"><bold>0.9344</bold>
</td>
</tr>
<tr>
<td valign="top" align="left">PE+Batch_Norm+con_ANA+BLOCKConv</td>
<td valign="top" align="center">0.0534</td>
<td valign="top" align="center">0.9242</td>
<td valign="top" align="center">0.9442</td>
<td valign="top" align="center">0.2049</td>
<td valign="top" align="center">0.8433</td>
<td valign="top" align="center">0.8839</td>
</tr>
<tr>
<td valign="top" align="left">PE+Layer_Norm+con_ANA+BLOCKConv</td>
<td valign="top" align="center">0.0610</td>
<td valign="top" align="center">0.9197</td>
<td valign="top" align="center">0.9328</td>
<td valign="top" align="center">0.2031</td>
<td valign="top" align="center">0.8340</td>
<td valign="top" align="center">0.8581</td>
</tr>
<tr>
<td valign="top" align="left">PE+HLA_Norm+add_ANA+BLOCKConv</td>
<td valign="top" align="center">0.0781</td>
<td valign="top" align="center">0.9038</td>
<td valign="top" align="center">0.9291</td>
<td valign="top" align="center">0.2274</td>
<td valign="top" align="center">0.8014</td>
<td valign="top" align="center">0.8230</td>
</tr>
<tr>
<td valign="top" align="left">PE+HLA_onehot+con_ANA+BLOCKConv</td>
<td valign="top" align="center">0.1042</td>
<td valign="top" align="center">0.8467</td>
<td valign="top" align="center">0.8835</td>
<td valign="top" align="center">0.2625</td>
<td valign="top" align="center">0.7637</td>
<td valign="top" align="center">0.7784</td>
</tr>
<tr>
<td valign="top" align="left">PE+BLOCKConv</td>
<td valign="top" align="center">0.2427</td>
<td valign="top" align="center">0.8046</td>
<td valign="top" align="center">0.8476</td>
<td valign="top" align="center">0.4691</td>
<td valign="top" align="center">0.5745</td>
<td valign="top" align="center">0.5872</td>
</tr>
<tr>
<td valign="top" align="left">PE+HLA_Norm+con_ANA+Conv</td>
<td valign="top" align="center">0.0578</td>
<td valign="top" align="center">0.8656</td>
<td valign="top" align="center">0.9103</td>
<td valign="top" align="center">0.2128</td>
<td valign="top" align="center">0.7877</td>
<td valign="top" align="center">0.8237</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>PE refers to the general peptide embedding, Batch_Norm, Layer_Norm, and HLA_Norm refer to the different HLA normalization methods described in section 2.4, while con_ANA is used to refer to the concatenate peptide_Embedding and HLA_Embedding header to get a [26&#xd7;128] matrix for the following step calculation, add_ANA refers to peptide_Embedding, and HLA_Embedding is processed by direct addition, which is mentioned in Section 2.5.</p>
</fn>
<fn>
<p>Bold means highlight superiority of our model.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_3">
<title>3.3 FIONA-P Favors Balanced Positive and Negative MHC-II Peptide Presentation Data</title>
<p>In a natural environment, the proportion of presented antigens compared to non-presented peptides degraded by protease is relatively low; therefore, the unbalanced data amount should theoretically and more faithfully reflect the actual situation. However, the unbalanced data amount of positive and negative samples is a great challenge to the construction and optimization of the deep learning model. Here, we selected a specific number of samples from multiple negative samples generated by the method mentioned in Section 2.2 to build 2 datasets with relatively balanced and unbalanced positive and negative ratios (positive data to negative data = 1:1 and 1:5, respectively) to compare the influence to our model FIONA-P. As shown in <xref ref-type="fig" rid="f5"><bold>Figure&#xa0;5</bold></xref>, FIONA-P has a better performance for balanced datasets, especially in terms of the performance of PR, which has a pronounced degradation if unbalanced data are used.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Influence of balanced and unbalanced data ratios on FIONA-P. <bold>(A, B)</bold> are the ROC (receiver operating characteristic) curve and PR (precision and recall) curve of unbalanced data (AUC=0.90, PR=0.90), respectively, while <bold>(C, D)</bold> are the ROC curve and PR curve of balanced data (AUC=0.94, PR=0.95), respectively.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g005.tif"/>
</fig>
<p>Similarly, we also compared the performance of natively uneven immunogenicity data from the IEDB and artificial synthetic datasets after randomly reducing a portion of the positive data (positive data to negative data = 1.62:1 and 1:1, respectively) to test the performance of the FIONA-I model under such circumstances. The test results show only minor changes in terms of ROC and PR.</p>
</sec>
<sec id="s3_4">
<title>3.4 FIONA-P Achieves Comparable Performance</title>
<p>The IEDB benchmark dataset is often used to compare the performance of different binding prediction tools. However, these datasets are usually intracellular binding data rather than elution data. To test the ability of the presentation prediction of several existing MHC-II epitope tools [Maria, NetMHCIIpan4.0, BERTMHC (<xref ref-type="bibr" rid="B54">54</xref>), and MixMHC2pred (<xref ref-type="bibr" rid="B55">55</xref>)], we used an independent dataset from the University of T&#xfc;bingen (<xref ref-type="bibr" rid="B56">56</xref>) that contains 142,625 naturally eluted ligands from 29 tissues across 42 MHC-II subtypes (33 MHC-II subtypes in total after omitting the &#x3b1; chains of MHC, <xref ref-type="supplementary-material" rid="ST3"><bold>Supplementary Table&#xa0;3</bold></xref>). The independent dataset is deduplicated by sequence and the corresponding MHC-II subtypes compared with the training dataset. All the supported MHC-II subtypes that overlap the MHC-II subtypes of the independent dataset are tested. For all tools, our FIONA-P model achieved the best performance for 25 out of the 33 MHC-II subtypes, especially in subtypes with higher corresponding eluted peptides as shown in <xref ref-type="fig" rid="f6"><bold>Figure&#xa0;6</bold></xref>. Our model has shown a bit of advancement compared with MixMHC2 and great improvement compared with other tools. However, MixMHC2 only supports 38 MHC-II subtypes; thus, 3 of unsupported MHC-II subtypes have no available results in this comparison. Our model not only supports 65 MHC-II subtypes by direct training but is also able to predict the peptide presentation of corresponding untrained MHC-II subtypes by our new breakthrough HLA hierarchical encoding method. Since the number of supported MHC-II subtypes is also very important in epitope prediction, our model has greatly broadened the scope of available MHC-II subtypes.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Comparison of FIONA-P and other prediction tools on the presentation data of all available MHC-II subtypes. The black ones indicate that those MHC-II subtypes are not supported.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-888556-g006.tif"/>
</fig>
</sec>
<sec id="s3_5">
<title>3.5 FIONA-I Improves Positive Prediction Value of True Neoantigen Through Validation of Curated Neoantigen Dataset</title>
<p>As previously discussed, only a small proportion of peptides presented by APC can trigger the downstream immunogenicity of T cells, resulting in a fairly low false-positive rate (FPR), which is presumably one of the main reasons that cancer vaccines do not have enough clinical benefits since these vaccines cannot load sufficient epitopes to inhibit the immune escape of cancer cells given such high FPR.</p>
<p>We used a fully manually annotated neoantigen database, NEPdb (<xref ref-type="bibr" rid="B57">57</xref>), newly published in 2021 to demonstrate that our FIONA-I model substantially improves the positive predictive value (PPV) of neoantigen prediction compared to the antigen presentation model. All MHC-II neoantigen data entries containing DP, DQ, and DR alleles were retrieved from the NEPdb, which contains 182 positive and 3,508 negative epitopes across 31 different MHC-II subtypes (<xref ref-type="supplementary-material" rid="ST4"><bold>Supplementary Table&#xa0;4</bold></xref>). FIONA-I, FIONA-P, and other MHC-II epitope tools (Maria, NetMHCIIpan4.0, BERTMHC, and MixMHC2pred) are used to calculate the PPV with a default parameter setting. Maria/BERTMHC directly returns &#x2018;0&#x2019; for negative and &#x2018;1&#x2019; for positive, NetMHCIIpan4.0 and MixMHC2pred take top 10% peptides as positive; all the MHC-II subtypes that are not supported by these tools are neglected. As shown in <xref ref-type="table" rid="T2"><bold>Table&#xa0;2</bold></xref>, FIONA-I raises the PPV from 22.51% (mean PPV of FIONA-P, Maria, NetMHCIIpan4.0, BERTMHC, and MixMHC2pred) to 40.27%, obtaining a near doubling of the increasement. The results showed that FIONA-I could improve the PPV significantly and retain the sensitivity at 0.89, indicating that the immunogenicity model could greatly contribute to high-confidence neoantigen identification.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Results of immunogenicity prediction of MHC-II-restricted epitopes in terms of sensitivity, specificity, and positive predictive value (PPV).</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Tools</th>
<th valign="top" colspan="3" align="center">MHC-II Immunogenicity</th>
</tr>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center">PPV</th>
<th valign="top" align="center">Sensitivity</th>
<th valign="top" align="center">Specificity</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">FIONA-I</td>
<td valign="top" align="center"><bold>0.4027</bold>
</td>
<td valign="top" align="center"><bold>0.8846</bold>
</td>
<td valign="top" align="center"><bold>0.9319</bold>
</td>
</tr>
<tr>
<td valign="top" align="left">FIONA-P</td>
<td valign="top" align="center">0.2188</td>
<td valign="top" align="center">0.7340</td>
<td valign="top" align="center">0.8640</td>
</tr>
<tr>
<td valign="top" align="left">NetMHCIIpan 4.0</td>
<td valign="top" align="center">0.1295</td>
<td valign="top" align="center">0.9271</td>
<td valign="top" align="center">0.6767</td>
</tr>
<tr>
<td valign="top" align="left">BERTMHC</td>
<td valign="top" align="center">0.1683</td>
<td valign="top" align="center">0.8093</td>
<td valign="top" align="center">0.7925</td>
</tr>
<tr>
<td valign="top" align="left">Maria</td>
<td valign="top" align="center">0.3279</td>
<td valign="top" align="center">0.7846</td>
<td valign="top" align="center">0.9166</td>
</tr>
<tr>
<td valign="top" align="left">MixMHC2pred</td>
<td valign="top" align="center">0.2812</td>
<td valign="top" align="center">0.7425</td>
<td valign="top" align="center">0.9032</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold means highlight superiority of our model.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_6">
<title>3.6 Web Service</title>
<p>We developed a user-friendly web interface (<uri xlink:href="http://therarna.cn/fiona.html">http://therarna.cn/fiona.html</uri>), allowing visitors to quickly query whether peptides would be presented or able to trigger an immune reaction given the specific MHC-II subtypes.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>Since 2018, a couple of useful tools for antigen presentation prediction have been reported. For example, Gritstone has published its proprietary software for predicting the MHC-I epitopes presented on the cell surface; similarly, MARIA is capable of predicting the MHC-II epitopes presented on the cell surface. Both tools use the same underlying hypothesis: antigen processing by proteases, antigen abundance, and peptide&#x2013;MHC interaction are 3 important factors that participate in antigen presentation. Therefore, both tools involved these 3 factors into the algorithm training and by far outperformed early versions of NetMHCIIpan and MHCFlurry.</p>
<p>The 3-factor theory of antigen presentation actually makes good scientific sense: short peptides need to be cleaved from long peptides by proteases to enable their binding to MHC-I or MHC-II; antigen abundance measured by the mRNA level determines the amount of MHC ligands displayed on the cell surface to be recognized by a T-cell receptor, whereas peptide&#x2013;MHC interaction would tell which ligands are more favorable to be displayed by MHC. However, there might be better ways to integrate these 3 factors into antigen presentation prediction. For instance, in the algorithm structure of MARIA, antigen presentation is simplified to a cleavage score, peptide&#x2013;MHC interaction is simplified to an HLA-DR binding score, and antigen abundance is standardized as mRNA TPM (transcript per million); afterwards, the 3 types of data from different dimensions were put into the algorithm training. We are not saying &#x201c;that approach is not right,&#x201d; but we seriously want to discuss what a better model should be. Biologically, long peptide cleavage by proteases occurs before short peptides interact with MHC. Therefore, it is more reasonable to develop a tool to enumerate all short peptides generated from a long peptide by protease cleavage, and the pool of short peptides would be the input of the next-step antigen presentation prediction. Moreover, in the antigen presentation process, antigen abundance would no longer be a limiting factor once it exceeds a reasonable level, which has been proven in the neoantigen meta-analysis of TESLA (<xref ref-type="bibr" rid="B58">58</xref>). Therefore, we may use TPM&gt;35 proposed by TESLA as a cut-off point of antigen abundance. In other words, the antigens whose expression levels are above the cut-off point should be regarded as &#x201c;abundant&#x201d; to be presented. Furthermore, in a natural infection caused by an exogenous virus or bacteria, all pathogen-related antigens should be regarded as &#x201c;abundant,&#x201d; even though their expression levels could hardly be standardized as TPM.</p>
<p>Based on the mechanistic analysis above, antigen processing had better been analyzed with a separate upstream tool, whereas antigen abundance could be reasonably simplified to a criterion of TPM &gt;35. Therefore, our antigen presentation prediction tool focuses more on peptide&#x2013;MHC interaction. We would not recommend MARIA&#x2019;s approach of oversimplifying the peptide&#x2013;MHC interaction to a binding score because the amino sequences of peptide ligands as well as MHC complex may reveal important information relevant to the antigen presentation process. For example, previous studies confirmed that peptide-MHC binding affinity reflected as IC50 (nM) does not accurately reflect the stability of the peptide&#x2013;MHC complex. Thus, the sequences of peptide ligands would provide information in more than one dimension. Taking all the foresaid into account, our antigen presentation prediction tool involves the sequences of peptide ligands and MHC into deep learning and therefore avoids the issue of oversimplification. This could be a possible explanation that our model outperforms the well-known tools.</p>
<p>Compared to antigen presentation, predicting the immunogenicity of MHC ligands is more challenging due to the lack of powerful theories. As previously discussed, there is a scientifically sound 3-factor theory that explains the mechanism of antigen presentation, and this theory effectively guided the development of multiple prediction tools. In contrast, the root cause of immunogenicity is more difficult to interpret.</p>
<p>Immunogenicity is shaped by a T-cell-negative selection; thus, the real challenge of immunogenicity prediction is the limited understanding of the mechanism of a T-cell-negative selection. A T-cell-negative selection process in thymus removes T cells reactive to self-antigens from the T-cell repertoire and therefore provides protection against unwanted T-cell responses. A T-cell-negative selection determines which MHC ligands will NOT elicit an immune response, whereas other MHC ligands may still encounter the corresponding TCR in the T-cell repertoire.</p>
<p>So far, the T-cell-negative selection process is still a &#x201c;black-box,&#x201d; and there is no powerful theory that clearly interprets its delicate mechanism. Especially, no theory could define what factors constitute the &#x201c;sufficient condition&#x201d; to trigger a T-cell-negative selection. At least, self-antigen alone does not constitute the &#x201c;sufficient condition.&#x201d; A T-cell-negative selection does NOT remove all T cells that recognize the MHC ligands derived from self-antigens, and such complexity is endorsed by 2 facts in clinical studies:</p>
<list list-type="order">
<list-item>
<p>Self-reactive T cells are present in patients with autoimmune diseases (<xref ref-type="bibr" rid="B59">59</xref>).</p>
</list-item>
<list-item>
<p>A peptide vaccine could elicit T-cell responses against TAA in cancer patients (<xref ref-type="bibr" rid="B2">2</xref>).</p>
</list-item>
</list>
<p>The lack of a robust theory to interpret a T-cell-negative selection makes it challenging to predict immunogenicity. All software tools for predicting immunogenicity, including ours, are based on an empirical approach: the tools are trained with T-cell assay data that distinguish immunogenic peptides from non-immunogenic ones, matched with MHC subtypes. Of course, even an empirical approach could solve many problems. For example, our trained software could achieve PPV at 40.27% on an independent dataset. Nevertheless, the limitation of the empirical approach should not be forgotten: such methodology requires tremendous T-cell assay data to train a functional model. For those MHC subtypes that do not have many corresponding T-cell assay results, the empirical approach cannot be used. Based on our discussion above, a more accurate immunogenicity prediction tool would rely on the emergence of a more robust theory that interprets the mechanism of the T-cell- negative selection mechanism. By then, it might be possible to deduce the immunogenicity of peptide ligands based on the host&#x2019;s MHC genotype and proteome information.</p>
<p>Our study proposed a systematic workflow that could identify MHC-II restricted epitopes that can be presented on the cell surface and elicit immune responses. This tool could be of great usefulness for identifying potential epitopes from cancer neoantigens and paving the way of designing effective cancer therapeutic vaccines.</p>
</sec>
<sec id="s5" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="SM1"><bold>Supplementary Material</bold></xref>. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author Contributions</title>
<p>SX is the first author of this article. SX and CF designed the concept and experiments. XW performed the ablation experiments. SX and CF prepared the data for training. CF implemented the negative data generation algorithm. XW implemented the CNN model. CF prepared the IEDB data and plotted the final figures and tables. CF and XW performed statistical analysis. SX and CF wrote the paper with input from XW. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s7" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>Authors SX, XW, and CF were employed by the company Nanjing Chengshi BioTech (TheraRNA) Co., Ltd.</p>
</sec>
<sec id="s8" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="s9" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fonc.2022.888556/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fonc.2022.888556/full#supplementary-material</ext-link></p>
<p>All the training and validation datasets mentioned in this work including human MHC-II elution and immunogenicity data from IEDB plus individual datasets from T&#xfc;bingen and NEPdb are <xref ref-type="supplementary-material" rid="ST4"><bold>Supplementary Tables&#xa0;1&#x2013;4</bold></xref> respectively. Full table of initial table of peptide encoding and WebLogo of sequence motifs of top 10 MHC-II subtypes corresponding peptides is shown in <xref ref-type="supplementary-material" rid="SM1"><bold>Supplementary pdf file</bold></xref>.</p>
<supplementary-material xlink:href="Image_1.pdf" id="SM1" mimetype="application/pdf"/>
<supplementary-material xlink:href="Table_1.xlsx" id="ST1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
<supplementary-material xlink:href="Table_2.xlsx" id="ST2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
<supplementary-material xlink:href="Table_3.csv" id="ST3" mimetype="text/csv"/>
<supplementary-material xlink:href="DataSheet_2.zip" id="ST4" mimetype="application/zip"/>
<supplementary-material xlink:href="DataSheet_3.zip" id="ST5" mimetype="application/zip"/>

</sec>
<sec id="s10">
<title>Abbreviations</title>
<p>APC, antigen-presenting cell; MHC, major histocompatibility complex; HLA, human leukocyte antigen; CNN, convolutional neural network; CRC, colorectal cancer; TAA, tumor-associated antigens; TSA, tumor-specific antigens; AML, acute myeloid leukemia; DC, dendritic cell; TCR, T-cell receptor; ROC curve, receiver operating characteristic curve; AUC, area under the curve; PR curve, precision and recall curve.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Robbins</surname> <given-names>PF</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Y-C</given-names>
</name>
<name>
<surname>El-Gamil</surname> <given-names>M</given-names>
</name>
<name>
<surname>Li</surname> <given-names>YF</given-names>
</name>
<name>
<surname>Gross</surname> <given-names>C</given-names>
</name>
<name>
<surname>Gartner</surname> <given-names>J</given-names>
</name>
<etal/>
</person-group>. <article-title>Mining Exomic Sequencing Data to Identify Mutated Antigens Recognized by Adoptively Transferred Tumor-Reactive T Cells</article-title>. <source>Nat Med</source> (<year>2013</year>) <volume>19</volume>:<fpage>747</fpage>. doi: <pub-id pub-id-type="doi">10.1038/nm.3161</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Anguille</surname> <given-names>S</given-names>
</name>
<name>
<surname>Van de Velde</surname> <given-names>AL</given-names>
</name>
<name>
<surname>Smits</surname> <given-names>EL</given-names>
</name>
<name>
<surname>Van Tendeloo</surname> <given-names>VF</given-names>
</name>
<name>
<surname>Juliusson</surname> <given-names>G</given-names>
</name>
<name>
<surname>Cools</surname> <given-names>N</given-names>
</name>
<etal/>
</person-group>. <article-title>Dendritic Cell Vaccination as Postremission Treatment to Prevent or Delay Relapse in Acute Myeloid Leukemia</article-title>. <source>Blood</source> (<year>2017</year>) <volume>130</volume>:<page-range>1713&#x2013;21</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1182/blood-2017-04-780155</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>van Rooij</surname> <given-names>N</given-names>
</name>
<name>
<surname>van Buuren</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Philips</surname> <given-names>D</given-names>
</name>
<name>
<surname>Velds</surname> <given-names>A</given-names>
</name>
<name>
<surname>Toebes</surname> <given-names>M</given-names>
</name>
<name>
<surname>Heemskerk</surname> <given-names>B</given-names>
</name>
<etal/>
</person-group>. <article-title>Tumor Exome Analysis Reveals Neoantigen-Specific T-Cell Reactivity in an Ipilimumab-Responsive Melanoma</article-title>. <source>J Clin Oncol Off J Am Soc Clin Oncol</source> (<year>2013</year>) <volume>31</volume>:<elocation-id>e439-e442</elocation-id>. doi: <pub-id pub-id-type="doi">10.1200/JCO.2012.47.7521</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shrimali</surname> <given-names>RK</given-names>
</name>
<name>
<surname>Ahmad</surname> <given-names>S</given-names>
</name>
<name>
<surname>Verma</surname> <given-names>V</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>P</given-names>
</name>
<name>
<surname>Ananth</surname> <given-names>S</given-names>
</name>
<name>
<surname>Gaur</surname> <given-names>P</given-names>
</name>
<etal/>
</person-group>. <article-title>Concurrent PD-1 Blockade Negates the Effects of OX40 Agonist Antibody in Combination Immunotherapy Through Inducing T-Cell Apoptosis</article-title>. <source>Cancer Immunol Res</source> (<year>2017</year>) <volume>5</volume>:<page-range>755&#x2013;66</page-range>. doi: <pub-id pub-id-type="doi">10.1158/2326-6066.CIR-17-0292</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Parsons</surname> <given-names>BL</given-names>
</name>
</person-group>. <article-title>Many Different Tumor Types Have Polyclonal Tumor Origin: Evidence and Implications</article-title>. <source>Mutat Res Mutat Res</source> (<year>2008</year>) <volume>659</volume>:<page-range>232&#x2013;47</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.mrrev.2008.05.004</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X</given-names>
</name>
</person-group>. <article-title>PD-1 and Its Ligands are Important Immune Checkpoints in Cancer</article-title>. <source>Oncotarget</source> (<year>2017</year>) <volume>8</volume>:<fpage>2171</fpage>. doi: <pub-id pub-id-type="doi">10.18632/oncotarget.13895</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mellman</surname> <given-names>I</given-names>
</name>
<name>
<surname>Coukos</surname> <given-names>G</given-names>
</name>
<name>
<surname>Dranoff</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>Cancer Immunotherapy Comes of Age</article-title>. <source>Nature</source> (<year>2011</year>) <volume>480</volume>:<page-range>480&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nature10673</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dougan</surname> <given-names>M</given-names>
</name>
<name>
<surname>Dranoff</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>Immune Therapy for Cancer</article-title>. <source>Annu Rev Immunol</source> (<year>2009</year>) <volume>27</volume>:<fpage>83</fpage>&#x2013;<lpage>117</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.immunol.021908.132544</pub-id>
</citation>
</ref>
<ref id="B9">
<label>9</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>PP</given-names>
</name>
<name>
<surname>Yee</surname> <given-names>C</given-names>
</name>
<name>
<surname>Savage</surname> <given-names>PA</given-names>
</name>
<name>
<surname>Fong</surname> <given-names>L</given-names>
</name>
<name>
<surname>Brockstedt</surname> <given-names>D</given-names>
</name>
<name>
<surname>Weber</surname> <given-names>JS</given-names>
</name>
<etal/>
</person-group>. <article-title>Characterization of Circulating T Cells Specific for Tumor-Associated Antigens in Melanoma Patients</article-title>. <source>Nat Med</source> (<year>1999</year>) <volume>5</volume>:<page-range>677&#x2013;85</page-range>. doi: <pub-id pub-id-type="doi">10.1038/9525</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Criscitiello</surname> <given-names>C</given-names>
</name>
</person-group>. <article-title>Tumor-Associated Antigens in Breast Cancer</article-title>. <source>Breast Care</source> (<year>2012</year>) <volume>7</volume>:<page-range>262&#x2013;6</page-range>. doi: <pub-id pub-id-type="doi">10.1159/000342164</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Higgins</surname> <given-names>JP</given-names>
</name>
<name>
<surname>Bernstein</surname> <given-names>MB</given-names>
</name>
<name>
<surname>Hodge</surname> <given-names>JW</given-names>
</name>
</person-group>. <article-title>Enhancing Immune Responses to Tumor-Associated Antigens</article-title>. <source>Cancer Biol Ther</source> (<year>2009</year>) <volume>8</volume>:<page-range>1440&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.4161/cbt.8.15.9133</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>W</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>B</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>W</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>W</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J-Y</given-names>
</name>
</person-group>. <article-title>Autoantibodies to Tumor-Associated Antigens as Biomarkers in Cancer Immunodiagnosis</article-title>. <source>Autoimmun Rev</source> (<year>2011</year>) <volume>10</volume>:<page-range>331&#x2013;5</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.autrev.2010.12.002</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Linnebacher</surname> <given-names>M</given-names>
</name>
<name>
<surname>Gebert</surname> <given-names>J</given-names>
</name>
<name>
<surname>Rudy</surname> <given-names>W</given-names>
</name>
<name>
<surname>Woerner</surname> <given-names>S</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>YP</given-names>
</name>
<name>
<surname>Bork</surname> <given-names>P</given-names>
</name>
<etal/>
</person-group>. <article-title>Frameshift Peptide-Derived T-Cell Epitopes: A Source of Novel Tumor-Specific Antigens</article-title>. <source>Int J Cancer</source> (<year>2001</year>) <volume>93</volume>:<fpage>6</fpage>&#x2013;<lpage>11</lpage>. doi: <pub-id pub-id-type="doi">10.1002/ijc.1298</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Laumont</surname> <given-names>CM</given-names>
</name>
<name>
<surname>Vincent</surname> <given-names>K</given-names>
</name>
<name>
<surname>Hesnard</surname> <given-names>L</given-names>
</name>
<name>
<surname>Audemard</surname> <given-names>&#xc9;</given-names>
</name>
<name>
<surname>Bonneil</surname> <given-names>&#xc9;</given-names>
</name>
<name>
<surname>Laverdure</surname> <given-names>J-P</given-names>
</name>
<etal/>
</person-group>. <article-title>Noncoding Regions are the Main Source of Targetable Tumor-Specific Antigens</article-title>. <source>Sci Transl Med</source> (<year>2018</year>) <volume>10</volume>(<issue>470</issue>). doi: <pub-id pub-id-type="doi">10.1126/scitranslmed.aau5516</pub-id>
</citation>
</ref>
<ref id="B15">
<label>15</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Apavaloaei</surname> <given-names>A</given-names>
</name>
<name>
<surname>Hardy</surname> <given-names>M-P</given-names>
</name>
<name>
<surname>Thibault</surname> <given-names>P</given-names>
</name>
<name>
<surname>Perreault</surname> <given-names>C</given-names>
</name>
</person-group>. <article-title>The Origin and Immune Recognition of Tumor-Specific Antigens</article-title>. <source>Cancers</source> (<year>2020</year>) <volume>12</volume>:<fpage>2607</fpage>. doi: <pub-id pub-id-type="doi">10.3390/cancers12092607</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boon</surname> <given-names>T</given-names>
</name>
<name>
<surname>Coulie</surname> <given-names>PG</given-names>
</name>
<name>
<surname>Van den Eynde</surname> <given-names>B</given-names>
</name>
</person-group>. <article-title>Tumor Antigens Recognized by T Cells</article-title>. <source>Immunol Today</source> (<year>1997</year>) <volume>18</volume>:<page-range>267&#x2013;8</page-range>. doi: <pub-id pub-id-type="doi">10.1016/S0167-5699(97)80020-5</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Martin</surname> <given-names>SD</given-names>
</name>
<name>
<surname>Brown</surname> <given-names>SD</given-names>
</name>
<name>
<surname>Wick</surname> <given-names>DA</given-names>
</name>
<name>
<surname>Nielsen</surname> <given-names>JS</given-names>
</name>
<name>
<surname>Kroeger</surname> <given-names>DR</given-names>
</name>
<name>
<surname>Twumasi-Boateng</surname> <given-names>K</given-names>
</name>
<etal/>
</person-group>. <article-title>Low Mutation Burden in Ovarian Cancer May Limit the Utility of Neoantigen-Targeted Vaccines</article-title>. <source>PloS One</source> (<year>2016</year>) <volume>11</volume>:<elocation-id>e0155189</elocation-id>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0155189</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hilf</surname> <given-names>N</given-names>
</name>
<name>
<surname>Kuttruff-Coqui</surname> <given-names>S</given-names>
</name>
<name>
<surname>Frenzel</surname> <given-names>K</given-names>
</name>
<name>
<surname>Bukur</surname> <given-names>V</given-names>
</name>
<name>
<surname>Stevanovi&#x107;</surname> <given-names>S</given-names>
</name>
<name>
<surname>Gouttefangeas</surname> <given-names>C</given-names>
</name>
<etal/>
</person-group>. <article-title>Actively Personalized Vaccination Trial for Newly Diagnosed Glioblastoma</article-title>. <source>Nature</source> (<year>2019</year>) <volume>565</volume>:<page-range>240&#x2013;5</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41586-018-0810-y</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carreno</surname> <given-names>BM</given-names>
</name>
<name>
<surname>Magrini</surname> <given-names>V</given-names>
</name>
<name>
<surname>Becker-Hapak</surname> <given-names>M</given-names>
</name>
<name>
<surname>Kaabinejadian</surname> <given-names>S</given-names>
</name>
<name>
<surname>Hundal</surname> <given-names>J</given-names>
</name>
<name>
<surname>Petti</surname> <given-names>AA</given-names>
</name>
<etal/>
</person-group>. <article-title>A Dendritic Cell Vaccine Increases the Breadth and Diversity of Melanoma Neoantigen-Specific T Cells</article-title>. <source>Science</source> (<year>2015</year>) <volume>348</volume>:<page-range>803&#x2013;8</page-range>. doi: <pub-id pub-id-type="doi">10.1126/science.aaa3828</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ott</surname> <given-names>PA</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Keskin</surname> <given-names>DB</given-names>
</name>
<name>
<surname>Shukla</surname> <given-names>SA</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
<name>
<surname>Bozym</surname> <given-names>DJ</given-names>
</name>
<etal/>
</person-group>. <article-title>An Immunogenic Personal Neoantigen Vaccine for Patients With Melanoma</article-title>. <source>Nature</source> (<year>2017</year>) <volume>547</volume>:<page-range>217&#x2013;21</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nature22991</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kreiter</surname> <given-names>S</given-names>
</name>
<name>
<surname>Vormehr</surname> <given-names>M</given-names>
</name>
<name>
<surname>Van de Roemer</surname> <given-names>N</given-names>
</name>
<name>
<surname>Diken</surname> <given-names>M</given-names>
</name>
<name>
<surname>L&#xf6;wer</surname> <given-names>M</given-names>
</name>
<name>
<surname>Diekmann</surname> <given-names>J</given-names>
</name>
<etal/>
</person-group>. <article-title>Mutant MHC Class II Epitopes Drive Therapeutic Immune Responses to Cancer</article-title>. <source>Nature</source> (<year>2015</year>) <volume>520</volume>:<page-range>692&#x2013;6</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nature14426</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Keskin</surname> <given-names>DB</given-names>
</name>
<name>
<surname>Anandappa</surname> <given-names>AJ</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
<name>
<surname>Tirosh</surname> <given-names>I</given-names>
</name>
<name>
<surname>Mathewson</surname> <given-names>ND</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S</given-names>
</name>
<etal/>
</person-group>. <article-title>Neoantigen Vaccine Generates Intratumoral T Cell Responses in Phase Ib Glioblastoma Trial</article-title>. <source>Nature</source> (<year>2019</year>) <volume>565</volume>:<page-range>234&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41586-018-0792-9</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sahin</surname> <given-names>U</given-names>
</name>
<name>
<surname>Oehm</surname> <given-names>P</given-names>
</name>
<name>
<surname>Derhovanessian</surname> <given-names>E</given-names>
</name>
<name>
<surname>Jabulowsky</surname> <given-names>RA</given-names>
</name>
<name>
<surname>Vormehr</surname> <given-names>M</given-names>
</name>
<name>
<surname>Gold</surname> <given-names>M</given-names>
</name>
<etal/>
</person-group>. <article-title>An RNA Vaccine Drives Immunity in Checkpoint-Inhibitor-Treated Melanoma</article-title>. <source>Nature</source> (<year>2020</year>) <volume>585</volume>:<page-range>107&#x2013;12</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41586-020-2537-9</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sahin</surname> <given-names>U</given-names>
</name>
<name>
<surname>Derhovanessian</surname> <given-names>E</given-names>
</name>
<name>
<surname>Miller</surname> <given-names>M</given-names>
</name>
<name>
<surname>Kloke</surname> <given-names>B-P</given-names>
</name>
<name>
<surname>Simon</surname> <given-names>P</given-names>
</name>
<name>
<surname>L&#xf6;wer</surname> <given-names>M</given-names>
</name>
<etal/>
</person-group>. <article-title>Personalized RNA Mutanome Vaccines Mobilize Poly-Specific Therapeutic Immunity Against Cancer</article-title>. <source>Nature</source> (<year>2017</year>) <volume>547</volume>:<page-range>222&#x2013;6</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nature23003</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Walters</surname> <given-names>JN</given-names>
</name>
<name>
<surname>Ferraro</surname> <given-names>B</given-names>
</name>
<name>
<surname>Duperret</surname> <given-names>EK</given-names>
</name>
<name>
<surname>Kraynyak</surname> <given-names>KA</given-names>
</name>
<name>
<surname>Chu</surname> <given-names>J</given-names>
</name>
<name>
<surname>Saint-Fleur</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>A Novel DNA Vaccine Platform Enhances Neo-Antigen-Like T Cell Responses Against WT1 to Break Tolerance and Induce Anti-Tumor Immunity</article-title>. <source>Mol Ther</source> (<year>2017</year>) <volume>25</volume>:<page-range>976&#x2013;88</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.ymthe.2017.01.022</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Leet</surname> <given-names>DE</given-names>
</name>
<name>
<surname>Alles&#xf8;e</surname> <given-names>RL</given-names>
</name>
<name>
<surname>Oliveira</surname> <given-names>G</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S</given-names>
</name>
<name>
<surname>Luoma</surname> <given-names>AM</given-names>
</name>
<etal/>
</person-group>. <article-title>Personal Neoantigen Vaccines Induce Persistent Memory T Cell Responses and Epitope Spreading in Patients With Melanoma</article-title>. <source>Nat Med</source> (<year>2021</year>) <volume>27</volume>:<page-range>515&#x2013;25</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41591-020-01206-4</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Natali</surname> <given-names>PG</given-names>
</name>
<name>
<surname>Nicotra</surname> <given-names>MR</given-names>
</name>
<name>
<surname>Bigotti</surname> <given-names>A</given-names>
</name>
<name>
<surname>Venturo</surname> <given-names>I</given-names>
</name>
<name>
<surname>Slamon</surname> <given-names>DJ</given-names>
</name>
<name>
<surname>Fendly</surname> <given-names>BM</given-names>
</name>
<etal/>
</person-group>. <article-title>Expression of the P185 Encoded by HER2 Oncogene in Normal and Transformed Human Tissues</article-title>. <source>Int J Cancer</source> (<year>1990</year>) <volume>45</volume>:<page-range>457&#x2013;61</page-range>. doi: <pub-id pub-id-type="doi">10.1002/ijc.2910450314</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname> <given-names>W</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>L</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>J</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>J</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>T</given-names>
</name>
<name>
<surname>Pang</surname> <given-names>L</given-names>
</name>
<etal/>
</person-group>. <article-title>The Molecular Mechanism and Clinical Significance of LDHA in HER2-Mediated Progression of Gastric Cancer</article-title>. <source>Am J Transl Res</source> (<year>2018</year>) <volume>10</volume>:<fpage>2055</fpage>.</citation>
</ref>
<ref id="B29">
<label>29</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fukuda</surname> <given-names>S</given-names>
</name>
<name>
<surname>Pelus</surname> <given-names>LM</given-names>
</name>
</person-group>. <article-title>Survivin, a Cancer Target With an Emerging Role in Normal Adult Tissues</article-title>. <source>Mol Cancer Ther</source> (<year>2006</year>) <volume>5</volume>:<page-range>1087&#x2013;98</page-range>. doi: <pub-id pub-id-type="doi">10.1158/1535-7163.MCT-05-0375</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>D</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>C</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H</given-names>
</name>
</person-group>. <article-title>Survivin as a Novel Target Protein for Reducing the Proliferation of Cancer Cells</article-title>. <source>BioMed Rep</source> (<year>2018</year>) <volume>8</volume>:<fpage>399</fpage>&#x2013;<lpage>406</lpage>. doi: <pub-id pub-id-type="doi">10.3892/br.2018.1077</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>MHC Class II&#x2013;restricted Tumor Antigens Recognized by CD4+ T Cells: New Strategies for Cancer Vaccine Design</article-title>. <source>J Immunother</source> (<year>2001</year>) <volume>24</volume>:<fpage>195</fpage>&#x2013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1097/00002371-200105000-00002</pub-id>
</citation>
</ref>
<ref id="B32">
<label>32</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dolan</surname> <given-names>BP</given-names>
</name>
<name>
<surname>Gibbs</surname> <given-names>KD</given-names>
</name>
<name>
<surname>Ostrand-Rosenberg</surname> <given-names>S</given-names>
</name>
</person-group>. <article-title>Tumor-Specific CD4+ T Cells are Activated by &#x201c;Cross-Dressed&#x201d; Dendritic Cells Presenting Peptide-MHC Class II Complexes Acquired From Cell-Based Cancer Vaccines</article-title>. <source>J Immunol</source> (<year>2006</year>) <volume>176</volume>:<page-range>1447&#x2013;55</page-range>. doi: <pub-id pub-id-type="doi">10.4049/jimmunol.176.3.1447</pub-id>
</citation>
</ref>
<ref id="B33">
<label>33</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lybaert</surname> <given-names>L</given-names>
</name>
</person-group>. <article-title>Immunosurveillance and the Importance of CD4 T Cells in Developing Cancer Vaccines</article-title>. (<year>2021</year>).</citation>
</ref>
<ref id="B34">
<label>34</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oh</surname> <given-names>DY</given-names>
</name>
<name>
<surname>Kwek</surname> <given-names>SS</given-names>
</name>
<name>
<surname>Raju</surname> <given-names>SS</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T</given-names>
</name>
<name>
<surname>McCarthy</surname> <given-names>E</given-names>
</name>
<name>
<surname>Chow</surname> <given-names>E</given-names>
</name>
<etal/>
</person-group>. <article-title>Intratumoral CD4+ T Cells Mediate Anti-Tumor Cytotoxicity in Human Bladder Cancer</article-title>. <source>Cell</source> (<year>2020</year>) <volume>181</volume>:<page-range>1612&#x2013;25</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2020.05.017</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schumacher</surname> <given-names>TN</given-names>
</name>
<name>
<surname>Schreiber</surname> <given-names>RD</given-names>
</name>
</person-group>. <article-title>Neoantigens in Cancer Immunotherapy</article-title>. <source>Science</source> (<year>2015</year>) <volume>348</volume>:<fpage>69</fpage>&#x2013;<lpage>74</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.aaa4971</pub-id>
</citation>
</ref>
<ref id="B36">
<label>36</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yadav</surname> <given-names>M</given-names>
</name>
<name>
<surname>Jhunjhunwala</surname> <given-names>S</given-names>
</name>
<name>
<surname>Phung</surname> <given-names>QT</given-names>
</name>
<name>
<surname>Lupardus</surname> <given-names>P</given-names>
</name>
<name>
<surname>Tanguay</surname> <given-names>J</given-names>
</name>
<name>
<surname>Bumbaca</surname> <given-names>S</given-names>
</name>
<etal/>
</person-group>. <article-title>Predicting Immunogenic Tumour Mutations by Combining Mass Spectrometry and Exome Sequencing</article-title>. <source>Nature</source> (<year>2014</year>) <volume>515</volume>:<page-range>572&#x2013;6</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/nature14001</pub-id>
</citation>
</ref>
<ref id="B37">
<label>37</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rizvi</surname> <given-names>NA</given-names>
</name>
<name>
<surname>Hellmann</surname> <given-names>MD</given-names>
</name>
<name>
<surname>Snyder</surname> <given-names>A</given-names>
</name>
<name>
<surname>Kvistborg</surname> <given-names>P</given-names>
</name>
<name>
<surname>Makarov</surname> <given-names>V</given-names>
</name>
<name>
<surname>Havel</surname> <given-names>JJ</given-names>
</name>
<etal/>
</person-group>. <article-title>Mutational Landscape Determines Sensitivity to PD-1 Blockade in Non&#x2013;Small Cell Lung Cancer</article-title>. <source>Science</source> (<year>2015</year>) <volume>348</volume>:<page-range>124&#x2013;8</page-range>. doi: <pub-id pub-id-type="doi">10.1126/science.aaa1348</pub-id>
</citation>
</ref>
<ref id="B38">
<label>38</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nielsen</surname> <given-names>M</given-names>
</name>
<name>
<surname>Andreatta</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>NNAlign: A Platform to Construct and Evaluate Artificial Neural Network Models of Receptor&#x2013;Ligand Interactions</article-title>. <source>Nucleic Acids Res</source> (<year>2017</year>) <volume>45</volume>:<page-range>W344&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkx276</pub-id>
</citation>
</ref>
<ref id="B39">
<label>39</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jurtz</surname> <given-names>V</given-names>
</name>
<name>
<surname>Paul</surname> <given-names>S</given-names>
</name>
<name>
<surname>Andreatta</surname> <given-names>M</given-names>
</name>
<name>
<surname>Marcatili</surname> <given-names>P</given-names>
</name>
<name>
<surname>Peters</surname> <given-names>B</given-names>
</name>
<name>
<surname>Nielsen</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>NetMHCpan-4.0: Improved Peptide&#x2013;MHC Class I Interaction Predictions Integrating Eluted Ligand and Peptide Binding Affinity Data</article-title>. <source>J Immunol</source> (<year>2017</year>) <volume>199</volume>:<page-range>3360&#x2013;8</page-range>. doi: <pub-id pub-id-type="doi">10.4049/jimmunol.1700893</pub-id>
</citation>
</ref>
<ref id="B40">
<label>40</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Reynisson</surname> <given-names>B</given-names>
</name>
<name>
<surname>Alvarez</surname> <given-names>B</given-names>
</name>
<name>
<surname>Paul</surname> <given-names>S</given-names>
</name>
<name>
<surname>Peters</surname> <given-names>B</given-names>
</name>
<name>
<surname>Nielsen</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>NetMHCpan-4.1 and NetMHCIIpan-4.0: Improved Predictions of MHC Antigen Presentation by Concurrent Motif Deconvolution and Integration of MS MHC Eluted Ligand Data</article-title>. <source>Nucleic Acids Res</source> (<year>2020</year>) <volume>48</volume>:<page-range>W449&#x2013;54</page-range>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkaa379</pub-id>
</citation>
</ref>
<ref id="B41">
<label>41</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>O&#x2019;Donnell</surname> <given-names>TJ</given-names>
</name>
<name>
<surname>Rubinsteyn</surname> <given-names>A</given-names>
</name>
<name>
<surname>Laserson</surname> <given-names>U</given-names>
</name>
</person-group>. <article-title>MHCflurry 2.0: Improved Pan-Allele Prediction of MHC Class I-Presented Peptides by Incorporating Antigen Processing</article-title>. <source>Cell Syst</source> (<year>2020</year>) <volume>11</volume>:<fpage>42</fpage>&#x2013;<lpage>48.e7</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cels.2020.06.010</pub-id>
</citation>
</ref>
<ref id="B42">
<label>42</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boehm</surname> <given-names>KM</given-names>
</name>
<name>
<surname>Bhinder</surname> <given-names>B</given-names>
</name>
<name>
<surname>Raja</surname> <given-names>VJ</given-names>
</name>
<name>
<surname>Dephoure</surname> <given-names>N</given-names>
</name>
<name>
<surname>Elemento</surname> <given-names>O</given-names>
</name>
</person-group>. <article-title>Predicting Peptide Presentation by Major Histocompatibility Complex Class I: An Improved Machine Learning Approach to the Immunopeptidome</article-title>. <source>BMC Bioinf</source> (<year>2019</year>) <volume>20</volume>:<fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s12859-018-2561-z</pub-id>
</citation>
</ref>
<ref id="B43">
<label>43</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>B</given-names>
</name>
<name>
<surname>Khodadoust</surname> <given-names>MS</given-names>
</name>
<name>
<surname>Olsson</surname> <given-names>N</given-names>
</name>
<name>
<surname>Wagar</surname> <given-names>LE</given-names>
</name>
<name>
<surname>Fast</surname> <given-names>E</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>CL</given-names>
</name>
<etal/>
</person-group>. <article-title>Predicting HLA Class II Antigen Presentation Through Integrated Deep Learning</article-title>. <source>Nat Biotechnol</source> (<year>2019</year>) <volume>37</volume>:<page-range>1332&#x2013;43</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41587-019-0280-2</pub-id>
</citation>
</ref>
<ref id="B44">
<label>44</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Starr</surname> <given-names>TK</given-names>
</name>
<name>
<surname>Jameson</surname> <given-names>SC</given-names>
</name>
<name>
<surname>Hogquist</surname> <given-names>KA</given-names>
</name>
</person-group>. <article-title>Positive and Negative Selection of T Cells</article-title>. <source>Annu Rev Immunol</source> (<year>2003</year>) <volume>21</volume>:<page-range>139&#x2013;76</page-range>. doi: <pub-id pub-id-type="doi">10.1146/annurev.immunol.21.120601.141107</pub-id>
</citation>
</ref>
<ref id="B45">
<label>45</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blackman</surname> <given-names>M</given-names>
</name>
<name>
<surname>Kappler</surname> <given-names>J</given-names>
</name>
<name>
<surname>Marrack</surname> <given-names>P</given-names>
</name>
</person-group>. <article-title>The Role of the T Cell Receptor in Positive and Negative Selection of Developing T Cells</article-title>. <source>Science</source> (<year>1990</year>) <volume>248</volume>:<page-range>1335&#x2013;41</page-range>. doi: <pub-id pub-id-type="doi">10.1126/science.1972592</pub-id>
</citation>
</ref>
<ref id="B46">
<label>46</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Accolla</surname> <given-names>RS</given-names>
</name>
<name>
<surname>Tosi</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>Optimal MHC-II-Restricted Tumor Antigen Presentation to CD4+ T Helper Cells: The Key Issue for Development of Anti-Tumor Vaccines</article-title>. <source>J Transl Med</source> (<year>2012</year>) <volume>10</volume>:<fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi: <pub-id pub-id-type="doi">10.1186/1479-5876-10-154</pub-id>
</citation>
</ref>
<ref id="B47">
<label>47</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ibrahim</surname> <given-names>NAM</given-names>
</name>
<name>
<surname>Mansour</surname> <given-names>YSE</given-names>
</name>
</person-group>. <article-title>A Review on Anticancer Peptide-Based Vaccines: Advantages, Limitations, and Current Challenges</article-title>. <source>Indian J Drugs</source> (<year>2020</year>) <volume>8</volume>:<fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi: <pub-id pub-id-type="doi">10.5281/zenodo.3351737</pub-id>
</citation>
</ref>
<ref id="B48">
<label>48</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>J</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>B</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>W</given-names>
</name>
<name>
<surname>Su</surname> <given-names>Z</given-names>
</name>
<etal/>
</person-group>. <article-title>DeepHLApan: A Deep Learning Approach for Neoantigen Prediction Considering Both HLA-Peptide Binding and Immunogenicity</article-title>. <source>Front Immunol</source> (<year>2019</year>) <volume>10</volume>:<elocation-id>2559</elocation-id>. doi: <pub-id pub-id-type="doi">10.3389/fimmu.2019.02559</pub-id>
</citation>
</ref>
<ref id="B49">
<label>49</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>L</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>F</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>DeepNetBim: Deep Learning Model for Predicting HLA-Epitope Interactions Based on Network Analysis by Harnessing Binding and Immunogenicity Information</article-title>. <source>BMC Bioinf</source> (<year>2021</year>) <volume>22</volume>:<fpage>1</fpage>&#x2013;<lpage>16</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s12859-021-04155-y</pub-id>
</citation>
</ref>
<ref id="B50">
<label>50</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cui</surname> <given-names>C</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J</given-names>
</name>
<name>
<surname>Fagerberg</surname> <given-names>E</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>P-M</given-names>
</name>
<name>
<surname>Connolly</surname> <given-names>KA</given-names>
</name>
<name>
<surname>Damo</surname> <given-names>M</given-names>
</name>
<etal/>
</person-group>. <article-title>Neoantigen-Driven B Cell and CD4 T Follicular Helper Cell Collaboration Promotes Anti-Tumor CD8 T Cell Responses</article-title>. <source>Cell</source> (<year>2021</year>) <volume>184</volume>:<page-range>6101&#x2013;18</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2021.11.007</pub-id>
</citation>
</ref>
<ref id="B51">
<label>51</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vita</surname> <given-names>R</given-names>
</name>
<name>
<surname>Mahajan</surname> <given-names>S</given-names>
</name>
<name>
<surname>Overton</surname> <given-names>JA</given-names>
</name>
<name>
<surname>Dhanda</surname> <given-names>SK</given-names>
</name>
<name>
<surname>Martini</surname> <given-names>S</given-names>
</name>
<name>
<surname>Cantrell</surname> <given-names>JR</given-names>
</name>
<etal/>
</person-group>. <article-title>The Immune Epitope Database (IEDB): 2018 Update</article-title>. <source>Nucleic Acids Res</source> (<year>2019</year>) <volume>47</volume>:<page-range>D339&#x2013;43</page-range>. doi: <pub-id pub-id-type="doi">10.1093/nar/gky1006</pub-id>
</citation>
</ref>
<ref id="B52">
<label>52</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Santurkar</surname> <given-names>S</given-names>
</name>
<name>
<surname>Tsipras</surname> <given-names>D</given-names>
</name>
<name>
<surname>Ilyas</surname> <given-names>A</given-names>
</name>
<name>
<surname>M&#x105;dry</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>How Does Batch Normalization Help Optimization</article-title>?(<year>2018</year>) (Accessed <access-date>Proceedings of the 32nd international conference on neural information processing systems</access-date>).</citation>
</ref>
<ref id="B53">
<label>53</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ba</surname> <given-names>JL</given-names>
</name>
<name>
<surname>Kiros</surname> <given-names>JR</given-names>
</name>
<name>
<surname>Hinton</surname> <given-names>GE</given-names>
</name>
</person-group>. <article-title>Layer Normalization</article-title>. <source>ArXiv Prepr ArXiv160706450</source> (<year>2016</year>).</citation>
</ref>
<ref id="B54">
<label>54</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname> <given-names>J</given-names>
</name>
<name>
<surname>Bendjama</surname> <given-names>K</given-names>
</name>
<name>
<surname>Rittner</surname> <given-names>K</given-names>
</name>
<name>
<surname>Malone</surname> <given-names>B</given-names>
</name>
</person-group>. <article-title>BERTMHC: Improved MHC&#x2013;peptide Class II Interaction Prediction With Transformer and Multiple Instance Learning</article-title>. <source>Bioinformatics</source> (<year>2021</year>) <volume>37</volume>:<page-range>4172&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btab422</pub-id>
</citation>
</ref>
<ref id="B55">
<label>55</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moore</surname> <given-names>TV</given-names>
</name>
<name>
<surname>Nishimura</surname> <given-names>MI</given-names>
</name>
</person-group>. <article-title>Improved MHC II Epitope Prediction &#x2014; A Step Towards Personalized Medicine</article-title>. <source>Nat Rev Clin Oncol</source> (<year>2020</year>) <volume>17</volume>:<page-range>71&#x2013;2</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41571-019-0315-0</pub-id>
</citation>
</ref>
<ref id="B56">
<label>56</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marcu</surname> <given-names>A</given-names>
</name>
<name>
<surname>Bichmann</surname> <given-names>L</given-names>
</name>
<name>
<surname>Kuchenbecker</surname> <given-names>L</given-names>
</name>
<name>
<surname>Kowalewski</surname> <given-names>DJ</given-names>
</name>
<name>
<surname>Freudenmann</surname> <given-names>LK</given-names>
</name>
<name>
<surname>Backert</surname> <given-names>L</given-names>
</name>
<etal/>
</person-group>. <article-title>HLA Ligand Atlas: A Benign Reference of HLA-Presented Peptides to Improve T-Cell-Based Cancer Immunotherapy</article-title>. <source>J Immunother Cancer</source> (<year>2021</year>) <volume>9</volume>:<elocation-id>e002071</elocation-id>. doi: <pub-id pub-id-type="doi">10.1136/jitc-2020-002071</pub-id>
</citation>
</ref>
<ref id="B57">
<label>57</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xia</surname> <given-names>J</given-names>
</name>
<name>
<surname>Bai</surname> <given-names>P</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>W</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D</given-names>
</name>
<etal/>
</person-group>. <article-title>NEPdb: A Database of T-Cell Experimentally-Validated Neoantigens and Pan-Cancer Predicted Neoepitopes for Cancer Immunotherapy</article-title>. <source>Front Immunol</source> (<year>2021</year>) <volume>12</volume>:<elocation-id>644637</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fimmu.2021.644637</pub-id>
</citation>
</ref>
<ref id="B58">
<label>58</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wells</surname> <given-names>DK</given-names>
</name>
<name>
<surname>Buuren</surname> <given-names>MMv</given-names>
</name>
<name>
<surname>Dang</surname> <given-names>KK</given-names>
</name>
<name>
<surname>Hubbard-Lucey</surname> <given-names>VM</given-names>
</name>
<name>
<surname>Sheehan</surname> <given-names>KCF</given-names>
</name>
<name>
<surname>Campbell</surname> <given-names>KM</given-names>
</name>
<etal/>
</person-group>. <article-title>Key Parameters of Tumor Epitope Immunogenicity Revealed Through a Consortium Approach Improve Neoantigen Prediction</article-title>. <source>Cell</source> (<year>2020</year>) <volume>183</volume>:<fpage>818</fpage>&#x2013;<lpage>834.e13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cell.2020.09.015</pub-id>
</citation>
</ref>
<ref id="B59">
<label>59</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>C</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>H</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>R</given-names>
</name>
</person-group>. <article-title>The Emerging Epigenetic Role of CD8+T Cells in Autoimmune Diseases: A Systematic Review</article-title>. <source>Front Immunol</source> (<year>2019</year>) <volume>10</volume>:<elocation-id>856</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fimmu.2019.00856</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>