<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioinform.</journal-id>
<journal-title>Frontiers in Bioinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioinform.</abbrev-journal-title>
<issn pub-type="epub">2673-7647</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1198218</article-id>
<article-id pub-id-type="doi">10.3389/fbinf.2023.1198218</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioinformatics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Protein quality assessment with a loss function designed for high-quality decoys</article-title>
<alt-title alt-title-type="left-running-head">Roy and Ben-Hur</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fbinf.2023.1198218">10.3389/fbinf.2023.1198218</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Roy</surname>
<given-names>Soumyadip</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2074493/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Ben-Hur</surname>
<given-names>Asa</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1866258/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Department of Computer Science</institution>, <institution>Colorado State University</institution>, <addr-line>Fort Collins</addr-line>, <addr-line>CO</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/29610/overview">Mensur Dlakic</ext-link>, Montana State University, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/988612/overview">Hiroto Saigo</ext-link>, Kyushu University, Japan</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1702098/overview">Lim Heo</ext-link>, Michigan State University, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Asa Ben-Hur, <email>asa@colostate.edu</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>3</volume>
<elocation-id>1198218</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>03</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>09</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Roy and Ben-Hur.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Roy and Ben-Hur</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>
<bold>Motivation:</bold> The prediction of a protein 3D structure is essential for understanding protein function, drug discovery, and disease mechanisms; with the advent of methods like AlphaFold that are capable of producing very high-quality decoys, ensuring the quality of those decoys can provide further confidence in the accuracy of their predictions.</p>
<p>
<bold>Results:</bold> In this work, we describe Q<sub>
<italic>&#x3f5;</italic>
</sub>, a graph convolutional network (GCN) that utilizes a minimal set of atom and residue features as inputs to predict the global distance test total score (GDTTS) and local distance difference test (lDDT) score of a decoy. To improve the model&#x2019;s performance, we introduce a novel loss function based on the <italic>&#x3f5;</italic>-insensitive loss function used for SVM regression. This loss function is specifically designed for evaluating the characteristics of the quality assessment problem and provides predictions with improved accuracy over standard loss functions used for this task. Despite using only a minimal set of features, it matches the performance of recent state-of-the-art methods like DeepUMQA.</p>
<p>
<bold>Availability:</bold> The code for Q<sub>
<italic>&#x3f5;</italic>
</sub> is available at <ext-link ext-link-type="uri" xlink:href="https://github.com/soumyadip1997/qepsilon">https://github.com/soumyadip1997/qepsilon</ext-link>.</p>
</abstract>
<kwd-group>
<kwd>protein structure quality assessment</kwd>
<kwd>deep learning</kwd>
<kwd>graph convolutional networks</kwd>
<kwd>epsilon-insensitive loss function</kwd>
<kwd>critical assessment of structure prediction</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Protein Bioinformatics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Predicting a protein&#x2019;s 3D structure from its amino acid sequence has been an area of avid interest for many years (<xref ref-type="bibr" rid="B2">Al-Lazikani et al., 2001)</xref>. Recently, significant progress has been made in this field with the introduction of AlphaFold, a deep learning system that achieved remarkable accuracy in predicting protein structures (<xref ref-type="bibr" rid="B17">Jumper et al., 2021)</xref>. While experimental identification of native protein structures remains a time-consuming and costly process, computational methods have made it possible to generate thousands of tertiary structures, known as decoys, in a matter of hours (<xref ref-type="bibr" rid="B33">Shehu, 2015)</xref>. However, identifying the best structure remains a challenge. Therefore, it is necessary to employ a quality assessment stage to identify high-quality, near-native decoys among the generated decoys (<xref ref-type="bibr" rid="B1">Akhter et al., 2020)</xref>. This remains true even with AlphaFold&#x2019;s recent breakthrough performance (<xref ref-type="bibr" rid="B5">Chen et al., 2023)</xref>. Furthermore, with the subsequent availability of genome-wide predicted structures across many species (<xref ref-type="bibr" rid="B37">Varadi et al., 2022)</xref>, the quality assessment problem is as relevant as ever.</p>
<p>In this work, we address the decoy quality assessment problem with the help of graph convolutional networks (GCNs); we introduce a novel loss function inspired by the support vector regression, <italic>&#x3f5;</italic>-insensitive loss function, that is designed to take into account our intuition about what makes a good quality assessment predictor, namely, that it focuses on making correct predictions for those decoys that matter: decoys with high quality. We compare our method, called Q<sub>
<italic>&#x3f5;</italic>
</sub>, to other state-of-the-art methods and demonstrate that our method outperforms most of those methods while using only a very basic set of features computed from a decoy&#x2019;s sequence, without the need for engineered features.</p>
</sec>
<sec id="s2">
<title>2 Related work</title>
<p>Current techniques for quality assessment can be divided into two categories. One is single-model methods that operate on single structural models to estimate their quality (<xref ref-type="bibr" rid="B38">Wallner and Elofsson, 2003)</xref>. The second category consists of methods that use consistency among several candidates to estimate quality (<xref ref-type="bibr" rid="B22">Lundstr&#xf6;m et al., 2001)</xref>. Protein quality assessment methods have been evaluated in the Critical Assessment of Structure Prediction (CASP) competition (<xref ref-type="bibr" rid="B28">Moult et al., 1995)</xref> since CASP7. The CASP13 single-model methods, the focus of this work, performed comparably or better than consensus methods for the first time (<xref ref-type="bibr" rid="B6">Cheng et al., 2019)</xref>. A variety of single-model approaches have been proposed, and currently, machine learning-based methods dominate this area.</p>
<p>Until a few years ago, methods that use standard machine learning techniques with a large collection of engineered features computed from sequence and structure were the prevalent approaches for quality assessment. The ProQ series of methods (ProQ, ProQ2, ProQ3, and ProQ3D) (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref> used features such as the distribution of atom&#x2013;atom contacts, residue&#x2013;residue contacts, solvent accessibility, secondary structure, surface area, and evolutionary information. ProQ3 (<xref ref-type="bibr" rid="B35">Uziela et al., 2016)</xref> also incorporated features based on Rosetta energies. ProQ3D (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref> used the descriptors of ProQ3 as inputs in conjunction with a multi-layer perceptron and was one of the top performers of CASP13.</p>
<p>The current state-of-the-art method for quality assessment uses deep learning, including various types of 3D convolutional networks and graph neural networks, which have been demonstrated to be effective tools for modeling protein 3D structures (<xref ref-type="bibr" rid="B7">Derevyanko et al., 2018</xref>; <xref ref-type="bibr" rid="B11">Fout et al., 2017)</xref>. Deep convolutional networks as a tool for the representation of decoy structures were introduced by <xref ref-type="bibr" rid="B7">Derevyanko et al. (2018)</xref>. Their method, 3DCNN, used 3D convolutional networks applied to a volumetric representation of a decoy structure. The Ornate method by <xref ref-type="bibr" rid="B30">Pag&#xe8;s et al. (2019)</xref> improved upon 3DCNN by defining a canonical orientation for each residue. The GraphQA method by <xref ref-type="bibr" rid="B3">Baldassarre et al. (2020)</xref> employed a graph convolutional network with an extensive number of engineered features and achieved state-of-the-art performance on CASP13 decoys. <xref ref-type="bibr" rid="B5">Chen et al. (2023)</xref> used a graph neural network to estimate the accuracy of AlphaFold models, which is one of the current state-of-the-art methods, and improved on the results obtained with DeepAccNet by <xref ref-type="bibr" rid="B14">Hiranuma et al. (2021)</xref> while borrowing many ideas from its architecture. They used a combination of categorical loss and L2-loss on the lDDT scores to distinguish between decoys of varying quality levels. The DeepUMQA method uses 3D convolution over a collection of residue-level engineered features (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>, and its successor, DeepUMQA2 (<xref ref-type="bibr" rid="B21">Liu et al., 2023)</xref>, is also a state-of-the-art performer.</p>
<p>Most existing methods for quality assessment rely on engineered features. In contrast, our approach uses sequence embeddings computed using protein language models; convolutional layers applied to both atomic- and residue-level graphs are then used to put them in the context of the decoy structure. In combination with a novel loss function specifically designed for the quality assessment problem, our method can outperform the recent DeepUMQA method (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>.</p>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<sec id="s3-1">
<title>3.1 The quality assessment problem</title>
<p>Computational methods for predicting a protein&#x2019;s 3D structure produce large numbers of decoy conformations for a given target protein. In quality assessment, we seek to rank these decoys based on their similarity to the experimentally determined native structure. We address this as a regression problem: our method is designed to predict the global distance test total score (GDTTS) (<xref ref-type="bibr" rid="B39">Zemla, 2003)</xref> and the local distance difference test (lDDT) score (<xref ref-type="bibr" rid="B24">Mariani et al., 2013)</xref>, which are the official CASP scores for global-level decoy quality. While several recent methods were designed to predict the lDDT score (<xref ref-type="bibr" rid="B14">Hiranuma et al., 2021</xref>; <xref ref-type="bibr" rid="B5">Chen et al., 2023)</xref>, we used both scores to allow for direct comparison with GraphQA, which is the most similar approach to the method presented here and would allow us to compare with more recent QA methods like DeepUMQA and DeepUMQA2. GDTTS measures the percentage of residues in the superimposed predicted structure that are within a certain distance threshold of their corresponding residues in the true structure. lDDT score is a superposition-free score that represents the local distance difference among all atoms in a predicted structure, thereby providing an idea of the local quality of the predicted structure. Decoy structures with high GDTTS and lDDT score (<inline-formula id="inf1">
<mml:math id="m1">
<mml:mo>&#x3e;</mml:mo>
</mml:math>
</inline-formula>0.85) indicate that they closely resemble the native structure. In what follows, we describe Q<sub>
<italic>&#x3f5;</italic>
</sub>, a graph convolutional network that is trained on labeled decoy 3D structures, utilizing a basic set of features generated from atoms and residues using a combination of the L1-loss and a modification of the SVM regression <italic>&#x3f5;</italic>-insensitive loss function (<xref ref-type="bibr" rid="B8">Drucker et al., 1996)</xref>.</p>
</sec>
<sec id="s3-2">
<title>3.2 Atom- and residue-level graph convolution</title>
<p>Graph convolution is a powerful approach for representing protein 3D structures (<xref ref-type="bibr" rid="B11">Fout et al., 2017)</xref> and has proven its value for the quality assessment problem (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020</xref>). In order to enable us to forgo engineered features, we have chosen to represent the 3D structure of a decoy using dual graphs at the atom and residue levels (see <xref ref-type="fig" rid="F1">Figure 1</xref>). Each of the graphs is a nearest neighbor graph where a pair of nodes is connected by an edge if their distance in the structure is less than a given threshold, where 6&#xc5; was the selected value in our experiments, and the distance between residues is the minimum distance between their atoms. We used up to 20 nearest neighbors to define the edges in both the atom-level and the residue-level graphs.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Graph representation of a decoy structure. The structure of a decoy is represented using two graphs: one at the atomic level (left) and one at the residue level (right). Our graph convolution operation at the atom level differentiates between edges within a residue and edges across neighboring residues.</p>
</caption>
<graphic xlink:href="fbinf-03-1198218-g001.tif"/>
</fig>
<p>We perform graph convolution separately at the atom and residue levels. First, we describe the atom-level graph convolution (GCN<sub>atom</sub>). Each atom <italic>i</italic> is assigned a feature vector <inline-formula id="inf2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> that contains the features for layer <italic>l</italic> of graph convolution. The representation of a source atom <inline-formula id="inf3">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is updated based on its neighbors within the same residue <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and the neighbors across residues <inline-formula id="inf5">
<mml:math id="m5">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> according to<disp-formula id="e1">
<mml:math id="m6">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">L</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf6">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the weight matrix with respect to the source atom in layer <italic>l</italic>, <inline-formula id="inf7">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the weight matrix with respect to the neighboring atoms in layer <italic>l</italic> within the same residue as that of the source atom, <inline-formula id="inf8">
<mml:math id="m9">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the weight matrix with respect to the neighboring atoms in layer <italic>l</italic> that belong to a different residue than the source atom, and finally, <inline-formula id="inf9">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the bias in layer <italic>l</italic> for the atom-level GCN. The inputs to the atom-level convolution are derived from one-hot encoding of the atom type as described in the following sections.</p>
<p>In parallel to the atom-level convolution, we perform convolution over the residues that make up a decoy structure. This operation, denoted as GCN<sub>residue</sub>, is used to update the residue-level representation <inline-formula id="inf10">
<mml:math id="m11">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, which is the feature vector for residue <italic>i</italic> in layer <italic>l</italic> of the network. This operation is defined as follows:<disp-formula id="e2">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">L</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="script">R</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">R</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
</mml:mstyle>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf11">
<mml:math id="m13">
<mml:mi mathvariant="script">R</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the set of the neighboring residues of residue <italic>i</italic>, <inline-formula id="inf12">
<mml:math id="m14">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the weight matrix with respect to the source residue in layer <italic>l</italic>, <inline-formula id="inf13">
<mml:math id="m15">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the weight matrix with respect to the neighboring residues in layer <italic>l</italic>, and <inline-formula id="inf14">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the bias in layer <italic>l</italic>. The inputs to the residue-level convolution are embeddings computed using ProtTrans (<xref ref-type="bibr" rid="B9">Elnaggar et al., 2021)</xref> as described in the following sections.</p>
</sec>
<sec id="s3-3">
<title>3.3 Network architecture</title>
<p>The architecture for <italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub> includes four graph convolutional layers that aggregate information at the atomic level (<italic>GCN</italic>
<sub>
<italic>atom</italic>
</sub>) and four graph convolutional layers that pass information at the residue level (<italic>GCN</italic>
<sub>
<italic>residue</italic>
</sub>). To ensure model stability and generalization, we apply batch normalization (<xref ref-type="bibr" rid="B16">Ioffe and Szegedy, 2015)</xref> after each application of an activation function to normalize the activations across the nodes in the graph. To create a coherent representation at the residue level, we apply a maximum pooling operation to the output of the final layer of <italic>GCN</italic>
<sub>
<italic>atom</italic>
</sub>. The final residue-level representation is obtained by concatenating the output of the pooled atomic-level convolution and the output from the residue-level GCN. This concatenated output is passed through a multi-layer perceptron (MLP), which outputs a single output per residue of the decoy structure. The final output of the network, which is our predicted value of GDTTS or lDDT score, is then produced by averaging over the node-level scores. This process is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>
<italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub> model architecture illustrating how an input decoy structure is propagated through multiple graph convolutional layers (<italic>GCN</italic>
<sub>
<italic>atom</italic>
</sub> for the atom-level representation and <italic>GCN</italic>
<sub>
<italic>residue</italic>
</sub> for the residue-level representation of a protein); the outputs of the two sets of convolutional layers are concatenated and fed through a multi-layer perceptron (MLP) to generate local scores that are then averaged to compute the predicted GDTTS or lDDT score.</p>
</caption>
<graphic xlink:href="fbinf-03-1198218-g002.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Atom and residue features</title>
<p>Our method performs convolution at both the atom and residue levels. Here, we describe the features used at both levels.</p>
<sec id="s3-4-1">
<title>3.4.1 Atom features</title>
<p>We represent the atoms using one-hot encoding by grouping atoms into 11 different types (<xref ref-type="bibr" rid="B7">Derevyanko et al., 2018)</xref>. This grouping reflects both the type of atom (carbon, oxygen, or nitrogen) and its context within the residue (e.g., alpha carbon or the different group an atom belongs to). In doing so, we are able to incorporate important information of the atoms while also capturing the relationships between the atoms and their corresponding residues.</p>
</sec>
<sec id="s3-4-2">
<title>3.4.2 Residue features</title>
<p>We compute residue features by feeding the amino acid sequence of a decoy to the ProtTrans protein language model (<xref ref-type="bibr" rid="B9">Elnaggar et al., 2021)</xref>. ProtTrans embeddings provide a very useful representation of the amino acid sequence, capturing relationships between residues and their structural context (<xref ref-type="bibr" rid="B9">Elnaggar et al., 2021)</xref>. We take the embeddings from the last hidden state of the transformer attention stack of the ProtTrans model, with an output embedding of 1,024 dimensions, which serves as the input to the residue-level GCN.</p>
</sec>
</sec>
<sec id="s3-5">
<title>3.5 A modified <italic>&#x3f5;</italic>-insensitive loss</title>
<p>In this work, we address quality assessment as a regression problem with the objective of predicting GDTTS or lDDT score of a decoy. We propose a novel loss function that captures our desiderata for a quality assessment model: when it comes to poor decoys, we do not care about the accuracy of the prediction as long as we can differentiate it from a good decoy. On the other hand, the more accurate the decoy, the more accurate we want our prediction to be. This is especially important given the recent improvement in the quality of protein structure prediction methods. To achieve this goal, we modify the <italic>&#x3f5;</italic>-insensitive loss, which is the loss function employed in SVM regression (<xref ref-type="bibr" rid="B8">Drucker et al., 1996)</xref>, as follows:<disp-formula id="e3">
<mml:math id="m17">
<mml:mi>L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="italic">max</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where <italic>y</italic> and <italic>y</italic>&#x2032; are the true and predicted scores, respectively. As in the standard <italic>&#x3f5;</italic>-insensitive loss, this defines a tube of size <italic>&#x3f5;</italic> within which there is no penalty; outside the tube, the loss grows linearly as in the L1-loss, which is defined as <italic>L</italic>(<italic>y</italic>, <italic>y</italic>&#x2032;) &#x3d; &#x7c;<italic>y</italic> &#x2212; <italic>y</italic>&#x2032;&#x7c;. In our application, the size of the tube is a function <italic>&#x3f5;</italic>(<italic>y</italic>). In this work, we used a tube defined as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The motivation for the modified <italic>&#x3f5;</italic>-insensitive loss function is that the model should not try too hard to accurately fit poor-quality decoys where we do not need good accuracy anyhow. As decoy quality increases, models are trained to learn a fit that is much more accurate.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The modified <italic>&#x3f5;</italic>-insensitive loss uses a variable-sized band around the diagonal in which a predicted score is not penalized. The band becomes smaller as the GDTTS or lDDT score increases, reflecting our expectation for precise predictions for decoys that are closer to the native structure.</p>
</caption>
<graphic xlink:href="fbinf-03-1198218-g003.tif"/>
</fig>
</sec>
<sec id="s3-6">
<title>3.6 Network training</title>
<p>We have trained our network to predict GDTTS and lDDT score. For GDTTS prediction, we first pre-train Q<sub>
<italic>&#x3f5;</italic>
</sub> with the L1-loss for 50 epochs, followed by training with the modified <italic>&#x3f5;</italic>-insensitive loss for the next 10 epochs. To train the network with lDDT scores, we select the best model from GDTTS (&#x201c;best&#x201d; with respect to the validation set) and train it with the modified <italic>&#x3f5;</italic>-insensitive loss for another 50 epochs, keeping the same network architecture and hyperparameters.</p>
<p>The network was implemented in PyTorch (<xref ref-type="bibr" rid="B31">Paszke et al., 2019)</xref> and optimized using the Adam method (<xref ref-type="bibr" rid="B18">Kingma and Ba, 2014)</xref> with default parameters except for a learning rate of 0.001; training used a batch size of 70. Since our training set is highly imbalanced, i.e., contains very few high-quality decoys, we used the imbalanced sampler from the torchsampler package. During training, we monitored the loss over the validation set and used the model that gave the minimum loss. Our implementation used the PyTorch Lightning framework for training and testing and PyTorch Geometric (<xref ref-type="bibr" rid="B10">Fey and Lenssen, 2019)</xref> for performing graph convolution. Model selection was performed using the hyperparameters and values described in <xref ref-type="table" rid="T1">Table 1</xref>. We iterated over all parameters and, for each one, chose the value that gave the highest Pearson correlation coefficient on the validation set. Following model selection, training took approximately 42&#xa0;h on an NVIDIA RTX 3090 GPU.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Hyperparameter space. Model selection was performed based on performance on the validation set.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Hyperparameter</th>
<th align="center">Values</th>
<th align="center">Best</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Number of graph convolution layers</td>
<td align="center">2, 3, 4, 5, 6</td>
<td align="center">4</td>
</tr>
<tr>
<td align="center">Neighbor distance threshold</td>
<td align="center">4, 5, 6, 7, 8, 9</td>
<td align="center">6</td>
</tr>
<tr>
<td align="center">Maximum number of same residue atom neighbors</td>
<td align="center">10, 15, 20, 25</td>
<td align="center">20</td>
</tr>
<tr>
<td align="center">Maximum number of different residue atom neighbors</td>
<td align="center">10, 15, 20, 25</td>
<td align="center">20</td>
</tr>
<tr>
<td align="center">Maximum number of neighbors of a residue</td>
<td align="center">10, 15, 20, 25</td>
<td align="center">20</td>
</tr>
<tr>
<td align="center">Dropout rate for the graph convolution layers</td>
<td align="center">0, 0.1, 0.2, 0.3</td>
<td align="center">0.1</td>
</tr>
<tr>
<td align="center">Learning rate</td>
<td align="center">0.0001, 0.001, 0.01, 0.1</td>
<td align="center">0.001</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The &#x201c;Best&#x201d; column provides the chosen value for each hyperparameter.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3-7">
<title>3.7 Data</title>
<p>We collected decoys from CASP9 to CASP14 along with their labels from the CASP website (<xref ref-type="bibr" rid="B4">CASP, 2021)</xref>. We used CASP9&#x2013;CASP12 as our training and validation sets and CASP13 and CASP14 as our test sets (see <xref ref-type="table" rid="T2">Table 2</xref>). In order to match the decoys used in experiments performed by others, we created two separate datasets for GDDTS evaluation (CASP13 and CASP14) and two datasets for the evaluation of lDDT score prediction (CASP13 and CASP14).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Number of targets from CASP competitions in the training, validation, and testing data.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Mode</th>
<th align="center">CASP</th>
<th align="center">Target</th>
<th align="center">Decoy</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="center">Training data</td>
<td align="center">CASP9</td>
<td align="center">117</td>
<td align="center">31,863</td>
</tr>
<tr>
<td align="center">CASP10</td>
<td align="center">100</td>
<td align="center">23,755</td>
</tr>
<tr>
<td align="center">CASP11</td>
<td align="center">84</td>
<td align="center">15,573</td>
</tr>
<tr>
<td align="center">CASP12</td>
<td align="center">30</td>
<td align="center">5,351</td>
</tr>
<tr>
<td align="center">Validation data</td>
<td align="center">CASP12</td>
<td align="center">10</td>
<td align="center">1,338</td>
</tr>
<tr>
<td rowspan="2" align="center">Testing data (GDTTS)</td>
<td align="center">CASP13</td>
<td align="center">72</td>
<td align="center">34,654</td>
</tr>
<tr>
<td align="center">CASP14</td>
<td align="center">65</td>
<td align="center">38,293</td>
</tr>
<tr>
<td rowspan="3" align="center">Testing data (lDDT score)</td>
<td align="center">CASP13</td>
<td align="center">76</td>
<td align="center">10,739</td>
</tr>
<tr>
<td align="center">CASP14</td>
<td align="center">70</td>
<td align="center">10,380</td>
</tr>
<tr>
<td align="center">AlphaFold2 CASP15</td>
<td align="center">17</td>
<td align="center">85</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Two different CASP13 and CASP14 datasets, one for GDTTS evaluation and the other for lDDT score evaluation, are used to match decoys used in other publications.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>In CASP15, the focus shifted from predicting the accuracy of single-chain decoys to that of multi-chain complexes (<xref ref-type="bibr" rid="B19">Kryshtafovych et al., 2023)</xref>. However, some of the targets were composed of single chains, and we chose to focus on those targets in our evaluation, leading to a dataset with 17 targets.</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4 Results</title>
<p>We compare Q<sub>
<italic>&#x3f5;</italic>
</sub> with other methods that have either state-of-the-art or very good performance in CASP13 and CASP14. In our first set of experiments, we sought to compare our method with GraphQA, which uses a similar graph convolution architecture and was trained to predict GDTTS (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020)</xref>. The results in <xref ref-type="table" rid="T3">Table 3</xref> indicate that Q<sub>
<italic>&#x3f5;</italic>
</sub> outperforms GraphQA and several other recent methods trained to predict GDTTS despite not using engineered features; a detailed analysis of the contribution of the various components of the Q<sub>
<italic>&#x3f5;</italic>
</sub> architecture is described in an ablation study in the following section.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> and other methods in CASP13 and CASP14 GDTTS prediction.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Dataset</th>
<th align="center">Method</th>
<th align="center">
<italic>R</italic>
</th>
<th align="center">
<italic>R</italic>
<sub>
<italic>target</italic>
</sub>
</th>
<th align="center">
<italic>&#x3c1;</italic>
</th>
<th align="center">RMSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="5" align="center">CASP13</td>
<td align="center">Q<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">
<bold>0.90</bold>
</td>
<td align="center">
<bold>0.80</bold>
</td>
<td align="center">
<bold>0.89</bold>
</td>
<td align="center">
<bold>0.10</bold>
</td>
</tr>
<tr>
<td align="center">GraphQA (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020)</xref>
</td>
<td align="center">0.86</td>
<td align="center">0.78</td>
<td align="center">0.86</td>
<td align="center">0.13</td>
</tr>
<tr>
<td align="center">ModFOLD7_rank (<xref ref-type="bibr" rid="B25">McGuffin et al., 2019)</xref>
</td>
<td align="center">0.87</td>
<td align="center">0.74</td>
<td align="center">-</td>
<td align="center">0.16</td>
</tr>
<tr>
<td align="center">ProQ4 (<xref ref-type="bibr" rid="B15">Hurtado et al., 2018)</xref>
</td>
<td align="center">0.70</td>
<td align="center">0.66</td>
<td align="center">-</td>
<td align="center">0.18</td>
</tr>
<tr>
<td align="center">VoroMQA-A (<xref ref-type="bibr" rid="B29">Olechnovi&#x10d; and Venclovas, 2017)</xref>
</td>
<td align="center">0.66</td>
<td align="center">0.56</td>
<td align="center">-</td>
<td align="center">0.21</td>
</tr>
<tr>
<td align="center">CASP14</td>
<td align="center">Q<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">
<bold>0.81</bold>
</td>
<td align="center">
<bold>0.72</bold>
</td>
<td align="center">
<bold>0.82</bold>
</td>
<td align="center">
<bold>0.13</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The global Pearson correlation coefficient (<italic>R</italic>), Pearson correlation coefficient per target (<italic>R</italic>
<sub>target</sub>), and Spearman rank correlation between predicted and known GDTTS are provided. The best performance is highlighted in bold. Performance numbers for the other methods is quoted from <xref ref-type="bibr" rid="B3">Baldassarre et al. (2020)</xref>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The quality assessment community is transitioning to the use of the lDDT score, so we also compare Q<sub>
<italic>&#x3f5;</italic>
</sub> with more recent methods evaluated with lDDT. In this evaluation, the performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> was similar to that of DeepUMQA but outperformed by its successor, DeepUMQA2 (see <xref ref-type="table" rid="T4">Table 4</xref>). Results from EnQA (<xref ref-type="bibr" rid="B5">Chen et al., 2023)</xref>, whose performance was similar to that of DeepUMQA2, are also better than those of Q<sub>
<italic>&#x3f5;</italic>
</sub>. Both methods use more complex architectures and extensive engineered features; DeepUMQA2 also used evolutionary information, including structural features from homologous templates.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> and other methods in CASP13 and CASP14 with respect to lDDT scores.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Dataset</th>
<th align="center">Method</th>
<th align="center">
<italic>R</italic>
</th>
<th align="center">
<italic>&#x3c1;</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="8" align="center">CASP13</td>
<td align="center">Q<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.857</td>
<td align="center">
<bold>0.862</bold>
</td>
</tr>
<tr>
<td align="center">DeepUMQA2 (<xref ref-type="bibr" rid="B21">Liu et al., 2023)</xref>
</td>
<td align="center">
<bold>0.919</bold>
</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">DeepUMQA (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>
</td>
<td align="center">0.837</td>
<td align="center">0.804</td>
</tr>
<tr>
<td align="center">ModFOLD7_rank (<xref ref-type="bibr" rid="B23">Maghrabi and McGuffin, 2020)</xref>
</td>
<td align="center">0.826</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ3D (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref>
</td>
<td align="center">0.801</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ4 (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref>
</td>
<td align="center">0.777</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ2 (<xref ref-type="bibr" rid="B32">Ray et al., 2012)</xref>
</td>
<td align="center">0.715</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">VoroMQA-A (<xref ref-type="bibr" rid="B29">Olechnovi&#x10d; and Venclovas, 2017)</xref>
</td>
<td align="center">0.672</td>
<td align="center">-</td>
</tr>
<tr>
<td rowspan="9" align="center">CASP14</td>
<td align="center">Q<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.826</td>
<td align="center">
<bold>0.826</bold>
</td>
</tr>
<tr>
<td align="center">DeepUMQA2 (<xref ref-type="bibr" rid="B21">Liu et al., 2023)</xref>
</td>
<td align="center">
<bold>0.899</bold>
</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">DeepUMQA (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>
</td>
<td align="center">0.799</td>
<td align="center">0.736</td>
</tr>
<tr>
<td align="center">DeepAccNet (<xref ref-type="bibr" rid="B14">Hiranuma et al., 2021)</xref>
</td>
<td align="center">0.829</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ModFOLD8 (<xref ref-type="bibr" rid="B26">McGuffin et al., 2021)</xref>
</td>
<td align="center">0.629</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">GraphQA (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020)</xref>
</td>
<td align="center">0.706</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ3D (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref>
</td>
<td align="center">0.717</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ2 (<xref ref-type="bibr" rid="B32">Ray et al., 2012)</xref>
</td>
<td align="center">0.531</td>
<td align="center">-</td>
</tr>
<tr>
<td align="center">ProQ4 (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref>
</td>
<td align="center">0.547</td>
<td align="center">-</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The global Pearson correlation coefficient (<italic>R</italic>) and Spearman rank correlation between predicted and known lDDT scores are provided. The best performance is highlighted in bold. Performance figures for methods other than Q<sub>
<italic>&#x3f5;</italic>
</sub> are quoted from <xref ref-type="bibr" rid="B12">Guo et al. (2022)</xref> and <xref ref-type="bibr" rid="B21">Liu et al. (2023)</xref>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>To understand the contribution of the proposed modified <italic>&#x3f5;</italic>-insensitive loss to the performance of Q<sub>
<italic>&#x3f5;</italic>
</sub>, a scatter plot of true versus predicted GDTTS for the decoys in CASP13 and CASP14 is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. We observe that the modified <italic>&#x3f5;</italic>-insensitive loss leads to better learning of decoys of all quality levels compared to the L1-loss and leads to a pattern where the predictions are limited to a band around the true scores, which is a highly desirable property for a quality assessment method. It was interesting that the width of the band is similar across all quality levels, despite the loss having a variable width band compared to the original <italic>&#x3f5;</italic>-insensitive loss function.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Scatter plots comparing the true and predicted GDTTS for both CASP13 and CASP14 using L1-loss (bottom) and modified &#x03B5;-insensitive loss (top).</p>
</caption>
<graphic xlink:href="fbinf-03-1198218-g004.tif"/>
</fig>
<sec id="s4-1">
<title>4.1 Ablation study</title>
<p>To demonstrate the contribution of each of the major components of our method, we performed an ablation study with respect to GDTTS prediction, and its results are shown in <xref ref-type="table" rid="T5">Table 5</xref>. The first component we varied was the loss function. We observe that the pre-training with the L1-loss is key for the method&#x2019;s performance, serving to bootstrap the learning process. We also observe that performance dropped when using the original <italic>&#x3f5;</italic>-insensitive loss function, L1-loss, or L2-loss. This clearly shows the contribution of the proposed modification to the <italic>&#x3f5;</italic>-insensitive loss. Our next observation is that both the residue-level and atom-level convolutional blocks are crucial for the performance of the method. This is due to each of them providing different and complementary information. The residue-level blocks use ProtTrans embeddings, which have been documented to provide a variety of information regarding a residue&#x2019;s evolutionary history and structural context within the protein (<xref ref-type="bibr" rid="B9">Elnaggar et al., 2021)</xref>. The atom-level convolutional blocks provide a more fine-grained view of a decoy structure, complementing the information at the residue level.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Q<sub>
<italic>&#x3f5;</italic>
</sub> ablation study.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">
<italic>R</italic>
</th>
<th align="center">
<italic>R</italic>
<sub>
<italic>target</italic>
</sub>
</th>
<th align="center">
<italic>&#x3c1;</italic>
</th>
<th align="center">RMSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> (with atom and residue features, pre-trained with L1-loss and modified <italic>&#x3f5;</italic>-insensitive loss)</td>
<td align="center">
<bold>0.90</bold>
</td>
<td align="center">
<bold>0.80</bold>
</td>
<td align="center">
<bold>0.89</bold>
</td>
<td align="center">
<bold>0.11</bold>
</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> without modified <italic>&#x3f5;</italic>-insensitive loss</td>
<td align="center">0.75</td>
<td align="center">0.66</td>
<td align="center">0.69</td>
<td align="center">0.17</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> without L1-loss</td>
<td align="center">0.70</td>
<td align="center">0.59</td>
<td align="center">0.62</td>
<td align="center">0.20</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> with a constant <italic>&#x3f5;</italic> (0.2)</td>
<td align="center">0.63</td>
<td align="center">0.55</td>
<td align="center">0.66</td>
<td align="center">0.24</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> with only L2-loss</td>
<td align="center">0.65</td>
<td align="center">0.52</td>
<td align="center">0.56</td>
<td align="center">0.23</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> without residue features</td>
<td align="center">0.70</td>
<td align="center">0.65</td>
<td align="center">0.69</td>
<td align="center">0.19</td>
</tr>
<tr>
<td align="left">Q<sub>
<italic>&#x3f5;</italic>
</sub> without atom features</td>
<td align="center">0.79</td>
<td align="center">0.77</td>
<td align="center">0.76</td>
<td align="center">0.18</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Each of the major elements of Q<sub>
<italic>&#x3f5;</italic>
</sub> is removed, demonstrating that each of them provides a major contribution to the performance of the method.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-2">
<title>4.2 <italic>&#x3f5;</italic>-threshold selection</title>
<p>The modified <italic>&#x3f5;</italic>-insensitive loss has nine threshold parameters associated with the epsilon insensitive loss function, one for each bin of the prediction score. In our experiments, we have used the values shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. In order to determine that our initial choice was good, we ran an experiment where we varied all the values in a coordinated manner: we chose nine values lower or higher than the initial values (the columns low and high in <xref ref-type="table" rid="T6">Table 6</xref>). As shown in <xref ref-type="table" rid="T6">Table 6</xref>, reducing or increasing the values of all the thresholds in a coordinated fashion led to reduced accuracy on the validation set. As a sanity check, we verified that a similar decrease is observed on the test set as well.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Model selection over the <italic>&#x3f5;</italic> hyperparameter values.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Score ranges and results</th>
<th align="center">Low value</th>
<th align="center">Mid value</th>
<th align="center">High value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0&#x2013;0.1</td>
<td align="center">0.40</td>
<td align="center">0.45</td>
<td align="center">0.50</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.1&#x2013;0.2</td>
<td align="center">0.35</td>
<td align="center">0.40</td>
<td align="center">0.45</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.2&#x2013;0.3</td>
<td align="center">0.30</td>
<td align="center">0.35</td>
<td align="center">0.40</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.3&#x2013;0.4</td>
<td align="center">0.25</td>
<td align="center">0.30</td>
<td align="center">0.35</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.4&#x2013;0.5</td>
<td align="center">0.20</td>
<td align="center">0.25</td>
<td align="center">0.30</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.5&#x2013;0.6</td>
<td align="center">0.15</td>
<td align="center">0.2</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.6&#x2013;0.7</td>
<td align="center">0.10</td>
<td align="center">0.15</td>
<td align="center">0.20</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for 0.7&#x2013;0.8</td>
<td align="center">0.05</td>
<td align="center">0.1</td>
<td align="center">0.15</td>
</tr>
<tr>
<td align="center">
<italic>&#x3f5;</italic> for <inline-formula id="inf15">
<mml:math id="m18">
<mml:mo>&#x3e;</mml:mo>
</mml:math>
</inline-formula> 0.8</td>
<td align="center">0.005</td>
<td align="center">0.01</td>
<td align="center">0.015</td>
</tr>
<tr>
<td align="center">R on CASP12 (validation set) (GDTTS)</td>
<td align="center">0.84</td>
<td align="center">0.89</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="center">R on CASP12 (validation set) (lDDT score)</td>
<td align="center">0.81</td>
<td align="center">0.84</td>
<td align="center">0.77</td>
</tr>
<tr>
<td align="center">R on CASP13 (test set) (GDTTS)</td>
<td align="center">0.86</td>
<td align="center">0.90</td>
<td align="center">0.85</td>
</tr>
<tr>
<td align="center">R on CASP14 (test set) (GDTTS)</td>
<td align="center">0.80</td>
<td align="center">0.81</td>
<td align="center">0.79</td>
</tr>
<tr>
<td align="center">R on CASP13 (test set) (lDDT score)</td>
<td align="center">0.84</td>
<td align="center">0.86</td>
<td align="center">0.83</td>
</tr>
<tr>
<td align="center">R on CASP14 (test set) (lDDT score)</td>
<td align="center">0.82</td>
<td align="center">0.83</td>
<td align="center">0.80</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The top half shows the values of <italic>&#x3f5;</italic> for each score range. The lower half shows the performance for each combination of values (low/mid/high); R stands for the Pearson correlation coefficient. Results are shown for the validation set (first two rows) and the test set for both GDTTS and lDDT score.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-3">
<title>4.3 Local quality assessment with Q<sub>
<italic>&#x3f5;</italic>
</sub>
</title>
<p>In this section, we demonstrate the ability of Q<sub>
<italic>&#x3f5;</italic>
</sub> to make accurate predictions at the residue level, despite being trained only on global quality scores. This ability is a byproduct of the architecture of the network, where the global predicted score is an average of residue-level node summary scores (see <xref ref-type="fig" rid="F2">Figure 2</xref>). This forces the network to learn accurate local scores, as demonstrated in the results shown in <xref ref-type="table" rid="T7">Table 7</xref>. Similar to the global prediction problem, the performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> is between that of DeepUMQA and DeepUMQA2.</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> and other methods in CASP13 and CASP14 with respect to local lDDT scores.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Dataset</th>
<th align="center">Method</th>
<th align="center">
<italic>R</italic>
<sub>
<italic>local</italic>
</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="center">CASP13</td>
<td align="center">
<italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.80</td>
</tr>
<tr>
<td align="center">DeepUMQA2 (<xref ref-type="bibr" rid="B20">Liu et al., 2022)</xref>
</td>
<td align="center">
<bold>0.868</bold>
</td>
</tr>
<tr>
<td align="center">DeepUMQA (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>
</td>
<td align="center">0.766</td>
</tr>
<tr>
<td align="center">DeepAccNet (<xref ref-type="bibr" rid="B14">Hiranuma et al., 2021)</xref>
</td>
<td align="center">0.740</td>
</tr>
<tr>
<td rowspan="4" align="center">CASP14</td>
<td align="center">
<italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.76</td>
</tr>
<tr>
<td align="center">DeepUMQA2 (<xref ref-type="bibr" rid="B20">Liu et al., 2022)</xref>
</td>
<td align="center">
<bold>0.822</bold>
</td>
</tr>
<tr>
<td align="center">DeepUMQA (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>
</td>
<td align="center">0.680</td>
</tr>
<tr>
<td align="center">DeepAccNet (<xref ref-type="bibr" rid="B14">Hiranuma et al., 2021)</xref>
</td>
<td align="center">0.672</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The local Pearson correlation coefficient (<italic>R</italic>
<sub>
<italic>local</italic>
</sub>) between predicted and known local lDDT scores is provided. The best performance is highlighted in bold. Performance figures for methods other than Q<sub>
<italic>&#x3f5;</italic>
</sub> are quoted from <xref ref-type="bibr" rid="B20">Liu et al. (2022)</xref>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-4">
<title>4.4 Results on CAMEO decoys</title>
<p>For further validation of the performance of Q<sub>
<italic>&#x3f5;</italic>
</sub>, we evaluated its performance on decoys from the CAMEO evaluation project (<xref ref-type="bibr" rid="B13">Haas et al., 2018)</xref>. We downloaded decoys used from 13 May 2022 to 06 May 2023 and followed the same evaluation protocol used by CAMEO: the area under the ROC (AUROC) curve and area under the precision recall (AUPR) curve were calculated using a local lDDT score threshold of 0.6, and the obtained results are shown in <xref ref-type="table" rid="T8">Table 8</xref>. Again, we note that Q<sub>
<italic>&#x3f5;</italic>
</sub> was not trained on local scores (unlike the other methods) and yet is able to perform almost on par with DeepUMQA2. As mentioned previously, this can be traced to the fact that the global prediction score computed by Q<sub>
<italic>&#x3f5;</italic>
</sub> is evaluated by directly averaging local node summary scores, forcing those scores to reflect a local measure of quality.</p>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>Performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> and other methods on the CAMEO dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Dataset</th>
<th align="center">Method</th>
<th align="center">Model</th>
<th align="center">AUROC</th>
<th align="center">AUPR</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="center">CAMEO-QA</td>
<td align="center">Q<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">6,350</td>
<td align="center">0.93</td>
<td align="center">0.88</td>
</tr>
<tr>
<td align="center">DeepUMQA2 (<xref ref-type="bibr" rid="B21">Liu et al., 2023)</xref>
</td>
<td align="center">6,225</td>
<td align="center">
<bold>0.94</bold>
</td>
<td align="center">
<bold>0.89</bold>
</td>
</tr>
<tr>
<td align="center">ProQ3D_LDDT (<xref ref-type="bibr" rid="B36">Uziela et al., 2017)</xref>
</td>
<td align="center">6,498</td>
<td align="center">0.90</td>
<td align="center">0.81</td>
</tr>
<tr>
<td align="center">DeepUMQA (<xref ref-type="bibr" rid="B12">Guo et al., 2022)</xref>
</td>
<td align="center">6,247</td>
<td align="center">0.93</td>
<td align="center">0.86</td>
</tr>
<tr>
<td/>
<td align="center">ModFOLD9 (<xref ref-type="bibr" rid="B27">McGuffin et al., 2023)</xref>
</td>
<td align="center">6,498</td>
<td align="center">0.92</td>
<td align="center">0.87</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The best performance is highlighted in bold. All other results have been taken from the CAMEO website.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-5">
<title>4.5 Performance on AlphaFold2 decoys</title>
<p>In CASP14, AlphaFold2 provided, for the first time, decoys with near experimental resolution (<xref ref-type="bibr" rid="B34">Skolnick et al., 2021)</xref>, with a median GDTTS of 92.4, making it the first team to achieve the highest level of accuracy in CASP. We gathered the decoys submitted by the AlphaFold team (team no 427) from the CASP14 website and evaluated Q<sub>
<italic>&#x3f5;</italic>
</sub> on their decoys. We also ran AlphaFold2 version 2.3.1 on CASP15 single-chain targets. The results of this experiment are shown in <xref ref-type="table" rid="T9">Table 9</xref>. While AlphaFold2 provided better accuracy than our method, its value provided independent validation for the quality of AlphaFold2 predictions. EnQA (<xref ref-type="bibr" rid="B5">Chen et al., 2023)</xref> slightly improves on the quality of AlphaFold2 lDDT score estimates; however, it does so by using the AlphaFold2 scores as one of its features. Therefore, the results of the EnQA method are expected to be highly correlated with those of AlphaFold2 and less useful for independent verification of its predictions.</p>
<table-wrap id="T9" position="float">
<label>TABLE 9</label>
<caption>
<p>Performance of Q<sub>
<italic>&#x3f5;</italic>
</sub> and AlphaFold2 on AlphaFold2-generated decoys in CASP14 and CASP15.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Dataset</th>
<th align="center">Method</th>
<th align="center">
<italic>R</italic>
</th>
<th align="center">
<italic>R</italic>
<sub>
<italic>local</italic>
</sub>
</th>
<th align="center">
<italic>&#x3c1;</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="2" align="center">AlphaFold2-CASP14</td>
<td align="center">
<italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.772</td>
<td align="center">0.730</td>
<td align="center">0.832</td>
</tr>
<tr>
<td align="center">AlphaFold2</td>
<td align="center">
<bold>0.85</bold>
</td>
<td align="center">
<bold>0.792</bold>
</td>
<td align="center">
<bold>0.882</bold>
</td>
</tr>
<tr>
<td rowspan="2" align="center">AlphaFold2-CASP15</td>
<td align="center">
<italic>Q</italic>
<sub>
<italic>&#x3f5;</italic>
</sub>
</td>
<td align="center">0.64</td>
<td align="center">0.60</td>
<td align="center">0.60</td>
</tr>
<tr>
<td align="center">AlphaFold2</td>
<td align="center">
<bold>0.75</bold>
</td>
<td align="center">
<bold>0.72</bold>
</td>
<td align="center">
<bold>0.67</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The global Pearson correlation (<italic>R</italic>), local Pearson correlation (<italic>R</italic>
<sub>
<italic>local</italic>
</sub>), and Spearman rank correlation (<italic>&#x3c1;</italic>) between predicted and known local and global lDDT scores are provided. The best performance is highlighted in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion and future work</title>
<p>In this study, we proposed a novel loss function to enhance the performance of deep learning for quality assessment of decoy structures. Our approach performed close to other state-of-the-art methods while at the same time removing the need for engineered features computed from the protein structure, relying solely on features computed from the decoy sequence, demonstrating what is possible with a pure deep learning approach. These features were integrated using graph convolutional layers that operate at both the atom and residue levels, thereby improving the network&#x2019;s performance. The comparison of our approach with AlphaFold2 indicates there is a need for further research to provide accuracy estimates that improve on the local scores computed by AlphaFold2 in order to provide independent validation of the quality of its predicted structures.</p>
<p>Our approach can be extended in multiple ways. First, although it performs well in predicting local scores, the method is trained using only global quality scores. Joint learning of both global and local scores can potentially improve performance for both tasks. Second, we treated the prediction of GDDTS and lDDT score as independent tasks; there is a potential gain in addressing multiple quality scores at the same time (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020)</xref>. Finally, in this work, we chose to focus on the contribution of the loss function to method performance, so we used a relatively simple graph convolutional network similar to that used in GraphQA (<xref ref-type="bibr" rid="B3">Baldassarre et al., 2020)</xref>. Finally, we expect that the proposed loss function can be applied to regression problems, whose objective is to detect high-quality objects, and has the potential to be a useful addition to any deep learning toolbox.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary Material.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>This study was conceived by AB-H, and all the experiments were performed by SR. SR and AB-H analyzed the results and wrote the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This project was supported by NSF-ABI award #1564840.</p>
</sec>
<ack>
<p>The authors would like to thank Jianlin Cheng for fruitful discussions on the quality assessment problem.</p>
</ack>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akhter</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Chennupati</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Djidjev</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shehu</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Decoy selection for protein structure prediction via extreme gradient boosting and ranking</article-title>. <source>BMC Bioinformatics</source> <volume>21</volume>, <fpage>189</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3523-9</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Al-Lazikani</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Jung</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xiang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Honig</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Protein structure prediction</article-title>. <source>Curr. Opin. Chem. Biol.</source> <volume>5</volume>, <fpage>51</fpage>&#x2013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1016/s1367-5931(00)00164-2</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baldassarre</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Men&#xe9;ndez Hurtado</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Azizpour</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>GraphQA: protein model quality assessment using graph convolutional networks</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>360</fpage>&#x2013;<lpage>366</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa714</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="web">
<collab>CASP</collab> (<year>2021</year>). <article-title>CASP</article-title>. <comment>[Dataset]. Available at: <ext-link ext-link-type="uri" xlink:href="https://predictioncenter.org/download_area/">https://predictioncenter.org/download_area/</ext-link>.</comment>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Morehead</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>3D-equivariant graph neural networks for protein model quality assessment</article-title>. <source>Bioinformatics</source> <volume>39</volume>, <fpage>btad030</fpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btad030</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Choe</surname>
<given-names>M.-H.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>K.-S.</given-names>
</name>
<name>
<surname>Hou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Maghrabi</surname>
<given-names>A. H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Estimation of model accuracy in CASP13</article-title>. <source>Proteins Struct. Funct. Bioinforma.</source> <volume>87</volume>, <fpage>1361</fpage>&#x2013;<lpage>1377</lpage>. <pub-id pub-id-type="doi">10.1002/prot.25767</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Derevyanko</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Grudinin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lamoureux</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep convolutional networks for quality assessment of protein folds</article-title>. <source>Bioinformatics</source> <volume>34</volume>, <fpage>4046</fpage>&#x2013;<lpage>4053</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty494</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Drucker</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Burges</surname>
<given-names>C. J. C.</given-names>
</name>
<name>
<surname>Kaufman</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Smola</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vapnik</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>1996</year>). &#x201c;<article-title>Support vector regression machines</article-title>,&#x201d; in <source>Advances in neural information processing systems</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Mozer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Petsche</surname>
<given-names>T.</given-names>
</name>
</person-group> (<publisher-name>MIT Press</publisher-name>), <volume>9</volume>, <fpage>155</fpage>&#x2013;<lpage>161</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Elnaggar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Heinzinger</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dallago</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rehawi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>ProtTrans: toward understanding the language of life through self-supervised learning</article-title>. <source>IEEE Trans. pattern analysis Mach. Intell.</source> <volume>44</volume>, <fpage>7112</fpage>&#x2013;<lpage>7127</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2021.3095381</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fey</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lenssen</surname>
<given-names>J. E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fast graph representation learning with PyTorch Geometric</article-title>. <comment>
<italic>arXiv preprint arXiv:1903.02428</italic>
</comment>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Fout</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Byrd</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shariat</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ben-Hur</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Protein interface prediction using graph convolutional networks</source>, <volume>30</volume>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>S.-S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>X.-G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>G.-J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>DeepUMQA: ultrafast shape recognition-based protein model quality assessment using deep learning</article-title>. <source>Bioinformatics</source> <volume>38</volume>, <fpage>1895</fpage>&#x2013;<lpage>1903</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btac056</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Barbato</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Behringer</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Studer</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Roth</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bertoni</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Continuous Automated Model EvaluatiOn (CAMEO) complementing the critical assessment of structure prediction in CASP12</article-title>. <source>Proteins Struct. Funct. Bioinforma.</source> <volume>86</volume>, <fpage>387</fpage>&#x2013;<lpage>398</lpage>. <pub-id pub-id-type="doi">10.1002/prot.25431</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hiranuma</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Baek</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Anishchenko</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Dauparas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Baker</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Improved protein structure refinement guided by deep learning based accuracy estimation</article-title>. <source>Nat. Commun.</source> <volume>12</volume>, <fpage>1340</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-21511-x</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hurtado</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Uziela</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep transfer learning in the assessment of the quality of protein models</article-title>. <comment>
<italic>arXiv preprint arXiv:1804.06281</italic>
</comment>.</citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ioffe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Batch normalization: accelerating deep network training by reducing internal covariate shift</article-title>,&#x201d; in <source>International conference on machine learning</source> (<publisher-name>pmlr</publisher-name>), <fpage>448</fpage>&#x2013;<lpage>456</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumper</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Evans</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Figurnov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ronneberger</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Highly accurate protein structure prediction with alphafold</article-title>. <source>Nature</source> <volume>596</volume>, <fpage>583</fpage>&#x2013;<lpage>589</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Adam: A method for stochastic optimization</article-title>. <comment>
<italic>arXiv preprint arXiv:1412.6980</italic>
</comment>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kryshtafovych</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Antczak</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Szachniuk</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zok</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kretsch</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Rangan</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2023</year>). <article-title>New prediction categories in casp15</article-title>. <source>Proteins Struct. Funct. Bioinforma.</source> <pub-id pub-id-type="doi">10.1002/prot.26515</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Improved model quality assessment using sequence and structural information by enhanced deep neural networks</article-title>. <source>Briefings Bioinforma.</source> <volume>24</volume>, <fpage>bbac507</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbac507</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Improved model quality assessment using sequence and structural information by enhanced deep neural networks</article-title>. <source>Briefings Bioinforma.</source> <volume>24</volume>, <fpage>bbac507</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbac507</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lundstr&#xf6;m</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rychlewski</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bujnicki</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Pcons: A neural-network&#x2013;based consensus predictor that improves fold recognition</article-title>. <source>Protein Sci.</source> <volume>10</volume>, <fpage>2354</fpage>&#x2013;<lpage>2362</lpage>. <pub-id pub-id-type="doi">10.1110/ps.08501</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maghrabi</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>McGuffin</surname>
<given-names>L. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Estimating the quality of 3D protein models using the ModFOLD7 server</article-title>. <source>Protein Struct. Predict.</source> <volume>2165</volume>, <fpage>69</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-0716-0708-4_4</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mariani</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Biasini</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Barbato</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Schwede</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests</article-title>. <source>Bioinformatics</source> <volume>29</volume>, <fpage>2722</fpage>&#x2013;<lpage>2728</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt473</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McGuffin</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Adiyaman</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Maghrabi</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>Shuid</surname>
<given-names>A. N.</given-names>
</name>
<name>
<surname>Brackenridge</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Nealon</surname>
<given-names>J. O.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>IntFOLD: an integrated web resource for high performance protein structure and function prediction</article-title>. <source>Nucleic acids Res.</source> <volume>47</volume>, <fpage>W408</fpage>&#x2013;<lpage>W413</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz322</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McGuffin</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Aldowsari</surname>
<given-names>F. M.</given-names>
</name>
<name>
<surname>Alharbi</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Adiyaman</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>ModFOLD8: accurate global and local quality estimates for 3D protein models</article-title>. <source>Nucleic acids Res.</source> <volume>49</volume>, <fpage>W425</fpage>&#x2013;<lpage>W430</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkab321</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McGuffin</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Edmunds</surname>
<given-names>N. S.</given-names>
</name>
<name>
<surname>Genc</surname>
<given-names>A. G.</given-names>
</name>
<name>
<surname>Alharbi</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Salehe</surname>
<given-names>B. R.</given-names>
</name>
<name>
<surname>Adiyaman</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Prediction of protein structures, functions and interactions using the IntFOLD7, MultiFOLD and ModFOLDdock servers</article-title>. <source>Nucleic Acids Res.</source>, <fpage>gkad297</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkad297</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moult</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Pedersen</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Judson</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Fidelis</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>A large-scale experiment to assess protein structure prediction methods</article-title>. <source>Proteins</source>. <comment>[Dataset]</comment>. <pub-id pub-id-type="doi">10.1002/prot.340230303</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olechnovi&#x10d;</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Venclovas</surname>
<given-names>&#x10c;.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>VoroMQA: assessment of protein structure quality using interatomic contact areas</article-title>. <source>Proteins Struct. Funct. Bioinforma.</source> <volume>85</volume>, <fpage>1131</fpage>&#x2013;<lpage>1145</lpage>. <pub-id pub-id-type="doi">10.1002/prot.25278</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pag&#xe8;s</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Charmettant</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Grudinin</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Protein model quality assessment using 3D oriented convolutional neural networks</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>3313</fpage>&#x2013;<lpage>3319</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz122</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Paszke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gross</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Massa</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lerer</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bradbury</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chanan</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>PyTorch: an imperative style, high-performance deep learning library</article-title>. <source>Adv. neural Inf. Process. Syst.</source> <volume>32</volume>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ray</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lindahl</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Wallner</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Improved model quality assessment using ProQ2</article-title>. <source>BMC Bioinforma.</source> <volume>13</volume>, (<issue>224</issue>). <pub-id pub-id-type="doi">10.1186/1471-2105-13-224</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shehu</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>A review of evolutionary algorithms for computing functional conformations of protein molecules</article-title>. <source>Computer-Aided Drug Discov.</source>, <fpage>31</fpage>&#x2013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1007/7653_2015_47</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Skolnick</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>AlphaFold 2: why it works and its implications for understanding the relationships of protein sequence, structure, and function</article-title>. <source>J. Chem. Inf. Model.</source> <volume>61</volume>, <fpage>4827</fpage>&#x2013;<lpage>4831</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jcim.1c01114</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Uziela</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wallner</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>ProQ3: improved model quality assessments using Rosetta energy terms</article-title>. <source>Sci. Rep.</source> <volume>6</volume>, <fpage>33509</fpage>&#x2013;<lpage>33510</lpage>. <pub-id pub-id-type="doi">10.1038/srep33509</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Uziela</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Menendez Hurtado</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wallner</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>ProQ3D: improved model quality assessments using deep learning</article-title>. <source>Bioinformatics</source> <volume>33</volume>, <fpage>1578</fpage>&#x2013;<lpage>1580</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btw819</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Varadi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Anyango</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Deshpande</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nair</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Natassia</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yordanova</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>AlphaFold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models</article-title>. <source>Nucleic acids Res.</source> <volume>50</volume>, <fpage>D439</fpage>&#x2013;<lpage>D444</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkab1061</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wallner</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Elofsson</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Can correct protein models be identified?</article-title> <source>Protein Sci.</source> <volume>12</volume>, <fpage>1073</fpage>&#x2013;<lpage>1086</lpage>. <pub-id pub-id-type="doi">10.1110/ps.0236803</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zemla</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>LGA: a method for finding 3D similarities in protein structures</article-title>. <source>Nucleic Acids Res.</source> <volume>31</volume>, <fpage>3370</fpage>&#x2013;<lpage>3374</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkg571</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>