<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioinform.</journal-id>
<journal-title>Frontiers in Bioinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioinform.</abbrev-journal-title>
<issn pub-type="epub">2673-7647</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fbinf.2021.657529</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioinformatics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deriving and Using Descriptors of Elementary Functions in Rational Protein Design</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Yin</surname> <given-names>Melvin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Goncearenco</surname> <given-names>Alexander</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/983053/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Berezovsky</surname> <given-names>Igor N.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/163898/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Bioinformatics Institute, Agency for Science, Technology, and Research (A<sup>&#x0002A;</sup>STAR)</institution>, <addr-line>Singapore</addr-line>, <country>Singapore</country></aff>
<aff id="aff2"><sup>2</sup><institution>National Center for Biotechnology Information, National Institute of Health (NIH)</institution>, <addr-line>Bethesda, MD</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Biological Sciences (DBS), National University of Singapore (NUS)</institution>, <addr-line>Singapore</addr-line>, <country>Singapore</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Michael Gromiha, Indian Institute of Technology Madras, India</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Selvaraj Samuel, Bharathidasan University, India; Kumar Yugandhar, Cornell University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Igor N. Berezovsky <email>igorb&#x00040;bii.a-star.edu.sg</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Protein Bioinformatics, a section of the journal Frontiers in Bioinformatics</p></fn>
<fn fn-type="other" id="fn002"><p>&#x02020;These authors have contributed equally to this work</p></fn></author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>04</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>1</volume>
<elocation-id>657529</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>01</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>03</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Yin, Goncearenco and Berezovsky.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Yin, Goncearenco and Berezovsky</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>The rational design of proteins with desired functions requires a comprehensive description of the functional building blocks. The evolutionary conserved functional units constitute nature&#x00027;s toolbox; however, they are not readily available to protein designers. This study focuses on protein units of subdomain size that possess structural properties and amino acid residues sufficient to carry out elementary reactions in the catalytic mechanisms. The interactions within such elementary functional loops (ELFs) and the interactions with the surrounding protein scaffolds constitute the descriptor of elementary function. The computational approach to deriving descriptors directly from protein sequences and structures and applying them in rational design was implemented in a proof-of-concept DEFINED-PROTEINS software package. Once the descriptor is obtained, the ELF can be fitted into existing or novel scaffolds to obtain the desired function. For instance, the descriptor may be used to determine the necessary spatial restraints in a fragment-based grafting protocol. We illustrated the approach by applying it to well-known cases of ELFs, including phosphate-binding P-loop, diphosphate-binding glycine-rich motif, and calcium-binding EF-hand motif, which could be used to jumpstart templates for user applications. The DEFINED-PROTEINS package is available for free at <ext-link ext-link-type="uri" xlink:href="https://github.com/MelvinYin/Defined_Proteins">https://github.com/MelvinYin/Defined_Proteins</ext-link>.</p></abstract>
<kwd-group>
<kwd>protein function</kwd>
<kwd>protein design</kwd>
<kwd>elementary functional loops</kwd>
<kwd>elementary function</kwd>
<kwd>descriptor of the elementary function</kwd>
<kwd>DEFINED-PROTEINS software package</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="62"/>
<page-count count="11"/>
<word-count count="6678"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Contemporary views of the enzymatic functions are dominated by either consideration of functional domains or the catalytic active sites as their minimal structural/functional units (Marchler-Bauer et al., <xref ref-type="bibr" rid="B40">2015</xref>; Finn et al., <xref ref-type="bibr" rid="B20">2016</xref>; Trudeau and Tawfik, <xref ref-type="bibr" rid="B59">2019</xref>). The relationships in protein function evolution, however, are far more complex than the sequence-based models can describe (Nath et al., <xref ref-type="bibr" rid="B43">2014</xref>; Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B27">2015</xref>; Aziz et al., <xref ref-type="bibr" rid="B3">2016</xref>; Romero Romero et al., <xref ref-type="bibr" rid="B46">2016</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>). In many cases, the closely related structurally similar folds can carry completely different biochemical functions, while, on the other hand, the same function can be performed by many different protein folds (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B27">2015</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). According to the Enzyme Commission number (EC) nomenclature (Bairoch, <xref ref-type="bibr" rid="B5">2000</xref>), the current number of enzymatic functions reaches up to 5,000. The number of underlying biochemical mechanisms (Holliday et al., <xref ref-type="bibr" rid="B30">2012</xref>), however, does not exceed 500; moreover, the number of elementary chemical reactions (Holliday et al., <xref ref-type="bibr" rid="B31">2005</xref>) is less than 50. The differences in the order of magnitude between the numbers of enzymatic functions, biochemical mechanisms, and elementary chemical reactions prompt one to consider the biochemical function as a combination of elementary ones. Also, the domains themselves had to evolve from some primitive forms (Berezovsky, <xref ref-type="bibr" rid="B7">2003</xref>, <xref ref-type="bibr" rid="B8">2019</xref>; Nath et al., <xref ref-type="bibr" rid="B43">2014</xref>; Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B27">2015</xref>; Aziz et al., <xref ref-type="bibr" rid="B3">2016</xref>; Romero Romero et al., <xref ref-type="bibr" rid="B46">2016</xref>, <xref ref-type="bibr" rid="B47">2018</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>,<xref ref-type="bibr" rid="B11">b</xref>). It has been hypothesized that the first group of enzymatic domains emerged as combinations of prebiotic (Romero Romero et al., <xref ref-type="bibr" rid="B46">2016</xref>) ring-like peptides with simple chemical transformation (Trifonov et al., <xref ref-type="bibr" rid="B58">2001</xref>; Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B27">2015</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). Closed loops of preferential 25 to 35-amino acid residue size, a universal basic element of soluble proteins, are the descendants of the prebiotic peptides (Berezovsky et al., <xref ref-type="bibr" rid="B9">2000</xref>, <xref ref-type="bibr" rid="B10">2017a</xref>,<xref ref-type="bibr" rid="B11">b</xref>; Berezovsky, <xref ref-type="bibr" rid="B7">2003</xref>) determined by the polymer nature of polypeptide chains (Yamakawa and Stockmayer, <xref ref-type="bibr" rid="B61">1972</xref>; Shimada and Yamakawa, <xref ref-type="bibr" rid="B55">1984</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B9">2000</xref>, <xref ref-type="bibr" rid="B10">2017a</xref>; Orevi et al., <xref ref-type="bibr" rid="B44">2013</xref>; Jacob et al., <xref ref-type="bibr" rid="B36">2018</xref>).</p>
<p>Previous studies showed that biochemical functions can be represented as a combination of the elementary ones, provided by the elementary functional loops (EFLs), which are closed loops with specific signatures that perform elementary steps of biochemical transformations (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B24">2010</xref>, <xref ref-type="bibr" rid="B25">2011</xref>, <xref ref-type="bibr" rid="B26">2012</xref>, <xref ref-type="bibr" rid="B27">2015</xref>). We suggest that EFLs can be considered as potential elementary units in the design of biochemical functions, which requires an exhaustive description of their characteristics that are important for building the required catalytic site in the environment of the particular protein fold. Even though EFLs perform only elementary steps of the biochemical reactions, the relationship between their sequences, structures, and functions is complex. For example, CxxC motifs (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B25">2011</xref>) are known to be involved in a variety of functions, such as metal or metal-containing cofactor binding or redox reactions. Consequently, structures of the EFLs containing this signature in many folds differ significantly (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B26">2012</xref>; Zheng et al., <xref ref-type="bibr" rid="B62">2016</xref>), depending on both the structural environment in the protein and its overall biochemical function. The interactions between the EFL and the substrate and between the EFL and the rest of the structure will also depend on the fold and its function (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). Therefore, the descriptor of EFL has to capture structure- and function-dependent interaction propensities.</p>
<p>The protein design adventure (Das and Baker, <xref ref-type="bibr" rid="B18">2008</xref>) started more than 50 years ago from the general protein folding problem (Dill and MacCallum, <xref ref-type="bibr" rid="B19">2012</xref>) formulated in terms of polymer and statistical physics (Shakhnovich and Gutin, <xref ref-type="bibr" rid="B54">1993</xref>) of biomolecules (Sali et al., <xref ref-type="bibr" rid="B51">1994</xref>; Shakhnovich, <xref ref-type="bibr" rid="B53">2006</xref>), in terms of statistical predictions of structures from the sequence (Sippl, <xref ref-type="bibr" rid="B57">1990</xref>; Crippen, <xref ref-type="bibr" rid="B17">1996</xref>), and as an inverse protein-folding problem (Rooman et al., <xref ref-type="bibr" rid="B49">1990</xref>; Bowie et al., <xref ref-type="bibr" rid="B15">1991</xref>; Rooman and Wodak, <xref ref-type="bibr" rid="B50">1995</xref>) of finding the sequence that can be threaded into the certain fold structure. Current progress in the evolutionary-inspired fragment-based (Hocker, <xref ref-type="bibr" rid="B29">2014</xref>) and <italic>de novo</italic> (Huang et al., <xref ref-type="bibr" rid="B32">2016a</xref>; Silva et al., <xref ref-type="bibr" rid="B56">2019</xref>) design is described in several original works (Brunette et al., <xref ref-type="bibr" rid="B16">2015</xref>) and reviews (Lechner et al., <xref ref-type="bibr" rid="B39">2018</xref>; Baker, <xref ref-type="bibr" rid="B6">2019</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). It is further facilitated by advances in machine learning and artificial intelligence, as well as by the quality and quantity of high-throughput sequence and structural data available, leading to significant improvements in the performance of computational approaches (Senior et al., <xref ref-type="bibr" rid="B52">2020</xref>). Despite significant progress in the computational design of protein structures, the journey toward solving the great challenge of the <italic>de novo</italic> design of protein functions is, as of yet, at its very beginning (Huang et al., <xref ref-type="bibr" rid="B32">2016a</xref>; Lechner et al., <xref ref-type="bibr" rid="B39">2018</xref>; Baker, <xref ref-type="bibr" rid="B6">2019</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). Although the repertoire of conserved continuous functional units is available on the sequence level, a more comprehensive characterization is required to define spatial and interaction restraints (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). The computational framework presented in this study facilitates the derivation of the descriptor of elementary function and conceptualizes the objective function for protein engineering and design applications using the descriptor. It merges structure, sequence, and interaction features important for defining elementary functions on a residue level. In protein design, descriptors of elementary functional units serve as off-the-shelf building blocks, for instance, in protein grafting, while the objective function optimizes the choice of such building blocks from a library, considering their geometry and interactions with the protein scaffold, particularly in the key catalytic or binding residues.</p>
<p>We illustrate this approach by calculating the descriptors for three ubiquitous ELFs: the calcium-binding EF-hand, the phosphate-binding in mononucleotide-containing ligands, and the phosphate-binding in dinucleotide-containing ligands, such as ATP, nicotinamide&#x02013;adenine&#x02013;dinucleotide (NAD), and NAD phosphate (NADP), in a variety of structural scaffolds. We also model a hypothetical grafting experiment by swapping the EFLs and EFL-derived chimeras among the scaffolds.</p></sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and Methods</title>
<p>The proof-of-concept computational framework is aimed at, first, derivation of the descriptor of elementary function and, second, application of descriptors to design proteins with desired structures and functions. A descriptor represents a set of characteristics of the EFL (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B24">2010</xref>, <xref ref-type="bibr" rid="B25">2011</xref>, <xref ref-type="bibr" rid="B26">2012</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>), including the position-specific information on the sequence, and several structural features encoded as probabilistic distributions (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). An elementary function is defined as the smallest structural unit sufficient to carry out an elementary reaction in a biochemical transformation. Depending on the protein engineering task, it may be needed to introduce or replace an elementary function in an existing protein of interest or build and design a protein with the required structure and function <italic>de novo</italic> (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). The flowchart in <xref ref-type="supplementary-material" rid="SM3">Supplementary Figure 1</xref> illustrates the sequence of steps described below.</p>
<sec>
<title>Deriving the Descriptor</title>
<p>Motivated by the biophysical constraints of the polypeptide chain (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B24">2010</xref>, <xref ref-type="bibr" rid="B25">2011</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>), the procedure starts from 30-residue long seed sequence fragments of the functional loops represented by a gapless multiple sequence alignment and a position-specific scoring matrix (PSSM) profile. The sequence profile is then iteratively scanned against the non-redundant UniRef database (Hunter et al., <xref ref-type="bibr" rid="B35">2009</xref>) with an expectation&#x02013;maximization (EM)-like algorithm (Goncearenco and Berezovsky, <xref ref-type="bibr" rid="B24">2010</xref>, <xref ref-type="bibr" rid="B25">2011</xref>) converging to an expanded sequence profile of the descriptor with a functional signature (Berezovsky et al., <xref ref-type="bibr" rid="B12">2003a</xref>,<xref ref-type="bibr" rid="B13">b</xref>). The corresponding structures characterizing the functional loop are then obtained by looking for profile matches against the sequences in the Protein Data Bank (Berman et al., <xref ref-type="bibr" rid="B14">2000</xref>; wwPDB consortium, <xref ref-type="bibr" rid="B60">2019</xref>), followed by extraction and encoding of the structural feature in the form of parametrized probability distribution functions. Structural and functional annotations are extracted from the corresponding databases, such as Uniprot, MaCiE, and Conserved Domains Database (CDD) (Bairoch, <xref ref-type="bibr" rid="B5">2000</xref>; Andreini et al., <xref ref-type="bibr" rid="B2">2009</xref>; Fischer et al., <xref ref-type="bibr" rid="B21">2010</xref>; Holliday et al., <xref ref-type="bibr" rid="B30">2012</xref>; Marchler-Bauer et al., <xref ref-type="bibr" rid="B41">2013</xref>; Akiva et al., <xref ref-type="bibr" rid="B1">2014</xref>; Furnham et al., <xref ref-type="bibr" rid="B22">2014</xref>). The structural features include dihedral angles, Van der Waals (VdW) interactions, and hydrogen bonds (H-bonds). Intra-EFL interactions and interactions between the functional loop and the rest of the structure are encoded separately. Thus, a descriptor contains information about the immediate environment of the functional loop in all protein scaffolds and enzymatic functions where it was encountered.</p></sec>
<sec>
<title>Objective Function for Protein Engineering and Design</title>
<p>In protein engineering and <italic>de novo</italic> design, the descriptors of elementary function need to be integrated into a given structural scaffold. Once the elementary function that is required to be incorporated into the protein is selected, the sequence, structure, and interactions that would fit best into the scaffold have to be determined. The objective function scores how well a given structural loop fits in and should be maximized to obtain the best matching implementation of the descriptor with an assumption that the native structure has the best fit. Essentially, the score represents the joint likelihood of all amino acid positions in the grafted loop with respect to distributions parametrized in the descriptor with N-residue positions and M features: <inline-formula><mml:math id="M1"><mml:mtext>&#x000A0;</mml:mtext><mml:mi>F</mml:mi><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:msubsup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. The weight that is given to a residue position <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow></mml:msubsup></mml:math></inline-formula> reflects the relative degree of conservation of features in each position: <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>j</mml:mtext><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>j</mml:mtext><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:math></inline-formula>. Descriptor features <italic>j</italic> enumerate along a sequence signature (<italic>j</italic> &#x0003D; <italic>a</italic>), dihedral angles (<italic>j</italic> &#x0003D; <italic>d</italic>), H-bonds (<italic>j</italic> &#x0003D; <italic>h</italic>), and vdW interactions (<italic>j</italic> &#x0003D; <italic>v</italic>), with the corresponding scores <italic>S</italic><sub>ij</sub> and weights <italic>w</italic><sub>ij</sub>. Score <italic>S</italic><sub><italic>i,a</italic></sub> is the log-odds score for amino acid substitution according to BLOSUM62 (Henikoff and Henikoff, <xref ref-type="bibr" rid="B28">1993</xref>). The weight of the residue feature is given by the relative frequency of the two most frequent residues in the sequence profile <inline-formula><mml:math id="M4"><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext><mml:mo>,</mml:mo><mml:mtext>a&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>argma</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mover class="msup"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mn>20</mml:mn></mml:mrow></mml:mover><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>, where w<sub>k</sub> is amino acid frequency. Dihedral angles are first clustered as two-dimensional vector quantities using the EM algorithm as implemented in scikit-learn, and their weights and fitting scores are derived from parameters in the trained model with the following equation:</p>
<p><inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext><mml:mo>,</mml:mo><mml:mtext>d</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mo class="qopname">erf</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo class="qopname">&#x02211;</mml:mo><mml:msup><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msup><mml:mtext>CD</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>erf</italic> refers to the error function, D = X - Y is the displacement of the compared dihedral angle points, normalized by the median of all points in the descriptor, expressed as phi-psi angles, and <italic>C</italic> is a matrix proportional to the degree of confidence that the point belongs to the distribution. Given the precision matrix &#x0039B;&#x0003D;&#x003C3;<sup>&#x02212;2</sup> and posterior probability matrix <italic>P</italic> obtained from the clustered model, C = &#x0039B; P.</p>
<p>The isotropic VdW distribution score, S<sub>i,v</sub>, is based on counting the number of VdW contacts with non-hydrogen atoms within the 5-&#x000C5; radius. Hydrogen bonds are additionally split into acceptors and donors: S<sub>i,hA</sub> and S<sub>i,hD</sub>, respectively. They are scalar quantities and are measured as the number of donor H-bonds present at the residue position. For H-bonds, the effective radius is 3.5 &#x000C5;. The score is the ratio between the absolute difference in the number of bonds (&#x003B4;) to the higher number of contacts for either structure (VdW interactions and H-bonds): <inline-formula><mml:math id="M6"><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>d</mml:mtext><mml:mo>,</mml:mo><mml:mtext>v</mml:mtext><mml:mo>,</mml:mo><mml:mtext>hA</mml:mtext><mml:mo>,</mml:mo><mml:mtext>hD</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>abs</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>&#x003B4;</mml:mtext></mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mtext>max</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>&#x003B4;</mml:mtext></mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula>. The weight of the feature is the SD of the spread of a half-normal distribution fitted onto the data <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mtext>i&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>b</mml:mtext><mml:mo>,</mml:mo><mml:mtext>c</mml:mtext><mml:mo>,</mml:mo><mml:mtext>d</mml:mtext><mml:mo>,</mml:mo><mml:mtext>e</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi></mml:math></inline-formula>. Each feature weight is further tuned by an empirically derived scalar factor to account for differences in the spread of absolute values returned by the scoring functions.</p></sec></sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Deriving the Descriptors of EFs</title>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> contains examples of some of the structural features for three descriptors: the calcium-binding Ca<sup>2&#x0002B;</sup>-binding helix&#x02013;loop&#x02013;helix EF-hand motif (Gifford et al., <xref ref-type="bibr" rid="B23">2007</xref>) (with the characteristic signature DxDxD, <xref ref-type="fig" rid="F1">Figure 1A</xref>), the glycine-rich motif of the phosphate-binding in dinucleotide-containing ligands (with the characteristic signature GxGxxG, <xref ref-type="fig" rid="F1">Figure 1B</xref>), and the phosphate-binding P-loop in nucleotide-containing ligands (with the characteristic signature GxxGxG, <xref ref-type="fig" rid="F1">Figure 1C</xref>). For EF-hand (Gifford et al., <xref ref-type="bibr" rid="B23">2007</xref>), dihedral angles on amino acid positions before and after the calcium-binding sites form tight clusters, i.e., the backbone is structurally conserved. Donor-acceptor hydrogen bond pairs spaced three to four residue positions apart are also present in the first and last 10 residues of the EF-section (<xref ref-type="fig" rid="F1">Figure 1A</xref>), representing, together with conserved dihedral angles, two &#x003B1;-helices in the structural motif. Distributions of VdW contacts in the EF-hand motif show a large proportion of the conserved intrinsic contacts flanking the &#x003B1;-helices, whereas the central flexible link between them forms vital external contacts with the rest of the fold. The same patterns of the internal contacts are observed in the second halves of the GxGxxG and GxxGxG motif structures also containing the &#x003B1;-helices. The &#x003B2;-strands in these loops, as expected, interact more strongly with the rest of the structure via VdW interactions (<xref ref-type="supplementary-material" rid="SM4">Supplementary Figure 2</xref>) and H-bonds (<xref ref-type="supplementary-material" rid="SM5">Supplementary Figures 3</xref>, <xref ref-type="supplementary-material" rid="SM6">4</xref>). In addition to the interactions with ligands, the substrate-binding residues in the descriptor are characterized by multiple interactions with the fold.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>(A)</bold> Descriptors for Ca<sup>2&#x0002B;</sup>-binding helix-loop-helix EF-hand motif (Gifford et al., <xref ref-type="bibr" rid="B23">2007</xref>), <bold>(B)</bold> phosphate-binding loop in dinucleotide-containing ligands, and <bold>(C)</bold> phosphate-binding P-loop. Charts in the left column show per-residues numbers of Van der Waals (VdW) interactions and hydrogen bonds (H-bonds); contact maps in the right column show the total number of VdW contacts in H-bonds in all structures used for the derivation of corresponding descriptors. The central column contains examples of structures used for derivation of corresponding descriptors (represented in the form of structurally aligned segments) along with the logos representing the position-specific matrices of derived descriptors. The numbering of residues is sequential and follows positions in the logo.</p></caption>
<graphic xlink:href="fbinf-01-657529-g0001.tif"/>
</fig></sec>
<sec>
<title>Swapping the EFLs Within the Same Structural/Functional Context</title>
<p>The objective function was calibrated and exemplified on three motifs: (i) the EF-hand from <italic>Bos taurus</italic> calcium-binding protein structure (PDBID 1A29) with the Ca<sup>2&#x0002B;</sup> ligand; (ii) the phosphate-binding motif (GxGxxG) from <italic>Equus caballus</italic> oxidoreductase (PDBID 1A71) with dinucleotide-containing NAD ligand; and (iii) the phosphate-binding motif (GxxGxG) from <italic>Methanocaldococcus jannaschii</italic> ABC transporter (PDBID 1G6H) with nucleotide-containing an ADP ligand. These motifs were replaced with the realization of the corresponding descriptors (see <xref ref-type="supplementary-material" rid="SM7">Supplementary Figure 5</xref> and explanations in the figure caption).</p>
<p><xref ref-type="supplementary-material" rid="SM8">Supplementary Figure 6</xref> illustrates how objective function can be used to obtain the best descriptor realization in the corresponding engineering or design tasks. First, a segment of seven consecutive residues (half of the typical functional signature; Berezovsky et al., <xref ref-type="bibr" rid="B10">2017a</xref>) with the highest cumulative fitting score is determined. Then, this segment is extended with residues that contribute the highest scores to the fitting function. There might be insertions/deletions in the descriptor realization, or the reference could be undefined as in the case of missing data, with disordered or structurally unresolved regions in the query structure. In such a case, the algorithm first scores all valid residue positions and finds the appropriate segments to merge if the best match does not come from a singular structure. Next, adjacent segments are extended, possibly from both ends in the case of a gap, to fill in the uncharacterized positions. This returns the best-effort re-engineered structure that partially relies on the input structure and fills in the rest, where data are missing. <xref ref-type="supplementary-material" rid="SM8">Supplementary Figure 6</xref> exemplifies the case of descriptor realization that builds the functional loops out of two segments, providing the most optimal score.</p></sec>
<sec>
<title>Swapping EFLs Between Different Functions and Structures in the Cross-Validation Experiment</title>
<p>To assess the robustness of the derived descriptors and versatility of the fitting function, we selected seven structures with the phosphate-binding functionality, but of different folds, origins, and ligands: NADP-binding <italic>Thermotoga maritima</italic> lactate dehydrogenase (1A5Z) and <italic>Methylobacterium extorquens AM1</italic> methylene-tetrahydromethanopterin dehydrogenase (1LUA), AMP-binding <italic>Escherichia coli</italic> MoeB-MoaD protein complex (1JWB), flavin&#x02013;adenine dinucleotide (FAD)-binding <italic>Klebsiella pneumoniae</italic> udp-galactopyranose mutase (2BI7) and <italic>Thermus thermophilus GidA-related protein</italic>, and both binding sites for FAD and NADP in <italic>E. coli</italic> CoA reductase (1PS9) and human dihydrolipoamide dehydrogenase (1ZMC).</p>
<p>The donor functional loop that had to be transplanted is an elementary function of phosphate binding in dinucleotide-containing ligands (GxGxxG). We derived its descriptor from the dataset that excluded the aforementioned structures. Seven proteins representing variations in the structure and sequences of the phosphate-binding functional loops (both in nucleotide- and dinucleotide-containing ligands) were used as the targets or acceptors for the EFL-replacement procedure, using the above descriptor. <xref ref-type="fig" rid="F2">Figure 2</xref> illustrates the fit between the re-engineered functional loop into the original structure, which binds various ligands, including AMP, FAD, and NAD(P). Additionally, recombinant functional loops can be built from the best matching segments of multiple structures. Taking 1PS9 structure as an example, its NADP binding site is replaced with the loops from structures 1KF6 from <italic>E. coli</italic> quinol-fumarate reductase; 1VRP from <italic>C. sarcosine</italic> oxidase; 6GNC from <italic>Clostridium acetobutylicum</italic> thioredoxin reductase; and 5ER0 from <italic>Lactobacillus</italic> oxidase. All of these structures contain FAD-binding sites replacing the original NADP-binding site.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Realizations of the descriptor of Gly-rich signature of the phosphate binding in dinucleotide-containing ligands. <bold>(A&#x02013;G)</bold> PDB IDs of corresponding structures are: 1A5Z, 1JWB, 1LUA, 1PS9, 1ZMC, 2BI7, and 2CUL, respectively.</p></caption>
<graphic xlink:href="fbinf-01-657529-g0002.tif"/>
</fig></sec>
<sec>
<title>Proof-of-Principle Grafting Experiment Using Descriptors of the Phosphate-Binding EFLs With GxGxxG and GxxGxG Signatures</title>
<p>To further illustrate the utility of descriptors of elementary functions and the potential of the DEFINED-PROTEINS package in protein design applications, we set up initial conditions for the grafting procedure where we replaced the original functional loop in a given protein with the non-native one representing a different elementary function. We recruited two distinct, but not opposite, elementary functions, for which we have already derived descriptors: binding of the phosphate in dinucleotide-(GxGxxG) and nucleotide-containing P-loop (GxxGxG) ligands (Zheng et al., <xref ref-type="bibr" rid="B62">2016</xref>). Thereby, we cross-grafted nucleotide-binding function into protein folds with native dinucleotide-ligand binding (<xref ref-type="fig" rid="F3">Figures 3A&#x02013;C</xref> and <xref ref-type="supplementary-material" rid="SM9">Supplementary Figures 7A&#x02013;C</xref>), and cross-grafted dinucleotide-binding function into folds with the native mononucleotide binding ability (<xref ref-type="fig" rid="F3">Figures 3D&#x02013;F</xref> and <xref ref-type="supplementary-material" rid="SM9">Supplementary Figures 7D&#x02013;F</xref>), respectively. Although both are elementary functions of the phosphate binding, the latter belong to different dinucleotide- and nucleotide-containing ligands that determine the corresponding diversity of protein functions that these elementary functions evolved into.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>(A&#x02013;F)</bold> Cross-grafting of functional loops of the phosphate binding in dinucleotide (GxGxxG) and nucleotide-containing (GxxGxG) ligands. <bold>(A&#x02013;C)</bold> Grafting of the phosphate-binding signature in dinucleotide-containing ligands (GxGxxG) in proteins with P-loop (GxxGxG) elementary function; recombinant realizations of descriptors are shown in proteins with PDB IDs: 1H5Y, 1BWV, and 1O5K. <bold>(D&#x02013;F)</bold> Recombinant realizations of the descriptor of P-loop (GxxGxG) elementary function in proteins (1SKY, 1II2, and 1NI3, respectively) with elementary functional loops of the phosphate-binding in dinucleotide-containing ligands (GxGxxG). The original structures are shown in green.</p></caption>
<graphic xlink:href="fbinf-01-657529-g0003.tif"/>
</fig>
<p>The signatures of EFLs and the most frequent interactions with distinct parts of ligands were discussed elsewhere (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>), showing both conservatism and diversity depending on the position in the functional loop. From the structural perspective, both EFLs have similar secondary structures of &#x003B2;-turn-&#x003B1;-helix composition and architecture. The latter results in a high number of intrinsic H-bonds and VdW contacts in its second &#x003B1;-helical part, while &#x003B2;-strand elements of both loops form more contacts with the surrounding structure. At the same time, the dinucleotide-binding-GxGxxG-loop is a more compact structure by itself, with more intrinsic contacts between its &#x003B1; and &#x003B2; elements. We replaced the original functional segments with recombinant (<xref ref-type="fig" rid="F3">Figure 3</xref>) and deconvoluted single-loop (<xref ref-type="supplementary-material" rid="SM9">Supplementary Figure 7</xref>) functional loops sampled from the corresponding descriptors. In all of the cases presented in <xref ref-type="fig" rid="F3">Figure 3</xref> and <xref ref-type="supplementary-material" rid="SM9">Supplementary Figure 7</xref> replacements were done with the highest-scoring matches.</p>
<p><xref ref-type="fig" rid="F3">Figures 3A&#x02013;C</xref> and <xref ref-type="supplementary-material" rid="SM9">Supplementary Figures 7A&#x02013;C</xref> show results of grafting of the P-loop (GxxGxG) descriptor in places of Gly-rich signature (GxGxxG) of the phosphate-binding in dinucleotide-containing ligands. While obtaining a good match between the original and replacement loops in both unblended single-loop and recombinant cases, the latter was allowed to obtain a better per-residue fit between the original and replacement loops. Recombinant loop replacements consist of several fragments (typically, two to three), covering most of the loop (see examples in <xref ref-type="fig" rid="F4">Figure 4</xref> and <xref ref-type="supplementary-material" rid="SM10">Supplementary Figure 8</xref>) with higher scores in the functional positions at the expense of non-function-bearing positions in the secondary structure elements. It agrees with less conserved dihedral angles in the turn segments of the loops that contain functional signatures, which also result in different angles between the secondary structure elements of the loop and difference in the intra-loop contacts in GxxGxG and GxGxxG loops.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Alignments for recombinant cross-grafting of the phosphate-binding loop in dinucleotide (GxGxxG) and nucleotide-containing (GxxGxG) ligands. <bold>(A)</bold> Replacement of the P-loop elementary function in 1NI3 and <bold>(B)</bold> the phosphate-binding signature in dinucleotide-containing ligands in 1BWV. The thick, thin, and arrow sections represent the beta sheet, turn section, and alpha helix secondary structures, respectively.</p></caption>
<graphic xlink:href="fbinf-01-657529-g0004.tif"/>
</fig>
<p><xref ref-type="fig" rid="F4">Figure 4</xref> and <xref ref-type="supplementary-material" rid="SM10">Supplementary Figure 8</xref> show a comparison of secondary structure alignments in several examples of the single-loop and recombinant replacements. The secondary structure is affected by the environment (Minor and Kim, <xref ref-type="bibr" rid="B42">1996</xref>), pointing to the need for adjustments and optimization in length and location after the original descriptor realization is placed instead of the natural ELF. Overall, examples of structural replacements (<xref ref-type="fig" rid="F3">Figures 3</xref> and <xref ref-type="supplementary-material" rid="SM9">Supplementary Figure 7</xref>) and alignments (<xref ref-type="fig" rid="F4">Figure 4</xref> and <xref ref-type="supplementary-material" rid="SM10">Supplementary Figure 8</xref>) show that realizations of descriptors can be used as a starting point for further optimization of positions and interactions involving functional residues and structures of elementary functional loops to obtain required modifications/design in the context of new protein structures and functions.</p>
<p>It is important to note that, in general, scores obtained with the objective function depend on the sequence/structure characteristics of the elementary function and the type of procedure in which the descriptor is used. The major contributors to the objective function score are sequence conservation and dihedral angles, with weights 0.62/0.52 and 0.3/0.38 for the GxxGxG/GxGxxG signatures, respectively. Weights for VdW interactions and H-bonds are smaller (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref> legend for details): the VdW interactions are weak though omnipresent, whereas H-bonds between the loops and rest of the folds are rare, making the weights of both smaller. Nevertheless, contributions of all characteristics to the weight for each individual position are calculated as a sum of their weights normalized by the total feature weights across all residue positions. The score of the reference state for the descriptor, which can be indicative of the latter and can be used as a guideline in design, should be obtained for each descriptor. The straightforward way is to perform the cross-validation experiment, which assesses the robustness of the descriptor and provides the score that can be used as a ground state score. <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1B</xref> contains averaged scores obtained in the cross-validation experiment (<xref ref-type="fig" rid="F2">Figure 2</xref>), which can be used as a reference for the engineering and design of corresponding descriptors in other folds. As an example, <xref ref-type="supplementary-material" rid="SM2">Supplementary Table 2B</xref> shows scores in case of cross-grafting of descriptors of the phosphate-binding signatures in the nucleotide- (GxxGxG) and dinucleotide-containing (GxGxxG) ligands, revealing the deviation from scores in cross-validation that can further increase in case of <italic>de novo</italic> design of functions based on the descriptors. The utility of averaged (<xref ref-type="supplementary-material" rid="SM1">Supplementary Tables 1A</xref>, <xref ref-type="supplementary-material" rid="SM2">2A</xref>) and per-residue scores is in the information on the relative conservatism (importance) of positions in the descriptor and of the overall match between its realization and the rest of the fold and its capacity to contribute to engineered/designed function. For example, depending on the requirements on the interactions within the ELF and between the loop and the rest of the fold, H-bonds can be introduced in positions with a conservatism level allowing to do so, but the same holds for changing other characteristics.</p></sec></sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>There are two major tasks, engineering/modification and <italic>de novo</italic> design, in which descriptors of elementary functions can be used (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). The former is a modification of the natural protein function by replacing one or several natural functional loops (elementary functions) in the protein with other elementary function(s) encoded in the corresponding descriptor(s). The goal of this engineering effort can be to change a substrate specificity to modify a biochemical function and/or mechanism of the enzyme (Babbitt et al., <xref ref-type="bibr" rid="B4">1996</xref>; Pegg et al., <xref ref-type="bibr" rid="B45">2006</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). The quest on <italic>de novo</italic> design of the protein with required function can be set by providing the sequence, structure, or the sequence&#x02013;structure combination (Leaver-Fay et al., <xref ref-type="bibr" rid="B38">2011</xref>; Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). The specific task that DEFINED-PROTEINS presented in this study addresses the finding of a structural segment according to the functional descriptor, which will fit best into the original fold, providing required functional and stabilizing interactions with the rest of the fold. Ultimately, DEFINED-PROTEINS is aiming to build catalytic sites with desired activity and interactions while maintaining the overall fold structure and stability.</p>
<p>Steady progress in the computational design of new topologies and functions (Huang et al., <xref ref-type="bibr" rid="B34">2014</xref>, <xref ref-type="bibr" rid="B33">2016b</xref>; King et al., <xref ref-type="bibr" rid="B37">2015</xref>) and even a stronger drive toward <italic>de novo</italic> protein design (Huang et al., <xref ref-type="bibr" rid="B32">2016a</xref>; Lechner et al., <xref ref-type="bibr" rid="B39">2018</xref>; Baker, <xref ref-type="bibr" rid="B6">2019</xref>; Silva et al., <xref ref-type="bibr" rid="B56">2019</xref>) prompt the use of the basic units that would possess all traits determined by the polymer nature of proteins (Yamakawa and Stockmayer, <xref ref-type="bibr" rid="B61">1972</xref>; Shimada and Yamakawa, <xref ref-type="bibr" rid="B55">1984</xref>; Berezovsky et al., <xref ref-type="bibr" rid="B9">2000</xref>, <xref ref-type="bibr" rid="B10">2017a</xref>; Orevi et al., <xref ref-type="bibr" rid="B44">2013</xref>; Jacob et al., <xref ref-type="bibr" rid="B36">2018</xref>), their evolutionary history, and requirements on the structural stability and dynamics, as well as show the required functional activity (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>; Romero-Romero et al., <xref ref-type="bibr" rid="B48">2021</xref>). The concept of the elementary function standard/descriptor allows one to consider individual steps of biochemical functions provided by physics-based and evolutionary selected ELFs (Berezovsky, <xref ref-type="bibr" rid="B8">2019</xref>). An exhaustive description of elementary functions including all sequence, structure, and functional information that can be used in the design of new biochemical functions consisting of different combinations of elementary ones. Ultimately, the set of descriptors should, as exhaustively as possible, represent the diversity of sequence signatures that perform this elementary function, as well as the diversity of their structural implementations in different protein folds. In the combinations of descriptors into the desired biochemical function, the observables of the parameters of descriptors will be the result of the interference between their distributions and the type of the final structure that carries the function, the type of the overall transformation, interactions with other descriptors involved in the construction, and interactions with the substrate. Ideally, it should be possible to use descriptors of elementary functional units to build the geometry, the set of interactions, and the environment necessary for specified biochemical function. Then, the library of these units is supposed to be used in design efforts such as, the Rosetta enzyme design protocol (Leaver-Fay et al., <xref ref-type="bibr" rid="B38">2011</xref>), coupled with which the placement and refinement of a DEFINED-PROTEINS functional unit into a designable scaffold, which would include modeling of intra-loop and loop-fold interactions and energy minimization toward stable structure with required dynamics, would be enabled.</p></sec>
<sec sec-type="conclusions" id="s5">
<title>Conclusions</title>
<p>The DEFINED-PROTEINS is a proof-of-concept implementation that provides the tools to: (i) derive descriptor of elementary functions of interest directly from protein structures and (ii) apply descriptors in computational engineering and design. The software is available as a Python package that allows for integration in custom protein design workflows. We exemplified the derivation of the descriptor on three elementary functions: the calcium-binding EF-hand (Gifford et al., <xref ref-type="bibr" rid="B23">2007</xref>), the glycine-rich motif of the phosphate-binding in dinucleotide-containing ligands, and the mononucleotide phosphate-binding P-loop. We also assessed the robustness of derived descriptors and the objective function by replacing phosphate-binding EFLs in seven proteins with different functions derived on the set of proteins excluding the above seven proteins. Finally, we demonstrated a proof-of-principle grafting experiment by cross-replacing functional loops between P-loop containing proteins and those with the elementary function of the phosphate-binding in dinucleotide-containing ligands.</p></sec>
<sec sec-type="data-availability-statement" id="s6">
<title>Data Availability Statement</title>
<p>The software was built primarily using Python3 (<ext-link ext-link-type="uri" xlink:href="http://www.python.org">http://www.python.org</ext-link>), C with an optional MPI requirement, and C&#x0002B;&#x0002B;. Django (<ext-link ext-link-type="uri" xlink:href="https://djangoproject.com">https://djangoproject.com</ext-link>) was used as the web framework with embedded interactive Bokeh plots (<ext-link ext-link-type="uri" xlink:href="https://bokeh.org/">https://bokeh.org/</ext-link>). Docker container is available for OS-independent deployment. The software is BSD-licensed.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>IB: conceptualization, supervision, project administration, and funding acquisition. AG and IB: methodology. AG, MY, and IB: investigation, formal analysis, and writing&#x02014;review and editing. MY: software and visualization. MY and IB: writing&#x02014;original draft. All authors contributed to the article and approved the submitted version.</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</body>
<back>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fbinf.2021.657529/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fbinf.2021.657529/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table_2.docx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_1.PDF" id="SM3" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_2.PDF" id="SM4" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_3.PDF" id="SM5" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_4.PDF" id="SM6" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_5.PDF" id="SM7" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_6.PDF" id="SM8" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_7.PDF" id="SM9" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image_8.PDF" id="SM10" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Akiva</surname> <given-names>E.</given-names></name> <name><surname>Brown</surname> <given-names>S.</given-names></name> <name><surname>Almonacid</surname> <given-names>D. E.</given-names></name> <name><surname>Barber</surname> <given-names>A. E.</given-names> <suffix>2nd</suffix></name> <name><surname>Custer</surname> <given-names>A. F.</given-names></name> <name><surname>Hicks</surname> <given-names>M. A.</given-names></name> <name><surname>Huang</surname> <given-names>C. C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>The structure-function linkage database</article-title>. <source>Nucleic Acids Res.</source> <volume>42</volume>, <fpage>D521</fpage>&#x02013;<lpage>D530</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkt1130</pub-id><pub-id pub-id-type="pmid">24271399</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andreini</surname> <given-names>C.</given-names></name> <name><surname>Bertini</surname> <given-names>I.</given-names></name> <name><surname>Cavallaro</surname> <given-names>G.</given-names></name> <name><surname>Holliday</surname> <given-names>G. L.</given-names></name> <name><surname>Thornton</surname> <given-names>J. M.</given-names></name></person-group> (<year>2009</year>). <article-title>Metal-MACiE: a database of metals involved in biological catalysis</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2088</fpage>&#x02013;<lpage>2089</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp256</pub-id><pub-id pub-id-type="pmid">19369503</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aziz</surname> <given-names>M. F.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>K.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>The early history and emergence of molecular functions and modular scale-free network behavior</article-title>. <source>Sci. Rep.</source> <volume>6</volume>:<fpage>25058</fpage>. <pub-id pub-id-type="doi">10.1038/srep25058</pub-id><pub-id pub-id-type="pmid">27121452</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Babbitt</surname> <given-names>P. C.</given-names></name> <name><surname>Hasson</surname> <given-names>M. S.</given-names></name> <name><surname>Wedekind</surname> <given-names>J. E.</given-names></name> <name><surname>Palmer</surname> <given-names>D. R.</given-names></name> <name><surname>Barrett</surname> <given-names>W. C.</given-names></name> <name><surname>Reed</surname> <given-names>G. H.</given-names></name> <etal/></person-group>. (<year>1996</year>). <article-title>The enolase superfamily: a general strategy for enzyme-catalyzed abstraction of the alpha-protons of carboxylic acids</article-title>. <source>Biochemistry</source> <volume>35</volume>, <fpage>16489</fpage>&#x02013;<lpage>16501</lpage>. <pub-id pub-id-type="doi">10.1021/bi9616413</pub-id><pub-id pub-id-type="pmid">8987982</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bairoch</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>The ENZYME database</article-title> in <source>Nucleic Acids Res.</source> <volume>28</volume>, <fpage>304</fpage>&#x02013;<lpage>305</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.304</pub-id><pub-id pub-id-type="pmid">10592255</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>What has <italic>de novo</italic> protein design taught us about protein folding and biophysics?</article-title> <source>Protein Sci.</source> <volume>28</volume>, <fpage>678</fpage>&#x02013;<lpage>683</lpage>. <pub-id pub-id-type="doi">10.1002/pro.3588</pub-id><pub-id pub-id-type="pmid">30746840</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2003</year>). <article-title>Discrete structure of van der Waals domains in globular proteins</article-title>. <source>Protein engineering</source> <volume>16</volume>, <fpage>161</fpage>&#x02013;<lpage>167</lpage>. <pub-id pub-id-type="doi">10.1093/proeng/gzg026</pub-id><pub-id pub-id-type="pmid">12702795</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2019</year>). <article-title>Towards descriptor of elementary functions for protein design</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>58</volume>, <fpage>159</fpage>&#x02013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2019.06.010</pub-id><pub-id pub-id-type="pmid">31352188</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Grosberg</surname> <given-names>A. Y.</given-names></name> <name><surname>Trifonov</surname> <given-names>E. N.</given-names></name></person-group> (<year>2000</year>). <article-title>Closed loops of nearly standard size: common basic element of protein structure</article-title>. <source>FEBS Lett.</source> <volume>466</volume>, <fpage>283</fpage>&#x02013;<lpage>286</lpage>. <pub-id pub-id-type="doi">10.1016/S0014-5793(00)01091-7</pub-id><pub-id pub-id-type="pmid">10682844</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Guarnera</surname> <given-names>E.</given-names></name> <name><surname>Zheng</surname> <given-names>Z.</given-names></name></person-group> (<year>2017a</year>). <article-title>Basic units of protein structure, folding, and function</article-title>. <source>Progr. Biophys. Mol. Biol.</source> <volume>128</volume>, <fpage>85</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.pbiomolbio.2016.09.009</pub-id><pub-id pub-id-type="pmid">27697476</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Guarnera</surname> <given-names>E.</given-names></name> <name><surname>Zheng</surname> <given-names>Z.</given-names></name> <name><surname>Eisenhaber</surname> <given-names>B.</given-names></name> <name><surname>Eisenhaber</surname> <given-names>F.</given-names></name></person-group> (<year>2017b</year>). <article-title>Protein function machinery: from basic structural units to modulation of activity</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>42</volume>, <fpage>67</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2016.10.021</pub-id><pub-id pub-id-type="pmid">27865209</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Kirzhner</surname> <given-names>A.</given-names></name> <name><surname>Kirzhner</surname> <given-names>V. M.</given-names></name> <name><surname>Rosenfeld</surname> <given-names>V. R.</given-names></name> <name><surname>Trifonov</surname> <given-names>E. N.</given-names></name></person-group> (<year>2003a</year>). <article-title>Protein sequences yield a proteomic code</article-title>. <source>J. Biomol. Struct. Dyn.</source> <volume>21</volume>, <fpage>317</fpage>&#x02013;<lpage>325</lpage>. <pub-id pub-id-type="doi">10.1080/07391102.2003.10506928</pub-id><pub-id pub-id-type="pmid">14616028</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Kirzhner</surname> <given-names>A.</given-names></name> <name><surname>Kirzhner</surname> <given-names>V. M.</given-names></name> <name><surname>Trifonov</surname> <given-names>E. N.</given-names></name></person-group> (<year>2003b</year>). <article-title>Spelling protein structure</article-title>. <source>J. Biomol. Struct. Dyn.</source> <volume>21</volume>, <fpage>327</fpage>&#x02013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1080/07391102.2003.10506929</pub-id><pub-id pub-id-type="pmid">14616029</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berman</surname> <given-names>H. M.</given-names></name> <name><surname>Westbrook</surname> <given-names>J.</given-names></name> <name><surname>Feng</surname> <given-names>Z.</given-names></name> <name><surname>Gilliland</surname> <given-names>G.</given-names></name> <name><surname>Bhat</surname> <given-names>T. N.</given-names></name> <name><surname>Weissig</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2000</year>). <article-title>The protein data bank</article-title>. <source>Nucleic Acids Res.</source> <volume>28</volume>, <fpage>235</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.235</pub-id><pub-id pub-id-type="pmid">10592235</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bowie</surname> <given-names>J. U.</given-names></name> <name><surname>Luthy</surname> <given-names>R.</given-names></name> <name><surname>Eisenberg</surname> <given-names>D.</given-names></name></person-group> (<year>1991</year>). <article-title>A method to identify protein sequences that fold into a known three-dimensional structure</article-title>. <source>Science</source> <volume>253</volume>, <fpage>164</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1126/science.1853201</pub-id><pub-id pub-id-type="pmid">1853201</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunette</surname> <given-names>T. J.</given-names></name> <name><surname>Parmeggiani</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>P. S.</given-names></name> <name><surname>Bhabha</surname> <given-names>G.</given-names></name> <name><surname>Ekiert</surname> <given-names>D. C.</given-names></name> <name><surname>Tsutakawa</surname> <given-names>S. E.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Exploring the repeat protein universe through computational protein design</article-title>. <source>Nature</source> <volume>528</volume>, <fpage>580</fpage>&#x02013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1038/nature16162</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crippen</surname> <given-names>G. M.</given-names></name></person-group> (<year>1996</year>). <article-title>Failures of inverse folding and threading with gapped alignment</article-title>. <source>Proteins</source> <volume>26</volume>, <fpage>167</fpage>&#x02013;<lpage>171</lpage>. <pub-id pub-id-type="doi">10.1002/(SICI)1097-0134(199610)26:2&#x0003C;167::AID-PROT6&#x0003E;3.0.CO;2-D</pub-id><pub-id pub-id-type="pmid">8916224</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Das</surname> <given-names>R.</given-names></name> <name><surname>Baker</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Macromolecular modeling with rosetta</article-title>. <source>Annu. Rev. Biochem.</source> <volume>77</volume>, <fpage>363</fpage>&#x02013;<lpage>382</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.biochem.77.062906.171838</pub-id><pub-id pub-id-type="pmid">18410248</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dill</surname> <given-names>K. A.</given-names></name> <name><surname>MacCallum</surname> <given-names>J. L.</given-names></name></person-group> (<year>2012</year>). <article-title>The protein-folding problem, 50 years on</article-title>. <source>Science</source> <volume>338</volume>, <fpage>1042</fpage>&#x02013;<lpage>1046</lpage>. <pub-id pub-id-type="doi">10.1126/science.1219021</pub-id><pub-id pub-id-type="pmid">23180855</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Finn</surname> <given-names>R. D.</given-names></name> <name><surname>Coggill</surname> <given-names>P.</given-names></name> <name><surname>Eberhardt</surname> <given-names>R. Y.</given-names></name> <name><surname>Eddy</surname> <given-names>S. R.</given-names></name> <name><surname>Mistry</surname> <given-names>J.</given-names></name> <name><surname>Mitchell</surname> <given-names>A. L.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>The Pfam protein families database: towards a more sustainable future</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>D279</fpage>&#x02013;<lpage>D285</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1344</pub-id><pub-id pub-id-type="pmid">26673716</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fischer</surname> <given-names>J. D.</given-names></name> <name><surname>Holliday</surname> <given-names>G. L.</given-names></name> <name><surname>Thornton</surname> <given-names>J. M.</given-names></name></person-group> (<year>2010</year>). <article-title>The CoFactor database: organic cofactors in enzyme catalysis</article-title>. <source>Bioinformatics</source> <volume>26</volume>, <fpage>2496</fpage>&#x02013;<lpage>2497</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq442</pub-id><pub-id pub-id-type="pmid">20679331</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Furnham</surname> <given-names>N.</given-names></name> <name><surname>Holliday</surname> <given-names>G. L.</given-names></name> <name><surname>de Beer</surname> <given-names>T. A.</given-names></name> <name><surname>Jacobsen</surname> <given-names>J. O.</given-names></name> <name><surname>Pearson</surname> <given-names>W. R.</given-names></name> <name><surname>Thornton</surname> <given-names>J. M.</given-names></name></person-group> (<year>2014</year>). <article-title>The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes</article-title>. <source>Nucleic Acids Res.</source> <volume>42</volume>, <fpage>D485</fpage>&#x02013;<lpage>D489</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkt1243</pub-id><pub-id pub-id-type="pmid">24319146</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gifford</surname> <given-names>J. L.</given-names></name> <name><surname>Walsh</surname> <given-names>M. P.</given-names></name> <name><surname>Vogel</surname> <given-names>H. G.</given-names></name></person-group> (<year>2007</year>). <article-title>Structures and metal-ion-binding properties of the Ca<sup>2&#x0002B;</sup>-binding helix&#x02013;loop&#x02013;helix EF-hand motifs</article-title>. <source>Biochem. J</source>. <volume>405</volume>, <fpage>199</fpage>&#x02013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1042/BJ20070255</pub-id><pub-id pub-id-type="pmid">17590154</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2010</year>). <article-title>Prototypes of elementary functional loops unravel evolutionary connections between protein functions</article-title>. <source>Bioinformatics</source> <volume>26</volume>, <fpage>i497</fpage>&#x02013;<lpage>i503</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq374</pub-id><pub-id pub-id-type="pmid">20823313</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2011</year>). <article-title>Computational reconstruction of primordial prototypes of elementary functional loops in modern proteins</article-title>. <source>Bioinformatics</source> <volume>27</volume>, <fpage>2368</fpage>&#x02013;<lpage>2375</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btr396</pub-id><pub-id pub-id-type="pmid">21724592</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2012</year>). <article-title>Exploring the evolution of protein function in Archaea</article-title>. <source>BMC Evol. Biol.</source> <volume>12</volume>:<fpage>75</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-12-75</pub-id><pub-id pub-id-type="pmid">22646318</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2015</year>). <article-title>Protein function from its emergence to diversity in contemporary proteins</article-title>. <source>Phys. Biol.</source> <volume>12</volume>:<fpage>045002</fpage>. <pub-id pub-id-type="doi">10.1088/1478-3975/12/4/045002</pub-id><pub-id pub-id-type="pmid">26057563</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henikoff</surname> <given-names>S.</given-names></name> <name><surname>Henikoff</surname> <given-names>J. G.</given-names></name></person-group> (<year>1993</year>). <article-title>Performance evaluation of amino acid substitution matrices</article-title>. <source>Proteins</source> <volume>17</volume>, <fpage>49</fpage>&#x02013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1002/prot.340170108</pub-id><pub-id pub-id-type="pmid">8234244</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hocker</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>Design of proteins from smaller fragments-learning from evolution</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>27</volume>, <fpage>56</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2014.04.007</pub-id><pub-id pub-id-type="pmid">24865156</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holliday</surname> <given-names>G. L.</given-names></name> <name><surname>Andreini</surname> <given-names>C.</given-names></name> <name><surname>Fischer</surname> <given-names>J. D.</given-names></name> <name><surname>Rahman</surname> <given-names>S. A.</given-names></name> <name><surname>Almonacid</surname> <given-names>D. E.</given-names></name> <name><surname>Williams</surname> <given-names>S. T.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>MACiE: exploring the diversity of biochemical reactions</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>D783</fpage>&#x02013;<lpage>D789</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr799</pub-id><pub-id pub-id-type="pmid">22058127</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holliday</surname> <given-names>G. L.</given-names></name> <name><surname>Bartlett</surname> <given-names>G. J.</given-names></name> <name><surname>Almonacid</surname> <given-names>D. E.</given-names></name> <name><surname>O&#x00027;Boyle</surname> <given-names>N. M.</given-names></name> <name><surname>Murray-Rust</surname> <given-names>P.</given-names></name> <name><surname>Thornton</surname> <given-names>J. M.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>MACiE: a database of enzyme reaction mechanisms</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>4315</fpage>&#x02013;<lpage>4316</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti693</pub-id><pub-id pub-id-type="pmid">16188925</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>P. S.</given-names></name> <name><surname>Boyken</surname> <given-names>S. E.</given-names></name> <name><surname>Baker</surname> <given-names>D.</given-names></name></person-group> (<year>2016a</year>). <article-title>The coming of age of <italic>de novo</italic> protein design</article-title>. <source>Nature</source> <volume>537</volume>, <fpage>320</fpage>&#x02013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1038/nature19946</pub-id><pub-id pub-id-type="pmid">27629638</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>P. S.</given-names></name> <name><surname>Feldmeier</surname> <given-names>K.</given-names></name> <name><surname>Parmeggiani</surname> <given-names>F.</given-names></name> <name><surname>Velasco</surname> <given-names>D. A. F.</given-names></name> <name><surname>Hocker</surname> <given-names>B.</given-names></name> <name><surname>Baker</surname> <given-names>D.</given-names></name></person-group> (<year>2016b</year>). <article-title><italic>De novo</italic> design of a four-fold symmetric TIM-barrel protein with atomic-level accuracy</article-title>. <source>Nat. Chem. Biol.</source> <volume>12</volume>, <fpage>29</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1038/nchembio.1966</pub-id><pub-id pub-id-type="pmid">26595462</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>P. S.</given-names></name> <name><surname>Oberdorfer</surname> <given-names>G.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Pei</surname> <given-names>X. Y.</given-names></name> <name><surname>Nannenga</surname> <given-names>B. L.</given-names></name> <name><surname>Rogers</surname> <given-names>J. M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>High thermodynamic stability of parametrically designed helical bundles</article-title>. <source>Science</source> <volume>346</volume>, <fpage>481</fpage>&#x02013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1126/science.1257481</pub-id><pub-id pub-id-type="pmid">25342806</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hunter</surname> <given-names>S.</given-names></name> <name><surname>Apweiler</surname> <given-names>R.</given-names></name> <name><surname>Attwood</surname> <given-names>T. K.</given-names></name> <name><surname>Bairoch</surname> <given-names>A.</given-names></name> <name><surname>Bateman</surname> <given-names>A.</given-names></name> <name><surname>Binns</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>InterPro: the integrative protein signature database</article-title>. <source>Nucleic Acids Res.</source> <volume>37</volume>, <fpage>D211</fpage>&#x02013;<lpage>D215</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn785</pub-id><pub-id pub-id-type="pmid">18940856</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jacob</surname> <given-names>M. H.</given-names></name> <name><surname>D&#x00027;Souza</surname> <given-names>R. N.</given-names></name> <name><surname>Schwarzlose</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>F.</given-names></name> <name><surname>Haas</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Method-unifying view of loop-formation kinetics in peptide and protein folding</article-title>. <source>J. Phys. Chem. B</source> <volume>122</volume>, <fpage>4445</fpage>&#x02013;<lpage>4456</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jpcb.8b00879</pub-id><pub-id pub-id-type="pmid">29617564</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>King</surname> <given-names>I. C.</given-names></name> <name><surname>Gleixner</surname> <given-names>J.</given-names></name> <name><surname>Doyle</surname> <given-names>L.</given-names></name> <name><surname>Kuzin</surname> <given-names>A.</given-names></name> <name><surname>Hunt</surname> <given-names>J. F.</given-names></name> <name><surname>Xiao</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Precise assembly of complex beta sheet topologies from <italic>de novo</italic> designed building blocks</article-title>. <source>Elife</source> <volume>4</volume>:<fpage>e53865</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.11012.020</pub-id><pub-id pub-id-type="pmid">26650357</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leaver-Fay</surname> <given-names>A.</given-names></name> <name><surname>Tyka</surname> <given-names>M.</given-names></name> <name><surname>Lewis</surname> <given-names>S. M.</given-names></name> <name><surname>Lange</surname> <given-names>O. F.</given-names></name> <name><surname>Thompson</surname> <given-names>J.</given-names></name> <name><surname>Thompson</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>ROSETTA3: an object-oriented software suite for the simulation and design of macromolecules</article-title>. <source>Methods Enzymol.</source> <volume>487</volume>, <fpage>545</fpage>&#x02013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-381270-4.00019-6</pub-id><pub-id pub-id-type="pmid">21187238</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lechner</surname> <given-names>H.</given-names></name> <name><surname>Ferruz</surname> <given-names>N.</given-names></name> <name><surname>Hocker</surname> <given-names>B.</given-names></name></person-group> (<year>2018</year>). <article-title>Strategies for designing non-natural enzymes and binders</article-title>. <source>Curr. Opin. Chem. Biol.</source> <volume>47</volume>, <fpage>67</fpage>&#x02013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.1016/j.cbpa.2018.07.022</pub-id><pub-id pub-id-type="pmid">30248579</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marchler-Bauer</surname> <given-names>A.</given-names></name> <name><surname>Derbyshire</surname> <given-names>M. K.</given-names></name> <name><surname>Gonzales</surname> <given-names>N. R.</given-names></name> <name><surname>Lu</surname> <given-names>S.</given-names></name> <name><surname>Chitsaz</surname> <given-names>F.</given-names></name> <name><surname>Geer</surname> <given-names>L. Y.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>CDD: NCBI&#x00027;s conserved domain database</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>D222</fpage>&#x02013;<lpage>D226</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku1221</pub-id><pub-id pub-id-type="pmid">25414356</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marchler-Bauer</surname> <given-names>A.</given-names></name> <name><surname>Zheng</surname> <given-names>C.</given-names></name> <name><surname>Chitsaz</surname> <given-names>F.</given-names></name> <name><surname>Derbyshire</surname> <given-names>M. K.</given-names></name> <name><surname>Geer</surname> <given-names>L. Y.</given-names></name> <name><surname>Geer</surname> <given-names>R. C.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>CDD: conserved domains and protein three-dimensional structure</article-title>. <source>Nucleic Acids Res.</source> <volume>41</volume>, <fpage>D348</fpage>&#x02013;<lpage>352</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks1243</pub-id><pub-id pub-id-type="pmid">23197659</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Minor</surname> <given-names>D. L.</given-names> <suffix>Jr.</suffix></name> <name><surname>Kim</surname> <given-names>P. S.</given-names></name></person-group> (<year>1996</year>). <article-title>Context-dependent secondary structure formation of a designed protein sequence</article-title>. <source>Nature</source> <volume>380</volume>, <fpage>730</fpage>&#x02013;<lpage>734</lpage>. <pub-id pub-id-type="doi">10.1038/380730a0</pub-id><pub-id pub-id-type="pmid">8614471</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nath</surname> <given-names>N.</given-names></name> <name><surname>Mitchell</surname> <given-names>J. B.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>The natural history of biocatalytic mechanisms</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003642</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003642</pub-id><pub-id pub-id-type="pmid">24874434</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orevi</surname> <given-names>T.</given-names></name> <name><surname>Rahamim</surname> <given-names>G.</given-names></name> <name><surname>Hazan</surname> <given-names>G.</given-names></name> <name><surname>Amir</surname> <given-names>D.</given-names></name> <name><surname>Haas</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>The loop hypothesis: contribution of early formed specific non-local interactions to the determination of protein folding pathways</article-title>. <source>Biophys. Rev.</source> <volume>5</volume>, <fpage>85</fpage>&#x02013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1007/s12551-013-0113-3</pub-id><pub-id pub-id-type="pmid">28510159</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pegg</surname> <given-names>S. C.</given-names></name> <name><surname>Brown</surname> <given-names>S. D.</given-names></name> <name><surname>Ojha</surname> <given-names>S.</given-names></name> <name><surname>Seffernick</surname> <given-names>J.</given-names></name> <name><surname>Meng</surname> <given-names>E. C.</given-names></name> <name><surname>Morris</surname> <given-names>J. H.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Leveraging enzyme structure-function relationships for functional inference and experimental design: the structure-function linkage database</article-title>. <source>Biochemistry</source> <volume>45</volume>, <fpage>2545</fpage>&#x02013;<lpage>2555</lpage>. <pub-id pub-id-type="doi">10.1021/bi052101l</pub-id><pub-id pub-id-type="pmid">16489747</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero Romero</surname> <given-names>M. L.</given-names></name> <name><surname>Rabin</surname> <given-names>A.</given-names></name> <name><surname>Tawfik</surname> <given-names>D. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Functional proteins from short peptides: dayhoff&#x00027;s hypothesis Turns 50</article-title>. <source>Angew. Chem. Int. Ed. Engl.</source> <volume>55</volume>, <fpage>15966</fpage>&#x02013;<lpage>15971</lpage>. <pub-id pub-id-type="doi">10.1002/anie.201609977</pub-id><pub-id pub-id-type="pmid">27865046</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero Romero</surname> <given-names>M. L.</given-names></name> <name><surname>Yang</surname> <given-names>F.</given-names></name> <name><surname>Lin</surname> <given-names>Y. R.</given-names></name> <name><surname>Toth-Petroczy</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name> <name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Simple yet functional phosphate-loop proteins</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>115</volume>, <fpage>E11943</fpage>&#x02013;<lpage>E11950</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1812400115</pub-id><pub-id pub-id-type="pmid">30504143</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero-Romero</surname> <given-names>S.</given-names></name> <name><surname>Kordes</surname> <given-names>S.</given-names></name> <name><surname>Michel</surname> <given-names>F.</given-names></name> <name><surname>Hocker</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Evolution, folding, and design of TIM barrels and related proteins</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>68</volume>, <fpage>94</fpage>&#x02013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2020.12.007</pub-id><pub-id pub-id-type="pmid">33453500</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rooman</surname> <given-names>M. J.</given-names></name> <name><surname>Rodriguez</surname> <given-names>J.</given-names></name> <name><surname>Wodak</surname> <given-names>S. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Relations between protein sequence and structure and their significance</article-title>. <source>J. Mol. Biol.</source> <volume>213</volume>, <fpage>337</fpage>&#x02013;<lpage>350</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80195-0</pub-id><pub-id pub-id-type="pmid">2342111</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rooman</surname> <given-names>M. J.</given-names></name> <name><surname>Wodak</surname> <given-names>S. J.</given-names></name></person-group> (<year>1995</year>). <article-title>Are database-derived potentials valid for scoring both forward and inverted protein folding?</article-title> <source>Protein Eng.</source> <volume>8</volume>, <fpage>849</fpage>&#x02013;<lpage>858</lpage>. <pub-id pub-id-type="doi">10.1093/protein/8.9.849</pub-id><pub-id pub-id-type="pmid">8746722</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sali</surname> <given-names>A.</given-names></name> <name><surname>Shakhnovich</surname> <given-names>E.</given-names></name> <name><surname>Karplus</surname> <given-names>M.</given-names></name></person-group> (<year>1994</year>). <article-title>How does a protein fold?</article-title> <source>Nature</source> <volume>369</volume>, <fpage>248</fpage>&#x02013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.1038/369248a0</pub-id><pub-id pub-id-type="pmid">7710478</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Senior</surname> <given-names>A. W.</given-names></name> <name><surname>Evans</surname> <given-names>R.</given-names></name> <name><surname>Jumper</surname> <given-names>J.</given-names></name> <name><surname>Kirkpatrick</surname> <given-names>J.</given-names></name> <name><surname>Sifre</surname> <given-names>L.</given-names></name> <name><surname>Green</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Improved protein structure prediction using potentials from deep learning</article-title>. <source>Nature</source> <volume>577</volume>, <fpage>706</fpage>&#x02013;<lpage>710</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-1923-7</pub-id><pub-id pub-id-type="pmid">31942072</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shakhnovich</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>Protein folding thermodynamics and dynamics: where physics, chemistry, and biology meet</article-title>. <source>Chem. Rev.</source> <volume>106</volume>, <fpage>1559</fpage>&#x02013;<lpage>1588</lpage>. <pub-id pub-id-type="doi">10.1021/cr040425u</pub-id><pub-id pub-id-type="pmid">16683745</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shakhnovich</surname> <given-names>E. I.</given-names></name> <name><surname>Gutin</surname> <given-names>A. M.</given-names></name></person-group> (<year>1993</year>). <article-title>Engineering of stable and fast-folding sequences of model proteins</article-title>. <source>Proc. Natl. Acad. Sci. U.S. A.</source> <volume>90</volume>, <fpage>7195</fpage>&#x02013;<lpage>7199</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.90.15.7195</pub-id><pub-id pub-id-type="pmid">8346235</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shimada</surname> <given-names>J.</given-names></name> <name><surname>Yamakawa</surname> <given-names>H.</given-names></name></person-group> (<year>1984</year>). <article-title>Ring-closure probabilities for twisted wormlike chains. Application to DNA</article-title>. <source>Macromolecules</source> <volume>17</volume>, <fpage>689</fpage>&#x02013;<lpage>698</lpage>. <pub-id pub-id-type="doi">10.1021/ma00134a028</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silva</surname> <given-names>D. A.</given-names></name> <name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Ulge</surname> <given-names>U. Y.</given-names></name> <name><surname>Spangler</surname> <given-names>J. B.</given-names></name> <name><surname>Jude</surname> <given-names>K. M.</given-names></name> <name><surname>Labao-Almeida</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title><italic>De novo</italic> design of potent and selective mimics of IL-2 and IL-15</article-title>. <source>Nature</source> <volume>565</volume>, <fpage>186</fpage>&#x02013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0830-7</pub-id><pub-id pub-id-type="pmid">30626941</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sippl</surname> <given-names>M. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Calculation of conformational ensembles from potentials of mean force. An approach to the knowledge-based prediction of local structures in globular proteins</article-title>. <source>J. Mol. Biol.</source> <volume>213</volume>, <fpage>859</fpage>&#x02013;<lpage>883</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80269-4</pub-id><pub-id pub-id-type="pmid">2359125</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trifonov</surname> <given-names>E. N.</given-names></name> <name><surname>Kirzhner</surname> <given-names>A.</given-names></name> <name><surname>Kirzhner</surname> <given-names>V. M.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2001</year>). <article-title>Distinct stages of protein evolution as suggested by protein sequence analysis</article-title>. <source>J. Mol. Evol.</source> <volume>53</volume>, <fpage>394</fpage>&#x02013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1007/s002390010229</pub-id><pub-id pub-id-type="pmid">11675599</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trudeau</surname> <given-names>D. L.</given-names></name> <name><surname>Tawfik</surname> <given-names>D. S.</given-names></name></person-group> (<year>2019</year>). <article-title>Protein engineers turned evolutionists-the quest for the optimal starting point</article-title>. <source>Curr. Opin. Biotechnol.</source> <volume>60</volume>, <fpage>46</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.copbio.2018.12.002</pub-id><pub-id pub-id-type="pmid">30611116</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><collab>wwPDB consortium</collab></person-group> (<year>2019</year>). <article-title>Protein Data Bank: the single global archive for 3D macromolecular structure data</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>D520</fpage>&#x02013;<lpage>D528</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky949</pub-id><pub-id pub-id-type="pmid">30357364</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamakawa</surname> <given-names>H.</given-names></name> <name><surname>Stockmayer</surname> <given-names>W. H.</given-names></name></person-group> (<year>1972</year>). <article-title>Statistical mechanics of wormlike chains. II. Excluded volume effects</article-title>. <source>J. Chem. Phys.</source> <volume>57</volume>, <fpage>2843</fpage>&#x02013;<lpage>2854</lpage>. <pub-id pub-id-type="doi">10.1063/1.1678675</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>Z.</given-names></name> <name><surname>Goncearenco</surname> <given-names>A.</given-names></name> <name><surname>Berezovsky</surname> <given-names>I. N.</given-names></name></person-group> (<year>2016</year>). <article-title>Nucleotide binding database NBDB&#x02013;a collection of sequence motifs with specific protein-ligand interactions</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>D301</fpage>&#x02013;<lpage>D307</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1124</pub-id><pub-id pub-id-type="pmid">26507856</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This financial support provided by the Biomedical Research Council, via Agency for Science, Technology, and Research (A<sup>&#x0002A;</sup>STAR), is greatly appreciated. AG was supported in part by the Intramural Research Programs of the National Library of Medicine, National Institutes of Health.</p>
</fn>
</fn-group>
</back>
</article>