<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Immunol.</journal-id>
<journal-title>Frontiers in Immunology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Immunol.</abbrev-journal-title>
<issn pub-type="epub">1664-3224</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fimmu.2016.00681</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Immunology</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>BRILIA: Integrated Tool for High-Throughput Annotation and Lineage Tree Assembly of B-Cell Repertoires</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Lee</surname> <given-names>Donald W.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/383293"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Khavrutskii</surname> <given-names>Ilja V.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/393043"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wallqvist</surname> <given-names>Anders</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/187800"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bavari</surname> <given-names>Sina</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/265272"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Cooper</surname> <given-names>Christopher L.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Chaudhury</surname> <given-names>Sidhartha</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/289524"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Biotechnology HPC Software Applications Institute (BHSAI), Telemedicine and Advanced Technology Research Center, U.S. Army Medical Research and Materiel Command</institution>, <addr-line>Fort Detrick, MD</addr-line>, <country>USA</country></aff>
<aff id="aff2"><sup>2</sup><institution>Molecular and Translational Sciences, U.S. Army Medical Research Institute of Infectious Diseases</institution>, <addr-line>Frederick, MD</addr-line>, <country>USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Harry W. Schroeder, University of Alabama at Birmingham, USA</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Alex Rosenberg, University of Rochester Medical Center, USA; Johannes Tr&#x000FC;ck, University Children&#x02019;s Hospital Zurich, Switzerland</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Sidhartha Chaudhury, <email>sidhartha.chaudhury.civ&#x00040;mail.mil</email></corresp>
<fn fn-type="other" id="fn002"><p>Specialty section: This article was submitted to B Cell Biology, a section of the journal Frontiers in Immunology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>7</volume>
<elocation-id>681</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>12</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Lee, Khavrutskii, Wallqvist, Bavari, Cooper and Chaudhury.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Lee, Khavrutskii, Wallqvist, Bavari, Cooper and Chaudhury</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The somatic diversity of antigen-recognizing B-cell receptors (BCRs) arises from Variable (V), Diversity (D), and Joining (J) (VDJ) recombination and somatic hypermutation (SHM) during B-cell development and affinity maturation. The VDJ junction of the BCR heavy chain forms the highly variable complementarity determining region 3 (CDR3), which plays a critical role in antigen specificity and binding affinity. Tracking the selection and mutation of the CDR3 can be useful in characterizing humoral responses to infection and vaccination. Although tens to hundreds of thousands of unique BCR genes within an expressed B-cell repertoire can now be resolved with high-throughput sequencing, tracking SHMs is still challenging because existing annotation methods are often limited by poor annotation coverage, inconsistent SHM identification across the VDJ junction, or lack of B-cell lineage data. Here, we present B-cell repertoire inductive lineage and immunosequence annotator (BRILIA), an algorithm that leverages repertoire-wide sequencing data to globally improve the VDJ annotation coverage, lineage tree assembly, and SHM identification. On benchmark tests against simulated human and mouse BCR repertoires, BRILIA correctly annotated germline and clonally expanded sequences with 94 and 70% accuracy, respectively, and it has a 90% SHM-positive prediction rate in the CDR3 of heavily mutated sequences; these are substantial improvements over existing methods. We used BRILIA to process BCR sequences obtained from splenic germinal center B cells extracted from C57BL/6 mice. BRILIA returned robust B-cell lineage trees and yielded SHM patterns that are consistent across the VDJ junction and agree with known biological mechanisms of SHM. By contrast, existing BCR annotation tools, which do not account for repertoire-wide clonal relationships, systematically underestimated both the size of clonally related B-cell clusters and yielded inconsistent SHM frequencies. We demonstrate BRILIA&#x02019;s utility in B-cell repertoire studies related to VDJ gene usage, mechanisms for adenosine mutations, and SHM hot spot motifs. Furthermore, we show that the complete gene usage annotation and SHM identification across the entire CDR3 are essential for studying the B-cell affinity maturation process through immunosequencing methods.</p>
</abstract>
<kwd-group>
<kwd>B-cell receptor (BCR)</kwd>
<kwd>repertoire</kwd>
<kwd>annotation</kwd>
<kwd>lineage</kwd>
<kwd>VDJ</kwd>
<kwd>somatic hypermutation (SHM)</kwd>
</kwd-group>
<contract-sponsor id="cn01">Medical Research and Materiel Command<named-content content-type="fundref-id">10.13039/100000182</named-content></contract-sponsor>
<contract-sponsor id="cn02">U.S. Department of Defense<named-content content-type="fundref-id">10.13039/100000005</named-content></contract-sponsor>
<contract-sponsor id="cn03">U.S. Army Medical Research Institute of Infectious Diseases<named-content content-type="fundref-id">10.13039/100010201</named-content></contract-sponsor>
<contract-sponsor id="cn04">Defense Threat Reduction Agency<named-content content-type="fundref-id">10.13039/100000774</named-content></contract-sponsor>
<counts>
<fig-count count="9"/>
<table-count count="3"/>
<equation-count count="3"/>
<ref-count count="99"/>
<page-count count="18"/>
<word-count count="12443"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<title>Introduction</title>
<p>B cells synthesize transmembrane proteins called B-cell receptors (BCRs) that recognize foreign antigens. The binding of BCRs with an antigen activates B cell clonal expansion and somatic hypermutation (SHM), which increases the likelihood of synthesizing high-affinity BCRs that are later secreted as antibodies into the blood [see Ref. (<xref ref-type="bibr" rid="B1">1</xref>) for a review on affinity maturation]. Understanding how antigen-specific antibodies are produced can aid the development of effective vaccines, for instance, by showing which BCR genes become enriched in vaccinated subjects (<xref ref-type="bibr" rid="B2">2</xref>&#x02013;<xref ref-type="bibr" rid="B6">6</xref>) or by measuring the extent of affinity maturation in response to vaccination. Although thousands of BCR sequences from B cell repertoires can be obtained with high-throughput sequencing (<xref ref-type="bibr" rid="B7">7</xref>), tracking SHM among maturing B cells remains a challenge (<xref ref-type="bibr" rid="B8">8</xref>). The standard approach for processing BCR sequences is to first annotate the Variable (V), Diversity (D), and Joining (J) segments of the BCR gene, then cluster clonally related sequences, and finally construct a lineage tree for each cluster (<xref ref-type="bibr" rid="B8">8</xref>&#x02013;<xref ref-type="bibr" rid="B11">11</xref>). Each step is typically carried out using separate algorithms, which can yield results that are at odds with the biological mechanisms that underlie SHM or affinity maturation. Therefore, developing algorithms that provide VDJ annotations that agree with clonal expansion and SHM behaviors is critical for accurately characterizing B-cell repertoires.</p>
<p>The BCR consists of a heavy chain and a light chain. During B-cell development, functional BCR genes for the heavy and light chain are formed <italic>via</italic> the recombination of V, D, and J gene segments within the chromosomal DNA, mediated by recombinase enzymes RAG1 and RAG2 (<xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B13">13</xref>). The heavy chain of the BCR uses the V, D, and J segments, and the light chain uses a separate set of only V and J segments, giving rise to a baseline combinatorial diversity of BCRs. The VDJ junction of the heavy chain forms the highly variable complementarity determining region 3 (CDR3) loop structure that plays a critical role in antigen recognition (<xref ref-type="bibr" rid="B14">14</xref>). Further diversity is introduced in the VDJ junction at the joining regions between the gene segments through deletions of gene edges by nucleases (<xref ref-type="bibr" rid="B15">15</xref>), creation of palindromic sequences called palindromic nucleotides (P-nts), and insertions of non-templated nts (N-nts) by terminal deoxynucleotidyl transferase (TDT) (<xref ref-type="bibr" rid="B16">16</xref>&#x02013;<xref ref-type="bibr" rid="B18">18</xref>) [see Ref. (<xref ref-type="bibr" rid="B19">19</xref>) for a review on VDJ recombination]. A contiguous sequence of P- and N-nts is referred to as an N region. A final level of BCR diversity is introduced through SHM that is mediated by deaminases and error-prone DNA repair enzymes, which can obscure the original VDJ genes [see Ref. (<xref ref-type="bibr" rid="B20">20</xref>) for a review on SHM]. The combinatorial diversity of germline gene recombination, the variation that is introduced by insertions and deletions in the VDJ junction, and the subsequent accumulation of SHMs during affinity maturation pose major challenges to BCR gene annotation.</p>
<p>High-throughput sequencing focused on the heavy chain CDR3 is becoming a common tool for rapidly characterizing the entire B-cell repertoire from a single sample&#x02014;an inexpensive alternative to more costly single-cell sequencing approaches (<xref ref-type="bibr" rid="B21">21</xref>). For a given individual, the number of unique B cells is estimated to be greater than 10<sup>7</sup> (<xref ref-type="bibr" rid="B22">22</xref>), and analysis of antigen-specific B cells may require sequencing of up to 10<sup>4</sup> or 10<sup>5</sup> unique B cells (<xref ref-type="bibr" rid="B23">23</xref>). The Illumina deep-sequencing technology is capable of providing sufficient sequencing depth to capture this repertoire in its entirety for sequence read lengths of &#x0007E;150&#x02009;bp (<xref ref-type="bibr" rid="B24">24</xref>). However, the relatively short reads, which capture the complete CDR3 sequence but exclude much of the V and J regions (including CDR1 and CDR2), present additional challenges to BCR gene annotation and SHM characterization.</p>
<p>Most existing annotation algorithms use a sequence alignment-based approach to resolve the V, D, and J gene segments within a given BCR sequence. The most widely used algorithms is ImMunoGeneTics (IMGT)&#x02019;s VQUEST with JunctionAnalysis (VQUEST&#x02009;&#x0002B;&#x02009;JA) (<xref ref-type="bibr" rid="B25">25</xref>&#x02013;<xref ref-type="bibr" rid="B28">28</xref>), which finds annotations for the V, J, and D genes (in this order) that maximize the sequence alignment scores with respect to a database of unmutated, or germline, sequences. VQUEST&#x02009;&#x0002B;&#x02009;JA also use the conserved 104Cys and 118Trp/Phe residues surrounding the CDR3 [which are residues numbered according to IMGT&#x02019;s unique numbering system (<xref ref-type="bibr" rid="B29">29</xref>)] to fine-tune the annotations (<xref ref-type="bibr" rid="B28">28</xref>). Examples of other algorithms that use similar alignment-based annotation methods include IgBlast (<xref ref-type="bibr" rid="B30">30</xref>), SoDA (<xref ref-type="bibr" rid="B31">31</xref>), JointML (<xref ref-type="bibr" rid="B32">32</xref>), JOINSOLVER (<xref ref-type="bibr" rid="B33">33</xref>), VDJSeq-Solver (<xref ref-type="bibr" rid="B34">34</xref>), MiXCR (<xref ref-type="bibr" rid="B35">35</xref>), IMSEQ (<xref ref-type="bibr" rid="B36">36</xref>), and IgSQUEAL (<xref ref-type="bibr" rid="B37">37</xref>). A different strategy, hidden Markov modeling (HMM), uses statistical models and probability matrices to calculate the most likely series of events leading to each sequence. Algorithms that use HMM are SoDA2 (<xref ref-type="bibr" rid="B38">38</xref>), iHHMune-align (<xref ref-type="bibr" rid="B39">39</xref>), JointHMM (<xref ref-type="bibr" rid="B32">32</xref>), and <italic>partis</italic> (<xref ref-type="bibr" rid="B40">40</xref>). Most BCR annotation algorithms provide high accuracy in annotating V and J segments, but struggle to annotate the N and D segments that define the critical CDR3. The D genes are especially difficult to annotate because they are short in length (e.g., 9&#x02013;18 nt for mice) and can be inverted (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B41">41</xref>), severely truncated, blended with the N regions, and highly mutated. Thus, an effective method for annotating the D gene is to first find the least mutated or most ancestral BCR sequence among the clonally related BCR sequences. However, employing this strategy requires that B-cell lineages be determined concurrently with VDJ annotations.</p>
<p>B-cell lineages are typically determined separately from BCR annotation, despite their common biological basis. A common way to identify clonally related sequences is to cluster sequences with the same V(D)J annotation and CDR3 length and with a high level of BCR nt sequence similarity based on a Hamming distance cutoff (<xref ref-type="bibr" rid="B11">11</xref>, <xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B36">36</xref>, <xref ref-type="bibr" rid="B42">42</xref>). Although annotation-free clustering has been used (<xref ref-type="bibr" rid="B43">43</xref>), determining lineage trees per cluster is difficult without the full VDJ gene annotations. After clustering, lineage trees are assembled per cluster by using algorithms such as MEGA5 (<xref ref-type="bibr" rid="B44">44</xref>), PHLYPIS (<xref ref-type="bibr" rid="B45">45</xref>), ImmuniTree (<xref ref-type="bibr" rid="B46">46</xref>), or IgTree (<xref ref-type="bibr" rid="B47">47</xref>). The assumptions underlying the annotation, clustering, and tree assembly algorithms may be mutually inconsistent, leading to issues such as inadvertent segmentation of long lineage trees owing to divergent VDJ annotations or assembly of binary trees that do not properly reflect B-cell clonal expansion. B-cell lineage trees can be highly branched because a group of identical B cells can give rise to multiple lineages when undergoing SHM. Integrated software applications, such as Change-O (<xref ref-type="bibr" rid="B10">10</xref>) and RevertToGermline&#x02009;&#x0002B;&#x02009;AnnotateTree (<xref ref-type="bibr" rid="B48">48</xref>), streamline the process of assembling lineage trees that are consistent with the annotations, but in these methods, the lineage tree information is not used to improve annotations.</p>
<p>Evaluating the performance of BCR annotation tools is challenging because the true VDJ annotations are not known in real-life BCR sequencing data. As a result, most annotation tools are benchmarked against simulated BCR repertoires created either through in-house simulations or tools such as IgSimulator (<xref ref-type="bibr" rid="B49">49</xref>). However, when processing our own BCR data sets, existing annotation algorithms had difficulty yielding consistent nt substitution frequencies across all V, D, and J segments. There is no biological basis for this inconsistency&#x02014;SHM-inducing enzymes that are responsible for the substitution patterns, such as activation-induced cytidine deaminase (commonly referred to as AID) (<xref ref-type="bibr" rid="B50">50</xref>&#x02013;<xref ref-type="bibr" rid="B53">53</xref>), do not necessarily discriminate between V, D, and J segments. Therefore, we used the correlation between the SHM nt substitution patterns of the V segment and those of the DJ segments as a proxy for the overall annotation quality of real-life BCR repertoires.</p>
<p>In this study, we present B-cell repertoire inductive lineage and immunosequence annotator (BRILIA), a BCR annotation algorithm that concurrently annotates genes, clusters sequences, and assembles lineage trees. BRILIA refines annotations by exploiting mechanistic biases in SHM patterns, N region nt compositions, and directionality of N region synthesis by TDT on the coding versus non-coding DNA strand. We benchmarked BRILIA by processing short 125-bp sequences from simulated human and mouse BCR repertoires and real-life repertoire data obtained from splenic germinal center B cells isolated from C57BL/6 mice. BRILIA identified more highly branched lineage trees and obtained more consistent SHM patterns across the VDJ segments when compared to currently available methods. We demonstrate how BRILIA annotations and lineage trees can be applied in research on affinity maturation, VDJ gene usage frequencies, and SHM mechanisms and hot spot motifs.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2-1">
<title>Obtaining the Database for VDJ Germline Genes</title>
<p>Human and mouse VDJ germline genes were downloaded from the international IMGT database (<uri xlink:href="http://www.imgt.org">http://www.imgt.org</uri>) (<xref ref-type="bibr" rid="B54">54</xref>&#x02013;<xref ref-type="bibr" rid="B61">61</xref>). When annotating C57BL/6 mouse data sets, only genes obtained from the same mouse strain were kept to prevent strain bias of VDJ gene alleles (<xref ref-type="bibr" rid="B41">41</xref>). Pseudogenes, which are germline genes with stop codons or frame shift mutations, were included in the database since their role in producing functional VDJ is still debated (<xref ref-type="bibr" rid="B62">62</xref>, <xref ref-type="bibr" rid="B63">63</xref>). However, only annotations without any stop codons and frame shift errors, referred to as productive VDJ junctions, were analyzed at the end. We included inverted D genes by default because they have been observed occasionally (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B33">33</xref>, <xref ref-type="bibr" rid="B64">64</xref>, <xref ref-type="bibr" rid="B65">65</xref>). Users are given the option to disallow inverted D matching because it may not significantly improve the annotation results (<xref ref-type="bibr" rid="B66">66</xref>). We used the IMGT gene nomenclature, but added an &#x0201C;r&#x0201D; before the family name (e.g., rIGHD01-1&#x0002A;01) for the inverted D genes.</p>
</sec>
<sec id="S2-2">
<title>Simulating BCR Repertoires for Benchmarking Annotation Algorithms</title>
<p>To benchmark our annotation method, we simulated a BCR repertoire so that the true annotations are known. The purpose of this simulated repertoire is to gauge the ability of an algorithm to identify the actual VDJ genes and SHM events and not necessarily to simulate the actual usage frequencies of VDJ genes. Because different genes are used with lesser or higher frequencies than others in real-life repertoires (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B33">33</xref>), replicating this feature in simulated repertoires would have limited the VDJ combinations that were tested. The simulated sequences preserved other details such as the frequent C&#x02009;&#x02794;&#x02009;T and G&#x02009;&#x02794;&#x02009;A mutations mediated by AID (<xref ref-type="bibr" rid="B50">50</xref>&#x02013;<xref ref-type="bibr" rid="B53">53</xref>), preferential occurrence of A mutations over T mutations (referred to as strand-bias mutations) (<xref ref-type="bibr" rid="B67">67</xref>, <xref ref-type="bibr" rid="B68">68</xref>), and biased N region nt compositions (<xref ref-type="bibr" rid="B69">69</xref>, <xref ref-type="bibr" rid="B70">70</xref>). We refer to the N region between the V and D segments as N<sub>VD</sub> and that between D and J segments as N<sub>DJ</sub>. Only productive VDJ junctions were generated to ensure a fair comparison of each algorithm&#x02019;s core functions. A total of 1,000 unmutated, or &#x0201C;germline,&#x0201D; sequences were generated. To simulate SHM, five descendant sequences were generated for each germline sequence by mutating 5&#x02009;nt at non-repeating locations, five times. The combined germline and mutated sequences (a total of 6,000 sequences) were labeled as &#x0201C;clonally expanded&#x0201D; sequences. More details about the BCR simulations are provided in the Supplementary Material. Simulated human and mouse BCR sequences can be found in Datasheets S1 and S2 in Supplementary Material, respectively.</p>
</sec>
<sec id="S2-3">
<title>Extracting BCR Sequences from C57BL/6 Mice Spleens</title>
<p>Germinal center B cells were isolated from wild-type C57BL/6 mice purchased from Jackson Labs. In brief, single-cell suspensions from homogenized spleens were washed using FACS buffer (phosphate-buffered saline, 0.5% bovine serum albumin, and 2&#x02009;mM ethylenediaminetetraacetic acid; Corning, Sigma), lysed with red blood cell buffer (Sigma), and then counterstained with the following B-cell antibodies: B220, IgM, IgD, IgG1, CD38, CD138, and GL-7 (BD Biosciences). All samples were Fc-blocked (anti-CD16/CD32) and stained to evaluate viability (live/dead aqua, Invitrogen) prior to antibody staining. Highly purified (&#x0003E;&#x02009;90%) Spt-GC B cells (B220<sup>&#x0002B;</sup>GL7<sup>&#x0002B;</sup>CD95<sup>&#x0002B;</sup>CD38low) were then isolated using cell sorting on a BD Aria II. Subsequently, purified Spt-GC B cells were pelleted by centrifugation, snap frozen, and then shipped to Adaptive Biotechnologies for DNA extraction and next-generation sequencing of the murine VDJ loci. BCR sequence information was obtained from amplicons beginning within FR3 of the V gene and ending just 3&#x02032; of the complete VDJ junction. Each sequence was trimmed to 125&#x02009;bp maintaining the last 3&#x02009;nt of the sequence codes for the conserved 118Trp (TGG) of the CDR3. For each unique sequence, template counts were also provided, which reflect the number of B cells that had a copy of a particular BCR gene (<xref ref-type="bibr" rid="B71">71</xref>). Animal work was conducted under a United States Army Medical Research Institute of Infectious Diseases (USAMRIID) Institutional Animal Care and Use Committee-approved protocol in compliance with the US Animal Welfare Act, Public Health Service Policy, and other federal statutes and regulations relating to animals and experiments involving animals. The facility in which this research was conducted (USAMRIID) is accredited by the Association for Assessment and Accreditation of Laboratory Animal Care, International and adheres to principles stated in the Guide for the Care and Use of Laboratory Animals, National Research Council, 2011. Real-life BCR sequences used in this work can be found in Datasheet S3 in Supplementary Material.</p>
</sec>
</sec>
<sec id="S3">
<title>Stepwise Procedures</title>
<sec id="S3-1">
<title>BRILIA Algorithm Overview</title>
<p>A flowchart of our BRILIA annotation algorithm is shown in Figure <xref ref-type="fig" rid="F1">1</xref>, along with an example of the process and rationale behind each major step. The key features of our annotation strategy are the alignment strategy that accommodates variable SHM rates per sequence, preservation of nts during alignment that prevent &#x0201C;no D&#x0201D; results, lineage-based clustering, unification of annotations within a cluster, D inverse searches, and refinement steps for D and N regions. BRILIA is written in MATLAB (MathWorks), and all source codes and input files used in this study are available on request.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>BRILIA flowchart with examples and rationales for each step</bold>. For the sample sequences in the middle, the V, N<sub>VD</sub>, D, N<sub>DJ</sub>, and J segments are separated by a space, where a double space indicates a lack of N region (e.g., N<sub>DJ</sub> is absent initially). A lowercase letter is either an N nucleotide (nt) or a mismatched nt with respect to the germline genes, and a bolded letter is a consensus mismatched nt.</p></caption>
<graphic xlink:href="fimmu-07-00681-g001.tif"/>
</fig>
</sec>
<sec id="S3-2">
<title>Defining the Alignment Scoring Method</title>
<p>The initial VDJ annotations rely on aligning sequences to the germline sequences and maximizing the total alignment score for the VDJ segments. We used a custom alignment scoring method defined as
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mtext>Score</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:mrow><mml:msubsup><mml:mi>M</mml:mi><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
where C<italic><sub>i</sub></italic> and M<italic><sub>j</sub></italic> are the numbers of consecutively matched and mismatched nts, respectively, in a segment. For example, if the result of a sequence alignment is &#x0201C;<monospace>AGtTTcC</monospace>,&#x0201D; where lowercase letters represent mismatched nts, the alignment score would be computed as 2<sup>2</sup> &#x02212;&#x02009;1<sup>2</sup> &#x0002B;&#x02009;2<sup>2</sup> &#x02212;&#x02009;1<sup>2</sup> &#x0002B;&#x02009;1<sup>2</sup> &#x0003D;&#x02009;7. To prevent SHM from severely affecting the alignment score, we add a mismatch leniency rule so that point mutations (but not consecutively mismatched nts) can elongate consecutively matched segments. By using the above example and allowing for one mismatch, the score would now be computed as (2&#x02009;&#x0002B;&#x02009;0&#x02009;&#x0002B;&#x02009;2)<sup>2</sup> &#x02212;&#x02009;1<sup>2</sup> &#x0002B;&#x02009;1<sup>2</sup> &#x0003D;&#x02009;16. Note that despite having another point mutation near the 3&#x02032; end, we elongate segments near the 5&#x02032; end first because SHM occurs more frequently near the 5&#x02032; end (<xref ref-type="bibr" rid="B72">72</xref>).</p>
<p>For the V segment, we set the mismatch leniency to be as high as 15% of the V segment length, although the actually mutation% is often less. For the D and J segments, we set a mismatch leniency rule to have the same mutation% as what was found for the V segment, although the D and J segments in the CDR3 can accumulate more mutations (<xref ref-type="bibr" rid="B73">73</xref>). Since the first objective of BRILIA is to identify the least mutated sequence, setting a higher mutation% for the D and J segments is not necessary.</p>
</sec>
<sec id="S3-3">
<title>Step 1: Matching VDJ Genes and Correcting V Gene Indel Errors</title>
<p>For a given VDJ sequence, we matched the V gene first, but with the condition that the last 9&#x02009;nt were preserved for matching the D and J genes. For instance, given 125&#x02009;nt in a sequence, only the first 116&#x02009;nt would be used to determine a V gene. This nt preservation step prevents &#x02018;overmatching&#x02019; a V gene such that the J and D genes cannot be resolved. We also corrected for V gene insertions/deletions (indels) that occurred before the 104Cys, because indels here are likely caused by sequencing errors (<xref ref-type="bibr" rid="B39">39</xref>). We did not correct for indels in the CDR3 because indels can be caused by real VDJ recombination events.</p>
<p>Once a V gene match was found, we preserved 3&#x02009;nt to the right of the V gene segment and then determined the J gene with the remaining nts. For instance, if the first 100 of 125&#x02009;nt were matched to a V gene, then nts 101 to 103 were preserved, while the last 22&#x02009;nt were used to match the J gene. After determining a J gene, all remaining nts were used to match the D gene. Any nts not assigned to a V, D, or J gene were treated as P-nts) and/or non-templated (N) nts, which were then assigned to their respective N<sub>VD</sub> or N<sub>DJ</sub> region.</p>
</sec>
<sec id="S3-4">
<title>Step 2: Assembling Lineage Trees, Clustering Sequences, and Unifying VDJ Annotations</title>
<p>We next clustered the sequences by using lineage trees. Sequences with the same CDR3 lengths and VJ gene family numbers were clustered before constructing lineage trees because the latter process is more computationally expensive. Sequences were considered related to each other if they were within a certain sequence similarity distance. We used a custom distance metric, referred to here as the SHM distance, which resembles the Hamming distance but includes the following adjustments:
<list list-type="simple">
<list-item><label>(1)</label> <p>Consecutively mismatches M number of nts add M<sup>2</sup> to the SHM distance instead of merely M. This increases the distance between clonally unrelated sequences that may have similar VDJ genes but slightly different N regions.</p></list-item>
<list-item><label>(2)</label> <p>Frequently observed [C&#x02009;&#x02794;&#x02009;T, G&#x02009;&#x02794;&#x02009;A, A&#x02009;&#x02794;&#x02009;G, A&#x02009;&#x02794;&#x02009;T] mutations reduce the SHM distance by 0.5 units per mutation; less frequently observed [T&#x02009;&#x02794;&#x02009;C, A&#x02009;&#x02794;&#x02009;C] mutations have no effect; and all other rarer mutations increase the SHM distance by 0.5 units. These adjustments create asymmetry in distances between two sequences, which helps with determining parent&#x02013;child relationships.</p></list-item>
</list></p>
<p>Example: If Seq1&#x02009;&#x0003D;&#x02009;<monospace>ACGCTT</monospace> and Seq2&#x02009;&#x0003D;&#x02009;<monospace>AttgTT</monospace>, then the SHM distance is 3<sup>2</sup> &#x0002B;&#x02009;[&#x02212;0.5&#x02009;&#x0002B;&#x02009;0.5&#x02009;&#x0002B;&#x02009;0.5]&#x02009;&#x0003D;&#x02009;9.5, assuming Seq1 is the parent. If Seq2 is assumed to be the parent, the SHM distance is 3<sup>2</sup> &#x0002B;&#x02009;[0&#x02009;&#x0002B;&#x02009;0.5&#x02009;&#x0002B;&#x02009;0.5]&#x02009;&#x0003D;&#x02009;10. In this case, we would assume Seq1 is the parent of Seq2.</p>
<p>Parent&#x02013;child sequence relationships were determined within each cluster by using a nearest-distance method. The initial linkages generate cyclic dependencies (e.g., Seq1&#x02009;&#x02794;&#x02009;Seq2&#x02009;&#x02794;&#x02009;Seq1) because the root has not yet been assigned. For each independent tree cluster, the root sequence was determined as that which is involved in the cyclic dependency and has the smallest total SHM distance to all other sequences in that cluster. Any ties in the root sequence determination were broken by assigning the sequence with the highest VDJ alignment scores as the root. In an iterative process, the root of each small cluster was linked to any sequence in another cluster, as long as it did not exceed the SHM distance cutoff equal to 3% of the sequence length. Note that this cutoff distance can be adjusted by the user.</p>
<p>Finally, we defined a BRILIA cluster as a group of sequences that shared a common root sequence. The VDJ annotations and N region demarcations for each cluster were unified to match those of the root sequence. The lineage tree was rerooted only if another sequence served as a better root sequence based on a closer distance to the predicted germline genes. Hereafter, any annotation refinements to a cluster were applied to all sequences in the cluster.</p>
</sec>
<sec id="S3-5">
<title>Step 3: Refining D Annotations within a Cluster</title>
<p>After unifying the VDJ annotations per cluster in the prior step, annotation errors can become more apparent. A common indication of annotation error is when all sequences have a mismatched nt in the CDR3 that does not match with the germline sequence, which we will refer to as a consensus mismatched nt (see Figure <xref ref-type="fig" rid="F1">1</xref>, bold letters in example sequences). If a consensus mismatched nt was present in the framework region of the V gene (Vframe), then we assumed that the consensus mismatched nts in the CDR3 were a byproduct of real SHMs. Otherwise, we assumed that the consensus mismatched nts occurred by a suboptimal annotation and attempted to remove them by refining the D gene alignment. Since changing the D gene results also changes the N regions, we had to determine whether the resulting N region compositions agreed with the actions of TDT, which prefers to add A and G (<xref ref-type="bibr" rid="B69">69</xref>, <xref ref-type="bibr" rid="B70">70</xref>, <xref ref-type="bibr" rid="B74">74</xref>). We computed the probability that an N region is created by TDT relative to purely random nt insertion, denoted as <italic>P<sub>TDT</sub></italic>, using the equation below.</p>
<p><disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mtext>TDT</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x0220F;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>0.25</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x0220F;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
where L is the length of the N region and <italic>P<sub><italic>X(j)</italic></sub></italic> is the probability that TDT adds nt <italic>X</italic> (A, C, G, T) at position <italic>j</italic> in the N region. We assumed that nts that are not associated with TDT activity has an occurrence probability of 0.25, whereas nts added by TDT have the following occurrence probabilities: <italic>P<sub>A</sub></italic>&#x02009;&#x0003D;&#x02009;0.25, <italic>P<sub>C</sub></italic>&#x02009;&#x0003D;&#x02009;0.08, <italic>P<sub>G</sub></italic>&#x02009;&#x0003D;&#x02009;0.60, and <italic>P<sub>T</sub></italic>&#x02009;&#x0003D;&#x02009;0.07 (see Figure S2 in Supplementary Material). These probability values were obtained from the N regions of our data set from mice, after converting these regions to their complement sequences if there were more CT content than AG content, which should better capture the TDT-mediated DNA elongation patterns (<xref ref-type="bibr" rid="B69">69</xref>) (see <xref ref-type="supplementary-material" rid="S9">Supplementary Material</xref>). We note that the <italic>P<sub>X</sub></italic> values are reported for healthy mice, and we do not expect these to vary much across subjects unless there are abnormal conditions [e.g., nt pool imbalance (<xref ref-type="bibr" rid="B75">75</xref>)].</p>
<p>We next calculated a custom N region likelihood score, or <italic>Nscore</italic>, by using the following equation:
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mrow><mml:mtext>Nscore</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mtext>TDT</mml:mtext></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p>The <italic>Nscore</italic> was calculated for both the normal and complement sequences of each N region, and only the higher score was retained. A different D gene annotation was accepted only if it increased the sum of the VDJ alignment scores and the Nscores for N<sub>VD</sub> and N<sub>DJ</sub>. If a consensus mismatch persisted, then we evaluated whether this was caused by the incorrect demarcation of the N regions, as discussed next.</p>
</sec>
<sec id="S3-6">
<title>Step 4: Refining N Regions within a Cluster</title>
<p>Improper demarcation of N regions can also cause consecutively mismatched nts to exist in the CDR3 (marked as bold lower case letters in a sequence alignment), which can be fixed by redefining where the VDJ gene segments are. For any three consecutively mismatched nts near gene segment edges, we automatically reassigned the edges to the N regions because such events are likely caused by annotation errors. For example, if a V gene ended with &#x0201C;5&#x02032;<monospace>-TG<bold>agg</bold>GG</monospace>,&#x0201D; then &#x0201C;<monospace>agggg</monospace>&#x0201D; was automatically added to the N<sub>VD</sub> region. For all other cases, we checked whether the gene segment edges had compositions that reflected TDT-mediated nt insertion. Several examples cases are provided below.</p>
<list list-type="bullet">
<list-item><p>If no N region is initially present and trimming would create one, then we checked whether <italic>P<sub>TDT</sub></italic> of the trimmed nts was &#x0003E;0.50. For instance, if a V gene ended with &#x0201C;<monospace>5&#x02032;-TGCA<bold>g</bold>GG</monospace>,&#x0201D; then <italic>P<sub>TDT</sub></italic> for &#x0201C;<monospace><bold>g</bold>GG</monospace>&#x0201D; is 0.93 and therefore, &#x0201C;<monospace>ggg</monospace>&#x0201D; became the N<sub>VD</sub> region. If a V gene ended with &#x0201C;<monospace>5&#x02032;-CA<bold>t</bold>ATC</monospace>,&#x0201D; then <italic>P<sub>TDT</sub></italic> for &#x0201C;<monospace><bold>t</bold>ATC</monospace>&#x0201D; is 0.40, and therefore, no trimming was performed.</p></list-item>
<list-item><p>If an N region is initially present, then we calculate whether adding the edge nts to the N region would increase <italic>P<sub>TDT</sub></italic>. For instance, if the V gene ended with &#x0201C;<monospace>5&#x02032;-TGCA<bold>g</bold>GG</monospace>&#x0201D; and the N<sub>DV</sub> region was &#x0201C;<monospace>ccc</monospace>,&#x0201D; adding &#x0201C;<monospace><bold>g</bold>GG</monospace>&#x0201D; to &#x0201C;<monospace>ccc</monospace>&#x0201D; would have created an unfavorable &#x0201C;<monospace>gggccc</monospace>&#x0201D; in the N<sub>VD</sub> region with a reduction in P<sub>TDT</sub> from 0.93 to 0.31; hence, no trimming was performed.</p></list-item>
</list>
</sec>
</sec>
<sec id="S4">
<title>Results</title>
<sec id="S4-1">
<title>VDJ Annotation of Simulated BCR Repertoire Data</title>
<p>Obtaining accurate gene annotations is essential to measuring gene usage frequency (<xref ref-type="bibr" rid="B76">76</xref>), tracking affinity maturation and selection pressure (<xref ref-type="bibr" rid="B77">77</xref>&#x02013;<xref ref-type="bibr" rid="B79">79</xref>), and studying SHM-associated enzyme activities (<xref ref-type="bibr" rid="B51">51</xref>&#x02013;<xref ref-type="bibr" rid="B53">53</xref>, <xref ref-type="bibr" rid="B80">80</xref>&#x02013;<xref ref-type="bibr" rid="B83">83</xref>). Because the true accuracy of VDJ annotation cannot be determined when using real-life BCR repertoires, we created a synthetic repertoire for benchmarking purposes. We compared our annotations with those of two recently updated algorithms, VQUEST&#x02009;&#x0002B;&#x02009;JA (<xref ref-type="bibr" rid="B25">25</xref>&#x02013;<xref ref-type="bibr" rid="B28">28</xref>) and <italic>partis</italic> (<xref ref-type="bibr" rid="B40">40</xref>). If an algorithm suggested multiple VDJ annotations, we retained only the first suggestion.</p>
<p>We compared how well the algorithms could obtain an exact match to the actual gene (up to the gene allele number) and also a degenerate match to any gene name that contains &#x02265;98% of the same nts (Tables <xref ref-type="table" rid="T1">1</xref> and <xref ref-type="table" rid="T2">2</xref>). An example of a degenerate match is when &#x0201C;ATTAACTA&#x0201D; of IGHD1-1&#x0002A;01 was used generate a BCR sequence and the annotation suggested IGHD1-1&#x0002A;02, which has the same nts. Tables <xref ref-type="table" rid="T1">1</xref> and <xref ref-type="table" rid="T2">2</xref> compare the gene matching performances of the three algorithms on human and mouse sequences, respectively, for both germline and clonally expanded sequences. However, we did not annotate the mouse repertoire with <italic>partis</italic> because it was intended for annotating human BCR genes only (<xref ref-type="bibr" rid="B40">40</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Annotation accuracy in simulated human B-cell receptor (BCR) repertoire</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center"/>
<th valign="top" align="center" colspan="2">B-cell repertoire inductive lineage and immunosequence annotator (BRILIA)</th>
<th valign="top" align="center" colspan="2">VQUEST &#x0002B; JA</th>
<th valign="top" align="center" colspan="2"><italic>Partis</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" valign="middle" rowspan="8"><bold>Germline</bold> 1,000 sequences</td>
<td align="left" valign="top">V exact match</td>
<td align="center" valign="top">826</td>
<td align="center" valign="top">83%</td>
<td align="center" valign="top">393</td>
<td align="center" valign="top">39%</td>
<td align="center" valign="top">481</td>
<td align="center" valign="top">48%</td>
</tr>
<tr>
<td align="left" valign="top">D exact match</td>
<td align="center" valign="top">830</td>
<td align="center" valign="top">83%</td>
<td align="center" valign="top">613</td>
<td align="center" valign="top">61%</td>
<td align="center" valign="top">660</td>
<td align="center" valign="top">66%</td>
</tr>
<tr>
<td align="left" valign="top">J exact match</td>
<td align="center" valign="top">837</td>
<td align="center" valign="top">84%</td>
<td align="center" valign="top">530</td>
<td align="center" valign="top">53%</td>
<td align="center" valign="top">686</td>
<td align="center" valign="top">69%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ exact match</td>
<td align="center" valign="top">562</td>
<td align="center" valign="top">56%</td>
<td align="center" valign="top">139</td>
<td align="center" valign="top">14%</td>
<td align="center" valign="top">231</td>
<td align="center" valign="top">23%</td>
</tr>
<tr>
<td align="left" valign="top">V degen match</td>
<td align="center" valign="top">1,000</td>
<td align="center" valign="top">100%</td>
<td align="center" valign="top">862</td>
<td align="center" valign="top">86%</td>
<td align="center" valign="top">969</td>
<td align="center" valign="top">97%</td>
</tr>
<tr>
<td align="left" valign="top">D degen match</td>
<td align="center" valign="top">956</td>
<td align="center" valign="top">96%</td>
<td align="center" valign="top">808</td>
<td align="center" valign="top">81%</td>
<td align="center" valign="top">890</td>
<td align="center" valign="top">89%</td>
</tr>
<tr>
<td align="left" valign="top">J degen match</td>
<td align="center" valign="top">985</td>
<td align="center" valign="top">99%</td>
<td align="center" valign="top">843</td>
<td align="center" valign="top">84%</td>
<td align="center" valign="top">938</td>
<td align="center" valign="top">94%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ degen match</td>
<td align="center" valign="top">941</td>
<td align="center" valign="top">94%</td>
<td align="center" valign="top">655</td>
<td align="center" valign="top">66%</td>
<td align="center" valign="top">857</td>
<td align="center" valign="top">86%</td>
</tr>
<tr>
<td align="center" valign="middle" rowspan="8"><bold>Clonally expanded</bold> 6,000 sequences</td>
<td align="left" valign="top">V exact match</td>
<td align="center" valign="top">4,045</td>
<td align="center" valign="top">67%</td>
<td align="center" valign="top">1,928</td>
<td align="center" valign="top">32%</td>
<td align="center" valign="top">2,442</td>
<td align="center" valign="top">41%</td>
</tr>
<tr>
<td align="left" valign="top">D exact match</td>
<td align="center" valign="top">4,359</td>
<td align="center" valign="top">73%</td>
<td align="center" valign="top">2,990</td>
<td align="center" valign="top">50%</td>
<td align="center" valign="top">3,419</td>
<td align="center" valign="top">57%</td>
</tr>
<tr>
<td align="left" valign="top">J exact match</td>
<td align="center" valign="top">4,548</td>
<td align="center" valign="top">76%</td>
<td align="center" valign="top">2,783</td>
<td align="center" valign="top">46%</td>
<td align="center" valign="top">3,777</td>
<td align="center" valign="top">63%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ exact match</td>
<td align="center" valign="top">2,291</td>
<td align="center" valign="top">38%</td>
<td align="center" valign="top">508</td>
<td align="center" valign="top">8%</td>
<td align="center" valign="top">936</td>
<td align="center" valign="top">16%</td>
</tr>
<tr>
<td align="left" valign="top">V degen match</td>
<td align="center" valign="top">5,159</td>
<td align="center" valign="top">86%</td>
<td align="center" valign="top">4,570</td>
<td align="center" valign="top">76%</td>
<td align="center" valign="top">5,272</td>
<td align="center" valign="top">88%</td>
</tr>
<tr>
<td align="left" valign="top">D degen match</td>
<td align="center" valign="top">4,979</td>
<td align="center" valign="top">83%</td>
<td align="center" valign="top">3,893</td>
<td align="center" valign="top">65%</td>
<td align="center" valign="top">4,575</td>
<td align="center" valign="top">76%</td>
</tr>
<tr>
<td align="left" valign="top">J degen match</td>
<td align="center" valign="top">5,477</td>
<td align="center" valign="top">91%</td>
<td align="center" valign="top">4,546</td>
<td align="center" valign="top">76%</td>
<td align="center" valign="top">5,195</td>
<td align="center" valign="top">87%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ degen match</td>
<td align="center" valign="top">4,059</td>
<td align="center" valign="top">68%</td>
<td align="center" valign="top">2,670</td>
<td align="center" valign="top">45%</td>
<td align="center" valign="top">3,827</td>
<td align="center" valign="top">64%</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Annotation accuracy in simulated mouse (C57BL/6) B-cell receptor (BCR) repertoire</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center"/>
<th valign="top" align="center" colspan="2">B-cell repertoire inductive lineage and immunosequence annotator (BRILIA)</th>
<th valign="top" align="center" colspan="2">VQUEST &#x0002B; JA</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" valign="middle" rowspan="8"><bold>Germline</bold> 1,000 sequences</td>
<td align="left" valign="top">V exact match</td>
<td align="center" valign="top">970</td>
<td align="center" valign="top">97%</td>
<td align="center" valign="top">808</td>
<td align="center" valign="top">81%</td>
</tr>
<tr>
<td align="left" valign="top">D exact match</td>
<td align="center" valign="top">796</td>
<td align="center" valign="top">80%</td>
<td align="center" valign="top">382</td>
<td align="center" valign="top">38%</td>
</tr>
<tr>
<td align="left" valign="top">J exact match</td>
<td align="center" valign="top">965</td>
<td align="center" valign="top">97%</td>
<td align="center" valign="top">558</td>
<td align="center" valign="top">56%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ exact match</td>
<td align="center" valign="top">751</td>
<td align="center" valign="top">75%</td>
<td align="center" valign="top">184</td>
<td align="center" valign="top">18%</td>
</tr>
<tr>
<td align="left" valign="top">V degen match</td>
<td align="center" valign="top">998</td>
<td align="center" valign="top">100%</td>
<td align="center" valign="top">942</td>
<td align="center" valign="top">94%</td>
</tr>
<tr>
<td align="left" valign="top">D degen match</td>
<td align="center" valign="top">947</td>
<td align="center" valign="top">95%</td>
<td align="center" valign="top">833</td>
<td align="center" valign="top">83%</td>
</tr>
<tr>
<td align="left" valign="top">J degen match</td>
<td align="center" valign="top">993</td>
<td align="center" valign="top">99%</td>
<td align="center" valign="top">918</td>
<td align="center" valign="top">92%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ degen match</td>
<td align="center" valign="top">938</td>
<td align="center" valign="top">94%</td>
<td align="center" valign="top">753</td>
<td align="center" valign="top">75%</td>
</tr>
<tr>
<td align="center" valign="middle" rowspan="8"><bold>Clonally expanded</bold> 6,000 sequences</td>
<td align="left" valign="top">V exact match</td>
<td align="center" valign="top">5,271</td>
<td align="center" valign="top">88%</td>
<td align="center" valign="top">4,425</td>
<td align="center" valign="top">74%</td>
</tr>
<tr>
<td align="left" valign="top">D exact match</td>
<td align="center" valign="top">4,088</td>
<td align="center" valign="top">68%</td>
<td align="center" valign="top">1,721</td>
<td align="center" valign="top">29%</td>
</tr>
<tr>
<td align="left" valign="top">J exact match</td>
<td align="center" valign="top">5,604</td>
<td align="center" valign="top">93%</td>
<td align="center" valign="top">3,244</td>
<td align="center" valign="top">54%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ exact match</td>
<td align="center" valign="top">3,428</td>
<td align="center" valign="top">57%</td>
<td align="center" valign="top">7,65</td>
<td align="center" valign="top">13%</td>
</tr>
<tr>
<td align="left" valign="top">V degen match</td>
<td align="center" valign="top">5,474</td>
<td align="center" valign="top">91%</td>
<td align="center" valign="top">5,214</td>
<td align="center" valign="top">87%</td>
</tr>
<tr>
<td align="left" valign="top">D degen match</td>
<td align="center" valign="top">4,910</td>
<td align="center" valign="top">82%</td>
<td align="center" valign="top">3,878</td>
<td align="center" valign="top">65%</td>
</tr>
<tr>
<td align="left" valign="top">J degen match</td>
<td align="center" valign="top">5,777</td>
<td align="center" valign="top">96%</td>
<td align="center" valign="top">5,275</td>
<td align="center" valign="top">88%</td>
</tr>
<tr>
<td align="left" valign="top">VDJ degen match</td>
<td align="center" valign="top">4,351</td>
<td align="center" valign="top">73%</td>
<td align="center" valign="top">3,129</td>
<td align="center" valign="top">52%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The annotation accuracy of the germline sequences reflects how well each algorithm works in the best-case scenario where SHM does not obscure VDJ genes. This accuracy is important because if a germline sequence exists within a cluster of clonally related sequences, then the corresponding annotation will be applied to all members in the cluster. BRILIA provided the highest accuracy of matching, followed closely by <italic>partis</italic> and VQUEST&#x02009;&#x0002B;&#x02009;JA (Tables <xref ref-type="table" rid="T1">1</xref> and <xref ref-type="table" rid="T2">2</xref>, &#x0201C;Germline Sequence&#x0201D; rows).</p>
<p>The annotation accuracy of clonally expanded sequences reflects how well each algorithm can annotate sequences that have undergone extensive SHM. For the simulated human BCR sequences, BRILIA achieved 83% D gene degenerate matching accuracy, compared to 65% by VQUEST&#x02009;&#x0002B;&#x02009;JA and 76% by <italic>partis</italic> (Table <xref ref-type="table" rid="T1">1</xref>, &#x0201C;Clonally Expanded&#x0201D; rows). For the simulated mouse BCR sequences, BRILIA achieved 82% degenerate D gene matching accuracy, compared to 65% by VQUEST&#x02009;&#x0002B;&#x02009;JA (Table <xref ref-type="table" rid="T2">2</xref>, &#x0201C;Clonally Expanded&#x0201D; rows). The degenerate V and J gene annotation accuracies are comparable across all algorithms, as might be expected given the relatively long lengths of the V gene segments and the limited number of J germline genes. It is important to note that the overall V and J annotation accuracies presented here are lower than those obtained in previously published other benchmark tests (<xref ref-type="bibr" rid="B37">37</xref>, <xref ref-type="bibr" rid="B40">40</xref>); this is because previous benchmarks used the full VDJ segments (&#x0007E;400&#x02009;bp), while we used much shorter sequences (125&#x02009;bp) typical of CDR3-focused next-generation sequencing (<xref ref-type="bibr" rid="B71">71</xref>).</p>
</sec>
<sec id="S4-2">
<title>SHM Identification Accuracy of Simulated BCR Repertoire Data</title>
<p>In addition to correctly annotating the VDJ genes, it is important to accurately identify SHMs within the CDR3. For each algorithm, and for sequences grouped by the same number of simulated SHMs, we first computed the accuracy of determining mutated and unmutated nts in the CDR3 (Table <xref ref-type="table" rid="T3">3</xref>). For all algorithms, accuracy decreased as sequences accumulated more SHMs, as expected, but BRILIA retained the highest accuracy, followed by <italic>partis</italic> and closely by VQUEST&#x02009;&#x0002B;&#x02009;JA. We also computed the positive prediction rate of mutations, which reflects how many of the predicted SHMs were true. BRILIA retained the highest positive prediction rates, followed by VQUEST&#x02009;&#x0002B;&#x02009;JA and closely by <italic>partis</italic>.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p><bold>Accuracy (ACC) and positive prediction rate (PPR) of identifying SHMs in the complementarity determining region 3 (CDR3) from the simulated clonally expanded BCRs</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="left">Lineage generation</th>
<th valign="top" align="center">&#x00023; of SHMs in 125-bp BCR sequence</th>
<th valign="top" align="center">TP</th>
<th valign="top" align="center">TN</th>
<th valign="top" align="center">FP</th>
<th valign="top" align="center">FN</th>
<th valign="top" align="center">ACC</th>
<th valign="top" align="center">PPR</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" valign="middle" rowspan="6"><bold>BRILIA human</bold></td>
<td align="left" valign="top">Germline</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">51,544</td>
<td align="center" valign="top">1,031</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">1st</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">1,230</td>
<td align="center" valign="top">50,072</td>
<td align="center" valign="top">629</td>
<td align="center" valign="top">644</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.66</td>
</tr>
<tr>
<td align="left" valign="top">2nd</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">2,433</td>
<td align="center" valign="top">48,146</td>
<td align="center" valign="top">676</td>
<td align="center" valign="top">1,320</td>
<td align="center" valign="top">0.96</td>
<td align="center" valign="top">0.78</td>
</tr>
<tr>
<td align="left" valign="top">3rd</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">3,373</td>
<td align="center" valign="top">46,141</td>
<td align="center" valign="top">755</td>
<td align="center" valign="top">2,306</td>
<td align="center" valign="top">0.94</td>
<td align="center" valign="top">0.82</td>
</tr>
<tr>
<td align="left" valign="top">4th</td>
<td align="center" valign="top">20</td>
<td align="center" valign="top">4,376</td>
<td align="center" valign="top">44,310</td>
<td align="center" valign="top">633</td>
<td align="center" valign="top">3,256</td>
<td align="center" valign="top">0.93</td>
<td align="center" valign="top">0.87</td>
</tr>
<tr>
<td align="left" valign="top">5th</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">5,588</td>
<td align="center" valign="top">42,370</td>
<td align="center" valign="top">612</td>
<td align="center" valign="top">4,005</td>
<td align="center" valign="top">0.91</td>
<td align="center" valign="top">0.90</td>
</tr>
<tr>
<td align="center" valign="middle" rowspan="6"><bold>VQUEST&#x02009;&#x0002B;&#x02009;JA human</bold></td>
<td align="left" valign="top">Germline</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">46,724</td>
<td align="center" valign="top">502</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">1st</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">913</td>
<td align="center" valign="top">44,296</td>
<td align="center" valign="top">812</td>
<td align="center" valign="top">758</td>
<td align="center" valign="top">0.97</td>
<td align="center" valign="top">0.53</td>
</tr>
<tr>
<td align="left" valign="top">2nd</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">1,636</td>
<td align="center" valign="top">41,973</td>
<td align="center" valign="top">844</td>
<td align="center" valign="top">1,645</td>
<td align="center" valign="top">0.95</td>
<td align="center" valign="top">0.66</td>
</tr>
<tr>
<td align="left" valign="top">3rd</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">2,146</td>
<td align="center" valign="top">39,634</td>
<td align="center" valign="top">919</td>
<td align="center" valign="top">2,745</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">0.70</td>
</tr>
<tr>
<td align="left" valign="top">4th</td>
<td align="center" valign="top">20</td>
<td align="center" valign="top">2,365</td>
<td align="center" valign="top">36,937</td>
<td align="center" valign="top">1,010</td>
<td align="center" valign="top">4,007</td>
<td align="center" valign="top">0.89</td>
<td align="center" valign="top">0.70</td>
</tr>
<tr>
<td align="left" valign="top">5th</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">2,324</td>
<td align="center" valign="top">33,983</td>
<td align="center" valign="top">1,114</td>
<td align="center" valign="top">5,410</td>
<td align="center" valign="top">0.85</td>
<td align="center" valign="top">0.68</td>
</tr>
<tr>
<td align="center" valign="middle" rowspan="6"><bold><italic>partis</italic> human</bold></td>
<td align="left" valign="top">Germline</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">50,685</td>
<td align="center" valign="top">654</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">1st</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">1,160</td>
<td align="center" valign="top">48,182</td>
<td align="center" valign="top">1,149</td>
<td align="center" valign="top">665</td>
<td align="center" valign="top">0.96</td>
<td align="center" valign="top">0.50</td>
</tr>
<tr>
<td align="left" valign="top">2nd</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">2,325</td>
<td align="center" valign="top">45,946</td>
<td align="center" valign="top">1,695</td>
<td align="center" valign="top">1,334</td>
<td align="center" valign="top">0.94</td>
<td align="center" valign="top">0.58</td>
</tr>
<tr>
<td align="left" valign="top">3rd</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">3,505</td>
<td align="center" valign="top">43,251</td>
<td align="center" valign="top">2,228</td>
<td align="center" valign="top">2,001</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">0.61</td>
</tr>
<tr>
<td align="left" valign="top">4th</td>
<td align="center" valign="top">20</td>
<td align="center" valign="top">4,655</td>
<td align="center" valign="top">40,405</td>
<td align="center" valign="top">2,828</td>
<td align="center" valign="top">2,686</td>
<td align="center" valign="top">0.89</td>
<td align="center" valign="top">0.62</td>
</tr>
<tr>
<td align="left" valign="top">5th</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">5,648</td>
<td align="center" valign="top">37,713</td>
<td align="center" valign="top">3,154</td>
<td align="center" valign="top">3,441</td>
<td align="center" valign="top">0.87</td>
<td align="center" valign="top">0.64</td>
</tr>
<tr>
<td align="center" valign="middle" rowspan="6"><bold>BRILIA mouse</bold></td>
<td align="left" valign="top">Germline</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">42,144</td>
<td align="center" valign="top">804</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">1st</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">976</td>
<td align="center" valign="top">41,062</td>
<td align="center" valign="top">410</td>
<td align="center" valign="top">500</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.70</td>
</tr>
<tr>
<td align="left" valign="top">2nd</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">1,976</td>
<td align="center" valign="top">39,518</td>
<td align="center" valign="top">443</td>
<td align="center" valign="top">1,011</td>
<td align="center" valign="top">0.97</td>
<td align="center" valign="top">0.82</td>
</tr>
<tr>
<td align="left" valign="top">3rd</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">2,746</td>
<td align="center" valign="top">37,891</td>
<td align="center" valign="top">527</td>
<td align="center" valign="top">1,784</td>
<td align="center" valign="top">0.95</td>
<td align="center" valign="top">0.84</td>
</tr>
<tr>
<td align="left" valign="top">4th</td>
<td align="center" valign="top">20</td>
<td align="center" valign="top">3,531</td>
<td align="center" valign="top">36,427</td>
<td align="center" valign="top">421</td>
<td align="center" valign="top">2,569</td>
<td align="center" valign="top">0.93</td>
<td align="center" valign="top">0.89</td>
</tr>
<tr>
<td align="left" valign="top">5th</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">4,579</td>
<td align="center" valign="top">34,951</td>
<td align="center" valign="top">374</td>
<td align="center" valign="top">3,044</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">0.92</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="6"><bold>VQUEST&#x02009;&#x0002B;&#x02009;JA mouse</bold></td>
<td align="left" valign="top">Germline</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">38,602</td>
<td align="center" valign="top">209</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">1st</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">743</td>
<td align="center" valign="top">36,843</td>
<td align="center" valign="top">459</td>
<td align="center" valign="top">592</td>
<td align="center" valign="top">0.97</td>
<td align="center" valign="top">0.62</td>
</tr>
<tr>
<td align="left" valign="top">2nd</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">1,407</td>
<td align="center" valign="top">35,471</td>
<td align="center" valign="top">522</td>
<td align="center" valign="top">1,315</td>
<td align="center" valign="top">0.95</td>
<td align="center" valign="top">0.73</td>
</tr>
<tr>
<td align="left" valign="top">3rd</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">1,864</td>
<td align="center" valign="top">33,978</td>
<td align="center" valign="top">639</td>
<td align="center" valign="top">2,270</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">0.74</td>
</tr>
<tr>
<td align="left" valign="top">4th</td>
<td align="center" valign="top">20</td>
<td align="center" valign="top">2,176</td>
<td align="center" valign="top">32,294</td>
<td align="center" valign="top">680</td>
<td align="center" valign="top">3,328</td>
<td align="center" valign="top">0.90</td>
<td align="center" valign="top">0.76</td>
</tr>
<tr>
<td align="left" valign="top">5th</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">2,298</td>
<td align="center" valign="top">30,090</td>
<td align="center" valign="top">739</td>
<td align="center" valign="top">4,346</td>
<td align="center" valign="top">0.86</td>
<td align="center" valign="top">0.76</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Accuracy is computed as (TP&#x02009;&#x0002B;&#x02009;TN)/(TP&#x02009;&#x0002B;&#x02009;TN&#x02009;&#x0002B;&#x02009;FN&#x02009;&#x0002B;&#x02009;FP), where T&#x02009;&#x0003D;&#x02009;true, P&#x02009;&#x0003D;&#x02009;positive, F&#x02009;&#x0003D;&#x02009;false, and N&#x02009;&#x0003D;&#x02009;negative. PPR is computed as TP/(TP&#x02009;&#x0002B;&#x02009;FP). The nucleotides in the CDR3 are pooled together based on the sequence lineage (e.g., 1st descendants of the Germline B cells), prior to determining the SHM accuracy statistics. The total nucleotide count differs among algorithms because the annotation coverage differs</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>We next measured how well each BCR annotation method can identify the frequency in which one nt (X<sub>0</sub>) mutates to another nt (X<sub>1</sub>), which we will refer as the SHM propensity. The SHM propensities are not uniform, and certain X<sub>0</sub>&#x02009;&#x02794;&#x02009;X<sub>1</sub> mutations occur more frequently than others. For instance, the C&#x02009;&#x02794;&#x02009;T mutations (and the complement G&#x02009;&#x02794;&#x02009;A mutations) occur frequently because they are triggered by activation-induced cytidine deaminase (AID), which initiates the C&#x02009;&#x02794;&#x02009;U&#x02009;&#x02794;&#x02009;T substitutions (<xref ref-type="bibr" rid="B50">50</xref>&#x02013;<xref ref-type="bibr" rid="B53">53</xref>). Although deaminases are known to act on specific nt sequence motifs, called hot spots, there is no evidence that they can discriminate between the V, D, and J segments of the BCR. Therefore, we expect high-quality SHM identification to show SHM propensities that are (1) consistent across the entire VDJ junction and (2) agree with deaminase-mediated nt substitution patterns.</p>
<p>Figure <xref ref-type="fig" rid="F2">2</xref> shows the correlation between SHM propensities for the V and DJ segments, as predicted by each method for the simulated human and mouse BCRs. All methods yielded SHM rates that are highly correlated across the V and DJ segments, as shown by the Pearson correlation coefficient (<italic>R<sub>corr</sub></italic>) being close to 1. However, VQUEST&#x02009;&#x0002B;&#x02009;JA and <italic>partis</italic> tended to underpredict well-known SHM propensities (i.e., C&#x02009;&#x02794;&#x02009;T, G&#x02009;&#x02794;&#x02009;A, and A&#x02009;&#x02794;&#x02009;G) in the DJ segments, as shown by the reduced linear regression line slope (<italic>Slope</italic>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Somatic hypermutation (SHM) propensity correlations between the V and DJ segments for simulated (A) human and (B) mouse B-cell receptor sequences</bold>. The combination of color and shape of a data point represents a SHM propensity or the mutation frequency of nucleotide (nt) X<sub>0</sub> to nt X<sub>1</sub>. The <italic>x</italic>- and <italic>y</italic> axes show the normalized mutation frequencies (e.g., <italic>P<sub>A&#x02794;T</sub></italic>&#x02009;&#x0002B;&#x02009;<italic>P<sub>A&#x02794;C</sub></italic>&#x02009;&#x0002B;&#x02009;<italic>P<sub>A&#x02794;G</sub></italic>&#x02009;&#x0003D;&#x02009;1) for the V and DJ segments, respectively. <italic>R<sub>corr</sub></italic> is the Pearson correlation coefficient, while <italic>Slope</italic> is the slope of the linear regression line.</p></caption>
<graphic xlink:href="fimmu-07-00681-g002.tif"/>
</fig>
<p>Assessing annotation quality for real-life BCR repertoires is difficult because the true gene annotations are unknown. The correlation of SHM propensities across the V and DJ segments provides an alternate measure of SHM identification accuracy that does not rely on knowing the true VDJ gene assignments. One could assess the quality of VDJ gene annotations indirectly by testing if SHM propensities are consistent across the VDJ junction. Given the common biological basis for SHM across the VDJ junction, high-quality annotations should yield measures of <italic>R<sub>corr</sub></italic> and <italic>Slope</italic> that approach the value of 1. However, the proper generation of these correlation metrics in real BCR repertoire data requires the determination of B-cell lineages because SHM propensities should be identified with respect to a parent&#x02013;child relationship between a pair of sequences, not against a predicted germline sequence, where inherited mutations from a previous generation would be treated as new independent mutation events. For our simulated sequences, determining lineages was not as critical because mutations did not occur more than once in the same place; in other words, the SHM propensities computed from germline&#x02013;child sequences would be similar to those from parent&#x02013;child sequences. In real-life BCR sequences where multiple mutations can occur in the same position, the SHM propensities computed from parent&#x02013;child versus germline&#x02013;child sequence pairs will differ.</p>
</sec>
<sec id="S4-3">
<title>SHM Identification Accuracy of Real-Life Mice BCR Repertoire Data</title>
<p>To test how well BRILIA performs on real-life data sets, we sequenced and analyzed 12,300 unique BCR gene sequences collected from the spleen germinal centers of C57BL/6 mice. It is important to note that these mice were not immunized, and thus, the B cells isolated from the spleen likely developed within spontaneously-formed germinal centers (Spt-GCs). Although the exact cause of Spt-GC formation is unclear, it is thought to arise for a range of reasons, from autoimmunity (<xref ref-type="bibr" rid="B84">84</xref>, <xref ref-type="bibr" rid="B85">85</xref>) to bacterial infection or escape (<xref ref-type="bibr" rid="B86">86</xref>). Previous studies have suggested that Spt-GCs resemble immunization-induced GCs and spontaneous GC B cells undergo some degree of affinity maturation, including accumulation of SHMs and class switching (<xref ref-type="bibr" rid="B84">84</xref>).</p>
<p>We compared BRILIA with a method of processing BCR sequence data that entails grouping sequences with the same VDJ annotations and CDR3 lengths returned by VQUEST&#x02009;&#x0002B;&#x02009;JA, followed by an additional clustering step based on a sequence similarity cutoff distance (hereafter the &#x0201C;Standard&#x0201D; method). For the standard method, we used the same lineage-tree based clustering step as that used by BRILIA; the main differences between the standard and BRILIA methods lie in the alignment algorithms and annotation-based clustering step that occurs prior to the lineage-based clustering step.</p>
<p>We first describe the traditional approach of showing the SHM level of repertoires, i.e., to count the number of mutated nts in a sequence with respect to the germline sequence (Figure <xref ref-type="fig" rid="F3">3</xref>A). Both the standard method and BRILIA appear to return similar SHM frequencies. However, the correlation of SHM propensities between the V and DJ segments differ significantly between the two methods (Figures <xref ref-type="fig" rid="F3">3</xref>B,C). The standard method tended to underpredict SHMs in the DJ segments, and the T&#x02009;&#x02794;&#x02009;X mutation frequencies of the V segment show no correlation with those of the DJ segments (Figure <xref ref-type="fig" rid="F3">3</xref>B). In contrast, the same correlation plot based on BRILIA annotations show a high correlation between the V and DJ segments (Figure <xref ref-type="fig" rid="F3">3</xref>C). These results suggest that while BRILIA and the standard method estimate a similar <italic>number</italic> of SHMs, BRILIA is more accurate in identifying the SHM positions themselves.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Comparison of somatic hypermutation (SHM) identification in real-life C57BL/6 B-cell receptor repertoires between the standard method and BRILIA</bold>. <bold>(A)</bold> Frequency distribution of SHMs per sequence predicted for all sequences in relation to their corresponding cluster&#x02019;s germline sequence. <bold>(B)</bold> SHM propensity correlation returned by the standard method. Note that SHMs were determined for parent&#x02013;child sequence pairs and not germline&#x02013;child sequence pairs. <bold>(C)</bold> SHM propensity correlation returned by BRILIA.</p></caption>
<graphic xlink:href="fimmu-07-00681-g003.tif"/>
</fig>
</sec>
<sec id="S4-4">
<title>VDJ Gene Usages for Real-Life Mice BCR Repertoire Data</title>
<p>Tracking VDJ gene usage is relevant for combinatory gene usage frequency studies (<xref ref-type="bibr" rid="B33">33</xref>, <xref ref-type="bibr" rid="B76">76</xref>, <xref ref-type="bibr" rid="B87">87</xref>). Figure <xref ref-type="fig" rid="F4">4</xref> shows the VDJ gene usage frequencies in terms of both overall gene family usage frequency (top and right bar charts) and frequency of VD and DJ pairs (scatter plots). Although the standard and BRILIA methods predicted similar usage frequencies for V and J gene families (Figures <xref ref-type="fig" rid="F4">4</xref>A,B, top bar charts), their predictions for the D gene families differed substantially. BRILIA results show that IGHD4 is used twice as much as IGHD3 and that IGHD2 is used 20% more than IGHD1. In contrast, the standard method results show that IGHD1 and IGHD2 occur at similar frequencies, whereas the same applies to IGHD3 and IGHD4. Differences in D gene usages appear to arise from differences in clustering (see next section). BRILIA also returned a small number of inverted D gene annotations, although these occur less frequently than normal D gene annotations. Manual inspection of genes with inverted D annotations revealed that most were inherently ambiguous sequences, and disallowing inverted D annotations did not improve the alignment scores or correlation metrics (data not shown).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>V, D, and J gene usage frequencies</bold>. Frequency distributions of individual VDJ gene families, and VD and DJ pairs as determined by <bold>(A)</bold> the standard method and <bold>(B)</bold> BRILIA.</p></caption>
<graphic xlink:href="fimmu-07-00681-g004.tif"/>
</fig>
</sec>
<sec id="S4-5">
<title>Clustering Results for Real-Life Mice BCR Repertoire Data</title>
<p>B-cell lineage clustering can be used to describe the breadth and extent of affinity maturation and identify promising B-cell clonal lines for further study. Here, a cluster is a set of clonally related BCR sequences, as defined by the standard or BRILIA annotation method. We compared the cluster sizes and counts returned by the standard and BRILIA methods. In Figures <xref ref-type="fig" rid="F5">5</xref>A,B, we compared how many clusters of one method were associated with the cluster of the other method, where an associated cluster shares at least one BCR sequence. We found that typically, a single BRILIA cluster is associated with a given standard cluster (Figure <xref ref-type="fig" rid="F5">5</xref>A), but that the converse is not true&#x02014;multiple standard clusters are often associated with a given BRILIA cluster. These findings suggest that many standard clusters are a subset of BRILIA clusters.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>Comparison of cluster counts and sizes between annotations made using the standard method and BRILIA</bold>. <bold>(A)</bold> Number of BRILIA clusters that are associated (Assoc.) with each standard cluster, where associated clusters share at least one B-cell receptor sequence. The red dots represent clusters whose corresponding lineage trees are shown in Figure <xref ref-type="fig" rid="F6">6</xref>. <bold>(B)</bold> Number of standard clusters that are associated with each BRILIA cluster. <bold>(C)</bold> Largest BRILIA cluster size associated with each standard cluster. The dotted diagonal line (<italic>y</italic>&#x02009;&#x0003D;&#x02009;<italic>x</italic>) highlights differences in the associated cluster sizes between the two methods. <bold>(D)</bold> Largest standard cluster size associated with each BRILIA cluster.</p></caption>
<graphic xlink:href="fimmu-07-00681-g005.tif"/>
</fig>
<p>In Figures <xref ref-type="fig" rid="F5">5</xref>C,D, we compare the differences in cluster sizes from one method&#x02019;s cluster in relation to the other method&#x02019;s largest associated cluster. We found that BRILIA clusters are larger than their associated standard clusters (Figure <xref ref-type="fig" rid="F5">5</xref>C), while standard clusters are generally smaller than their associated BRILIA clusters (Figure <xref ref-type="fig" rid="F5">5</xref>D). In summary, BRILIA clusters are systematically larger than their standard cluster counterparts, and thus, we expect to see more complex lineage trees based on BRILIA annotations.</p>
</sec>
<sec id="S4-6">
<title>BRILIA Preserves Diverse Lineage Trees with High CDR3 Mutations</title>
<p>Different clustering results can ultimately translate to different lineage trees and interpretations of how affinity maturation has progressed within a clonal group. As an example, we compared lineage trees from the case where a large BRILIA cluster was represented by 15 separate standard clusters (red circles in Figure <xref ref-type="fig" rid="F5">5</xref>). Trees were drawn so that each unique BCR sequence was a circle, whose size reflected the sequence template count and whose color corresponded to a unique CDR3 sequence (Figure <xref ref-type="fig" rid="F6">6</xref>). We would expect trees in which clones with high template counts coincide with branch points because highly proliferating B cells are more likely to undergo SHM, generating diverse lineages. The largest tree given by the standard method shows general features of such a tree (Figure <xref ref-type="fig" rid="F6">6</xref>A), although the second standard tree (of size&#x02009;&#x0003D;&#x02009;16) displays an unlikely scenario in which sequences are interlinked without an expanded B-cell clone and with a high CDR3 variability within the small cluster.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Differences in lineage trees and somatic hypermutation (SHM) frequencies between the associated standard and BRILIA clusters from the example in Figure <xref ref-type="fig" rid="F5">5</xref></bold>. <bold>(A)</bold> Lineage trees are assembled from standard clusters that are subsets of a larger associated BRILIA cluster. The <italic>x</italic>-axis shows the absolute SHM distance, where the difference in SHM values between parent and child sequence is the SHM distance between the two sequences. Each dot color corresponds to a unique CDR3 sequence, and the dot size is scaled proportional to the sequence template count relative to the total template count within each lineage tree. The SHM distance is calculated based on the comparison of two 125-nucleotide sequences. Note that six single-member clusters are not drawn. <bold>(B)</bold>&#x02009;Lineage tree of a large BRILIA cluster that encompasses standard clusters. <bold>(C)</bold> Mutation frequencies of the V gene framework and CDR3 predicted by the two methods.</p></caption>
<graphic xlink:href="fimmu-07-00681-g006.tif"/>
</fig>
<p>The corresponding BRILIA tree had leaves ending with low template&#x02013;count clones that usually stemmed from larger clones (Figure <xref ref-type="fig" rid="F6">6</xref>B), and the tree was deeper and wider than the trees returned by the standard method. For the all BCR sequences in this example, BRILIA predicted a higher number of accumulated mutations in the CDR3 than when using the standard method (Figure <xref ref-type="fig" rid="F6">6</xref>C). This representative example illustrates how combining lineage tree assembly and gene annotation can result in substantially larger, richer B-cell lineage trees that are biologically plausible. It also shows how standard annotation methods can systematically underestimate the extent of SHM. While such clonal families make up a small percentage of the overall B cell repertoire, they may play a disproportionately important role in antibody responses to infection because they represent the most affinity-matured members of the repertoire.</p>
</sec>
<sec id="S4-7">
<title>Insights into SHM Mechanisms</title>
<p>Proper mutation annotations can help to validate proposed mechanisms of SHM. Currently, the C&#x02009;&#x02794;&#x02009;T and G&#x02009;&#x02794;&#x02009;A mutation rates can be explained by AID-mediated deamination of C that creates C&#x02009;&#x02794;&#x02009;U mutations, which triggers MSH2/6, polymerase &#x003B7;, and uracil DNA glycosylase to fix U:G mismatches [recently reviewed by Casellas et al. (<xref ref-type="bibr" rid="B88">88</xref>)]. AID recognizes certain 3- or 4-nt long sequences called hot spots (<xref ref-type="bibr" rid="B51">51</xref>, <xref ref-type="bibr" rid="B53">53</xref>, <xref ref-type="bibr" rid="B81">81</xref>, <xref ref-type="bibr" rid="B82">82</xref>); hence, one could identify SHMs caused by AID if the mutations occur at the signature hot spots. However, the A and T mutations occur at different 2-nt long hot spots (<xref ref-type="bibr" rid="B80">80</xref>), suggesting an alternate mechanism of mutation. There are two hotly debated theories of the mechanism underlying A and T mutations (<xref ref-type="bibr" rid="B89">89</xref>). One theory proposes that adenosine deaminase that acts on RNA (ADAR) converts adenosine to inosine, which occurs at a different hot spot and introduces A mutations mostly on the coding DNA strand (<xref ref-type="bibr" rid="B50">50</xref>, <xref ref-type="bibr" rid="B67">67</xref>, <xref ref-type="bibr" rid="B89">89</xref>). The other theory assumes that the AID-induced U:G mismatch triggers multiple DNA repair enzymes to eventually introduce A:T mutations nearby (<xref ref-type="bibr" rid="B51">51</xref>, <xref ref-type="bibr" rid="B90">90</xref>, <xref ref-type="bibr" rid="B91">91</xref>). Past studies identified SHM by comparing the V segment to the predicted germline V gene (<xref ref-type="bibr" rid="B92">92</xref>, <xref ref-type="bibr" rid="B93">93</xref>); in contrast, here, we show SHM across the entire VDJ junction and based on inferred parent&#x02013;child sequence relationships that better reflect the true nucleotide substitution frequencies. We present three findings on the mechanisms underlying SHM based on our analysis of the mice BCR sequencing data.</p>
<sec id="S4-7-1">
<title>Molecular Mechanisms for A Mutations</title>
<p>The mutation frequencies (Figure <xref ref-type="fig" rid="F7">7</xref>A) confirm the frequent C&#x02009;&#x02794;&#x02009;T and G&#x02009;&#x02794;&#x02009;A mutations generated by AID (<xref ref-type="bibr" rid="B50">50</xref>&#x02013;<xref ref-type="bibr" rid="B53">53</xref>, <xref ref-type="bibr" rid="B83">83</xref>), and also the higher A mutations over T mutations that reflect what is known as strand-biased mutations (<xref ref-type="bibr" rid="B67">67</xref>, <xref ref-type="bibr" rid="B68">68</xref>, <xref ref-type="bibr" rid="B94">94</xref>). The mutation frequencies of A follow the trend of G&#x02009;&#x0003E;&#x02009;T&#x02009;&#x0003E;&#x02009;C. To our best knowledge, this A mutation trend has not been previously discussed, and only the A&#x02009;&#x02794;&#x02009;G mutation was proposed to result from an inosine (I) intermediate during ADAR-mediated mutations [i.e., A&#x02009;&#x02794;&#x02009;I&#x02009;&#x02794;&#x02009;G (<xref ref-type="bibr" rid="B80">80</xref>)]. Interestingly, the A mutation trend coincides with that of I:X base-pairing free energy measurements (<xref ref-type="bibr" rid="B95">95</xref>). The Gibbs free energies of I:C, I:A, and I:G base pairs within a short dsDNA segment are &#x02212;8.8, &#x02212;7.5, and &#x02212;6.3&#x02009;kcal/mol, respectively (<xref ref-type="bibr" rid="B95">95</xref>); that is, inosine most closely resembles G, then T, and finally C. The explanation for frequent A mutations over T mutations is still being sought; the transcription of the BCR gene may provide an opportunity for A mutations in one DNA strand (<xref ref-type="bibr" rid="B67">67</xref>, <xref ref-type="bibr" rid="B80">80</xref>, <xref ref-type="bibr" rid="B94">94</xref>, <xref ref-type="bibr" rid="B96">96</xref>, <xref ref-type="bibr" rid="B97">97</xref>). There is a competing hypothesis that A:T mutations are the result of an AID-triggered patchwork repair process around C&#x02009;&#x02794;&#x02009;U mutation sites (<xref ref-type="bibr" rid="B51">51</xref>, <xref ref-type="bibr" rid="B90">90</xref>, <xref ref-type="bibr" rid="B91">91</xref>). In this case, given that C:G mutations would induce A:T mutations, we would expect a correlation between C:G mutations (<italic>CG<sub>mut</sub></italic>) and A:T mutations (<italic>AT<sub>mut</sub></italic>); this does not appear to be the case (Figure <xref ref-type="fig" rid="F7">7</xref>B).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Somatic hypermutation (SHM) frequencies returned by BRILIA, for the purpose of evaluating SHM mechanistic models</bold>. <bold>(A)</bold> Cumulative frequency of SHM propensities for VDJ segments, excluding N regions. X<sub>0</sub> is the parent nt and X<sub>1</sub> is the child nucleotide. <bold>(B)</bold> The [A&#x02009;&#x02794;&#x02009;G&#x02009;&#x0002B;&#x02009;T&#x02009;&#x02794;&#x02009;C] mutation frequency (<italic>AT<sub>mut</sub></italic>) normalized by the total A&#x02009;&#x0002B;&#x02009;T content (<italic>AT<sub>tot</sub></italic>), plotted against the [C&#x02009;&#x02794;&#x02009;T&#x02009;&#x0002B;&#x02009;G&#x02009;&#x02794;&#x02009;A] mutation frequency (<italic>CG<sub>mut</sub></italic>) normalized by the total C&#x02009;&#x0002B;&#x02009;G content in the VDJ segments (<italic>CG<sub>tot</sub></italic>). The dotted red line, which depicts a circle with its center at the origin and a radius of 0.06, marks the mutation rate that captures 90% of the mutated sequences.</p></caption>
<graphic xlink:href="fimmu-07-00681-g007.tif"/>
</fig>
</sec>
<sec id="S4-7-2">
<title>Hot Spot Motifs</title>
<p>AID has been shown to mutate Cs near a WR<underline>C</underline>Y (<xref ref-type="bibr" rid="B82">82</xref>), WG<underline>C</underline>W (<xref ref-type="bibr" rid="B51">51</xref>), WR<underline>C</underline>H (<xref ref-type="bibr" rid="B81">81</xref>), or WR<underline>C</underline> (<xref ref-type="bibr" rid="B53">53</xref>) hot spot motif, where W&#x02009;&#x0003D;&#x02009;A/T, R&#x02009;&#x0003D;&#x02009;A/G, Y&#x02009;&#x0003D;&#x02009;C/T, or H&#x02009;&#x0003D;&#x02009;A/C/T. We evaluated the composition of nts around mutated Cs in our data set and found that it initially agreed with a WG<underline>C</underline>W hot spot for AID (Figure <xref ref-type="fig" rid="F8">8</xref>A); however, we found that for any C, regardless of mutation, the&#x02009;&#x0002B;&#x02009;1 position consistently contained Ws (Figure <xref ref-type="fig" rid="F8">8</xref>B). Hence, the WG<underline>C</underline>W (and potentially WR<underline>C</underline>Y and WR<underline>C</underline>H) hot spots predicted by others may be simply arise from the fact that the &#x0002B;1&#x02009;nt is biased toward a certain nt depending on the host species (<xref ref-type="bibr" rid="B53">53</xref>). Overall, we found that C mutations prefer the WG<underline>C</underline> motif, which is a subset of the WR<underline>C</underline> hot spot predicted by <italic>in vitro</italic> studies of AID mechanisms (<xref ref-type="bibr" rid="B53">53</xref>). The hot spot for G mutations is the complement sequence, <underline>G</underline>CW. Meanwhile, mutations of A have been previously shown to occur at W<underline>A</underline> hot spots (<xref ref-type="bibr" rid="B80">80</xref>). In support of this, our predicted hot spots for A and T mutations are T<underline>A</underline> and <underline>T</underline>A, respectively.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Somatic hypermutation (SHM) hot spot analysis using BRILIA annotations</bold>. <bold>(A)</bold> Evaluation of nucleotide (nt) compositions near only mutated nts, which are at the 0 positions. The negative and positive positions are nts toward the 5&#x02032; and 3&#x02032; sides, respectively, of the 0 position nt. The nt color codes are A&#x02009;&#x0003D;&#x02009;red, C&#x02009;&#x0003D;&#x02009;green, G&#x02009;&#x0003D;&#x02009;blue, and T&#x02009;&#x0003D;&#x02009;gray. <bold>(B)</bold> Evaluation of nt compositions of all nts, regardless of whether they mutated.</p></caption>
<graphic xlink:href="fimmu-07-00681-g008.tif"/>
</fig>
</sec>
<sec id="S4-7-3">
<title>V Gene Mutations as a Proxy for CDR3 Mutations</title>
<p>Past studies often used the V gene mutation rates as a proxy for the CDR3 mutation rates (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>, <xref ref-type="bibr" rid="B98">98</xref>) because resolving the germline D genes was difficult, especially when repertoire-wide sequencing data were unavailable. We investigated the correlation between mutations in the V gene framework (Vframe) and CDR3. Although SHMs are likely unfavorable in the Vframe region since it encodes conserved structural areas of the BCR (<xref ref-type="bibr" rid="B99">99</xref>), we still expected some level of silent mutations to correlate with SHMs in the CDR3. Results from BRILIA and the standard annotation methods show that there is a lack of correlation between SHM rates in the Vframe and CDR3 (Figure <xref ref-type="fig" rid="F9">9</xref>), suggesting that Vframe SHMs are a poor proxy for CDR3 SHMs. Given that the CDR3 region has the highest sequence diversity and typically accommodates the most SHM (<xref ref-type="bibr" rid="B14">14</xref>, <xref ref-type="bibr" rid="B99">99</xref>), these findings demonstrate that it is critical to measure SHM frequency across the entire CDR3 to accurately assess the overall degree of SHM and affinity maturation in B cells.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>Somatic hypermutations (SHM) in CDR3 and V framework regions</bold>. Comparison of the mutations accumulated in the CDR3 versus V framework (Vframe) regions, as determined by <bold>(A)</bold> the standard method and <bold>(B)</bold> BRILIA.</p></caption>
<graphic xlink:href="fimmu-07-00681-g009.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="S5" sec-type="discussion">
<title>Discussion</title>
<p>We presented a novel approach to annotating and analyzing BCR sequencing data that leveraged repertoire-wide B-cell lineage information. By using simulated BCR sequencing data, BRILIA performed substantially better than existing method in annotating the D gene and identifying SHMs. We also showed that the identified SHMs from existing algorithms often provided biologically implausible results, such as inconsistent nt substitution frequencies between the V and the DJ segments. Finally, we applied BRILIA to real-life BCR sequencing data from splenic germinal center B cells of C57BL/6 mice. BRILIA yielded larger, more complex B-cell lineage trees compared to other methods.</p>
<p>Unlike common methods of determining SHM by comparing germline-child sequence pairs, BRILIA calculates SHM frequencies based on inferred parent&#x02013;child relationships across the entire repertoire. The resulting SHM identification provided a more accurate description of SHM patterns, which was used to evaluate hypotheses related to SHM mechanisms, SHM hot spots, and extent of affinity maturation. Our results showed that there was a distinct order to A mutations (G&#x02009;&#x0003E;&#x02009;T&#x02009;&#x0003E;&#x02009;C), which supports the theory of an ADAR-based mutation mechanism <italic>via</italic> an inosine intermediate (<xref ref-type="bibr" rid="B80">80</xref>). Furthermore, we found that the hot spot motif associated with C mutations could most simply be described as WG<underline>C</underline>, which agrees with <italic>in vitro</italic> experiments on AID (<xref ref-type="bibr" rid="B53">53</xref>) and suggests that other more complex hot spots might be the result of intrinsic nucleotide position biases irrespective of mutations. Finally, we showed that SHM frequency in the V gene, a common proxy for overall SHM frequency in many B cell repertoire studies, was a poor predictor of SHM frequency in the highly variable CDR3. These findings highlight the importance of using repertoire-based, full VDJ annotations to evaluate the extent of affinity maturation of B-cell repertoires.</p>
<sec id="S5-1">
<title>BRILIA Helps Separate Real BCR Genes from Those Created by Sequencing Error</title>
<p>A persistent issue with analyzing high-throughput sequencing data is separating real sequences from those created by sequencing error. The ImmuniTree (<xref ref-type="bibr" rid="B46">46</xref>) and IMSEQ (<xref ref-type="bibr" rid="B36">36</xref>) algorithms address this issue, but completely removing sequencing errors is not always feasible. We expect that BCR sequences generated by error will most likely have low template counts and be assigned as &#x0201C;leaves&#x0201D; in the lineage tree. If the goal is to identify real BCR genes, and especially those from clonally expanded B cells, then this can be achieved by looking for sequences with higher-than-background template counts and are designated as lineage tree &#x0201C;nodes.&#x0201D; BRILIA helps identify lineage tree nodes and clonally expanded B cells by outputting the number of descendants associated with each sequence.</p>
</sec>
<sec id="S5-2">
<title>Limitations of BRILIA</title>
<p>The consolidation of lineage trees, clustering, and annotation into a single algorithm makes BRILIA a powerful tool for immunosequencing research. However, limitations also stem from this strength, in that the cluster-based annotation scheme can underperform if the sequences are clustered incorrectly or if the root sequence is not correctly identified. For instance, BRILIA is not fully immune to accidentally grouping clonally unrelated sequences into the same cluster if a path is available or if the cutoff distance is set too large. We are investigating strategies to automatically determine the cutoff point and allow for variable cutoffs among different clusters. In addition, there are certain VDJ recombination events that BRILIA does not account for, including double D insertions [which creates VDDJ junctions (<xref ref-type="bibr" rid="B65">65</xref>)] and lack of D usage (which creates VJ junctions).</p>
<p>If, after the annotation process, multiple VDJ annotations are suggested, then BRILIA removes only pseudogene suggestions. Additional calculations to remove or prioritize the remaining degenerate annotations are not performed as this may bias the repertoire-wide VDJ gene usage frequencies. Processing longer sequences can help to reduce the occurrence of degenerate solutions.</p>
</sec>
<sec id="S5-3">
<title>BRILIA Processing Speed</title>
<p>BRILIA can process large volumes of BCR sequences within a reasonable amount of time, even while determining lineage relationships among B cells. By using a 3.4&#x02009;GHz quad-core processor with 16 GB of memory, BRILIA required 400&#x02009;s to process our repertoire data with 12,300 sequences (or 31&#x02009;ms per each 125-bp sequence). The overall computation time can be further reduced by splitting annotation jobs across more processors.</p>
</sec>
<sec id="S5-4">
<title>BRILIA Input and Output Files</title>
<p>Although we focused on short 125-bp sequences in this study, BRILIA can process longer sequences that extend the full length of the V and J segments, including the CDR1, CDR2, FR1, FR2, and constant regions. If sequences contain the constant region attached to the J segment, BRILIA will trim the constant region. The input files for BRILIA are currently fasta, fastq, csv, xls, and xlsx files containing the sample name and sequence data. To supply the template count data for plotting lineage trees (as shown in Figure <xref ref-type="fig" rid="F6">6</xref>), tabulated data formats are preferred. Datasheets S1&#x02013;S3 Supplementary Material show both example input and output files. BRILIA assumes that the input files contain contiguous sequences and not raw pair-end sequence reads. We recommend performing basic sequence formatting before running BRILIA to ensure most sequences are in the positive sense direction and span the VDJ junction.</p>
</sec>
<sec id="S5-5">
<title>Concluding Remarks</title>
<p>In conclusion, we have demonstrated the ability of BRILIA to predict consistent SHM rates across the VDJ segments and its ability to identify clonally related sequences. These capabilities have wide utility for research related to tracking B-cell affinity maturation in a range of areas of research from infection and vaccination to autoimmune disorders and cancer. BRILIA is a powerful resource for processing and analyzing BCR sequences, and we are currently developing a publicly accessible web-based server for it. The BRILIA source code can be found in GitHub at <uri xlink:href="https://github.com/BHSAI/BRILIA">https://github.com/BHSAI/BRILIA</uri>, and the stand-alone executable version is available on request. Please contact the corresponding author for technical support or further information.</p>
</sec>
</sec>
<sec id="S6" sec-type="author-contributor">
<title>Author Contributions</title>
<p>DL, SC, and IK conceived the project idea, interpreted the results, and wrote the manuscript. SC, CC, SB, and AW acquired project funding. SC and AW oversaw the project. CC performed mouse experiments and reviewed the manuscript. DL developed BRILIA and processed the sequencing data.</p>
</sec>
<sec id="S7">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to thank Sabrina Stronsky and Jackie Benko of USAMRIID for assistance with animal work and cell sorting of GC B cells. The authors thank Dr. Tatsuya Omaya for reviewing and editing this manuscript.</p>
</ack>
<sec id="S8">
<title>Funding</title>
<p>Support for this research was provided by the Military Infectious Diseases Research Program of the United States (US) Army Medical Research and Materiel Command and the US Department of Defense (DoD) High-Performance Computing Modernization Program. This research was in part funded by grants provided to USAMRIID by the US Department of Defense&#x02019;s Defense Threat Reduction Agency (DTRA). The opinions and assertions contained herein are the private views of the authors and are not to be construed as official or as reflecting the views of the US Army or the US DoD. This paper has been approved for public release with unlimited distribution.</p>
</sec>
<sec id="S9" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at <uri xlink:href="http://journal.frontiersin.org/article/10.3389/fimmu.2016.00681/full&#x00023;supplementary-material">http://journal.frontiersin.org/article/10.3389/fimmu.2016.00681/full&#x00023;supplementary-material</uri>.</p>
<supplementary-material xlink:href="Presentation_1.PDF" id="SM1" mimetype="applicationn/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_1.XLSX" id="SM2" mimetype="applicationn/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_2.XLSX" id="SM3" mimetype="applicationn/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_3.XLSX" id="SM4" mimetype="applicationn/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<sec id="S10">
<title>Abbreviations</title>
<p>ADAR, adenosine deaminase acting on RNA; AID, activation-induced deaminase; bp, base pair; BCR, B-cell receptor; BRILIA, B-cell repertoire inductive lineage and immunosequence annotator; CDR, complementarity-determining region; HMM, hidden Markov model; IMGT, ImMunoGeneTics; nt(s), nucleotide(s); SHM, somatic hypermutation; TDT, terminal deoxynucleotidyl transferase.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><label>1</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Victora</surname> <given-names>GD</given-names></name> <name><surname>Nussenzweig</surname> <given-names>MC</given-names></name></person-group>. <article-title>Germinal centers</article-title>. <source>Annu Rev Immunol</source> (<year>2012</year>) <volume>30</volume>:<fpage>429</fpage>&#x02013;<lpage>57</lpage>.<pub-id pub-id-type="doi">10.1146/annurev-immunol-020711-075032</pub-id><pub-id pub-id-type="pmid">22224772</pub-id></citation></ref>
<ref id="B2"><label>2</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>N</given-names></name> <name><surname>He</surname> <given-names>J</given-names></name> <name><surname>Weinstein</surname> <given-names>JA</given-names></name> <name><surname>Penland</surname> <given-names>L</given-names></name> <name><surname>Sasaki</surname> <given-names>S</given-names></name> <name><surname>He</surname> <given-names>XS</given-names></name> <etal/></person-group> <article-title>Lineage structure of the human antibody repertoire in response to influenza vaccination</article-title>. <source>Sci Transl Med</source> (<year>2013</year>) <volume>5</volume>(<issue>171</issue>):<fpage>ra19</fpage>&#x02013;<lpage>171</lpage>.<pub-id pub-id-type="doi">10.1126/scitranslmed.3004794</pub-id><pub-id pub-id-type="pmid">23390249</pub-id></citation></ref>
<ref id="B3"><label>3</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Legutki</surname> <given-names>JB</given-names></name> <name><surname>Johnston</surname> <given-names>SA</given-names></name></person-group>. <article-title>Immunosignatures can predict vaccine efficacy</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2013</year>) <volume>110</volume>(<issue>46</issue>):<fpage>18614</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1309390110</pub-id><pub-id pub-id-type="pmid">24167296</pub-id></citation></ref>
<ref id="B4"><label>4</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galson</surname> <given-names>JD</given-names></name> <name><surname>Clutterbuck</surname> <given-names>EA</given-names></name> <name><surname>Tr&#x000FC;ck</surname> <given-names>J</given-names></name> <name><surname>Ramasamy</surname> <given-names>MN</given-names></name> <name><surname>M&#x000FC;nz</surname> <given-names>M</given-names></name> <name><surname>Fowler</surname> <given-names>A</given-names></name> <etal/></person-group> <article-title>BCR repertoire sequencing: different patterns of B-cell activation after two meningococcal vaccines</article-title>. <source>Immunol Cell Biol</source> (<year>2015</year>) <volume>93</volume>(<issue>10</issue>):<fpage>885</fpage>&#x02013;<lpage>95</lpage>.<pub-id pub-id-type="doi">10.1038/icb.2015.57</pub-id><pub-id pub-id-type="pmid">25976772</pub-id></citation></ref>
<ref id="B5"><label>5</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galson</surname> <given-names>JD</given-names></name> <name><surname>Tr&#x000FC;ck</surname> <given-names>J</given-names></name> <name><surname>Fowler</surname> <given-names>A</given-names></name> <name><surname>Clutterbuck</surname> <given-names>EA</given-names></name> <name><surname>M&#x000FC;nz</surname> <given-names>M</given-names></name> <name><surname>Cerundolo</surname> <given-names>V</given-names></name> <etal/></person-group> <article-title>Analysis of B cell repertoire dynamics following hepatitis b vaccination in humans, and enrichment of vaccine-specific antibody sequences</article-title>. <source>EBioMedicine</source> (<year>2015</year>) <volume>2</volume>(<issue>12</issue>):<fpage>2070</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1016/j.ebiom.2015.11.034</pub-id><pub-id pub-id-type="pmid">26844287</pub-id></citation></ref>
<ref id="B6"><label>6</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tr&#x000FC;ck</surname> <given-names>J</given-names></name> <name><surname>Ramasamy</surname> <given-names>MN</given-names></name> <name><surname>Galson</surname> <given-names>JD</given-names></name> <name><surname>Rance</surname> <given-names>R</given-names></name> <name><surname>Parkhill</surname> <given-names>J</given-names></name> <name><surname>Lunter</surname> <given-names>G</given-names></name> <etal/></person-group> <article-title>Identification of antigen-specific B cell receptor sequences using public repertoire analysis</article-title>. <source>J Immunol</source> (<year>2015</year>) <volume>194</volume>(<issue>1</issue>):<fpage>252</fpage>&#x02013;<lpage>61</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1401405</pub-id><pub-id pub-id-type="pmid">25392534</pub-id></citation></ref>
<ref id="B7"><label>7</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robins</surname> <given-names>H</given-names></name></person-group>. <article-title>Immunosequencing: applications of immune repertoire deep sequencing</article-title>. <source>Curr Opin Immunol</source> (<year>2013</year>) <volume>25</volume>(<issue>5</issue>):<fpage>646</fpage>&#x02013;<lpage>52</lpage>.<pub-id pub-id-type="doi">10.1016/j.coi.2013.09.017</pub-id><pub-id pub-id-type="pmid">24140071</pub-id></citation></ref>
<ref id="B8"><label>8</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yaari</surname> <given-names>G</given-names></name> <name><surname>Kleinstein</surname> <given-names>SH</given-names></name></person-group>. <article-title>Practical guidelines for B-cell receptor repertoire sequencing analysis</article-title>. <source>Genome Med</source> (<year>2015</year>) <volume>7</volume>(<issue>1</issue>):<fpage>1</fpage>&#x02013;<lpage>14</lpage>.<pub-id pub-id-type="doi">10.1186/s13073-015-0243-2</pub-id><pub-id pub-id-type="pmid">26589402</pub-id></citation></ref>
<ref id="B9"><label>9</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kepler</surname> <given-names>TB</given-names></name></person-group>. <article-title>Reconstructing a B-cell clonal lineage. I. Statistical inference of unobserved ancestors</article-title>. <source>F1000Res</source> (<year>2013</year>) <volume>2</volume>:<fpage>103</fpage>.<pub-id pub-id-type="doi">10.12688/f1000research.2-103.v1</pub-id></citation></ref>
<ref id="B10"><label>10</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gupta</surname> <given-names>NT</given-names></name> <name><surname>Vander Heiden</surname> <given-names>JA</given-names></name> <name><surname>Uduman</surname> <given-names>M</given-names></name> <name><surname>Gadala-Maria</surname> <given-names>D</given-names></name> <name><surname>Yaari</surname> <given-names>G</given-names></name> <name><surname>Kleinstein</surname> <given-names>SH</given-names></name></person-group>. <article-title>Change-O: a toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data</article-title>. <source>Bioinformatics</source> (<year>2015</year>) <volume>31</volume>(<issue>20</issue>):<fpage>3356</fpage>&#x02013;<lpage>8</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btv359</pub-id><pub-id pub-id-type="pmid">26069265</pub-id></citation></ref>
<ref id="B11"><label>11</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stern</surname> <given-names>JN</given-names></name> <name><surname>Yaari</surname> <given-names>G</given-names></name> <name><surname>Vander Heiden</surname> <given-names>JA</given-names></name> <name><surname>Church</surname> <given-names>G</given-names></name> <name><surname>Donahue</surname> <given-names>WF</given-names></name> <name><surname>Hintzen</surname> <given-names>RQ</given-names></name> <etal/></person-group> <article-title>B cells populating the multiple sclerosis brain mature in the draining cervical lymph nodes</article-title>. <source>Sci Transl Med</source> (<year>2014</year>) <volume>6</volume>(<issue>248</issue>):<fpage>ra107</fpage>&#x02013;<lpage>248</lpage>.<pub-id pub-id-type="doi">10.1126/scitranslmed.3008879</pub-id><pub-id pub-id-type="pmid">25100741</pub-id></citation></ref>
<ref id="B12"><label>12</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schatz</surname> <given-names>DG</given-names></name> <name><surname>Oettinger</surname> <given-names>MA</given-names></name> <name><surname>Baltimore</surname> <given-names>D</given-names></name></person-group>. <article-title>The V(D)J recombination activating gene, <italic>Rag-1</italic></article-title>. <source>Cell</source> (<year>1989</year>) <volume>59</volume>(<issue>6</issue>):<fpage>1035</fpage>&#x02013;<lpage>48</lpage>.<pub-id pub-id-type="doi">10.1016/0092-8674(89)90760-5</pub-id></citation></ref>
<ref id="B13"><label>13</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oettinger</surname> <given-names>MA</given-names></name> <name><surname>Schatz</surname> <given-names>DG</given-names></name> <name><surname>Gorka</surname> <given-names>C</given-names></name> <name><surname>Baltimore</surname> <given-names>D</given-names></name></person-group>. <article-title>RAG-1 and RAG-2, adjacent genes that synergistically activate V(D)J recombination</article-title>. <source>Science</source> (<year>1990</year>) <volume>248</volume>(<issue>4962</issue>):<fpage>1517</fpage>.<pub-id pub-id-type="doi">10.1126/science.2360047</pub-id><pub-id pub-id-type="pmid">2360047</pub-id></citation></ref>
<ref id="B14"><label>14</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>JL</given-names></name> <name><surname>Davis</surname> <given-names>MM</given-names></name></person-group>. <article-title>Diversity in the CDR3 region of VH is sufficient for most antibody specificities</article-title>. <source>Immunity</source> (<year>2000</year>) <volume>13</volume>(<issue>1</issue>):<fpage>37</fpage>&#x02013;<lpage>45</lpage>.<pub-id pub-id-type="doi">10.1016/S1074-7613(00)00006-6</pub-id></citation></ref>
<ref id="B15"><label>15</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Y</given-names></name> <name><surname>Pannicke</surname> <given-names>U</given-names></name> <name><surname>Schwarz</surname> <given-names>K</given-names></name> <name><surname>Lieber</surname> <given-names>MR</given-names></name></person-group>. <article-title>Hairpin opening and overhang processing by an artemis/dna-dependent protein kinase complex in nonhomologous end joining and V(D)J recombination</article-title>. <source>Cell</source> (<year>2002</year>) <volume>108</volume>(<issue>6</issue>):<fpage>781</fpage>&#x02013;<lpage>94</lpage>.<pub-id pub-id-type="doi">10.1016/S0092-8674(02)00671-2</pub-id><pub-id pub-id-type="pmid">11955432</pub-id></citation></ref>
<ref id="B16"><label>16</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Komori</surname> <given-names>T</given-names></name> <name><surname>Okada</surname> <given-names>A</given-names></name> <name><surname>Stewart</surname> <given-names>V</given-names></name> <name><surname>Alt</surname> <given-names>FW</given-names></name></person-group>. <article-title>Lack of N regions in antigen receptor variable region genes of TdT-deficient lymphocytes</article-title>. <source>Science</source> (<year>1993</year>) <volume>261</volume>(<issue>5125</issue>):<fpage>1171</fpage>&#x02013;<lpage>5</lpage>.<pub-id pub-id-type="doi">10.1126/science.8356451</pub-id><pub-id pub-id-type="pmid">8356451</pub-id></citation></ref>
<ref id="B17"><label>17</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilfillan</surname> <given-names>S</given-names></name> <name><surname>Dierich</surname> <given-names>A</given-names></name> <name><surname>Lemeur</surname> <given-names>M</given-names></name> <name><surname>Benoist</surname> <given-names>C</given-names></name> <name><surname>Mathis</surname> <given-names>D</given-names></name></person-group>. <article-title>Mice lacking TdT: mature animals with an immature lymphocyte repertoire</article-title>. <source>Science</source> (<year>1993</year>) <volume>261</volume>(<issue>5125</issue>):<fpage>1175</fpage>&#x02013;<lpage>8</lpage>.<pub-id pub-id-type="doi">10.1126/science.8356452</pub-id><pub-id pub-id-type="pmid">8356452</pub-id></citation></ref>
<ref id="B18"><label>18</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Motea</surname> <given-names>EA</given-names></name> <name><surname>Berdis</surname> <given-names>AJ</given-names></name></person-group>. <article-title>Terminal deoxynucleotidyl transferase: the story of a misguided DNA polymerase</article-title>. <source>Biochim Biophys Acta</source> (<year>2010</year>) <volume>1804</volume>(<issue>5</issue>):<fpage>1151</fpage>&#x02013;<lpage>66</lpage>.<pub-id pub-id-type="doi">10.1016/j.bbapap.2009.06.030</pub-id><pub-id pub-id-type="pmid">19596089</pub-id></citation></ref>
<ref id="B19"><label>19</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schatz</surname> <given-names>DG</given-names></name> <name><surname>Ji</surname> <given-names>Y</given-names></name></person-group>. <article-title>Recombination centres and the orchestration of V(D)J recombination</article-title>. <source>Nat Rev Immunol</source> (<year>2011</year>) <volume>11</volume>(<issue>4</issue>):<fpage>251</fpage>&#x02013;<lpage>63</lpage>.<pub-id pub-id-type="doi">10.1038/nri2941</pub-id><pub-id pub-id-type="pmid">21394103</pub-id></citation></ref>
<ref id="B20"><label>20</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noia</surname> <given-names>JMD</given-names></name> <name><surname>Neuberger</surname> <given-names>MS</given-names></name></person-group>. <article-title>Molecular mechanisms of antibody somatic hypermutation</article-title>. <source>Annu Rev Biochem</source> (<year>2007</year>) <volume>76</volume>(<issue>1</issue>):<fpage>1</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1146/annurev.biochem.76.061705.090740</pub-id><pub-id pub-id-type="pmid">17328676</pub-id></citation></ref>
<ref id="B21"><label>21</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeKosky</surname> <given-names>BJ</given-names></name> <name><surname>Ippolito</surname> <given-names>GC</given-names></name> <name><surname>Deschner</surname> <given-names>RP</given-names></name> <name><surname>Lavinder</surname> <given-names>JJ</given-names></name> <name><surname>Wine</surname> <given-names>Y</given-names></name> <name><surname>Rawlings</surname> <given-names>BM</given-names></name> <etal/></person-group> <article-title>High-throughput sequencing of the paired human immunoglobulin heavy and light chain repertoire</article-title>. <source>Nat Biotechnol</source> (<year>2013</year>) <volume>31</volume>(<issue>2</issue>):<fpage>166</fpage>.<pub-id pub-id-type="doi">10.1038/nbt.2492</pub-id><pub-id pub-id-type="pmid">23334449</pub-id></citation></ref>
<ref id="B22"><label>22</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeWitt</surname> <given-names>WS</given-names></name> <name><surname>Lindau</surname> <given-names>P</given-names></name> <name><surname>Snyder</surname> <given-names>TM</given-names></name> <name><surname>Sherwood</surname> <given-names>AM</given-names></name> <name><surname>Vignali</surname> <given-names>M</given-names></name> <name><surname>Carlson</surname> <given-names>CS</given-names></name> <etal/></person-group> <article-title>A public database of memory and naive B-Cell receptor sequences</article-title>. <source>PLoS One</source> (<year>2016</year>) <volume>11</volume>(<issue>8</issue>):<fpage>e0160853</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0160853</pub-id><pub-id pub-id-type="pmid">27513338</pub-id></citation></ref>
<ref id="B23"><label>23</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larimore</surname> <given-names>K</given-names></name> <name><surname>McCormick</surname> <given-names>MW</given-names></name> <name><surname>Robins</surname> <given-names>HS</given-names></name> <name><surname>Greenberg</surname> <given-names>PD</given-names></name></person-group>. <article-title>Shaping of human germline IgH repertoires revealed by deep sequencing</article-title>. <source>J Immunol</source> (<year>2012</year>) <volume>189</volume>(<issue>6</issue>):<fpage>3221</fpage>&#x02013;<lpage>30</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1201303</pub-id><pub-id pub-id-type="pmid">22865917</pub-id></citation></ref>
<ref id="B24"><label>24</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>D</given-names></name> <name><surname>Chen</surname> <given-names>C</given-names></name> <name><surname>Seely</surname> <given-names>EJ</given-names></name> <name><surname>Chen</surname> <given-names>S</given-names></name> <name><surname>Song</surname> <given-names>Y</given-names></name></person-group>. <article-title>High-throughput sequencing-based immune repertoire study during infectious disease</article-title>. <source>Front Immunol</source> (<year>2016</year>) <volume>7</volume>:<fpage>336</fpage>.<pub-id pub-id-type="doi">10.3389/fimmu.2016.00336</pub-id><pub-id pub-id-type="pmid">27630639</pub-id></citation></ref>
<ref id="B25"><label>25</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Chaume</surname> <given-names>D</given-names></name> <name><surname>Lefranc</surname> <given-names>M-P</given-names></name></person-group>. <article-title>IMGT/V-QUEST, an integrated software program for immunoglobulin and T cell receptor V&#x02013;J and V&#x02013;D&#x02013;J rearrangement analysis</article-title>. <source>Nucleic Acids Res</source> (<year>2004</year>) <volume>32</volume>(<issue>Web Server issue</issue>):<fpage>W435</fpage>&#x02013;<lpage>40</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkh412</pub-id></citation></ref>
<ref id="B26"><label>26</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brochet</surname> <given-names>X</given-names></name> <name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name></person-group>. <article-title>IMGT/V-QUEST: the highly customized and integrated system for IG and TR standardized V-J and V-D-J sequence analysis</article-title>. <source>Nucleic Acids Res</source> (<year>2008</year>) <volume>36</volume>(<issue>Suppl 2</issue>):<fpage>W503</fpage>&#x02013;<lpage>8</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkn316</pub-id><pub-id pub-id-type="pmid">18503082</pub-id></citation></ref>
<ref id="B27"><label>27</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S</given-names></name> <name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Miles</surname> <given-names>JJ</given-names></name> <name><surname>Alamyar</surname> <given-names>E</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Duroux</surname> <given-names>P</given-names></name> <etal/></person-group> <article-title>IMGT/HighV QUEST paradigm for T cell receptor IMGT clonotype diversity and next generation repertoire immunoprofiling</article-title>. <source>Nat Commun</source> (<year>2013</year>) <volume>4</volume>:<fpage>2333</fpage>.<pub-id pub-id-type="doi">10.1038/ncomms3333</pub-id><pub-id pub-id-type="pmid">23995877</pub-id></citation></ref>
<ref id="B28"><label>28</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yousfi Monod</surname> <given-names>M</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Chaume</surname> <given-names>D</given-names></name> <name><surname>Lefranc</surname> <given-names>M-P</given-names></name></person-group>. <article-title>IMGT/JunctionAnalysis: the first tool for the analysis of the immunoglobulin and T cell receptor complex V&#x02013;J and V&#x02013;D&#x02013;J JUNCTIONs</article-title>. <source>Bioinformatics</source> (<year>2004</year>) <volume>20</volume>(<issue>Suppl 1</issue>):<fpage>i379</fpage>&#x02013;<lpage>85</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/bth945</pub-id></citation></ref>
<ref id="B29"><label>29</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Pommi&#x000E9;</surname> <given-names>C</given-names></name> <name><surname>Kaas</surname> <given-names>Q</given-names></name> <name><surname>Duprat</surname> <given-names>E</given-names></name> <name><surname>Bosc</surname> <given-names>N</given-names></name> <name><surname>Guiraudou</surname> <given-names>D</given-names></name> <etal/></person-group> <article-title>IMGT unique numbering for immunoglobulin and T cell receptor variable domains and Ig superfamily V-like domains</article-title>. <source>Dev Comp Immunol</source> (<year>2003</year>) <volume>27</volume>(<issue>1</issue>):<fpage>55</fpage>&#x02013;<lpage>77</lpage>.<pub-id pub-id-type="doi">10.1016/S0145-305X(02)00039-3</pub-id></citation></ref>
<ref id="B30"><label>30</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ye</surname> <given-names>J</given-names></name> <name><surname>Ma</surname> <given-names>N</given-names></name> <name><surname>Madden</surname> <given-names>TL</given-names></name> <name><surname>Ostell</surname> <given-names>JM</given-names></name></person-group>. <article-title>IgBLAST: an immunoglobulin variable domain sequence analysis tool</article-title>. <source>Nucleic Acids Res</source> (<year>2013</year>) <volume>41</volume>(<issue>Web Server issue</issue>):<fpage>W34</fpage>&#x02013;<lpage>40</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkt382</pub-id><pub-id pub-id-type="pmid">23671333</pub-id></citation></ref>
<ref id="B31"><label>31</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Volpe</surname> <given-names>JM</given-names></name> <name><surname>Cowell</surname> <given-names>LG</given-names></name> <name><surname>Kepler</surname> <given-names>TB</given-names></name></person-group>. <article-title>SoDA: implementation of a 3D alignment algorithm for inference of antigen receptor recombinations</article-title>. <source>Bioinformatics</source> (<year>2006</year>) <volume>22</volume>(<issue>4</issue>):<fpage>438</fpage>&#x02013;<lpage>44</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btk004</pub-id><pub-id pub-id-type="pmid">16357034</pub-id></citation></ref>
<ref id="B32"><label>32</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohm-Laursen</surname> <given-names>L</given-names></name> <name><surname>Nielsen</surname> <given-names>M</given-names></name> <name><surname>Larsen</surname> <given-names>SR</given-names></name> <name><surname>Barington</surname> <given-names>T</given-names></name></person-group>. <article-title>No evidence for the use of DIR, D&#x02013;D fusions, chromosome 15 open reading frames or VHreplacement in the peripheral repertoire was found on application of an improved algorithm, JointML, to 6329 human immunoglobulin H rearrangements</article-title>. <source>Immunology</source> (<year>2006</year>) <volume>119</volume>(<issue>2</issue>):<fpage>265</fpage>&#x02013;<lpage>77</lpage>.<pub-id pub-id-type="doi">10.1111/j.1365-2567.2006.02431.x</pub-id></citation></ref>
<ref id="B33"><label>33</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Souto-Carneiro</surname> <given-names>MM</given-names></name> <name><surname>Longo</surname> <given-names>NS</given-names></name> <name><surname>Russ</surname> <given-names>DE</given-names></name> <name><surname>Sun</surname> <given-names>HW</given-names></name> <name><surname>Lipsky</surname> <given-names>PE</given-names></name></person-group>. <article-title>Characterization of the human Ig heavy chain antigen binding complementarity determining region 3 using a newly developed software algorithm, JOINSOLVER</article-title>. <source>J Immunol</source> (<year>2004</year>) <volume>172</volume>(<issue>11</issue>):<fpage>6790</fpage>&#x02013;<lpage>802</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.172.11.6790</pub-id></citation></ref>
<ref id="B34"><label>34</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paciello</surname> <given-names>G</given-names></name> <name><surname>Acquaviva</surname> <given-names>A</given-names></name> <name><surname>Pighi</surname> <given-names>C</given-names></name> <name><surname>Ferrarini</surname> <given-names>A</given-names></name> <name><surname>Macii</surname> <given-names>E</given-names></name> <name><surname>Zamo&#x02019;</surname> <given-names>A</given-names></name> <etal/></person-group> <article-title>VDJSeq-solver: in silico V(D)J recombination detection tool</article-title>. <source>PLoS One</source> (<year>2015</year>) <volume>10</volume>(<issue>3</issue>):<fpage>e0118192</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0118192</pub-id><pub-id pub-id-type="pmid">25799103</pub-id></citation></ref>
<ref id="B35"><label>35</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolotin</surname> <given-names>DA</given-names></name> <name><surname>Poslavsky</surname> <given-names>S</given-names></name> <name><surname>Mitrophanov</surname> <given-names>I</given-names></name> <name><surname>Shugay</surname> <given-names>M</given-names></name> <name><surname>Mamedov</surname> <given-names>IZ</given-names></name> <name><surname>Putintseva</surname> <given-names>EV</given-names></name> <etal/></person-group> <article-title>MiXCR: software for comprehensive adaptive immunity profiling</article-title>. <source>Nat Methods</source> (<year>2015</year>) <volume>12</volume>(<issue>5</issue>):<fpage>380</fpage>&#x02013;<lpage>1</lpage>.<pub-id pub-id-type="doi">10.1038/nmeth.3364</pub-id></citation></ref>
<ref id="B36"><label>36</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuchenbecker</surname> <given-names>L</given-names></name> <name><surname>Nienen</surname> <given-names>M</given-names></name> <name><surname>Hecht</surname> <given-names>J</given-names></name> <name><surname>Neumann</surname> <given-names>AU</given-names></name> <name><surname>Babel</surname> <given-names>N</given-names></name> <name><surname>Reinert</surname> <given-names>K</given-names></name> <etal/></person-group> <article-title>IMSEQ &#x02013; a fast and error aware approach to immunogenetic sequence analysis</article-title>. <source>Bioinformatics</source> (<year>2015</year>) <volume>31</volume>(<issue>18</issue>):<fpage>2963</fpage>&#x02013;<lpage>71</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btv309</pub-id></citation></ref>
<ref id="B37"><label>37</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frost</surname> <given-names>SD</given-names></name> <name><surname>Murrell</surname> <given-names>B</given-names></name> <name><surname>Hossain</surname> <given-names>AS</given-names></name> <name><surname>Silverman</surname> <given-names>GJ</given-names></name> <name><surname>Pond</surname> <given-names>SL</given-names></name></person-group>. <article-title>Assigning and visualizing germline genes in antibody repertoires</article-title>. <source>Philos Trans R Soc Lond B Biol Sci</source> (<year>2015</year>) <volume>370</volume>(<issue>1676</issue>):<fpage>20140240</fpage>.<pub-id pub-id-type="doi">10.1098/rstb.2014.0240</pub-id><pub-id pub-id-type="pmid">26194754</pub-id></citation></ref>
<ref id="B38"><label>38</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Munshaw</surname> <given-names>S</given-names></name> <name><surname>Kepler</surname> <given-names>TB</given-names></name></person-group>. <article-title>SoDA2: a hidden Markov model approach for identification of immunoglobulin rearrangements</article-title>. <source>Bioinformatics</source> (<year>2010</year>) <volume>26</volume>(<issue>7</issue>):<fpage>867</fpage>&#x02013;<lpage>72</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btq056</pub-id><pub-id pub-id-type="pmid">20147303</pub-id></citation></ref>
<ref id="B39"><label>39</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ga&#x000EB;ta</surname> <given-names>BA</given-names></name> <name><surname>Malming</surname> <given-names>HR</given-names></name> <name><surname>Jackson</surname> <given-names>KJ</given-names></name> <name><surname>Bain</surname> <given-names>ME</given-names></name> <name><surname>Wilson</surname> <given-names>P</given-names></name> <name><surname>Collins</surname> <given-names>AM</given-names></name></person-group>. <article-title>iHMMune-align: hidden Markov model-based alignment and identification of germline genes in rearranged immunoglobulin gene sequences</article-title>. <source>Bioinformatics</source> (<year>2007</year>) <volume>23</volume>(<issue>13</issue>):<fpage>1580</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btm147</pub-id><pub-id pub-id-type="pmid">17463026</pub-id></citation></ref>
<ref id="B40"><label>40</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ralph</surname> <given-names>DK</given-names></name> <name><surname>Matsen</surname> <given-names>FA</given-names> <suffix>IV</suffix></name></person-group>. <article-title>Consistency of VDJ rearrangement and substitution parameters enables accurate B cell receptor sequence annotation</article-title>. <source>PLoS Comput Biol</source> (<year>2016</year>) <volume>12</volume>(<issue>1</issue>): <fpage>e1004409</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pcbi.1004409</pub-id><pub-id pub-id-type="pmid">26751373</pub-id></citation></ref>
<ref id="B41"><label>41</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collins</surname> <given-names>AM</given-names></name> <name><surname>Wang</surname> <given-names>Y</given-names></name> <name><surname>Roskin</surname> <given-names>KM</given-names></name> <name><surname>Marquis</surname> <given-names>CP</given-names></name> <name><surname>Jackson</surname> <given-names>KJ</given-names></name></person-group>. <article-title>The mouse antibody heavy chain repertoire is germline-focused and highly variable between inbred strains</article-title>. <source>Philos Trans R Soc Lond B Biol Sci</source> (<year>2015</year>) <volume>370</volume>(<issue>1676</issue>):<fpage>20140236</fpage>.<pub-id pub-id-type="doi">10.1098/rstb.2014.0236</pub-id><pub-id pub-id-type="pmid">26194750</pub-id></citation></ref>
<ref id="B42"><label>42</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liberman</surname> <given-names>G</given-names></name> <name><surname>Benichou</surname> <given-names>J</given-names></name> <name><surname>Tsaban</surname> <given-names>L</given-names></name> <name><surname>Glanville</surname> <given-names>J</given-names></name> <name><surname>Louzoun</surname> <given-names>Y</given-names></name></person-group>. <article-title>Multi step selection in Ig H chains is initially focused on CDR3 and then on other CDR regions</article-title>. <source>Front Immunol</source> (<year>2013</year>) <volume>4</volume>:<fpage>274</fpage>.<pub-id pub-id-type="doi">10.3389/fimmu.2013.00274</pub-id><pub-id pub-id-type="pmid">24062742</pub-id></citation></ref>
<ref id="B43"><label>43</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giraud</surname> <given-names>M</given-names></name> <name><surname>Salson</surname> <given-names>M</given-names></name> <name><surname>Duez</surname> <given-names>M</given-names></name> <name><surname>Villenet</surname> <given-names>C</given-names></name> <name><surname>Quief</surname> <given-names>S</given-names></name> <name><surname>Caillault</surname> <given-names>A</given-names></name> <etal/></person-group> <article-title>Fast multiclonal clusterization of V (D) J recombinations from high-throughput sequencing</article-title>. <source>BMC Genomics</source> (<year>2014</year>) <volume>15</volume>(<issue>1</issue>):<fpage>409</fpage>.<pub-id pub-id-type="doi">10.1186/1471-2164-15-409</pub-id><pub-id pub-id-type="pmid">24885090</pub-id></citation></ref>
<ref id="B44"><label>44</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tamura</surname> <given-names>K</given-names></name> <name><surname>Peterson</surname> <given-names>D</given-names></name> <name><surname>Peterson</surname> <given-names>N</given-names></name> <name><surname>Stecher</surname> <given-names>G</given-names></name> <name><surname>Nei</surname> <given-names>M</given-names></name> <name><surname>Kumar</surname> <given-names>S</given-names></name></person-group>. <article-title>MEGA5: molecular evolutionary genetics analysis using maximum likelihood, evolutionary distance, and maximum parsimony methods</article-title>. <source>Mol Biol Evol</source> (<year>2011</year>) <volume>28</volume>(<issue>10</issue>):<fpage>2731</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1093/molbev/msr121</pub-id><pub-id pub-id-type="pmid">21546353</pub-id></citation></ref>
<ref id="B45"><label>45</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Felesenstein</surname> <given-names>J</given-names></name></person-group>. <article-title>PHYLIP-phylogeny inference package (version 3.2)</article-title>. <source>Cladistics</source> (<year>1989</year>) <volume>5</volume>:<fpage>163</fpage>&#x02013;<lpage>6</lpage>.</citation></ref>
<ref id="B46"><label>46</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sok</surname> <given-names>D</given-names></name> <name><surname>Laserson</surname> <given-names>U</given-names></name> <name><surname>Laserson</surname> <given-names>J</given-names></name> <name><surname>Liu</surname> <given-names>Y</given-names></name> <name><surname>Vigneault</surname> <given-names>F</given-names></name> <name><surname>Julien</surname> <given-names>JP</given-names></name> <etal/></person-group> <article-title>The effects of somatic hypermutation on neutralization and binding in the PGT121 family of broadly neutralizing HIV antibodies</article-title>. <source>PLoS Pathog</source> (<year>2013</year>) <volume>9</volume>(<issue>11</issue>):<fpage>e1003754</fpage>.<pub-id pub-id-type="doi">10.1371/journal.ppat.1003754</pub-id><pub-id pub-id-type="pmid">24278016</pub-id></citation></ref>
<ref id="B47"><label>47</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barak</surname> <given-names>M</given-names></name> <name><surname>Zuckerman</surname> <given-names>NS</given-names></name> <name><surname>Edelman</surname> <given-names>H</given-names></name> <name><surname>Unger</surname> <given-names>R</given-names></name> <name><surname>Mehr</surname> <given-names>R</given-names></name></person-group>. <article-title>IgTree&#x000A9;: creating immunoglobulin variable region gene lineage trees</article-title>. <source>J Immunol Methods</source> (<year>2008</year>) <volume>338</volume>(<issue>1&#x02013;2</issue>):<fpage>67</fpage>&#x02013;<lpage>74</lpage>.<pub-id pub-id-type="doi">10.1016/j.jim.2008.06.006</pub-id></citation></ref>
<ref id="B48"><label>48</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lees</surname> <given-names>WD</given-names></name> <name><surname>Shepherd</surname> <given-names>AJ</given-names></name></person-group>. <article-title>Utilities for high-throughput analysis of B-cell clonal lineages</article-title>. <source>J Immunol Res</source> (<year>2015</year>) <volume>2015</volume>:<fpage>323506</fpage>.<pub-id pub-id-type="doi">10.1155/2015/323506</pub-id><pub-id pub-id-type="pmid">26527585</pub-id></citation></ref>
<ref id="B49"><label>49</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Safonova</surname> <given-names>Y</given-names></name> <name><surname>Lapidus</surname> <given-names>A</given-names></name> <name><surname>Lill</surname> <given-names>J</given-names></name></person-group>. <article-title>IgSimulator: a versatile immunosequencing simulator</article-title>. <source>Bioinformatics</source> (<year>2015</year>) <volume>31</volume>(<issue>19</issue>):<fpage>3213</fpage>&#x02013;<lpage>5</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btv326</pub-id><pub-id pub-id-type="pmid">26007226</pub-id></citation></ref>
<ref id="B50"><label>50</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Longerich</surname> <given-names>S</given-names></name> <name><surname>Basu</surname> <given-names>U</given-names></name> <name><surname>Alt</surname> <given-names>F</given-names></name> <name><surname>Storb</surname> <given-names>U</given-names></name></person-group>. <article-title>AID in somatic hypermutation and class switch recombination [Current Opinion in Immunology 2006, 18:164&#x02013;174]</article-title>. <source>Curr Opin Immunol</source> (<year>2006</year>) <volume>18</volume>(<issue>6</issue>):<fpage>769</fpage>.<pub-id pub-id-type="doi">10.1016/j.coi.2006.09.017</pub-id></citation></ref>
<ref id="B51"><label>51</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maul</surname> <given-names>RW</given-names></name> <name><surname>Gearhart</surname> <given-names>PJ</given-names></name></person-group>. <article-title>Controlling somatic hypermutation in immunoglobulin variable and switch regions</article-title>. <source>Immunol Res</source> (<year>2010</year>) <volume>47</volume>(<issue>1</issue>):<fpage>113</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1007/s12026-009-8142-5</pub-id><pub-id pub-id-type="pmid">20082153</pub-id></citation></ref>
<ref id="B52"><label>52</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casali</surname> <given-names>P</given-names></name> <name><surname>Pal</surname> <given-names>Z</given-names></name> <name><surname>Xu</surname> <given-names>Z</given-names></name> <name><surname>Zan</surname> <given-names>H</given-names></name></person-group>. <article-title>DNA repair in antibody somatic hypermutation</article-title>. <source>Trends Immunol</source> (<year>2006</year>) <volume>27</volume>(<issue>7</issue>):<fpage>313</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1016/j.it.2006.05.001</pub-id></citation></ref>
<ref id="B53"><label>53</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pham</surname> <given-names>P</given-names></name> <name><surname>Bransteitter</surname> <given-names>R</given-names></name> <name><surname>Petruska</surname> <given-names>J</given-names></name> <name><surname>Goodman</surname> <given-names>MF</given-names></name></person-group>. <article-title>Processive AID-catalysed cytosine deamination on single-stranded DNA simulates somatic hypermutation</article-title>. <source>Nature</source> (<year>2003</year>) <volume>424</volume>(<issue>6944</issue>):<fpage>103</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1038/nature01760</pub-id><pub-id pub-id-type="pmid">12819663</pub-id></citation></ref>
<ref id="B54"><label>54</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name></person-group>. <article-title>IMGT, the international ImMunoGeneTics database</article-title>. <source>Nucleic Acids Res</source> (<year>2001</year>) <volume>29</volume>(<issue>1</issue>):<fpage>207</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1093/nar/29.1.207</pub-id><pub-id pub-id-type="pmid">11125093</pub-id></citation></ref>
<ref id="B55"><label>55</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name></person-group>. <article-title>IMGT, the international ImMunoGeneTics database<sup>&#x000AE;</sup></article-title>. <source>Nucleic Acids Res</source> (<year>2003</year>) <volume>31</volume>(<issue>1</issue>):<fpage>307</fpage>&#x02013;<lpage>10</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkg085</pub-id></citation></ref>
<ref id="B56"><label>56</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Ginestoux</surname> <given-names>C</given-names></name> <name><surname>Bodmer</surname> <given-names>J</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>W</given-names></name> <name><surname>Bontrop</surname> <given-names>R</given-names></name> <etal/></person-group> <article-title>IMGT, the international ImMunoGeneTics database</article-title>. <source>Nucleic Acids Res</source> (<year>1999</year>) <volume>27</volume>(<issue>1</issue>):<fpage>209</fpage>&#x02013;<lpage>12</lpage>.<pub-id pub-id-type="doi">10.1093/nar/27.1.209</pub-id><pub-id pub-id-type="pmid">9847182</pub-id></citation></ref>
<ref id="B57"><label>57</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Clement</surname> <given-names>O</given-names></name> <name><surname>Kaas</surname> <given-names>Q</given-names></name> <name><surname>Duprat</surname> <given-names>E</given-names></name> <name><surname>Chastellan</surname> <given-names>P</given-names></name> <name><surname>Coelho</surname> <given-names>I</given-names></name> <etal/></person-group> <article-title>IMGT-Choreography for immunogenetics and immunoinformatics</article-title>. <source>In Silico Biol</source> (<year>2005</year>) <volume>5</volume>(<issue>1</issue>):<fpage>45</fpage>&#x02013;<lpage>60</lpage>.<pub-id pub-id-type="pmid">15972004</pub-id></citation></ref>
<ref id="B58"><label>58</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Duroux</surname> <given-names>P</given-names></name> <name><surname>Jabado-Michaloud</surname> <given-names>J</given-names></name> <name><surname>Folch</surname> <given-names>G</given-names></name> <name><surname>Aouinti</surname> <given-names>S</given-names></name> <etal/></person-group> <article-title>IMGT<sup>&#x000AE;</sup>, the international ImMunoGeneTics information system<sup>&#x000AE;</sup> 25 years on</article-title>. <source>Nucleic Acids Res</source> (<year>2015</year>) <volume>43</volume>(<issue>D1</issue>):<fpage>D413</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gku1056</pub-id></citation></ref>
<ref id="B59"><label>59</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Ginestoux</surname> <given-names>C</given-names></name> <name><surname>Jabado-Michaloud</surname> <given-names>J</given-names></name> <name><surname>Folch</surname> <given-names>G</given-names></name> <name><surname>Bellahcene</surname> <given-names>F</given-names></name> <etal/></person-group> <article-title>IMGT<sup>&#x000AE;</sup>, the international ImMunoGeneTics information system<sup>&#x000AE;</sup></article-title>. <source>Nucleic Acids Res</source> (<year>2009</year>) <volume>37</volume>(<issue>Suppl 1</issue>):<fpage>D1006</fpage>&#x02013;<lpage>12</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkn838</pub-id></citation></ref>
<ref id="B60"><label>60</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lefranc</surname> <given-names>M-P</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Kaas</surname> <given-names>Q</given-names></name> <name><surname>Duprat</surname> <given-names>E</given-names></name> <name><surname>Jabado-Michaloud</surname> <given-names>J</given-names></name> <name><surname>Scaviner</surname> <given-names>D</given-names></name> <etal/></person-group> <article-title>IMGT, the international ImMunoGeneTics information system<sup>&#x000AE;</sup></article-title>. <source>Nucleic Acids Res</source> (<year>2005</year>) <volume>33</volume>(<issue>Suppl 1</issue>):<fpage>D593</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gki065</pub-id></citation></ref>
<ref id="B61"><label>61</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ruiz</surname> <given-names>M</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name> <name><surname>Ginestoux</surname> <given-names>C</given-names></name> <name><surname>Stoehr</surname> <given-names>P</given-names></name> <name><surname>Robinson</surname> <given-names>J</given-names></name> <name><surname>Bodmer</surname> <given-names>J</given-names></name> <etal/></person-group> <article-title>IMGT, the international ImMunoGeneTics database</article-title>. <source>Nucleic Acids Res</source> (<year>2000</year>) <volume>28</volume>(<issue>1</issue>):<fpage>219</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1093/nar/28.1.219</pub-id><pub-id pub-id-type="pmid">10592230</pub-id></citation></ref>
<ref id="B62"><label>62</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balakirev</surname> <given-names>ES</given-names></name> <name><surname>Ayala</surname> <given-names>FJ</given-names></name></person-group>. <article-title>Pseudogenes: are they &#x0201C;junk&#x0201D; or functional DNA?</article-title> <source>Annu Rev Genet</source> (<year>2003</year>) <volume>37</volume>:<fpage>123</fpage>.<pub-id pub-id-type="doi">10.1146/annurev.genet.37.040103.103949</pub-id></citation></ref>
<ref id="B63"><label>63</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schiff</surname> <given-names>C</given-names></name> <name><surname>Milili</surname> <given-names>M</given-names></name> <name><surname>Fougereau</surname> <given-names>M</given-names></name></person-group>. <article-title>Functional and pseudogenes are similarly organized and may equally contribute to the extensive antibody diversity of the IgVHII family</article-title>. <source>EMBO J</source> (<year>1985</year>) <volume>4</volume>(<issue>5</issue>):<fpage>1225</fpage>.<pub-id pub-id-type="pmid">3924600</pub-id></citation></ref>
<ref id="B64"><label>64</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sollbach</surname> <given-names>AE</given-names></name> <name><surname>Wu</surname> <given-names>GE</given-names></name></person-group>. <article-title>Inversions produced during V(D)J rearrangement at IgH, the immunoglobulin heavy-chain locus</article-title>. <source>Mol Cell Biol</source> (<year>1995</year>) <volume>15</volume>(<issue>2</issue>):<fpage>671</fpage>&#x02013;<lpage>81</lpage>.<pub-id pub-id-type="doi">10.1128/MCB.15.2.671</pub-id><pub-id pub-id-type="pmid">7823936</pub-id></citation></ref>
<ref id="B65"><label>65</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collins</surname> <given-names>AM</given-names></name> <name><surname>Ikutani</surname> <given-names>M</given-names></name> <name><surname>Puiu</surname> <given-names>D</given-names></name> <name><surname>Buck</surname> <given-names>GA</given-names></name> <name><surname>Nadkarni</surname> <given-names>A</given-names></name> <name><surname>Gaeta</surname> <given-names>B</given-names></name></person-group>. <article-title>Partitioning of rearranged Ig genes by mutation analysis demonstrates D-D fusion and V gene replacement in the expressed human repertoire</article-title>. <source>J Immunol</source> (<year>2004</year>) <volume>172</volume>(<issue>1</issue>):<fpage>340</fpage>&#x02013;<lpage>8</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.172.1.340</pub-id><pub-id pub-id-type="pmid">14688342</pub-id></citation></ref>
<ref id="B66"><label>66</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corbett</surname> <given-names>SJ</given-names></name> <name><surname>Tomlinson</surname> <given-names>IM</given-names></name> <name><surname>Sonnhammer</surname> <given-names>ELL</given-names></name> <name><surname>Buck</surname> <given-names>D</given-names></name> <name><surname>Winter</surname> <given-names>G</given-names></name></person-group>. <article-title>Sequence of the human immunoglobulin diversity (D) segment locus: a systematic analysis provides no evidence for the use of DIR segments, inverted D segments, &#x0201C;minor&#x0201D; D segments or D-D recombination1</article-title>. <source>J Mol Biol</source> (<year>1997</year>) <volume>270</volume>(<issue>4</issue>):<fpage>587</fpage>&#x02013;<lpage>97</lpage>.<pub-id pub-id-type="doi">10.1006/jmbi.1997.1141</pub-id></citation></ref>
<ref id="B67"><label>67</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>EJ</given-names></name></person-group>. <article-title>Mechanism of somatic hypermutation: critical analysis of strand biased mutation signatures at A: T and G: C base pairs</article-title>. <source>Mol Immunol</source> (<year>2009</year>) <volume>46</volume>(<issue>3</issue>):<fpage>305</fpage>&#x02013;<lpage>20</lpage>.<pub-id pub-id-type="doi">10.1016/j.molimm.2008.10.021</pub-id><pub-id pub-id-type="pmid">19062097</pub-id></citation></ref>
<ref id="B68"><label>68</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>EJ</given-names></name> <name><surname>Franklin</surname> <given-names>A</given-names></name> <name><surname>Blanden</surname> <given-names>RV</given-names></name></person-group>. <article-title>Genesis of the strand-biased signature in somatic hypermutation of rearranged immunoglobulin variable genes</article-title>. <source>Immunol Cell Biol</source> (<year>2004</year>) <volume>82</volume>(<issue>2</issue>):<fpage>209</fpage>&#x02013;<lpage>18</lpage>.<pub-id pub-id-type="doi">10.1046/j.0818-9641.2004.01224.x</pub-id><pub-id pub-id-type="pmid">15061776</pub-id></citation></ref>
<ref id="B69"><label>69</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kepler</surname> <given-names>TB</given-names></name> <name><surname>Borrero</surname> <given-names>M</given-names></name> <name><surname>Rugerio</surname> <given-names>B</given-names></name> <name><surname>McCray</surname> <given-names>SK</given-names></name> <name><surname>Clarke</surname> <given-names>SH</given-names></name></person-group>. <article-title>Interdependence of N nucleotide addition and recombination site choice in V(D)J rearrangement</article-title>. <source>J Immunol</source> (<year>1996</year>) <volume>157</volume>(<issue>10</issue>):<fpage>4451</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="pmid">8906821</pub-id></citation></ref>
<ref id="B70"><label>70</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gauss</surname> <given-names>GH</given-names></name> <name><surname>Lieber</surname> <given-names>MR</given-names></name></person-group>. <article-title>Mechanistic constraints on diversity in human V (D) J recombination</article-title>. <source>Mol Cell Biol</source> (<year>1996</year>) <volume>16</volume>(<issue>1</issue>):<fpage>258</fpage>&#x02013;<lpage>69</lpage>.<pub-id pub-id-type="doi">10.1128/MCB.16.1.258</pub-id><pub-id pub-id-type="pmid">8524303</pub-id></citation></ref>
<ref id="B71"><label>71</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carlson</surname> <given-names>CS</given-names></name> <name><surname>Emerson</surname> <given-names>RO</given-names></name> <name><surname>Sherwood</surname> <given-names>AM</given-names></name> <name><surname>Desmarais</surname> <given-names>C</given-names></name> <name><surname>Chung</surname> <given-names>MW</given-names></name> <name><surname>Parsons</surname> <given-names>JM</given-names></name> <etal/></person-group> <article-title>Using synthetic templates to design an unbiased multiplex PCR assay</article-title>. <source>Nat Commun</source> (<year>2013</year>) <volume>4</volume>:<fpage>2680</fpage>.<pub-id pub-id-type="doi">10.1038/ncomms3680</pub-id><pub-id pub-id-type="pmid">24157944</pub-id></citation></ref>
<ref id="B72"><label>72</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rada</surname> <given-names>C</given-names></name> <name><surname>Milstein</surname> <given-names>C</given-names></name></person-group>. <article-title>The intrinsic hypermutability of antibody heavy and light chain genes decays exponentially</article-title>. <source>EMBO J</source> (<year>2001</year>) <volume>20</volume>(<issue>16</issue>):<fpage>4570</fpage>&#x02013;<lpage>6</lpage>.<pub-id pub-id-type="doi">10.1093/emboj/20.16.4570</pub-id><pub-id pub-id-type="pmid">11500383</pub-id></citation></ref>
<ref id="B73"><label>73</label><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Shlomchik</surname> <given-names>MJ</given-names></name> <name><surname>Litwin</surname> <given-names>S</given-names></name> <name><surname>Weigert</surname> <given-names>M</given-names></name></person-group>. <article-title>The influence of somatic mutation on clonal expansion, in progress in immunology</article-title>. In: <person-group person-group-type="editor"><name><surname>Melchers</surname> <given-names>F</given-names></name></person-group>, editor. <conf-name>Vol. VII: Proceedings of the 7th International Congress Immunology Berlin 1989</conf-name>. <conf-loc>Berlin</conf-loc>: <conf-sponsor>Springer</conf-sponsor> (<year>1989</year>). p. <fpage>415</fpage>&#x02013;<lpage>23</lpage>.</citation></ref>
<ref id="B74"><label>74</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Basu</surname> <given-names>M</given-names></name> <name><surname>Hegde</surname> <given-names>MV</given-names></name> <name><surname>Modak</surname> <given-names>MJ</given-names></name></person-group>. <article-title>Synthesis of compositionally unique DNA by terminal deoxynucleotidyl transferase</article-title>. <source>Biochem Biophys Res Commun</source> (<year>1983</year>) <volume>111</volume>(<issue>3</issue>):<fpage>1105</fpage>&#x02013;<lpage>12</lpage>.<pub-id pub-id-type="doi">10.1016/0006-291X(83)91413-4</pub-id><pub-id pub-id-type="pmid">6301484</pub-id></citation></ref>
<ref id="B75"><label>75</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gangi-Peterson</surname> <given-names>L</given-names></name> <name><surname>Sorscher</surname> <given-names>DH</given-names></name> <name><surname>Reynolds</surname> <given-names>JW</given-names></name> <name><surname>Kepler</surname> <given-names>TB</given-names></name> <name><surname>Mitchell</surname> <given-names>BS</given-names></name></person-group>. <article-title>Nucleotide pool imbalance and adenosine deaminase deficiency induce alterations of N-region insertions during V(D)J recombination</article-title>. <source>J Clin Invest</source> (<year>1999</year>) <volume>103</volume>(<issue>6</issue>):<fpage>833</fpage>&#x02013;<lpage>41</lpage>.<pub-id pub-id-type="doi">10.1172/JCI4320</pub-id><pub-id pub-id-type="pmid">10079104</pub-id></citation></ref>
<ref id="B76"><label>76</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>J</given-names></name> <name><surname>Panavas</surname> <given-names>T</given-names></name> <name><surname>Thys</surname> <given-names>K</given-names></name> <name><surname>Aerssens</surname> <given-names>J</given-names></name> <name><surname>Naso</surname> <given-names>M</given-names></name> <name><surname>Fisher</surname> <given-names>J</given-names></name> <etal/></person-group> <article-title>IgG variable region and VH CDR3 diversity in unimmunized mice analyzed by massively parallel sequencing</article-title>. <source>Mol Immunol</source> (<year>2014</year>) <volume>57</volume>(<issue>2</issue>):<fpage>274</fpage>&#x02013;<lpage>83</lpage>.<pub-id pub-id-type="doi">10.1016/j.molimm.2013.09.008</pub-id><pub-id pub-id-type="pmid">24211535</pub-id></citation></ref>
<ref id="B77"><label>77</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kepler</surname> <given-names>TB</given-names></name> <name><surname>Munshaw</surname> <given-names>S</given-names></name> <name><surname>Wiehe</surname> <given-names>K</given-names></name> <name><surname>Zhang</surname> <given-names>R</given-names></name> <name><surname>Yu</surname> <given-names>JS</given-names></name> <name><surname>Woods</surname> <given-names>CW</given-names></name> <etal/></person-group> <article-title>Reconstructing a B-cell clonal lineage. II. mutation, selection, and affinity maturation</article-title>. <source>Front Immunol</source> (<year>2014</year>) <volume>5</volume>:<fpage>170</fpage>.<pub-id pub-id-type="doi">10.3389/fimmu.2014.00170</pub-id><pub-id pub-id-type="pmid">24795717</pub-id></citation></ref>
<ref id="B78"><label>78</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marshall</surname> <given-names>AJ</given-names></name> <name><surname>Paige</surname> <given-names>CJ</given-names></name> <name><surname>Wu</surname> <given-names>GE</given-names></name></person-group>. <article-title>V (H) repertoire maturation during B cell development in vitro: differential selection of Ig heavy chains by fetal and adult B cell progenitors</article-title>. <source>J Immunol</source> (<year>1997</year>) <volume>158</volume>(<issue>9</issue>):<fpage>4282</fpage>&#x02013;<lpage>91</lpage>.<pub-id pub-id-type="pmid">9126990</pub-id></citation></ref>
<ref id="B79"><label>79</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elhanati</surname> <given-names>Y</given-names></name> <name><surname>Murugan</surname> <given-names>A</given-names></name> <name><surname>Callan</surname> <given-names>CG</given-names> <suffix>Jr</suffix></name> <name><surname>Mora</surname> <given-names>T</given-names></name> <name><surname>Walczak</surname> <given-names>AM</given-names></name></person-group>. <article-title>Quantifying selection in immune receptor repertoires</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2014</year>) <volume>111</volume>(<issue>27</issue>):<fpage>9875</fpage>&#x02013;<lpage>80</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1409572111</pub-id><pub-id pub-id-type="pmid">24941953</pub-id></citation></ref>
<ref id="B80"><label>80</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>EJ</given-names></name> <name><surname>Lindley</surname> <given-names>RA</given-names></name> <name><surname>Wen</surname> <given-names>J</given-names></name> <name><surname>Weiller</surname> <given-names>GF</given-names></name></person-group>. <article-title>Computational analyses show A-to-G mutations correlate with nascent mRNA hairpins at somatic hypermutation hotspots</article-title>. <source>DNA Repair</source> (<year>2006</year>) <volume>5</volume>(<issue>11</issue>):<fpage>1346</fpage>&#x02013;<lpage>63</lpage>.<pub-id pub-id-type="doi">10.1016/j.dnarep.2006.06.002</pub-id><pub-id pub-id-type="pmid">16884961</pub-id></citation></ref>
<ref id="B81"><label>81</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogozin</surname> <given-names>IB</given-names></name> <name><surname>Diaz</surname> <given-names>M</given-names></name></person-group>. <article-title>Cutting edge: DGYW/WRCH is a better predictor of mutability at G:C bases in ig hypermutation than the widely accepted RGYW/WRCY motif and probably reflects a two-step activation-induced cytidine deaminase-triggered process</article-title>. <source>J Immunol</source> (<year>2004</year>) <volume>172</volume>(<issue>6</issue>):<fpage>3382</fpage>&#x02013;<lpage>4</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.172.6.3382</pub-id><pub-id pub-id-type="pmid">15004135</pub-id></citation></ref>
<ref id="B82"><label>82</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Foster</surname> <given-names>SJ</given-names></name> <name><surname>Dorner</surname> <given-names>T</given-names></name> <name><surname>Lipsky</surname> <given-names>PE</given-names></name></person-group>. <article-title>Somatic hypermutation of VkJk rearrangements: targeting of RGYW motifs on both DNA strands and preferential selection of mutated codons within RGYW motifs</article-title>. <source>Eur J Immunol</source> (<year>1999</year>) <volume>29</volume>:<fpage>4011</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1002/(SICI)1521-4141(199912)29:12&#x0003C;4011::AID-IMMU4011&#x0003E;3.0.CO;2-W</pub-id></citation></ref>
<ref id="B83"><label>83</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>NY</given-names></name> <name><surname>Wilson</surname> <given-names>K</given-names></name> <name><surname>Jared</surname> <given-names>M</given-names></name> <name><surname>Wilson</surname> <given-names>PC</given-names></name></person-group>. <article-title>Intricate targeting of immunoglobulin somatic hypermutation maximizes the efficiency of affinity maturation</article-title>. <source>J Exp Med</source> (<year>2005</year>) <volume>201</volume>(<issue>9</issue>):<fpage>1467</fpage>&#x02013;<lpage>78</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20042483</pub-id><pub-id pub-id-type="pmid">15867095</pub-id></citation></ref>
<ref id="B84"><label>84</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luzina</surname> <given-names>IG</given-names></name> <name><surname>Atamas</surname> <given-names>SP</given-names></name> <name><surname>Storrer</surname> <given-names>CE</given-names></name> <name><surname>daSilva</surname> <given-names>LC</given-names></name> <name><surname>Kelsoe</surname> <given-names>G</given-names></name> <name><surname>Papadimitriou</surname> <given-names>JC</given-names></name> <etal/></person-group> <article-title>Spontaneous formation of germinal centers in autoimmune mice</article-title>. <source>J Leukoc Biol</source> (<year>2001</year>) <volume>70</volume>(<issue>4</issue>):<fpage>578</fpage>&#x02013;<lpage>84</lpage>.<pub-id pub-id-type="pmid">11590194</pub-id></citation></ref>
<ref id="B85"><label>85</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Domeier</surname> <given-names>PP</given-names></name> <name><surname>Chodisetti</surname> <given-names>SB</given-names></name> <name><surname>Soni</surname> <given-names>C</given-names></name> <name><surname>Schell</surname> <given-names>SL</given-names></name> <name><surname>Elias</surname> <given-names>MJ</given-names></name> <name><surname>Wong</surname> <given-names>EB</given-names></name> <etal/></person-group> <article-title>IFN-&#x003B3; receptor and STAT1 signaling in B cells are central to spontaneous germinal center formation and autoimmunity</article-title>. <source>J Exp Med</source> (<year>2016</year>) <volume>213</volume>(<issue>5</issue>):<fpage>715</fpage>&#x02013;<lpage>32</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20151722</pub-id><pub-id pub-id-type="pmid">27069112</pub-id></citation></ref>
<ref id="B86"><label>86</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jones</surname> <given-names>L</given-names></name> <name><surname>Ho</surname> <given-names>WQ</given-names></name> <name><surname>Ying</surname> <given-names>S</given-names></name> <name><surname>Ramakrishna</surname> <given-names>L</given-names></name> <name><surname>Srinivasan</surname> <given-names>KG</given-names></name> <name><surname>Yurieva</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>A subpopulation of high IL-21-producing CD4&#x0002B; T cells in Peyer&#x02019;s patches is induced by the microbiota and regulates germinal centers</article-title>. <source>Sci Rep</source> (<year>2016</year>) <volume>6</volume>:<fpage>30784</fpage>.<pub-id pub-id-type="doi">10.1038/srep30784</pub-id></citation></ref>
<ref id="B87"><label>87</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glanville</surname> <given-names>J</given-names></name> <name><surname>Zhai</surname> <given-names>W</given-names></name> <name><surname>Berka</surname> <given-names>J</given-names></name> <name><surname>Telman</surname> <given-names>D</given-names></name> <name><surname>Huerta</surname> <given-names>G</given-names></name> <name><surname>Mehta</surname> <given-names>GR</given-names></name> <etal/></person-group> <article-title>Precise determination of the diversity of a combinatorial antibody library gives insight into the human immunoglobulin repertoire</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2009</year>) <volume>106</volume>(<issue>48</issue>):<fpage>20216</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.0909775106</pub-id><pub-id pub-id-type="pmid">19875695</pub-id></citation></ref>
<ref id="B88"><label>88</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casellas</surname> <given-names>R</given-names></name> <name><surname>Basu</surname> <given-names>U</given-names></name> <name><surname>Yewdell</surname> <given-names>WT</given-names></name> <name><surname>Chaudhuri</surname> <given-names>J</given-names></name> <name><surname>Robbiani</surname> <given-names>DF</given-names></name> <name><surname>Di Noia</surname> <given-names>JM</given-names></name></person-group>. <article-title>Mutations, kataegis and translocations in B cells: understanding AID promiscuous activity</article-title>. <source>Nat Rev Immunol</source> (<year>2016</year>) <volume>16</volume>(<issue>3</issue>):<fpage>164</fpage>&#x02013;<lpage>76</lpage>.<pub-id pub-id-type="doi">10.1038/nri.2016.2</pub-id><pub-id pub-id-type="pmid">26898111</pub-id></citation></ref>
<ref id="B89"><label>89</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>EJ</given-names></name></person-group>. <article-title>Somatic hypermutation in immunity and cancer: critical analysis of strand-biased and codon-context mutation signatures</article-title>. <source>DNA Repair</source> (<year>2016</year>) <volume>45</volume>:<fpage>1</fpage>&#x02013;<lpage>24</lpage>.<pub-id pub-id-type="doi">10.1016/j.dnarep.2016.07.001</pub-id><pub-id pub-id-type="pmid">27449479</pub-id></citation></ref>
<ref id="B90"><label>90</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>A</given-names></name> <name><surname>Li</surname> <given-names>Z</given-names></name> <name><surname>Lin</surname> <given-names>DP</given-names></name> <name><surname>Bardwell</surname> <given-names>PD</given-names></name> <name><surname>Iglesias-Ussel</surname> <given-names>MD</given-names></name> <name><surname>Edelmann</surname> <given-names>W</given-names></name> <etal/></person-group> <article-title>Msh2 ATPase activity is essential for somatic hypermutation at A-T basepairs and for efficient class switch recombination</article-title>. <source>J Exp Med</source> (<year>2003</year>) <volume>198</volume>(<issue>8</issue>):<fpage>1171</fpage>&#x02013;<lpage>8</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20030880</pub-id><pub-id pub-id-type="pmid">14568978</pub-id></citation></ref>
<ref id="B91"><label>91</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rada</surname> <given-names>C</given-names></name> <name><surname>Ehrenstein</surname> <given-names>MR</given-names></name> <name><surname>Neuberger</surname> <given-names>MS</given-names></name> <name><surname>Milstein</surname> <given-names>C</given-names></name></person-group>. <article-title>Hot spot focusing of somatic hypermutation in MSH2-deficient mice suggests two stages of mutational targeting</article-title>. <source>Immunity</source> (<year>1998</year>) <volume>9</volume>(<issue>1</issue>):<fpage>135</fpage>&#x02013;<lpage>41</lpage>.<pub-id pub-id-type="doi">10.1016/S1074-7613(00)80595-6</pub-id><pub-id pub-id-type="pmid">9697843</pub-id></citation></ref>
<ref id="B92"><label>92</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Milstein</surname> <given-names>C</given-names></name> <name><surname>Neuberger</surname> <given-names>MS</given-names></name> <name><surname>Staden</surname> <given-names>R</given-names></name></person-group>. <article-title>Both DNA strands of antibody genes are hypermutation targets</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>1998</year>) <volume>95</volume>(<issue>15</issue>):<fpage>8791</fpage>&#x02013;<lpage>4</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.95.15.8791</pub-id><pub-id pub-id-type="pmid">9671757</pub-id></citation></ref>
<ref id="B93"><label>93</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murray</surname> <given-names>F</given-names></name> <name><surname>Darzentas</surname> <given-names>N</given-names></name> <name><surname>Hadzidimitriou</surname> <given-names>A</given-names></name> <name><surname>Tobin</surname> <given-names>G</given-names></name> <name><surname>Boudjogra</surname> <given-names>M</given-names></name> <name><surname>Scielzo</surname> <given-names>C</given-names></name> <etal/></person-group> <article-title>Stereotyped patterns of somatic hypermutation in subsets of patients with chronic lymphocytic leukemia: implications for the role of antigen selection in leukemogenesis</article-title>. <source>Blood</source> (<year>2008</year>) <volume>111</volume>(<issue>3</issue>):<fpage>1524</fpage>&#x02013;<lpage>33</lpage>.<pub-id pub-id-type="doi">10.1182/blood-2007-07-099564</pub-id><pub-id pub-id-type="pmid">17959859</pub-id></citation></ref>
<ref id="B94"><label>94</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steele</surname> <given-names>EJ</given-names></name></person-group>. <article-title>Reflections on the state of play in somatic hypermutation</article-title>. <source>Mol Immunol</source> (<year>2008</year>) <volume>45</volume>(<issue>10</issue>):<fpage>2723</fpage>&#x02013;<lpage>6</lpage>.<pub-id pub-id-type="doi">10.1016/j.molimm.2008.02.002</pub-id><pub-id pub-id-type="pmid">18359085</pub-id></citation></ref>
<ref id="B95"><label>95</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>FH</given-names></name> <name><surname>Castro</surname> <given-names>MM</given-names></name> <name><surname>Aboul-ela</surname> <given-names>F</given-names></name> <name><surname>Tinoco</surname> <given-names>I</given-names> <suffix>Jr</suffix></name></person-group>. <article-title>Base pairing involving deoxyinosine: implications for probe design</article-title>. <source>Nucleic Acids Res</source> (<year>1985</year>) <volume>13</volume>(<issue>24</issue>):<fpage>8927</fpage>&#x02013;<lpage>38</lpage>.<pub-id pub-id-type="doi">10.1093/nar/13.24.8927</pub-id><pub-id pub-id-type="pmid">4080553</pub-id></citation></ref>
<ref id="B96"><label>96</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Franklin</surname> <given-names>A</given-names></name> <name><surname>Milburn</surname> <given-names>PJ</given-names></name> <name><surname>Blanden</surname> <given-names>RV</given-names></name> <name><surname>Steele</surname> <given-names>EJ</given-names></name></person-group>. <article-title>Special feature human DNA polymerase-&#x003B7;, an A-T mutator in somatic hypermutation of rearranged immunoglobulin genes, is a reverse transcriptase</article-title>. <source>Immunol Cell Biol</source> (<year>2004</year>) <volume>82</volume>(<issue>2</issue>):<fpage>219</fpage>&#x02013;<lpage>25</lpage>.<pub-id pub-id-type="doi">10.1046/j.0818-9641.2004.01221.x</pub-id></citation></ref>
<ref id="B97"><label>97</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tsuruoka</surname> <given-names>N</given-names></name> <name><surname>Arima</surname> <given-names>M</given-names></name> <name><surname>Yoshida</surname> <given-names>N</given-names></name> <name><surname>Okada</surname> <given-names>S</given-names></name> <name><surname>Sakamoto</surname> <given-names>A</given-names></name> <name><surname>Hatano</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>ADAR1 protein induces adenosine-targeted DNA mutations in senescent Bcl6 gene-deficient cells</article-title>. <source>J Biol Chem</source> (<year>2013</year>) <volume>288</volume>(<issue>2</issue>):<fpage>826</fpage>&#x02013;<lpage>36</lpage>.<pub-id pub-id-type="doi">10.1074/jbc.M112.365718</pub-id><pub-id pub-id-type="pmid">23209284</pub-id></citation></ref>
<ref id="B98"><label>98</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roy</surname> <given-names>B</given-names></name> <name><surname>Shukla</surname> <given-names>S</given-names></name> <name><surname>&#x00141;yszkiewicz</surname> <given-names>M</given-names></name> <name><surname>Krey</surname> <given-names>M</given-names></name> <name><surname>Viegas</surname> <given-names>N</given-names></name> <name><surname>D&#x000FC;ber</surname> <given-names>S</given-names></name> <etal/></person-group> <article-title>Somatic hypermutation in peritoneal B1b cells</article-title>. <source>Mol Immunol</source> (<year>2009</year>) <volume>46</volume>(<issue>8&#x02013;9</issue>):<fpage>1613</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1016/j.molimm.2009.02.026</pub-id><pub-id pub-id-type="pmid">19327839</pub-id></citation></ref>
<ref id="B99"><label>99</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gojobori</surname> <given-names>T</given-names></name> <name><surname>Nei</surname> <given-names>M</given-names></name></person-group>. <article-title>Relative contributions of germline gene variation and somatic mutation to immunoglobulin diversity in the mouse</article-title>. <source>Mol Biol Evol</source> (<year>1986</year>) <volume>3</volume>(<issue>2</issue>):<fpage>156</fpage>&#x02013;<lpage>67</lpage>.<pub-id pub-id-type="pmid">3444398</pub-id></citation></ref>
</ref-list>
</back>
</article>