<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Immunol.</journal-id>
<journal-title>Frontiers in Immunology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Immunol.</abbrev-journal-title>
<issn pub-type="epub">1664-3224</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fimmu.2017.01550</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Immunology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Local Chromatin Features Including PU.1 and IKAROS Binding and H3K4 Methylation Shape the Repertoire of Immunoglobulin Kappa Genes Chosen for V(D)J Recombination</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Matheson</surname> <given-names>Louise S.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/482627"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bolland</surname> <given-names>Daniel J.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/494231"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chovanec</surname> <given-names>Peter</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/493195"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Krueger</surname> <given-names>Felix</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/302456"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Andrews</surname> <given-names>Simon</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/132648"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Koohy</surname> <given-names>Hashem</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/188449"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Corcoran</surname> <given-names>Anne E.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/96329"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Nuclear Dynamics Programme, Babraham Institute</institution>, <addr-line>Cambridge</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff2"><sup>2</sup><institution>Bioinformatics Group, Babraham Institute</institution>, <addr-line>Cambridge</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Deborah K. Dunn-Walters, University of Surrey, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Albert Jeltsch, University of Stuttgart, Germany; Kim Good-Jacobson, Monash University, Australia</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Hashem Koohy, <email>hashem.koohy&#x00040;rdm.ox.ac.uk</email>; Anne E. Corcoran, <email>anne.corcoran&#x00040;babraham.ac.uk</email></corresp>
<fn fn-type="present-address" id="fn001"><p><sup>&#x02020;</sup>Present address: Hashem Koohy, MRC Human Immunology Unit, Weatherall Institute of Molecular Medicine, University of Oxford, Oxford, United Kingdom</p></fn>
<fn fn-type="other" id="fn002"><p>Specialty section: This article was submitted to B Cell Biology, a section of the journal Frontiers in Immunology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>11</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1550</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>09</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>10</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Matheson, Bolland, Chovanec, Krueger, Andrews, Koohy and Corcoran.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Matheson, Bolland, Chovanec, Krueger, Andrews, Koohy and Corcoran</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>V(D)J recombination is essential for the generation of diverse antigen receptor (AgR) repertoires. In B cells, immunoglobulin kappa (<italic>Ig&#x003BA;</italic>) light chain recombination follows immunoglobulin heavy chain (<italic>Igh</italic>) recombination. We recently developed the DNA-based VDJ-seq assay for the unbiased quantitation of <italic>Igh</italic> VH and DH repertoires. Integration of VDJ-seq data with genome-wide datasets revealed that two chromatin states at the recombination signal sequence (RSS) of VH genes are highly predictive of recombination in mouse pro-B cells. It is unknown whether local chromatin states contribute to V&#x003BA; gene choice during <italic>Ig&#x003BA;</italic> recombination. Here we adapt VDJ-seq to profile the <italic>Ig&#x003BA;</italic> V&#x003BA;J&#x003BA; repertoire and present a comprehensive readout in mouse pre-B cells, revealing highly variable V&#x003BA; gene usage. Integration with genome-wide datasets for histone modifications, DNase hypersensitivity, transcription factor binding and germline transcription identified PU.1 binding at the RSS, which was unimportant for <italic>Igh</italic>, as highly predictive of whether a V&#x003BA; gene will recombine or not, suggesting that it plays a binary, all-or-nothing role, priming genes for recombination. Thereafter, the frequency with which these genes recombine was shaped both by the presence and level of enrichment of several other chromatin features, including H3K4 methylation and IKAROS binding. Moreover, in contrast to the <italic>Igh</italic> locus, the chromatin landscape of the promoter, as well as of the RSS, contributes to V&#x003BA; gene recombination. Thus, multiple facets of local chromatin features explain much of the variation in V&#x003BA; gene usage. Together, these findings reveal shared and divergent roles for epigenetic features and transcription factors in AgR V(D)J recombination and provide avenues for further investigation of chromatin signatures that may underpin V(D)J-mediated chromosomal translocations.</p>
</abstract>
<kwd-group>
<kwd>V(D)J recombination</kwd>
<kwd>immunoglobulin kappa</kwd>
<kwd>epigenetic regulation</kwd>
<kwd>chromatin state</kwd>
<kwd>PU.1</kwd>
<kwd>IKAROS</kwd>
</kwd-group>
<contract-sponsor id="cn01">Biotechnology and Biological Sciences Research Council<named-content content-type="fundref-id">10.13039/501100000268</named-content></contract-sponsor>
<counts>
<fig-count count="9"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="82"/>
<page-count count="22"/>
<word-count count="15029"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<title>Introduction</title>
<p>V(D)J recombination enables the sequential rearrangement of variable (V), diversity (D) and joining (J) gene segments in B and T cell antigen receptor (AgR) loci. This mechanism, catalysed by the RAG recombinase complex which recognises the recombination signal sequence (RSS) of each gene segment, is the essential first step in the generation of diverse AgR repertoires, transforming a couple of hundred genes into millions of different antigen specificities (<xref ref-type="bibr" rid="B1">1</xref>). In B cells, the immunoglobulin heavy chain (<italic>Igh</italic>) locus recombines first, with D to J<sub>H</sub> recombination on both alleles preceding V<sub>H</sub> to DJ<sub>H</sub> recombination on one allele in pro-B cells (<xref ref-type="bibr" rid="B2">2</xref>). The joining of these genes is imprecise, due to exonuclease activity and the addition of non-templated nucleotides, partly mediated by terminal deoxynucleotidyl transferase (TdT), thereby enhancing diversity (<xref ref-type="bibr" rid="B3">3</xref>). Functional IgH chains are expressed on the cell surface with the surrogate light chain as the pre-B cell receptor. This promotes proliferation, differentiation to the pre-B cell stage and recombination of the immunoglobulin kappa light chain locus (<italic>Ig&#x003BA;</italic>) (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>).</p>
<p>The mouse <italic>Ig&#x003BA;</italic> locus, located on chromosome 6, is 3.2&#x02009;Mb in size and contains 162 V&#x003BA; genes, 5 J&#x003BA; genes and a single C&#x003BA; gene (<xref ref-type="bibr" rid="B6">6</xref>). In contrast to the <italic>Igh</italic> locus, in which all V<sub>H</sub> genes are in the same orientation, over half of the V&#x003BA; genes are in the reverse orientation with respect to the J&#x003BA; and C&#x003BA; genes (<xref ref-type="bibr" rid="B6">6</xref>), and their recombination leads to inversion rather than to the deletion of the intervening DNA. Whilst joining is still imprecise, light chain V-J junctions are much less diverse than <italic>Igh</italic> junctions since TdT is not expressed in pre-B cells (<xref ref-type="bibr" rid="B7">7</xref>, <xref ref-type="bibr" rid="B8">8</xref>) and exonuclease activity is reduced (<xref ref-type="bibr" rid="B9">9</xref>). Surface expression of IgH and Ig&#x003BA; together as the B cell receptor (BCR) allows selection that favours productive V&#x003BA;J&#x003BA; rearrangements and eliminates autoreactive BCRs. If necessary, recombination between the remaining upstream V&#x003BA; and downstream J&#x003BA; genes, termed receptor editing, is permitted (<xref ref-type="bibr" rid="B10">10</xref>); rearrangement of the second allele may also occur. The first recombination event at each allele is biased towards usage of the <italic>J&#x003BA;1</italic> gene, through suppression of DNA breaks at downstream J&#x003BA; genes (<xref ref-type="bibr" rid="B11">11</xref>).</p>
<p>The RAG recombinase-recruiting RSS of each V gene varies in quality, which can be quantified as the RSS information content (RIC) score, with a higher score theoretically more conducive to recombination (<xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B13">13</xref>). However, accumulating evidence shows that whilst the RIC score provides one layer of regulation, epigenetic features including H3K4 methylation also contribute to regulation of VDJ recombination (<xref ref-type="bibr" rid="B14">14</xref>&#x02013;<xref ref-type="bibr" rid="B20">20</xref>). Moreover, several transcription factors (TFs), including PAX5, IRF4, IKAROS, PU.1, E2A and P300, promote the activation and recombination of the <italic>Ig&#x003BA;</italic> locus. However, their specific contribution to shaping the repertoire is unclear (<xref ref-type="bibr" rid="B21">21</xref>&#x02013;<xref ref-type="bibr" rid="B29">29</xref>), and may include long-range or local V gene roles, or a combination thereof. Loss of CTCF or of its binding sites leads to increased transcription and usage of J-proximal V genes in the <italic>Igh</italic> and <italic>Ig&#x003BA;</italic> loci. This suggests a role in long-range looping of the locus, bringing the distal V genes into proximity with the (D)J region (<xref ref-type="bibr" rid="B30">30</xref>&#x02013;<xref ref-type="bibr" rid="B34">34</xref>). Deletion of PAX5 or YY1 also reduces distal V<sub>H</sub> gene recombination in the <italic>Igh</italic> locus (<xref ref-type="bibr" rid="B35">35</xref>, <xref ref-type="bibr" rid="B36">36</xref>). However, these general biases towards proximal recombination cannot explain why genes that are close to each other and similar in sequence can recombine at substantially different levels. A recent RNA-based, high-throughput analysis of the expressed mouse V&#x003BA; gene repertoire revealed that it was highly variable across the locus (<xref ref-type="bibr" rid="B37">37</xref>). Similarly, a DNA-based assay revealed diverse but variable V&#x003BA; gene usage in mouse splenic B cells (<xref ref-type="bibr" rid="B38">38</xref>); both studies also revealed that the V&#x003BA; repertoire for each J&#x003BA; gene differed. Highly represented V&#x003BA; genes in the RNA repertoire interact more frequently with <italic>Ig&#x003BA;</italic> enhancers compared to genes represented at low levels, and E2A has been implicated in orchestrating these interactions (<xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>). YY1 may direct the recombination of specific V&#x003BA; genes since expression of a YY1 mutant lacking a Polycomb Group binding domain resulted in a skewed repertoire in mouse pre-B cells (<xref ref-type="bibr" rid="B41">41</xref>), although a concomitant decrease in receptor editing may contribute to this finding. Thus, the features of the <italic>Ig&#x003BA;</italic> locus that determine the capacity of each V&#x003BA; gene to recombine, and the nature of their contribution, are poorly understood.</p>
<p>We recently developed the DNA-based VDJ-seq assay for unbiased high-throughput quantitation of <italic>Igh</italic> V<sub>H</sub> and D<sub>H</sub> repertoires and applied it to mouse bone marrow pro-B cells (<xref ref-type="bibr" rid="B14">14</xref>). By integrating our VDJ-seq data with genome-wide datasets for numerous histone modifications and TFs, we identified two mutually exclusive chromatin states: an architectural state, characterised by binding of CTCF and RAD21, and an enhancer state, characterised by binding of IRF4, PAX5 and YY1 and by histone modifications associated with enhancers and transcriptional activation. These chromatin states form at the RSS of V<sub>H</sub> genes, and both are highly predictive of active recombination (<xref ref-type="bibr" rid="B14">14</xref>). Moreover, they are enriched at non-canonical genome-wide binding sites for the recombinase enzymes that catalyse V(D)J recombination (<xref ref-type="bibr" rid="B42">42</xref>&#x02013;<xref ref-type="bibr" rid="B44">44</xref>), suggesting these states may also be permissive for the aberrant recombination events that underpin B cell leukaemias. The extent to which the chromatin signatures that underpin V(D)J recombination are shared between AgR loci is unknown. Moreover, whether a consensus signature exists that is predictive of susceptibility to aberrant recombination remains poorly understood.</p>
<p>Whilst the expressed V&#x003BA; gene repertoire has been quantified in pre-B cells (<xref ref-type="bibr" rid="B37">37</xref>), this does not accurately reflect the comparative frequency of recombination of each gene at the DNA level. This is because RNA quantity is an indirect measure of recombination frequency and is affected by factors that include different V&#x003BA; gene promoter strengths, the ratios of productive (in-frame):non-productive rearrangements, and transcript stabilities. For example, many recombined V&#x003BA; pseudogenes will not be detected in an RNA-based assay. However, 11 pseudogenes were detected in the DNA repertoire of splenic B cells (<xref ref-type="bibr" rid="B38">38</xref>), and the frequency of these events is much higher in pre-B cells, before non-functional rearrangements have been removed (<xref ref-type="bibr" rid="B45">45</xref>, <xref ref-type="bibr" rid="B46">46</xref>).</p>
<p>In this study, we adapt the VDJ-seq assay for the unbiased profiling and analysis of the V&#x003BA;J&#x003BA; repertoire, and present a comprehensive inventory of V&#x003BA; gene usage in mouse bone marrow pre-B cells, the pre-selection population of B cells in which recombination is taking place. We additionally build a novel two-step machine learning model to study the relationship between the primary V&#x003BA; gene repertoire and locus-wide profiles of chromatin features and transcription. Our results both confirm previous findings concerning the potential mechanisms that underpin recombination of the <italic>Ig&#x003BA;</italic> locus, and substantially advance our understanding of these regulatory mechanisms. We found that local chromatin features are highly predictive of whether a given V&#x003BA; gene is recombined or not, and of its recombination frequency. In contrast to the <italic>Igh</italic> locus, we observed that not only the RSS but also the V&#x003BA; gene promoter and its surrounding chromatin contribute to recombination frequency. We identified PU.1 binding at the RSS as a crucial feature in determining whether or not a V&#x003BA; gene will actively recombine, whilst IKAROS binding and H3K4 methylation are important in promoting a higher frequency of recombination. Moreover, whilst some local chromatin features that drive recombination are shared with the <italic>Igh</italic>, the regulatory mechanisms contributing to recombination of these two AgR loci are substantially different.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2-1">
<title>Primary Cells</title>
<p>C57BL/6 (WT) and <italic>Rag1<sup>&#x02013;/&#x02013;</sup>/VH81X</italic> mice were maintained in accordance with local and Home Office rules and ARRIVE guidelines under Project Licence 80/2529.</p>
<p>For each biological replicate, bone marrow from nine 6- to 8-week-old <italic>Rag1<sup>&#x02212;/&#x02212;</sup>/VH81X</italic> mice (<xref ref-type="bibr" rid="B47">47</xref>, <xref ref-type="bibr" rid="B48">48</xref>) or from fifteen 12-week-old male wild-type (WT) C57BL/6 mice was incubated with biotinylated antibodies against CD11B (MAC-1; ebioscience), Ly6G (Gr-1; ebioscience), Ly6C (Abd Serotec), TER119 (ebioscience), and CD3E (ebioscience) followed by incubation with streptavidin MACs beads (Miltenyi), to deplete macrophages, granulocytes, erythroid lineage, and T cells. Pre-B cells (surface IgM<italic><sup>&#x02212;</sup></italic>CD25<sup>&#x0002B;</sup>B220<sup>&#x0002B;</sup>CD19<sup>&#x0002B;</sup>) were then flow sorted on a BD FACSAria in the Babraham Institute Flow Cytometry facility. Antibodies used were CD45R BV421 (B220, RA3-6B2, Biolegend), CD19 PerCP-Cy5.5 (1D3, BD Pharmingen), CD25 APC (PC61.5, eBioscience), and IgM PE (eB121-15F9, eBioscience). Sort purities were all greater than 92%.</p>
</sec>
<sec id="S2-2">
<title>VDJ-Seq</title>
<sec id="S2-2-1">
<title>V&#x003BA;J&#x003BA;-Seq Assay</title>
<p>The VDJ-seq assay (<xref ref-type="bibr" rid="B14">14</xref>) was adapted for analysis of the <italic>Ig&#x003BA;</italic> repertoire (Supplementary Text S1 and Figure S1 in Supplementary Material). DNA was isolated from flow sorted pre-B cells using a DNeasy kit (Qiagen) and 10&#x02009;&#x000B5;g was sonicated to 400&#x02009;bp using a Covaris E220 sonicator. Except where AMPure XP beads were used, all following reactions were cleaned up by column purification (Qiagen QIAquick PCR purification kit). Fragmented DNA was end-repaired and A-tailed using standard protocols. Samples were divided in half and short asymmetric adaptors, including a molecular identifier and one of the two different anchor sequences, were ligated to both ends of all fragments (T4 DNA ligase, NEB; 16&#x000B0;C overnight); the two ligations were then pooled. Primer extension (8&#x02009;&#x000B5;l&#x02009;&#x000D7;&#x02009;50&#x02009;&#x000B5;l; 2&#x02009;U NEB Vent Exo-polymerase per reaction) using biotinylated primers that anneal downstream of all functional J&#x003BA; genes (J&#x003BA;1, 2, 4, and 5) allowed for the enrichment of fragments that contain a J&#x003BA; gene using streptavidin beads (My-one C1; Invitrogen), following the manufacturer&#x02019;s protocol with incubation overnight, rotating at room temperature (20&#x02009;&#x000B5;l beads per sample). After washing the beads, four cycles of PCR amplification off the beads were performed, using an Illumina PE1 primer corresponding to the long strand of the asymmetric adaptors, in combination with J&#x003BA;-specific PE2 primers (4&#x02009;&#x000B5;l&#x02009;&#x000D7;&#x02009;25&#x02009;&#x000B5;l; Pwo master mix, Roche). A second primer extension reaction (4&#x02009;&#x000B5;l&#x02009;&#x000D7;&#x02009;50&#x02009;&#x000B5;l) using biotinylated primers that anneal within intergenic regions upstream of each functional J&#x003BA; gene (i.e., present only when unrecombined) was then performed. Unrecombined sequences were removed using streptavidin beads with a 4&#x02009;h incubation at room temperature. The remaining DNA fragments, containing the V&#x003BA;-J&#x003BA; recombined sequences, were further enriched, with 11 additional PCR cycles using the same PE1/J&#x003BA;-PE2 primers as above. PCR products were cleaned up and small products removed, using AMPure XP beads (1&#x000D7;; Beckman Coulter). A final PCR amplification of five cycles was performed to add the flowcell-binding portions of the PE1 and PE2 adaptors, including Illumina Truseq bar codes within PE2. Final libraries were purified and size-selected by double-sided AMPure XP bead purification (0.5&#x000D7; followed by 1&#x000D7;), before quality control using a high sensitivity DNA assay on the Agilent Bioanalyzer, and KAPA qPCR (Illumina library quantification kit, KAPA Biosystems). Libraries were sequenced on the Illumina HiSeq, with 2&#x02009;&#x000D7;&#x02009;100bp paired end sequencing. Sequences of all oligonucleotides used and the cycling conditions are provided in Table S4 in Supplementary Material.</p>
</sec>
<sec id="S2-2-2">
<title>V&#x003BA;J&#x003BA;-Seq Pipeline</title>
<p>We adapted our Babraham LinkON pipeline (<xref ref-type="bibr" rid="B14">14</xref>)<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> for processing of V&#x003BA;J&#x003BA;-seq data. Briefly, sequences were demultiplexed based on Truseq barcodes and trimmed for adaptors and low quality (Phred&#x02009;&#x0003C;&#x02009;20) using TrimGalore version 0.3.8 (Babraham Bioinformatics<xref ref-type="fn" rid="fn2"><sup>2</sup></xref>). Due to the similarity of the J primers, the sequencing quality can drop at positions 3&#x02013;4 for the J reads (Read 2). The first four bases were, therefore, trimmed off all J read sequences (using the option&#x02014;clip r2 4 in Trim Galore). Chimaeric J reads produced through mis-priming of a J&#x003BA; gene with the incorrect J&#x003BA;-PE2 PCR primer were identified by examining the sequence immediately downstream of the primer binding sites, and a find-and-replace step was used to replace the incorrect J&#x003BA; primer sequences within commonly occurring chimaeras with the correct primer sequence. Thereafter, the J&#x003BA; primer (&#x0201C;bait&#x0201D;) sequences were used to assign each read to the corresponding J&#x003BA; gene. Any sequence without a bait was discarded. Sequences were further filtered to exclude any that had less than 20 bases downstream of the bait in Read 2, or that did not include one of the two anchor sequences following the molecular identifier in Read 1. The V end reads (Read 1), excluding the first 15 bases, which comprise the molecular identifier and anchor sequence, were aligned to the NCBIM37/mm<sup>9</sup> mouse genome assembly using Bowtie version 1.1.0 (<xref ref-type="bibr" rid="B49">49</xref>), discarding multi-mapping hits (options: &#x0201C;-m 1&#x02014;strata&#x02014;best&#x0201D;). The data were de-duplicated based on: the sequence of the molecular identifier (6N); the sequence downstream of the bait in Read 2, which includes the V&#x003BA;&#x02013;J&#x003BA; junctions; and the start position of the V read alignment. Any paired reads identical for all criteria were considered to be PCR duplicates, and only one was retained. Finally, aligned Read 1 sequences were produced as output as Bowtie mapped unique_V-BAM files. The entire pipeline is documented in more detail here: <uri xlink:href="https://github.com/FelixKrueger/BabrahamLinkON/blob/master/run_VkSk-Seq_pipeline.md">https://github.com/FelixKrueger/BabrahamLinkON/blob/master/run_VkSk-Seq_pipeline.md</uri>. Read counts for each stage of the pipeline are shown in Figure S2A in Supplementary Material.</p>
</sec>
</sec>
<sec id="S2-3">
<title>Quantification of V&#x003BA;-J&#x003BA; Recombination</title>
<p>BAM files of aligned V reads were loaded using default parameters into Seqmonk version 1.37.1 (Babraham Bioinformatics<xref ref-type="fn" rid="fn3"><sup>3</sup></xref>), a tool for the visualisation and analysis of next generation sequencing data. Based on examination of the reads, two V&#x003BA; gene annotations were corrected (Supplementary Text S2). V reads on the same strand as the gene were then counted within windows extending 750&#x02009;bp upstream (with respect to the gene orientation) from the 3&#x02032; end of each V&#x003BA; gene, for each J&#x003BA; gene separately. Read counts were normalised to the replicate with the median number of reads (replicate 1), either across all Js, or just for J&#x003BA;1-associated reads. For J&#x003BA;1-associated V sequences, the median of the normalised replicates for each V&#x003BA; gene was used as the recombination frequency in downstream analyses. The mappability (defined as the percentage of all possible 85&#x02009;bp sequences over a given window that can be mapped uniquely) over 350&#x02009;bp windows upstream of each V&#x003BA; gene 3&#x02032; end (within which the vast majority of V&#x003BA;J&#x003BA;-seq reads are localised) was calculated, and where stated, genes with low mappability (below 70%) were excluded. Only 11 genes fell into this category, 10 of which actively recombine, and only 4 of these had mappability below 60%. Quantitated data and other information relating to each V&#x003BA; gene is provided in Table S1 in Supplementary Material.</p>
</sec>
<sec id="S2-4">
<title>IMGT HighV-QUEST Analysis</title>
<p>To facilitate analysis by IMGT HighV-QUEST, a tool for the high-throughput analysis of V(D)J-recombined sequences (<xref ref-type="bibr" rid="B50">50</xref>), V and J sequences were first merged, and any gaps filled in, as described previously (<xref ref-type="bibr" rid="B14">14</xref>). <italic>Ig&#x003BA;</italic> rearrangements were then analysed by IMGT HighV-QUEST, with default parameters. Complementarity determining region 3 (CDR3) lengths and productive versus non-productive rearrangement data were obtained from the &#x0201C;Summary&#x0201D; file.</p>
</sec>
<sec id="S2-5">
<title>Definition of Active and Inactive Genes</title>
<p>We defined active genes as those that were significantly enriched (padj&#x02009;&#x0003C;&#x02009;0.01) for V&#x003BA;J&#x003BA;-seq reads compared to the V&#x003BA; region as a whole, using a binomial test. Thus, the probability was defined as the length of the window in which reads were quantified, as a fraction of the total length of the V&#x003BA; region, <italic>n</italic> as the rounded median of normalised read counts for each gene, and <italic>N</italic> as the rounded median of normalised read counts across the entire V&#x003BA; region.</p>
</sec>
<sec id="S2-6">
<title>ChIP Co-Localisation and Enrichment Analysis</title>
<p>ChIP peaks (including DHS) were called using MACS peak calling algorithm (<xref ref-type="bibr" rid="B51">51</xref>). MACS 1.4, which performs better with broad peaks was used for H3K27me3 with pvalue cutoff&#x02009;&#x0003D;&#x02009;1.00e&#x02212;02. For all the remaining features, MACS2 was used with pvalue cutoff&#x02009;&#x0003D;&#x02009;1.00e&#x02212;05. The number of peaks over the <italic>Ig&#x003BA;</italic> locus for each dataset is shown in Table S3 in Supplementary Material.</p>
<p>For each V&#x003BA; gene, we calculated the distance from the centre of the gene to the closest upstream and closest downstream (defined with respect to the gene orientation) peak summits. If a peak summit was located upstream of the gene centre, and within 1&#x02009;kb of the gene start site, it was labelled as promoter-associated; conversely any peaks with summits located downstream of the gene centre and within 1&#x02009;kb of the gene 3&#x02032; end were labelled as RSS-associated.</p>
<p>Relative enrichment for ChIP-seq datasets was calculated as the number of reads over a given window, relative to the average number of reads within windows of identical size across the entire <italic>Ig&#x003BA;</italic> locus.</p>
</sec>
<sec id="S2-7">
<title>Phylogenetic Analysis</title>
<p>A phylogenetic tree of C57BL/6 mouse V&#x003BA; gene germline sequences, based on NCBI Reference Sequence: NG_005612.1<xref ref-type="fn" rid="fn4"><sup>4</sup></xref> but with the corrected annotations detailed in Supplementary Text S2, was constructed using the <uri xlink:href="http://Phylogeny.fr">Phylogeny.fr</uri> tool, without including alignment curation (<xref ref-type="bibr" rid="B52">52</xref>, <xref ref-type="bibr" rid="B53">53</xref>). Multiple sequence alignment was performed with MUSCLE (<xref ref-type="bibr" rid="B54">54</xref>), and the maximum likelihood used for tree construction (<xref ref-type="bibr" rid="B55">55</xref>, <xref ref-type="bibr" rid="B56">56</xref>). The tree was visualised using the R package ggtree (<xref ref-type="bibr" rid="B57">57</xref>), and nodes comprising the majority of each V&#x003BA; gene family were collapsed.</p>
</sec>
<sec id="S2-8">
<title>Computational Approach</title>
<p>Our computational approach comprises an unsupervised and a supervised step. In the unsupervised step, we set out to interrogate the chromatin landscape of the <italic>Ig&#x003BA;</italic> locus through integration of the histone modification, DNase hypersensitivity (DHS), germline transcription, and TF-binding profiles included in this study. The supervised step is constructed in two layers: first, we train a Random Forest Classifier (RF-C) C(X) to predict whether a given gene is active or not; second, we construct a Random Forest Regression (RF-R) model R(X) to predict the frequency of recombination of a given active gene. In what follows, we describe both steps in more details.</p>
<sec id="S2-8-1">
<title>Chromatin Segmentation Analysis</title>
<p>For the supervised step, we used EpiCSeg (<xref ref-type="bibr" rid="B58">58</xref>), which combines the input features for the segmentation and characterisation of a context-specific chromatin landscape. EpiCSeg was originally developed to learn the epigenomic landscape from histone marks. However, we chose this algorithm instead of its commonly used counterparts, such as chromHMM (<xref ref-type="bibr" rid="B59">59</xref>), which works well for combinations of TFs and histone marks. This was because while the underlying mathematical modelling is very similar, it bypasses the binary mode of chromHMM and allows the user to proceed with continuous values, preventing loss of information and overfitting, which is particularly useful for analysis of a single large locus.</p>
<p>For this, we divided the locus into 200&#x02009;bp non-overlapping bins and calculated the enrichment of each feature over all bins using bedtools (<xref ref-type="bibr" rid="B60">60</xref>) multibam coverage function. As an input for EpiCSeg, we constructed a raw read-counts matrix X in which x_{ij} corresponds to the enrichment of feature i in bin number j. We ran EpiCSeg with varying numbers of states ranging from 3 to 15.</p>
<p>The A or E state were assigned to the promoter or RSS, if they overlapped a window extending from the centre of the gene to 500&#x02009;bp up- or downstream, respectively, with the exception that if the A/E state segment did not overlap with the gene start, and its centre was downstream of the gene centre, it would be assigned to the RSS but not the promoter, or vice versa.</p>
</sec>
</sec>
<sec id="S2-9">
<title>Random Forest Classification and Regression Models</title>
<p>For the unsupervised step, we first trained a RF-C to predict whether a V&#x003BA; gene is &#x0201C;active&#x0201D; or not. We then constructed a RF-R model to predict the recombination level of an active gene. We chose RF since it is generally accepted to be superior in tackling high dimensionality (relatively high number of features with low number of samples for the training) and co-linearity between the features (<xref ref-type="bibr" rid="B61">61</xref>).</p>
<p>Read counts for DHS-seq, ChIP-seq, and RNA-seq data were generated using Seqmonk, within four distinct windows for each V&#x003BA; gene: &#x0201C;promoter,&#x0201D; extending from 500&#x02009;bp upstream of the start of the gene to its centre; &#x0201C;RSS,&#x0201D; extending from the centre of the gene to 500&#x02009;bp downstream of its 3&#x02032; end; and &#x0201C;upstream&#x0201D; and &#x0201C;downstream&#x0201D; windows, extending from 500&#x02009;bp to 3&#x02009;kb up- or downstream of the gene start or end, respectively. In addition to these four windows for each of the genome-wide datasets, giving a total of 76 chromatin features, we also included three genetic features. These were: the RSS RIC score, the orientation (or strand) of the V&#x003BA; gene, and the distance from the V&#x003BA; gene to J&#x003BA;1. All of these features, except for the gene orientation, were projected between 0 and 1. These 79 features were considered as the explanatory variable for both the RF-C and RF-R.</p>
<p>The response variable for RF-C was the binary recombination classes (active and inactive), which were defined as described above. For RF-R, the log2-transformed median of the normalised recombination frequencies of active genes was used as the response variable.</p>
<p>Both the RF-C and RF-R approaches were performed with 10-fold cross-validation: 10% of genes were assigned to the test set each time, with every gene included in a test set exactly once. The number of trees generated for each fold was 1,000. For the initial RF-C including all features, the number of features tried at each step was set to 20; for all other models default parameters were used. The average importance of each feature, and SE across the 10-folds, was recorded. For the classification model, the performance was assessed by calculating the percentage of correct predictions (accuracy) across all ten test sets: this was calculated overall, as well as for the active V&#x003BA; genes (giving a measure of the sensitivity with which we could identify an active gene) and inactive genes (which gives a measure of specificity). We also calculated the F1 score as a combined measure of sensitivity and specificity. To assess the performance of the regression model, we used the root mean squared error (RMSE) for the predicted recombination frequencies compared to the observed values across all ten test sets. The RMSE gives a measure of the SD in errors, thus 68% of our predictions are expected to have an error within this range. Since our recombination frequencies are log2-transformed, an RMSE of <italic>x</italic> corresponds to a 2<italic><sup>x</sup></italic>-fold difference between the predicted recombination frequency and the observed recombination frequency.</p>
<p>For model selection, we focussed on the 16 most important features from the initial classification or regression model. We trained RF classification or regression models, with 10-fold cross-validation, for all possible combinations of the respective 16 features. These models were then compared using the performance metrics described above. Our analysis was performed using the R package randomForest (<xref ref-type="bibr" rid="B62">62</xref>).</p>
</sec>
<sec id="S2-10">
<title>Data Availability</title>
<p>Publicly available genome-wide datasets analysed during this study are available in the GEO repository; details including accession numbers are listed in Table <xref ref-type="table" rid="T1">1</xref>. All were downloaded from GEO as raw short-read files (SRA) and realigned to NCBIM37/mm9 using Bowtie (<xref ref-type="bibr" rid="B49">49</xref>) or Bowtie 2 (<xref ref-type="bibr" rid="B63">63</xref>). The V&#x003BA;J&#x003BA;-seq datasets generated in this study are available in the GEO repository with accession number GSE101606.<xref ref-type="fn" rid="fn5"><sup>5</sup></xref> Some of the quantitated data from this study is also provided in Table S1 in Supplementary Material.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Publicly available next generation sequencing datasets utilised in our study.</p></caption>
<table frame="hsides" rules="rows">
<thead>
<tr>
<th valign="top" align="left">Feature</th>
<th valign="top" align="left">Reference</th>
<th valign="top" align="left">Series</th>
<th valign="top" align="left">Accession(s)</th>
<th valign="top" align="left">Bowtie settings and notes</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Nuclear RNA</td>
<td align="left" valign="top">Bolland et al. (<xref ref-type="bibr" rid="B14">14</xref>)</td>
<td align="left" valign="top">GSE80155</td>
<td align="left" valign="top">GSM2113570</td>
<td align="left" valign="top">bowtie -n 0&#x02009;-m 1&#x02014;best&#x02014;strata&#x02014; maxins 1,000</td>
</tr>
<tr>
<td align="left" valign="top">DHS (DNaseI hypersensitivity)</td>
<td align="left" valign="top">Revilla-i-Domingo et al. (<xref ref-type="bibr" rid="B64">64</xref>)</td>
<td align="left" valign="top">GSE38046</td>
<td align="left" valign="top">GSM932968</td>
<td align="left" valign="top">Expt. 8,439 (1). bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">H3K4me1</td>
<td align="left" valign="top">Revilla-i-Domingo et al. (<xref ref-type="bibr" rid="B64">64</xref>)</td>
<td align="left" valign="top">GSE38046</td>
<td align="left" valign="top">GSM932934</td>
<td align="left" valign="top">Expt. 8,666 (1). bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">H3K4me2</td>
<td align="left" valign="top">Lin et al. (<xref ref-type="bibr" rid="B39">39</xref>)</td>
<td align="left" valign="top">GSE40173</td>
<td align="left" valign="top">GSM987804</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">H3K4me3</td>
<td align="left" valign="top">Bolland et al. (<xref ref-type="bibr" rid="B14">14</xref>)</td>
<td align="left" valign="top">GSE80155</td>
<td align="left" valign="top">GSM2113571, GSM2113573</td>
<td align="left" valign="top">bowtie -n 0&#x02009;-m 1&#x02014;best&#x02014;strata&#x02014;maxins 700</td>
</tr>
<tr>
<td align="left" valign="top">H3K9ac</td>
<td align="left" valign="top">Revilla-i-Domingo et al. (<xref ref-type="bibr" rid="B64">64</xref>)</td>
<td align="left" valign="top">GSE38046</td>
<td align="left" valign="top">GSM932943, GSM932944, GSM932945, GSM932946</td>
<td align="left" valign="top">Expts. 8,108, 8,113. bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">H3K27me3</td>
<td align="left" valign="top">Revilla-i-Domingo et al. (<xref ref-type="bibr" rid="B64">64</xref>)</td>
<td align="left" valign="top">GSE38046</td>
<td align="left" valign="top">GSM932947, GSM932948, GSM932949, GSM932950, GSM932951</td>
<td align="left" valign="top">Expts. 8,111, 8,116. bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">CTCF</td>
<td align="left" valign="top">Ebert et al. (<xref ref-type="bibr" rid="B65">65</xref>)</td>
<td align="left" valign="top">GSE27214</td>
<td align="left" valign="top">GSM672401</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">RAD21</td>
<td align="left" valign="top">Ebert et al. (<xref ref-type="bibr" rid="B65">65</xref>)</td>
<td align="left" valign="top">GSE27214</td>
<td align="left" valign="top">GSM672403</td>
<td align="left" valign="top">Bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">P300</td>
<td align="left" valign="top">Lin et al. (<xref ref-type="bibr" rid="B39">39</xref>)</td>
<td align="left" valign="top">GSE40173</td>
<td align="left" valign="top">GSM987808</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">PAX5</td>
<td align="left" valign="top">Revilla-i-Domingo et al. (<xref ref-type="bibr" rid="B64">64</xref>)</td>
<td align="left" valign="top">GSE38046</td>
<td align="left" valign="top">GSM932924</td>
<td align="left" valign="top">Expt. 8,417. bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">YY1</td>
<td align="left" valign="top">Medvedovic et al. (<xref ref-type="bibr" rid="B66">66</xref>)</td>
<td align="left" valign="top">GSE43008</td>
<td align="left" valign="top">GSM1145864</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">PU.1</td>
<td align="left" valign="top">Mullen et al. (<xref ref-type="bibr" rid="B67">67</xref>)</td>
<td align="left" valign="top">GSE21614</td>
<td align="left" valign="top">GSM539538</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">MED1</td>
<td align="left" valign="top">Whyte et al. (<xref ref-type="bibr" rid="B68">68</xref>)</td>
<td align="left" valign="top">GSE44288</td>
<td align="left" valign="top">GSM1038263</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014; strata</td>
</tr>
<tr>
<td align="left" valign="top">EBF1</td>
<td align="left" valign="top">Vilagos et al. (<xref ref-type="bibr" rid="B69">69</xref>)</td>
<td align="left" valign="top">GSE35857</td>
<td align="left" valign="top">GSM876622, GSM876623</td>
<td align="left" valign="top">bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">IRF4</td>
<td align="left" valign="top">Schwickert et al. (<xref ref-type="bibr" rid="B70">70</xref>)</td>
<td align="left" valign="top">GSE53595</td>
<td align="left" valign="top">GSM1296534</td>
<td align="left" valign="top">Bowtie2</td>
</tr>
<tr>
<td align="left" valign="top">E2A</td>
<td align="left" valign="top">Lin et al. (<xref ref-type="bibr" rid="B71">71</xref>)</td>
<td align="left" valign="top">GSE21978</td>
<td align="left" valign="top">GSM546523</td>
<td align="left" valign="top">Bowtie -m 1</td>
</tr>
<tr>
<td align="left" valign="top">BRG1</td>
<td align="left" valign="top">Bossen et al. (<xref ref-type="bibr" rid="B72">72</xref>)</td>
<td align="left" valign="top">GSE66978</td>
<td align="left" valign="top">GSM1635413, GSM1635414</td>
<td align="left" valign="top">Bowtie -m 1&#x02014;best&#x02014;strata</td>
</tr>
<tr>
<td align="left" valign="top">IKAROS</td>
<td align="left" valign="top">Bossen et al. (<xref ref-type="bibr" rid="B72">72</xref>)</td>
<td align="left" valign="top">GSE66978</td>
<td align="left" valign="top">GSM1635411, GSM1635414</td>
<td align="left" valign="top">Bowtie2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="S3">
<title>Results</title>
<sec id="S3-1">
<title>V&#x003BA;J&#x003BA;-Seq&#x02014;A High-Throughput Assay for Quantification of Recombined V&#x003BA; Gene Repertoires</title>
<p>To quantify the usage of V&#x003BA; genes in an unbiased way from DNA, we adapted our previously reported mouse <italic>Igh</italic> VDJ-seq assay (<xref ref-type="bibr" rid="B14">14</xref>) for the mouse <italic>Ig&#x003BA;</italic> locus (Figure S1 and Supplementary Text S1 in Supplementary Material). We generated V&#x003BA;J&#x003BA;-seq data for three biological replicates in wild-type (WT) bone marrow pre-B cells (B220<sup>&#x0002B;</sup>/CD19<sup>&#x0002B;</sup>/CD25<sup>&#x0002B;</sup>/IgM<sup>&#x02212;</sup>), and one replicate in pre-B cells from a <italic>Rag1</italic><sup>&#x02212;/&#x02212;</sup> mouse with a rearranged <italic>Igh</italic> transgene (<italic>VH81X</italic>) (<xref ref-type="bibr" rid="B47">47</xref>). These <italic>Rag1</italic><sup>&#x02212;/&#x02212;</sup><italic>/VH81X</italic> cells lack the RAG1 recombinase, precluding V(D)J recombination, but progression to the pre-B cell stage is permitted through expression of the <italic>VH81X</italic> transgene. Thus, they serve as a negative control in which reads mapping to V&#x003BA; genes, indicating a V&#x003BA;J&#x003BA; recombination event, should not be detected, giving a measure of the spurious incorporation of these reads into our libraries. Indeed, while 91.8% of unique, J&#x003BA; bait-associated reads for this library mapped upstream of unrecombined J&#x003BA; genes, only 9 reads (0.0002%) mapped to V&#x003BA; genes. Conversely, in WT pre-B cells over 30% of reads mapped to V&#x003BA; genes for all replicates. This equated to a total of 400&#x02013;530,000 unique V&#x003BA;J&#x003BA; recombined fragments for each replicate (Figure S2A in Supplementary Material), an order of magnitude greater than previous high-throughput assays of the V&#x003BA; gene repertoire (<xref ref-type="bibr" rid="B37">37</xref>, <xref ref-type="bibr" rid="B38">38</xref>). In the <italic>Rag1</italic><sup>&#x02212;/&#x02212;</sup><italic>/VH81X</italic> library, we noted a slight bias towards <italic>J&#x003BA;2</italic> (Figure S2B in Supplementary Material), suggesting minor preferential priming of the <italic>J&#x003BA;2</italic> gene. However, within the V&#x003BA;J&#x003BA; recombined fragments, over 35% of reads were associated with <italic>J&#x003BA;1</italic>, which usually recombines first (<xref ref-type="bibr" rid="B11">11</xref>), indicating that we are capturing a large proportion of the primary V&#x003BA; gene repertoire. Analysis of the sequences using IMGT/HighV-QUEST revealed that the ratio of productive:non-productive rearrangements was approximately 37:63 (Figure S2C in Supplementary Material), close to the expected two-thirds of non-productive rearrangements. Consistent with previous reports (<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B9">9</xref>, <xref ref-type="bibr" rid="B38">38</xref>), the vast majority of functional rearrangements had a CDR3 of 9 amino acids (Figure S2D in Supplementary Material).</p>
<p>Using normalised frequencies for each V&#x003BA;-J&#x003BA; gene combination, we clustered the dataset based on both V&#x003BA; and J&#x003BA; genes, taking each replicate separately. While V&#x003BA; genes with poor RIC scores clearly recombine infrequently, we did not observe any relationship between the repertoire of V&#x003BA; genes and their distance from, or orientation with respect to, the J&#x003BA; genes. For each J&#x003BA; gene, the three replicates were highly correlated (Pearson correlation coefficients &#x0003E;0.992) and clustered more closely to each other than to the repertoires of other J&#x003BA; genes, indicating that they are associated with distinct V&#x003BA; gene profiles (Figure <xref ref-type="fig" rid="F1">1</xref>). Notably, the pattern of recombination to <italic>J&#x003BA;1</italic> segregated from the other J&#x003BA; genes. This is consistent with the preferential usage of <italic>J&#x003BA;1</italic> in the generation of the primary repertoire, while the other J&#x003BA; genes are subsequently used for receptor editing (<xref ref-type="bibr" rid="B11">11</xref>). Since we aimed to assess the features driving recombination of the germline <italic>Ig&#x003BA;</italic> locus in the formation of the primary repertoire, we chose to focus on the <italic>J&#x003BA;1</italic> repertoire for further analyses (Table S1 in Supplementary Material).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Recombination frequencies for all J&#x003BA; genes and replicates. Heatmap showing log2-transformed recombination frequencies for each V&#x003BA;-J&#x003BA; combination across all replicates. Read counts for each replicate were first normalised to the replicate with the median number of reads aligning to V&#x003BA; genes. Each row represents a V&#x003BA; gene and each column represents an individual replicate for a J&#x003BA; gene. V&#x003BA; genes and J&#x003BA; gene replicates are clustered based on the similarity of their repertoires. The strand (&#x0002B;, pink; &#x02212;, blue), recombination signal sequence (RSS) Information Content (RIC) score (low&#x02009;&#x0003D;&#x02009;light; high&#x02009;&#x0003D;&#x02009;dark) and distance from J&#x003BA;1 (low&#x02009;&#x0003D;&#x02009;light; high&#x02009;&#x0003D;&#x02009;dark) for each V&#x003BA; gene are displayed on the left, and the colour of the V&#x003BA; gene label represents its family.</p></caption>
<graphic xlink:href="fimmu-08-01550-g001.tif"/>
</fig>
<p>The V&#x003BA;-J&#x003BA;1 repertoire varied widely across the locus, with no clear geographical pattern (Figure <xref ref-type="fig" rid="F2">2</xref>A). The RIC scores of V&#x003BA; genes from the same family were quite homogeneous, while their recombination frequencies could vary by more than 10-fold (Figures <xref ref-type="fig" rid="F2">2</xref>B,C). Comparing the families, we noted some patterns: for example, V&#x003BA;1 genes recombine quite frequently compared to several other families, even when their median RIC score was similar (e.g., V&#x003BA;2) or higher (e.g., V&#x003BA;4). For all genes with a RIC score &#x0003E;&#x02009;&#x02212;38.81, which are theoretically considered capable of recombination (<xref ref-type="bibr" rid="B12">12</xref>), usage of V&#x003BA; genes on the forward and reverse strands was not significantly different (Figure <xref ref-type="fig" rid="F2">2</xref>D), despite the significantly lower RIC scores of forward compared to reverse strand genes (Figure <xref ref-type="fig" rid="F2">2</xref>E). This contrasts with observations from an RNA-based assay (<xref ref-type="bibr" rid="B37">37</xref>) in which recombination to <italic>J&#x003BA;1</italic> was biased towards inversional rearrangements. Moreover, while 6 out of the 10 V&#x003BA; genes that were most frequently represented in their expressed <italic>J&#x003BA;1</italic> repertoire (<xref ref-type="bibr" rid="B37">37</xref>) are included in the top 20 of our V&#x003BA;-J&#x003BA;1 repertoire, 4 are not, and the DNA repertoire is not dominated by a small number of genes. Our assay also reveals numerous V&#x003BA; genes that are more highly represented at the DNA level, including 14 pseudogenes that were not detected in the RNA repertoire. Conversely, all V&#x003BA; genes present in the expressed repertoire were detected in our assay, albeit in some cases with very low read counts. This highlights the significant contribution of transcription and posttranscriptional processes to the expressed repertoire, which would confound the aim of this study to interrogate the pre-recombination chromatin state.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>V&#x003BA; recombination to J&#x003BA;1 varies widely across the locus. <bold>(A)</bold> Reads associated with J&#x003BA;1 for each V&#x003BA; gene were counted for all replicates, and normalised to the replicate with the median total number of J&#x003BA;1-associated V&#x003BA; reads. Bars represent the median normalised read count, while each replicate is displayed as a circle. Genes are arranged geographically, from the 5&#x02032; end of the locus (left) to the 3&#x02032; end (J-proximal end; right). Below, names and localisation on chromosome 6 of all actively recombining V&#x003BA; genes (Figure <xref ref-type="fig" rid="F3">3</xref>A) are shown, with genes on the forward strand (that recombine by deletion) displayed above the scale bar, and genes on the reverse strand (that recombine by inversion) displayed below. V&#x003BA; gene families are represented by colour, with pseudogenes (pg) displayed in grey. <bold>(B,C)</bold> Normalised median recombination frequencies to J&#x003BA;1 <bold>(B)</bold> and recombination signal sequence (RSS) Information Content (RIC) scores <bold>(C)</bold> are shown for all genes in each V&#x003BA; family. The 64 pseudogenes (PG) originate from 13 out of the 20 V&#x003BA; gene families. <bold>(D,E)</bold> Normalised median recombination frequencies <bold>(D)</bold> and RSS RIC scores <bold>(E)</bold> for genes on the forward and reverse strands. Only genes with a RIC score &#x0003E;&#x02009;&#x02212;38.81, which are considered theoretically capable of recombination, are included; <italic>p</italic>-values from two-sided <italic>t</italic>-tests are shown.</p></caption>
<graphic xlink:href="fimmu-08-01550-g002.tif"/>
</fig>
<p>To facilitate further investigation of the V&#x003BA;-J&#x003BA;1 repertoire, we performed a binomial test to distinguish V&#x003BA; genes that are significantly recombining (padj&#x02009;&#x0003C;&#x02009;0.01, Figure <xref ref-type="fig" rid="F3">3</xref>A; Table S1 in Supplementary Material). Out of 162 genes, 105 (64.8%) passed the binomial test and were labelled &#x0201C;active&#x0201D; to denote &#x0201C;actively recombining&#x0201D;; these genes were detected with a minimum of 59 reads, and included 15 pseudogenes, and are hereafter referred to as active genes. The remaining 57 genes had insufficient evidence of activity, with the median read count for each below 39, and were labelled &#x0201C;inactive.&#x0201D; These inactive genes included eight V&#x003BA; genes that are considered to be functional, suggesting that they contribute little to the primary repertoire. The usage of active genes was weakly correlated with the RIC score (<italic>R</italic>&#x02009;&#x0003D;&#x02009;0.42; Figure <xref ref-type="fig" rid="F3">3</xref>B); however, genes with a similar RIC score could recombine at markedly different frequencies. Moreover, some inactive genes have RIC scores that are comparable to those of active genes (Figure <xref ref-type="fig" rid="F3">3</xref>C). Importantly, a linear regression model revealed that only 17.7% of the variability in V&#x003BA; gene usage could be explained by RIC score alone (Figure <xref ref-type="fig" rid="F3">3</xref>B), highlighting the need to explore whether other mechanisms, such as chromatin features, contribute to shaping the repertoire.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Recombination signal sequence (RSS) Information Content (RIC) score can only partially explain variation in V&#x003BA; gene activity. <bold>(A)</bold> Distribution of V&#x003BA;-J&#x003BA;1 recombination frequencies for all 162 V&#x003BA; genes. A one-sided binomial test was used to gauge the significance of their recombination frequency, allowing each gene to be labelled as active (fdr adjusted <italic>p</italic>-value &#x0003C;0.01) or inactive. <bold>(B)</bold> Dependence of active genes&#x02019; recombination frequency on RSS RIC score. Linear regression model (dashed line) reveals that only 17.7% of the variation in recombination frequency of active genes can be explained by the RIC score. <bold>(C)</bold> RSS RIC scores of active and inactive V&#x003BA; genes; <italic>p</italic>-value from a two-sided Wilcoxon rank sum test.</p></caption>
<graphic xlink:href="fimmu-08-01550-g003.tif"/>
</fig>
</sec>
<sec id="S3-2">
<title>Chromatin Landscape of the <italic>Ig&#x003BA;</italic> Locus</title>
<sec id="S3-2-1">
<title>Colocalisation of Chromatin Features with V&#x003BA; Genes</title>
<p>In order to assess the contribution of chromatin features to V&#x003BA; gene recombination, we used published genome-wide datasets from mouse pro-B cell models that are developmentally stalled prior to recombination of the <italic>Igh</italic> locus (<xref ref-type="bibr" rid="B48">48</xref>, <xref ref-type="bibr" rid="B73">73</xref>). There are numerous pro-B cell datasets available, and the regulatory state of the <italic>Ig&#x003BA;</italic> locus has already begun to be established by this stage (<xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B74">74</xref>). Our analysis aims to determine the importance of these early regulatory events in priming the locus for recombination, thus shaping the primary repertoire. Moreover, <italic>Ig&#x003BA;</italic> locus gene-specific studies (<xref ref-type="bibr" rid="B75">75</xref>, <xref ref-type="bibr" rid="B76">76</xref>), as well as the small number of available pre-B cell datasets (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B77">77</xref>), revealed similar enrichment of CTCF, YY1, and histone H3 acetylation in pro-B and pre-B cells. The chromatin features we chose to assess included DHS, germline transcription, and ChIP for several histone modifications and TFs (Table <xref ref-type="table" rid="T1">1</xref>).</p>
<p>We first measured the distance from the centre of each V&#x003BA; gene to the summit of the closest peak for each DHS- and ChIP-seq dataset that had at least 35 peaks over the locus, both upstream (towards the promoter) and downstream (towards the RSS). Several TFs showed a bimodal distribution both up- and downstream of the V&#x003BA; genes. This was generally more pronounced for active V&#x003BA; genes, with peaks close to both promoters and RSSs (Figure <xref ref-type="fig" rid="F4">4</xref>A; Figure S3A in Supplementary Material). For some TFs, including PAX5 and IRF4, promoter-associated peaks were primarily located towards the 5&#x02032; end of the V&#x003BA; region, while PU.1 and IKAROS were also located close to promoters towards the 3&#x02032; end. Very few ChIP-seq peaks were found close to central V&#x003BA; gene promoters (Figure <xref ref-type="fig" rid="F4">4</xref>A). In contrast, RSS-associated peaks were located at V&#x003BA; genes throughout the locus. With the exception of PU.1, peaks were more frequently associated with V&#x003BA; gene promoters than with RSSs. We also noted that whilst RAD21 peaks directly mapped to only one promoter and one RSS, several peaks were located approximately 2&#x02009;kb upstream of V&#x003BA; gene promoters (Figure S3A in Supplementary Material).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Chromatin features associate with both the promoters and recombination signal sequences (RSSs) of active V&#x003BA; genes. <bold>(A)</bold> Scatter plots and density plots showing the log10-transformed distances of the closest ChIP-seq peaks both up- (&#x0002B;) and downstream (&#x02212;) from the centre of active (orange) and inactive (dark grey) V&#x003BA; genes. Yellow (promoter) and blue (RSS) shading indicates the range of distances within which 80% of the start and end sites of V&#x003BA; genes, respectively, are located. Grey shading indicates a distance of &#x0003E;1&#x02009;kb from the V&#x003BA; gene, based on the median V&#x003BA; gene length (525&#x02009;bp). <bold>(B)</bold> Average enrichment, relative to background, of chromatin features across all active (orange) and inactive (dark grey) V&#x003BA; genes. Genes have been scaled such that 0 and 100 represent the start and end of the gene, respectively. Yellow and blue shading indicates the location of the promoters and RSSs, respectively. <bold>(C,D)</bold> Median recombination frequencies and RSS Information Content (RIC) scores of active genes, excluding 10 genes that had &#x0003C;70% mappability. Genes are categorised based on the number of ChIP-seq peaks within 1&#x02009;kb of the gene <bold>(C)</bold>, or the localisation of those peaks upstream (promoter) or downstream (RSS) of the gene centre <bold>(D)</bold>. <italic>p</italic>-values for V&#x003BA;-J&#x003BA;1 recombination frequency are fdr adjusted, based on a two-sided Wilcoxon rank sum test. <italic>n</italic> values indicate the number of genes in each category.</p></caption>
<graphic xlink:href="fimmu-08-01550-g004.tif"/>
</fig>
<p>The localisation of TF peaks close to both the promoters and RSSs prompted us to examine the distribution of chromatin features over the V&#x003BA; genes in more detail, considering the overall enrichment of each feature without the threshold applied in peak calling. All TFs were found to be enriched over V&#x003BA; gene promoters, and most were enriched over RSSs, while the distribution of histone modifications was more variable (Figure <xref ref-type="fig" rid="F4">4</xref>B; Figure S3B in Supplementary Material). Importantly, with the exception of H3K27me3 and CTCF, the enrichment of all of these chromatin features was greater over active compared to inactive genes. Active genes with associated ChIP-seq peaks tended to recombine more frequently, despite having poorer quality RIC scores, than those without (Figure <xref ref-type="fig" rid="F4">4</xref>C), although this was not significant. Genes with peaks close to both the promoter and the RSS recombined with significantly greater frequency than genes with no associated peaks (<italic>p</italic>&#x02009;&#x0003D;&#x02009;0.0061; Wilcoxon rank sum test) or with peaks that were only associated with the promoter (<italic>p</italic>&#x02009;&#x0003D;&#x02009;0.0014) or RSS (<italic>p</italic>&#x02009;&#x0003D;&#x02009;0.0246; Figure <xref ref-type="fig" rid="F4">4</xref>D). This suggests that both the promoter and the RSS are important in facilitating efficient recombination. This is in contrast to the <italic>Igh</italic> locus, in which TF enrichment is almost exclusively confined to the V<sub>H</sub> gene RSSs (<xref ref-type="bibr" rid="B14">14</xref>), indicating that the mechanisms that regulate V(D)J recombination differ between V<sub>H</sub> and V&#x003BA;.</p>
</sec>
<sec id="S3-2-2">
<title>Chromatin Segmentation of the <italic>Ig&#x003BA;</italic> Locus</title>
<p>In order to shed further light on how chromatin features contribute to V&#x003BA; gene recombination, we investigated the regulatory landscape of the <italic>Ig&#x003BA;</italic> locus with EpiCSeg (<xref ref-type="bibr" rid="B58">58</xref>). This algorithm employs a multivariate Hidden Markov Model to integrate genome-wide datasets and segment a given genomic locus into characteristic chromatin states. We used read counts over 200&#x02009;bp bins covering the locus for each DHS- and ChIP-seq dataset as the input and ran the algorithm specifying an output of between 3 and 15 states.</p>
<p>Despite the complexity of the locus, we observed that within-class homogeneity and between-class heterogeneity is maximised with just three states (Figures <xref ref-type="fig" rid="F5">5</xref>A,B). This number and the characteristic attributes of the states are strikingly similar to our previous analysis of the <italic>Igh</italic> locus (<xref ref-type="bibr" rid="B14">14</xref>). Accordingly, we labelled these states as follows: a &#x0201C;Background&#x0201D; (Bg) state, which comprises most of the locus and shows little enrichment for any chromatin features; an &#x0201C;Architectural&#x0201D; (A) state, in which CTCF and RAD21 are enriched; and an &#x0201C;Enhancer&#x0201D; (E) state, which is enriched for several TFs and histone modifications, including PU.1, PAX5, IRF4, MED1, IKAROS, and BRG1 (Figures <xref ref-type="fig" rid="F5">5</xref>A,C). We note that our choice of three states is subjective. Running the algorithm with a higher number of states results in segregation into smaller sub-states, which display a low enrichment for a subset of the features enriched for in the A or E states, and are frequently adjacent to similar states (shown for 4&#x02013;8 states in Figure S4 in Supplementary Material). This suggests that they are not distinct from the A and E state, and that the three-state model is the most appropriate.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Chromatin state analysis of the <italic>Ig&#x003BA;</italic> locus. <bold>(A,B)</bold> Feature enrichment <bold>(A)</bold> and transition matrix <bold>(B)</bold> for states identified by the EpiCSeg algorithm when three states are specified. <bold>(C)</bold> Proportion of the <italic>Ig&#x003BA;</italic> locus that is in each state. State 1 was labelled as the &#x0201C;Enhancer&#x0201D; state (E State), state 2 as the &#x0201C;Architectural&#x0201D; state (A state), and state 3 as the &#x0201C;Background&#x0201D; state (Bg state). <bold>(D)</bold> Percentage of regions in the A, E, and Bg state, over and surrounding all V&#x003BA; genes, active and inactive V&#x003BA; genes separately, or random regions of equivalent size distribution. V&#x003BA; genes have been scaled to the median gene length (525&#x02009;bp), while distances surrounding the genes are not scaled. <bold>(E)</bold> Number of genes in each state, and the locations at which those states are present. <bold>(F,G)</bold> Enrichment relative to background of chromatin features characteristic of the E state over V&#x003BA; recombination signal sequences (RSSs) <bold>(F)</bold> and promoters <bold>(G)</bold> associated with each state. Fdr-adjusted <italic>p</italic>-values from a two-sided Wilcoxon rank sum test are shown for the difference in enrichment between A and E state-associated promoters or RSSs. All data are included for statistical testing, but to better visualise the data, some outliers are not displayed.</p></caption>
<graphic xlink:href="fimmu-08-01550-g005.tif"/>
</fig>
<p>When we examined the distribution of these states over the V&#x003BA; genes, we found that the E state was highly enriched over both the gene promoters and RSSs, but depleted elsewhere (Figure <xref ref-type="fig" rid="F5">5</xref>D). The A state displayed only slight enrichment over the V&#x003BA; genes, but was more broadly enriched in a region approximately 1&#x02013;3&#x02009;kb upstream of the genes. These patterns of enrichment were particularly striking when only actively recombining V&#x003BA; genes were considered, whilst inactive genes were almost exclusively enriched in the Bg state. We identified a total of 92 V&#x003BA; genes that were associated with only the E state, at the promoter of 26 genes, at the RSS of 35 genes, and at both of these regions of 31 genes (Figure <xref ref-type="fig" rid="F5">5</xref>E). Only 11 genes were associated exclusively with the A state, while 10 genes were associated with both the A and the E state. This distribution of states is in contrast to the <italic>Igh</italic> locus, in which we observed association only with the V<sub>H</sub> gene RSSs, and moreover, the two states were mutually exclusive, that is, no V<sub>H</sub> genes were associated with both the A and the E state (<xref ref-type="bibr" rid="B14">14</xref>). Features associated with the E state at the <italic>Ig&#x003BA;</italic> locus, including PAX5, PU.1, IRF4, IKAROS, MED1 BRG1, E2A, and H3K4me2, were more enriched over E state compared to A state RSSs and promoters (Figures <xref ref-type="fig" rid="F5">5</xref>F,G). Median enrichment of CTCF and RAD21 was higher over A state promoters and RSSs, although these differences were not significant (Figure S5 in Supplementary Material).</p>
<p>The distribution of these three states was significantly different over active versus inactive genes, with a much lower proportion of active genes in the Bg state. 83% of active V genes exhibited an E, A, or A/E chromatin state, compared with only 46% of inactive V genes (Figure <xref ref-type="fig" rid="F6">6</xref>A; Table S2 in Supplementary Material); the majority of these inactive genes had a RIC score below &#x02212;25 (the lowest score for any active gene). Moreover, active genes marked by the A and/or E state recombine with significantly greater frequency than active genes in the Bg state (Figure <xref ref-type="fig" rid="F6">6</xref>B). When we compared the localisation of these states, the median recombination frequency of genes marked by the A and/or E state at either the promoter or the RSS was higher than that of Bg genes. The highest frequency was observed for genes marked in both regions, which was significantly different from Bg genes (Figure <xref ref-type="fig" rid="F6">6</xref>C). This is consistent with our analyses of individual ChIP-seq peaks above, suggesting that the presence of an active chromatin state at the RSS is more important in facilitating high levels of recombination than is an active state at the promoter. However, active chromatin states at both locations is particularly conducive to recombination.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>V&#x003BA; gene chromatin state is associated with recombination frequency. <bold>(A)</bold> Number of active and inactive genes in each state. <italic>p-</italic>value based on a Fisher&#x02019;s exact test. <bold>(B,C)</bold> Violin plots with boxplot superimposed (black) showing median recombination frequencies of active genes in the Bg state compared to active genes associated with the A and/or E state. A and/or E state genes are considered altogether <bold>(B)</bold>, or categorised based on the localisation of the state to their promoter or recombination signal sequence (RSS) <bold>(C)</bold>. 10 genes that had &#x0003C;70% mappability were excluded. Fdr-adjusted <italic>p</italic>-values based on two-sided Wilcoxon rank sum test. <bold>(D)</bold> Proportion of each V&#x003BA; gene promoter and RSS (each window extending from the gene centre to 500&#x02009;bp up/downstream of the gene, respectively; median window size 762&#x02009;bp) and up- and downstream regions (from 500 to 3,500&#x02009;bp up/downstream of the gene) that are assigned to each of the three states. Each bar represents an individual V&#x003BA; gene. Note that up- and downstream windows represent a genomic region approximately four times the size of the promoter and RSS windows. Genes are ordered based on their recombination frequency, with highly recombining genes on the right (denoted by red shading on the recombination frequency scale). V&#x003BA; gene family and RSS Information Content (RIC) score are indicated to the left and above. <bold>(E)</bold> Top: phylogenetic tree of reference C57BL/6 V&#x003BA; gene sequences, collapsed at the nodes containing the majority of each V&#x003BA; gene family. The number of active genes in each family is shown. All nodes comprise a single V&#x003BA; gene family except for one node, the V&#x003BA;9 family branch, which also includes V&#x003BA;14-130. Bottom: proportion of active genes in each family that are associated with the A state, E state, or with both states at their promoters, RSSs and up/downstream regions [colours as in <bold>(A)</bold>].</p></caption>
<graphic xlink:href="fimmu-08-01550-g006.tif"/>
</fig>
<p>Since we had observed an enrichment of the A state upstream of active V&#x003BA; genes, we also undertook a broader analysis of the states located close to each individual V&#x003BA; gene. As expected from our earlier analyses (Figures <xref ref-type="fig" rid="F6">6</xref>A&#x02013;C), V&#x003BA; genes that recombine more frequently displayed a greater overlap with the E state in particular; this was apparent at both the promoter and the RSS (Figure <xref ref-type="fig" rid="F6">6</xref>D). Conversely, there was no clear relationship between recombination frequency and up- or downstream states, although this does not exclude a more complex influence of the surrounding chromatin on V&#x003BA; gene recombination.</p>
<p>We also compared the enrichment of chromatin states over a phylogenetic tree of V&#x003BA; gene families (Figure <xref ref-type="fig" rid="F6">6</xref>E). Considering only active genes, we observed some patterns in the state association of related families. For example, a large proportion of genes within the closely related V&#x003BA;4, V&#x003BA;17, V&#x003BA;8, and V&#x003BA;6 families (including in the two largest families, V&#x003BA;4 and V&#x003BA;8) are associated with the E state at their RSS; the A state was also frequently located upstream of these genes. However, unlike in the <italic>Igh</italic> locus (<xref ref-type="bibr" rid="B14">14</xref>), there was no clear evolutionary separation between genes associated with the A versus the E state. Indeed, the E state was also frequently associated with genes in the V&#x003BA;11, V&#x003BA;10, V&#x003BA;19, and V&#x003BA;12 families, which are closely related to each other but not to the families above.</p>
<p>We also asked whether the pre-recombination chromatin states at the V&#x003BA; gene promoters analysed here might in part explain the differences in the DNA V&#x003BA;-J&#x003BA;1 repertoire compared to the expressed repertoire (<xref ref-type="bibr" rid="B37">37</xref>), since these states may remain after recombination and additionally contribute to RNA expression. While 15 out of the top 20 most highly represented genes in the expressed repertoire were marked by either the A or E state at their promoters, 13 out of 20 V&#x003BA; genes that were present only in the DNA repertoire or that were highly represented in the DNA but had a low representation in the expressed repertoire, had promoters in the Bg state (Figure S6 in Supplementary Material). Thus, the chromatin state of the gene promoter can explain some of the differences in the repertoire. This underscores the value of using the DNA repertoire for these analyses, both to prevent the masking of the true recombination potential of each gene at the DNA level, and to ensure that conclusions drawn about the importance of chromatin features at the V&#x003BA; gene promoters prior to recombination are not confounded by differing expression levels.</p>
</sec>
</sec>
<sec id="S3-3">
<title>RSS RIC Score and PU.1 Binding Are Key Features That Distinguish Actively Recombining Genes from Inactive Genes</title>
<p>We next sought to understand how genetic and chromatin features regulate V&#x003BA; gene usage. Thus, we trained a RF-C to assess the power of each feature to correctly predict whether a gene is active or inactive; we have previously used this approach for analysis of <italic>Igh</italic> recombination (<xref ref-type="bibr" rid="B14">14</xref>). The RF-C takes features relating to each sample (here, each V&#x003BA; gene), and generates a large number of decision trees that vote on the response (here, V&#x003BA; gene recombination activity) for a training set of samples. At each step, one feature out of a random subset can be chosen, such that each tree will be unique. Feature importance is gauged by comparing the prediction accuracy for trees that include or exclude a given feature. The overall accuracy is then assessed by predicting the activity of an independent test set of genes. We chose the RF approach because it performs well both with a large number of features relative to the number of samples, and with highly correlated features (<xref ref-type="bibr" rid="B61">61</xref>).</p>
<p>We considered the signal intensity over the promoter separately from the signal intensity over the RSS, taking windows from the centre of each gene extending 500&#x02009;bp upstream (promoter) or downstream (RSS) of the gene. We also calculated the signal intensities in windows extending a further 2.5&#x02009;kb upstream of the promoter and downstream of the RSS (Figure <xref ref-type="fig" rid="F7">7</xref>A). In addition, we systematically gauged the contribution of the orientation of each V&#x003BA; gene and their genomic distance from <italic>J&#x003BA;1</italic> to facilitating recombination. Thus, the input for each gene included three genetic features: the RIC score, strand, and distance from <italic>J&#x003BA;1</italic>; and four separate features for each chromatin dataset (hereafter referred to as promoter, RSS, upstream and downstream).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Random forest classifier (RF-C) distinguishes active from inactive genes. <bold>(A)</bold> Locations of the four windows in which chromatin features were quantified as an input to the random forest models. <bold>(B)</bold> Relative importance of all features in distinguishing active from inactive genes, shown as the average out-of-bag gini impurity (a measure of the decrease in accuracy if the feature is excluded) in a RF-C, with 10-fold cross-validation. Error bars show the SEM. Inset: zoomed-in view of the features ranked from 2 to 30 in importance. Features are coloured based on whether they are a genetic feature, or the location in which the chromatin feature was measured. <bold>(C)</bold> Overall prediction accuracy of all RF-C models trained with all combinations of the 16 most important features in the initial RF-C. Colour denotes whether the RSS Information Content (RIC) score was included in the model: all models that included the RIC score were more accurate than those that did not. <bold>(D)</bold> Accuracy in correctly predicting active genes versus inactive genes in all RF-C models. <bold>(E)</bold> Overall prediction accuracy for models that included PU.1 binding at the RSS as a feature, compared to those that did not.</p></caption>
<graphic xlink:href="fimmu-08-01550-g007.tif"/>
</fig>
<p>A 10-fold cross-validation approach including all features revealed a mean prediction accuracy of 93.3% (SD 5.3%), with an F1 score of 0.949 (SD 0.0405). The prediction accuracy for active genes (97.2%, SD 6.2%) was better than that for inactive genes (85.7%, SD 14.5%), indicating a high sensitivity but slightly lower specificity in detecting active genes. The RIC score was by far the most important feature in distinguishing active from inactive genes (Figure <xref ref-type="fig" rid="F7">7</xref>B). The second most important feature was PU.1, which contributes to activation and recombination of the <italic>Ig&#x003BA;</italic> locus (<xref ref-type="bibr" rid="B21">21</xref>). Our model specifically suggests that while PU.1 binding at the promoter is of no consequence, its binding at the RSS is an important driver of recombination.</p>
<p>H3K4me2 enrichment within both the promoter and the RSS windows was also important for prediction accuracy, in addition to MED1 binding at the promoter and a number of E state-associated chromatin features at the RSS (Figure <xref ref-type="fig" rid="F7">7</xref>B). Notably, PAX5, CTCF, and RAD21, which are key drivers of <italic>Igh</italic> recombination (<xref ref-type="bibr" rid="B14">14</xref>), were not identified as important in promoting V&#x003BA; gene activity. Several up- or downstream features, such as DHS upstream and IRF4 binding downstream, were, however, ranked quite highly (Figure <xref ref-type="fig" rid="F7">7</xref>B), suggesting that in addition to the chromatin state of the V&#x003BA; genes themselves, the surrounding chromatin also influences the capacity of each gene to recombine.</p>
<p>Next, we performed a model selection analysis, considering all possible combinations of the 16 most important features. RF-C models that included the RIC score were all more accurate than those that did not (Figure <xref ref-type="fig" rid="F7">7</xref>C). This clear distinction was driven by the much lower prediction accuracy for inactive genes in RF-C models in which the RIC score was excluded (Figure <xref ref-type="fig" rid="F7">7</xref>D). There was also a striking contribution of PU.1 binding at the RSS to prediction accuracy, particularly in models that excluded the RIC score (Figure <xref ref-type="fig" rid="F7">7</xref>E); indeed, the bimodal distribution in prediction accuracy observed for models excluding the RIC score appears to be primarily dependent on this feature. This was apparent for both active and inactive genes (Figures S7A,B in Supplementary Material). We also observed a slight shift in prediction accuracy for models that included several other important features, such as IKAROS enrichment at the RSS (data not shown).</p>
</sec>
<sec id="S3-4">
<title>Chromatin Features Alone Are Highly Predictive of the Recombination Frequency of Active V&#x003BA; Genes</title>
<p>While the RF-C approach identifies the features that are most highly predictive of active V&#x003BA; gene recombination, and thus likely facilitate recombination, it does not directly show whether their levels contribute to the frequency with which an active gene is used. To address this, we trained a RF-R model for active V&#x003BA; genes, using the same set of features used for the classification model. This extends beyond the RF-C approach, giving a numerical prediction of the recombination frequency, with the log2-transformed recombination frequencies of active genes as the response variable. We used the RMSE to measure model performance. An initial RF-R using all features, with 10-fold cross-validation, achieved an RMSE of 1.57, indicating that 68% of the models&#x02019; predictions fall within 2<sup>1.57</sup>-fold (since recombination values are log2-transformed) of the observed recombination. This revealed the RIC score to be the most important feature for predicting the frequency of active V&#x003BA; gene recombination, as in the RF-C model (Figure <xref ref-type="fig" rid="F8">8</xref>A). However, there was far less separation between the importance of the RIC score and the other features. PU.1 binding at the RSS was much lower in importance compared to the RF-C model, while several features that were unimportant for classification contribute much more to the prediction of recombination frequency: these include PAX5 and CTCF binding upstream, and H3K4me3 and E2A binding at the RSS.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Random forest regression (RF-R) model for predicting recombination frequency of active V&#x003BA; genes. <bold>(A)</bold> Relative importance of all features in a RF-R model. Importance assessed by average out-of-bag node purity (a measure of the decrease in accuracy if the feature is excluded), with 10-fold cross validation. Error bars indicate the SEM. <bold>(B)</bold> Model selection, assessing all combinations of the 16 most important features from the initial RF-R model. Model performance is assessed as the root mean squared error (RMSE) across all test sets with 10-fold cross-validation. Colour denotes whether the RSS Information Content (RIC) score was included in the model as a feature. Top: scatter plot showing the RMSE of models with varying numbers of features; Bottom: density plot showing the distribution of the RMSE for all models that include or exclude the RIC score. <bold>(C,E)</bold> Observed versus predicted recombination frequencies across all test sets for the optimum RF-R model that included <bold>(C)</bold> or did not include <bold>(E)</bold> the RIC score as a feature, as assessed by RMSE. Observed&#x02009;&#x0003D;&#x02009;predicted is shown as a red line. <bold>(D,F)</bold> Frequency of inclusion of each feature in all models that included the RIC score and had an RMSE &#x0003C;1.39 <bold>(D)</bold>, or that excluded the RIC score and had an RMSE &#x0003C;1.48 <bold>(F)</bold>. <bold>(G,H)</bold> RMSE across all test sets for RF-R models that included H3K4me3 RSS <bold>(G)</bold> or MED1 promoter <bold>(H)</bold> as a feature compared to those that did not.</p></caption>
<graphic xlink:href="fimmu-08-01550-g008.tif"/>
</fig>
<p>We next performed a model selection analysis, considering the top 16 features, to identify a minimum subset that best predict V&#x003BA; gene usage. While several combinations of 6&#x02013;9 features have similarly low RMSE (Figure <xref ref-type="fig" rid="F8">8</xref>B), the minimum RMSE of 1.36 (Figure <xref ref-type="fig" rid="F8">8</xref>C) is achieved with a combination comprising: RIC score, H3K4me2, and H3K4me3 within the RSS window, IKAROS binding at both the promoter and the RSS, and PAX5 and PU.1 binding upstream. These seven features, in addition to H3K4me2 within the promoter window, were also the most highly represented in all models that had a low RMSE (&#x0003C;1.39), with the RIC score and H3K4me3 at the RSS being present in all models (Figure <xref ref-type="fig" rid="F8">8</xref>D). While this does not necessarily mean that these features are the most important individually (compare to Figure <xref ref-type="fig" rid="F8">8</xref>A), it suggests that together they are able to explain the largest proportion of the variability in our data. In models that excluded the RIC score, the best combination had an RMSE of 1.44 (Figure <xref ref-type="fig" rid="F8">8</xref>E), and comprised H3K4me3 at the RSS, IKAROS binding at both the promoter and the RSS, and PAX5 and PU.1 binding upstream, in addition to MED1 binding at the promoter. These six features, in addition to CTCF binding upstream, were the most frequently represented in models with low RMSE (&#x0003C;1.48) that excluded the RIC score, with MED1 binding to the promoter in all models (Figure <xref ref-type="fig" rid="F8">8</xref>F); this is particularly noteworthy since this feature was rarely present in the best combinations that included the RIC score. The inclusion of H3K4me3 at the RSS in a model was accompanied by a shift towards lower RMSE; this shift was more pronounced for the left-hand peak, corresponding to models that also include the RIC score (Figure <xref ref-type="fig" rid="F8">8</xref>G). Conversely, a shift towards lower RMSE for models that included MED1 binding at the promoter was only evident for the right-hand peak, representing models excluding the RIC score (Figure <xref ref-type="fig" rid="F8">8</xref>H). We also noted a shift towards a lower RMSE for models that included several other important features, including IKAROS binding at the promoter or RSS, or PAX5 or PU.1 binding upstream (Figure S7C in Supplementary Material).</p>
<p>These analyses revealed that we are able to predict the recombination frequency of V&#x003BA; genes from a combination of 6 or 7 features with a mean error rate of less than threefold, even when the RIC score is excluded (2.71-fold compared to 2.57-fold when including the RIC score). This model performance is highly significant when noting the variability of greater than 150-fold in recombination frequency across all active V&#x003BA; genes.</p>
<p>In order to further dissect the influence of the chromatin features shown to be important in our RF models, we split all active genes with a good RIC score (&#x0003E;&#x02212;14) into three groups based on their recombination frequency, and examined the enrichment of features of interest at the location(s) in which they were important (Figure <xref ref-type="fig" rid="F9">9</xref>). Genes that recombine more frequently tended to display higher levels of active histone modifications and TF binding at both the RSS (Figure <xref ref-type="fig" rid="F9">9</xref>A) and the promoter (Figure <xref ref-type="fig" rid="F9">9</xref>B). These trends were particularly striking for H3K4me2 (RSS and promoter windows) and IKAROS (RSS), in addition to IRF4 and E2A binding at the RSS, all of which displayed a significant positive relationship between recombination and enrichment. For PU.1 binding at the RSS, which was highly important in distinguishing active from inactive genes, the difference between genes that recombine at low and high frequency was more subtle (<italic>p</italic>&#x02009;&#x0003E;&#x02009;0.05), consistent with the lower importance of this feature for predicting the frequency of recombination. We also noted a subtle, but non-significant, trend towards higher enrichment of MED1 and IKAROS at the promoter, and H3K4me3 at the RSS, of more highly recombining genes: the importance of these features in the RF-R thus suggests that a more complex relationship exists between these features and recombination frequency. Conversely, we noted a slight, negative association between the recombination frequency and the binding of some TFs upstream of the gene, including PAX5, PU.1, and CTCF, with a significantly greater enrichment of PAX5 upstream of genes that recombine at a low level compared to those that recombine at a medium level (Figure <xref ref-type="fig" rid="F9">9</xref>C).</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Relationship between ChIP enrichment and recombination frequency for important RSS <bold>(A)</bold>, promoter <bold>(B)</bold>, and upstream <bold>(C)</bold> features in RF models. Enrichment of chromatin features over the locations in which they were found to be important (projected between 0 and 1 for each gene: identical to the input for RF-R models), for active genes with low (<italic>n</italic>&#x02009;&#x0003D;&#x02009;23; 117&#x02013;973 reads), medium (<italic>n</italic>&#x02009;&#x0003D;&#x02009;24; 1,006&#x02013;1,801 reads) and high (<italic>n</italic>&#x02009;&#x0003D;&#x02009;24; 1,880&#x02013;9,137 reads) relative frequency of recombination. Only genes with a high quality RIC score (&#x0003E;&#x02009;&#x02212;14) were considered. Fdr-adjusted <italic>p</italic>-values driven by two-sided Wilcoxon rank sum test. All data are included for statistical testing, but to better visualise the data, some outliers are not displayed.</p></caption>
<graphic xlink:href="fimmu-08-01550-g009.tif"/>
</fig>
</sec>
</sec>
<sec id="S4" sec-type="discussion">
<title>Discussion</title>
<p>We have adapted the VDJ-seq assay for the <italic>Ig&#x003BA;</italic> locus to quantitatively profile the V&#x003BA;-J&#x003BA; repertoire and to enable an in-depth analysis of the local drivers of recombination. Using cutting-edge random forest machine learning approaches to integrate genetic and chromatin features, we have distinguished genes that are actively recombining from those that are not, and have predicted the relative usage of active V&#x003BA; genes in primary recombination. We have found that local chromatin features, including PU.1 and IKAROS binding, and H3K4 methylation, explain much of the variation in recombination among V&#x003BA; genes.</p>
<p>The accuracy with which we can predict both V&#x003BA; gene activity and frequency of usage, even when the influence of the RIC score is excluded, is striking. Since we used pro-B cell genome-wide datasets, focussing on early events that prime the <italic>Ig&#x003BA;</italic> locus for recombination, the regulatory status of the <italic>Ig&#x003BA;</italic> locus may not fully reflect its state in pre-B cells immediately prior to V&#x003BA;-J&#x003BA;1 recombination. This suggests that early priming events are crucially important and that to a large extent, the recombination potential of each V&#x003BA; gene has been established by the pro-B cell stage. Nevertheless, we cannot exclude the possibility that features ranked unimportant here may become enriched at the <italic>Ig&#x003BA;</italic> locus later in development, or that additional pre-B cell specific features including IRF8, AIOLOS, and BRWD1 (<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B78">78</xref>, <xref ref-type="bibr" rid="B79">79</xref>) may play a local regulatory role in recombination. Profiling of the locus in a <italic>Rag<sup>&#x02212;/&#x02212;</sup></italic> model with a rearranged V<sub>H</sub>DJ<sub>H</sub> transgene would reveal pre-B cell developmental activation signatures that might predict V&#x003BA; gene usage with even greater accuracy.</p>
<p>Our recent analysis of the <italic>Igh</italic> V<sub>H</sub>DJ<sub>H</sub> repertoire (<xref ref-type="bibr" rid="B14">14</xref>) identified two mutually exclusive chromatin states, characterised by PAX5/IRF4 (enhancer/E state) or CTCF/RAD21 (architectural/A state) binding. Both localised exclusively to the RSSs of active V<sub>H</sub> genes, and their characteristic features were highly predictive of active recombination. We found striking similarities in the regulation of the <italic>Ig&#x003BA;</italic> locus, but with several important differences. While the two chromatin states at the <italic>Ig&#x003BA;</italic> locus were similar to those at the <italic>Igh</italic> locus, the E state predominates at V&#x003BA; genes. Moreover, the states were associated with both promoters and RSSs of V&#x003BA; genes; indeed both regions were represented within the most important features identified by our RF models. The individual chromatin features that we identified as important in driving <italic>Ig&#x003BA;</italic> recombination (e.g., PU.1) were also substantially different from those driving <italic>Igh</italic> recombination. Furthermore, our RF-R model, which was not used to assess <italic>Igh</italic> recombination, allowed us to take this analysis to the next level, giving a numerical prediction of recombination frequency. This approach allowed us to distinguish chromatin features that play a binary, all-or-nothing role in V&#x003BA; recombination from those that fine-tune the repertoire, shaping the frequency with which active genes will recombine.</p>
<p>Our RF-C model identified the RIC score as the most important feature for distinguishing actively recombining genes from inactive genes. Nevertheless, we achieved greater than 80% prediction accuracy based purely on chromatin features. This was primarily dependent on PU.1 binding at the RSS. PU.1 binding to the <italic>Ig&#x003BA;</italic> 3&#x02032; enhancer has been implicated in activation and recombination of the <italic>Ig&#x003BA;</italic> locus (<xref ref-type="bibr" rid="B23">23</xref>). In addition, Batista and colleagues observed frequent binding of PU.1 at V&#x003BA; RSSs and hypothesised that this may play a role in recruiting RAG enzymes for recombination (<xref ref-type="bibr" rid="B21">21</xref>). Our analysis provides direct mechanistic insight, revealing that PU.1 binding at the RSS of V&#x003BA; genes is a critical binary switch, which dictates whether that V&#x003BA; gene will recombine or not. Interestingly, PU.1 was not important in <italic>Igh</italic> V<sub>H</sub>DJ<sub>H</sub> recombination (<xref ref-type="bibr" rid="B14">14</xref>); conversely, CTCF, RAD21 and PAX5 binding, which were critical for <italic>Igh</italic> recombination, were unimportant in our <italic>Ig&#x003BA;</italic> RF-C model. We found IRF4 and DHS to be significant for recombination at both Ig loci. While H3K4 methylation featured strongly in both models, monomethylation was more prominent for V<sub>H</sub> genes in the <italic>Igh</italic> locus, while dimethylation was more important in shaping the V&#x003BA;J&#x003BA;1 repertoire. Thus, the local regulation of <italic>Igh</italic> and <italic>Ig&#x003BA;</italic> by histone modifications and TFs differs considerably. Features associated with the A state in particular appear to play a less significant role in priming the V&#x003BA; genes for recombination.</p>
<p>The RIC score was also the most important feature in our RF-R model. However, the distinction between it and the most important chromatin features was much lower than for the RF-C model, suggesting that similar to the <italic>Igh</italic> locus, the RIC score functions primarily as a binary switch that can be permissive or non-permissive to recombination.</p>
<p>Surprisingly, the influence of PU.1 binding to the RSS also appears to be binary: while it was key to the RF-C model, its importance in the RF-R model was much lower. Rather, the features with the greatest importance in predicting the frequency of recombination in the RF-R model included IKAROS, MED1, IRF4, and H3K4 di- and tri-methylation at the promoter and/or RSS. Moreover, at more frequently recombining genes, significant trends towards higher levels of enrichment were observed for H3K4me2, IKAROS, IRF4, and E2A. This suggests that both the binding and level of enrichment of these features are crucial for modulating the frequency with which each gene recombines, shaping the greater than 150-fold variation in active V&#x003BA; gene usage in the primary repertoire. While each of these have previously been implicated in promoting <italic>Ig&#x003BA;</italic> locus recombination (<xref ref-type="bibr" rid="B18">18</xref>, <xref ref-type="bibr" rid="B22">22</xref>, <xref ref-type="bibr" rid="B24">24</xref>, <xref ref-type="bibr" rid="B25">25</xref>, <xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B75">75</xref>), our findings provide mechanistic insight into their specific roles in shaping the repertoire through their localisation to individual V&#x003BA; genes. First, their locations at the promoter/RSS provide an additional layer of regulation beyond previously observed long-range interactions. Second, the correlation of higher levels of enrichment with higher recombination, measured here in bulk populations, suggests that these features localise to the relevant V&#x003BA; genes in a higher proportion of individual cells, or remain associated with these V&#x003BA; genes for longer, with a functional outcome of increased recombination.</p>
<p>A caveat of the RF approach is that it does not directly show how a given feature is related to V&#x003BA; gene recombination; the finite number of genes also means that over-fitting of the data could be a concern, although the use of 10-fold cross-validation mitigates this possibility. Nevertheless, the clear relationship that we observed between enrichment and recombination for several features, including IKAROS and IRF4, highlights the value of this approach in providing a shortlist of chromatin features that are potential drivers of recombination. The contribution of other features that did not display such a clear relationship with recombination, such as H3K4me3 at the RSS and MED1 binding at the promoter, will require further work to elucidate their roles. H3K4me3 binds and activates RAG2 (<xref ref-type="bibr" rid="B17">17</xref>&#x02013;<xref ref-type="bibr" rid="B20">20</xref>), suggesting a direct role in recruitment of the RAG complex. However, to our knowledge, this is the first time that MED1 has been implicated in recombination of the <italic>Ig&#x003BA;</italic> locus.</p>
<p>Notably, we did not observe a significant contribution of non-coding transcription to either RF model, suggesting that transcription does not play a predictive role in V&#x003BA; recombination. A previous study proposed that transcription causes the eviction of H2A/H2B around V&#x003BA; RSSs (<xref ref-type="bibr" rid="B75">75</xref>), and non-coding transcription has also been shown to mark recombinationally active domains of the V&#x003BA; region (<xref ref-type="bibr" rid="B16">16</xref>). Together, these findings suggest that, in common with histone H3 and H4 acetylation, non-coding transcription may play a priming role for all V&#x003BA; genes, setting the stage for the features we have described here to specifically activate V&#x003BA; genes for recombination with a range of frequencies.</p>
<p>We also identified the binding of PAX5, CTCF, and PU.1 upstream of V&#x003BA; genes as relatively important in predicting the recombination frequency of active genes, and observed subtle negative relationships between recombination frequency and enrichment of these features. While determining the mechanisms that underpin these relationships are beyond the scope of this study, it is noteworthy that PAX5 and CTCF have been implicated in long-range looping of the <italic>Igh</italic> and <italic>Igh</italic>/<italic>Ig&#x003BA;</italic> loci, respectively (<xref ref-type="bibr" rid="B30">30</xref>, <xref ref-type="bibr" rid="B32">32</xref>&#x02013;<xref ref-type="bibr" rid="B35">35</xref>, <xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B80">80</xref>), bringing V genes into proximity with the (D)J genes. Thus, it is tempting to speculate that the intergenic sites bound by these TFs might correspond to the anchors of these loops. Looping of the locus promotes the recombination of distal genes; however, genes located very close to loop anchors might be spatially constrained, disfavouring their recombination. Conversely, H3K4 methylation, E2A, and IKAROS binding have also been implicated in looping of the <italic>Ig&#x003BA;</italic> locus, but previous data, consistent with our RF models, suggest a positive correlation with recombination (<xref ref-type="bibr" rid="B39">39</xref>, <xref ref-type="bibr" rid="B40">40</xref>). Notably, while CTCF is generally located between the V&#x003BA; genes, the other features were localised to the genes themselves. Furthermore, CTCF-associated interactions were confined to the SIS regulatory element, a silencer located between the V&#x003BA; and J&#x003BA; genes (<xref ref-type="bibr" rid="B81">81</xref>, <xref ref-type="bibr" rid="B82">82</xref>), while loci marked by the other features also interacted with the <italic>Ig&#x003BA;</italic> enhancers. This suggests different classes of loops might exist, with CTCF and PAX5 required for the overall global architecture of the locus, bringing genes into the vicinity of the J&#x003BA; region. This might allow other TFs, such as E2A and IKAROS, to mediate the local clustering of active genes immediately adjacent to the J&#x003BA; genes and the recruitment of the <italic>Ig&#x003BA;</italic> enhancers to these genes. More detailed analysis of the three dimensional structure of the <italic>Ig&#x003BA;</italic> locus, and its relationship with the chromatin features implicated in its organisation, will be required to establish the validity of this hypothesis.</p>
<p>In addition to V(D)J recombination at the AgR loci, our findings also have implications for RAG-mediated off-target recombination events throughout the genome, which can lead to leukaemias. Promoter and enhancer signatures are associated with high frequency, genome-wide recruitment of the RAG complex (<xref ref-type="bibr" rid="B42">42</xref>&#x02013;<xref ref-type="bibr" rid="B44">44</xref>). Our previous study allowed us to refine those signatures by identifying that features of the A and E state are enriched at these sites (<xref ref-type="bibr" rid="B14">14</xref>). Here, we have identified additional candidates (PU.1, IKAROS) that may enhance predictive models of chromosomal translocation hotspots (<xref ref-type="bibr" rid="B43">43</xref>).</p>
<p>The findings reported here demonstrate that the mechanisms that regulate V&#x003BA; recombination differ substantially from those that regulate V<sub>H</sub> recombination. They also identify two distinct and crucial roles for chromatin features in regulating V&#x003BA; gene recombination. While PU.1 binding at the RSS plays a binary role in priming V&#x003BA; genes to recombine, the binding and variable enrichment of several other chromatin features, including H3K4 methylation, IKAROS binding at the RSS, and MED1 binding at the promoter, modulate the frequency with which each active gene recombines. Furthermore, inclusion of this canonical signature may refine prediction of genome-wide RAG1 binding sites susceptible to chromosomal translocation.</p>
</sec>
<sec id="S5">
<title>Ethics Statement</title>
<p>C57BL/6 (WT) and Rag1<sup>&#x02212;/&#x02212;</sup>/VH81X mice were maintained in accordance with Babraham Institute AWERB and Home Office rules and ARRIVE guidelines under Project Licence 80/2529.</p>
</sec>
<sec id="S6" sec-type="author-contributor">
<title>Author Contributions</title>
<p>LM developed the V&#x003BA;J&#x003BA;-seq assay reported here and generated V&#x003BA;J&#x003BA;-seq data. FK and SA adapted the Babraham LinkON pipeline for the Ig&#x003BA; locus. LM, DB, and PC refined the V&#x003BA;J&#x003BA;-seq protocol and analysis pipeline. HK pre-processed and Q.C. checked all NGS datasets. LM and HK performed computational and machine learning analyses. LM visualised the data and prepared the figures. LM, HK, and AC interpreted the results. LM and AC wrote the manuscript.</p>
</sec>
<sec id="S7">
<title>Conflict of Interest Statement</title>
<p>LM, DB, and AC are named inventors on a patent filed, &#x0201C;Covering the VDJ-seq technique: method of identifying VDJ recombination products.&#x0201D; (UK Patent Application No. GB1203720.6, filed March 2, 2012; PCT Patent Applic No. PCT/GB2013/05056, published September 6, 2013. National applications filed Europe, USA, Japan. US Publication number: 20150031042, publication date January 29, 2015.). All other authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>We thank Martin Turner, Mikhail Spivakov, and Peter Fraser for critical reading of the manuscript. Invaluable assistance was provided by Kristina Tabbada, Babraham Sequencing Facility, and Arthur Davis, Flow Facility.</p>
</ack>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was supported by the Biotechnology and Biological Sciences Research Council.</p></fn>
</fn-group>
<sec id="S8" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at <uri xlink:href="http://www.frontiersin.org/article/10.3389/fimmu.2017.01550/full&#x00023;supplementary-material">http://www.frontiersin.org/article/10.3389/fimmu.2017.01550/full&#x00023;supplementary-material</uri>.</p>
<supplementary-material xlink:href="Table_1.XLSX" id="SM1" mimetype="applicationn/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM2" mimetype="applicationn/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<sec id="S9">
<title>Abbreviations</title>
<p>AgR, antigen receptor; BCR, B cell receptor; D, Diversity; DHS, DNase hypersensitivity; <italic>Igh</italic>, Immunoglobulin heavy chain; <italic>Ig&#x003BA;</italic>, Immunoglobulin kappa light chain; J, joining; RF-C, random forest classification; RF-R, random forest regression; RSS, recombination signal sequence; RMSE, root mean squared error; RIC, RSS information content; TdT, terminal deoxynucleotidyl transferase; TF, transcription factor; V, variable; WT, wild type.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><label>1</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fugmann</surname> <given-names>SD</given-names></name> <name><surname>Lee</surname> <given-names>AI</given-names></name> <name><surname>Shockett</surname> <given-names>PE</given-names></name> <name><surname>Villey</surname> <given-names>IJ</given-names></name> <name><surname>Schatz</surname> <given-names>DG</given-names></name></person-group>. <article-title>The RAG proteins and V(D)J recombination: complexes, ends, and transposition</article-title>. <source>Annu Rev Immunol</source> (<year>2000</year>) <volume>18</volume>:<fpage>495</fpage>&#x02013;<lpage>527</lpage>.<pub-id pub-id-type="doi">10.1146/annurev.immunol.18.1.495</pub-id><pub-id pub-id-type="pmid">10837067</pub-id></citation></ref>
<ref id="B2"><label>2</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jung</surname> <given-names>D</given-names></name> <name><surname>Giallourakis</surname> <given-names>C</given-names></name> <name><surname>Mostoslavsky</surname> <given-names>R</given-names></name> <name><surname>Alt</surname> <given-names>FW</given-names></name></person-group>. <article-title>Mechanism and control of V(D)J recombination at the immunoglobulin heavy chain locus</article-title>. <source>Annu Rev Immunol</source> (<year>2006</year>) <volume>24</volume>:<fpage>541</fpage>&#x02013;<lpage>70</lpage>.<pub-id pub-id-type="doi">10.1146/annurev.immunol.23.021704.115830</pub-id><pub-id pub-id-type="pmid">16551259</pub-id></citation></ref>
<ref id="B3"><label>3</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benedict</surname> <given-names>CL</given-names></name> <name><surname>Gilfillan</surname> <given-names>S</given-names></name> <name><surname>Thai</surname> <given-names>TH</given-names></name> <name><surname>Kearney</surname> <given-names>JF</given-names></name></person-group>. <article-title>Terminal deoxynucleotidyl transferase and repertoire development</article-title>. <source>Immunol Rev</source> (<year>2000</year>) <volume>175</volume>:<fpage>150</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1111/j.1600-065X.2000.imr017518.x</pub-id><pub-id pub-id-type="pmid">10933600</pub-id></citation></ref>
<ref id="B4"><label>4</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hendriks</surname> <given-names>RW</given-names></name> <name><surname>Middendorp</surname> <given-names>S</given-names></name></person-group>. <article-title>The pre-BCR checkpoint as a cell-autonomous proliferation switch</article-title>. <source>Trends Immunol</source> (<year>2004</year>) <volume>25</volume>(<issue>5</issue>):<fpage>249</fpage>&#x02013;<lpage>56</lpage>.<pub-id pub-id-type="doi">10.1016/j.it.2004.02.011</pub-id></citation></ref>
<ref id="B5"><label>5</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herzog</surname> <given-names>S</given-names></name> <name><surname>Reth</surname> <given-names>M</given-names></name> <name><surname>Jumaa</surname> <given-names>H</given-names></name></person-group>. <article-title>Regulation of B-cell proliferation and differentiation by pre-B-cell receptor signalling</article-title>. <source>Nat Rev Immunol</source> (<year>2009</year>) <volume>9</volume>(<issue>3</issue>):<fpage>195</fpage>&#x02013;<lpage>205</lpage>.<pub-id pub-id-type="doi">10.1038/nri2491</pub-id><pub-id pub-id-type="pmid">19240758</pub-id></citation></ref>
<ref id="B6"><label>6</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brekke</surname> <given-names>KM</given-names></name> <name><surname>Garrard</surname> <given-names>WT</given-names></name></person-group>. <article-title>Assembly and analysis of the mouse immunoglobulin kappa gene sequence</article-title>. <source>Immunogenetics</source> (<year>2004</year>) <volume>56</volume>(<issue>7</issue>):<fpage>490</fpage>&#x02013;<lpage>505</lpage>.<pub-id pub-id-type="doi">10.1007/s00251-004-0659-0</pub-id><pub-id pub-id-type="pmid">15378297</pub-id></citation></ref>
<ref id="B7"><label>7</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>YS</given-names></name> <name><surname>Hayakawa</surname> <given-names>K</given-names></name> <name><surname>Hardy</surname> <given-names>RR</given-names></name></person-group>. <article-title>The regulated expression of B lineage associated genes during B cell differentiation in bone marrow and fetal liver</article-title>. <source>J Exp Med</source> (<year>1993</year>) <volume>178</volume>(<issue>3</issue>):<fpage>951</fpage>&#x02013;<lpage>60</lpage>.<pub-id pub-id-type="doi">10.1084/jem.178.3.951</pub-id><pub-id pub-id-type="pmid">8350062</pub-id></citation></ref>
<ref id="B8"><label>8</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Victor</surname> <given-names>KD</given-names></name> <name><surname>Vu</surname> <given-names>K</given-names></name> <name><surname>Feeney</surname> <given-names>AJ</given-names></name></person-group>. <article-title>Limited junctional diversity in kappa light chains. Junctional sequences from CD43&#x0002B;B220&#x0002B; early B cell progenitors resemble those from peripheral B cells</article-title>. <source>J Immunol</source> (<year>1994</year>) <volume>152</volume>(<issue>7</issue>):<fpage>3467</fpage>&#x02013;<lpage>75</lpage>.<pub-id pub-id-type="pmid">7511648</pub-id></citation></ref>
<ref id="B9"><label>9</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bertocci</surname> <given-names>B</given-names></name> <name><surname>De Smet</surname> <given-names>A</given-names></name> <name><surname>Berek</surname> <given-names>C</given-names></name> <name><surname>Weill</surname> <given-names>JC</given-names></name> <name><surname>Reynaud</surname> <given-names>CA</given-names></name></person-group>. <article-title>Immunoglobulin kappa light chain gene rearrangement is impaired in mice deficient for DNA polymerase mu</article-title>. <source>Immunity</source> (<year>2003</year>) <volume>19</volume>(<issue>2</issue>):<fpage>203</fpage>&#x02013;<lpage>11</lpage>.<pub-id pub-id-type="doi">10.1016/S1074-7613(03)00203-6</pub-id><pub-id pub-id-type="pmid">12932354</pub-id></citation></ref>
<ref id="B10"><label>10</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nemazee</surname> <given-names>D</given-names></name></person-group>. <article-title>Receptor editing in lymphocyte development and central tolerance</article-title>. <source>Nat Rev Immunol</source> (<year>2006</year>) <volume>6</volume>(<issue>10</issue>):<fpage>728</fpage>&#x02013;<lpage>40</lpage>.<pub-id pub-id-type="doi">10.1038/nri1939</pub-id><pub-id pub-id-type="pmid">16998507</pub-id></citation></ref>
<ref id="B11"><label>11</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vettermann</surname> <given-names>C</given-names></name> <name><surname>Timblin</surname> <given-names>GA</given-names></name> <name><surname>Lim</surname> <given-names>V</given-names></name> <name><surname>Lai</surname> <given-names>EC</given-names></name> <name><surname>Schlissel</surname> <given-names>MS</given-names></name></person-group>. <article-title>The proximal J kappa germline-transcript promoter facilitates receptor editing through control of ordered recombination</article-title>. <source>PLoS One</source> (<year>2015</year>) <volume>10</volume>(<issue>1</issue>):<fpage>e0113824</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0113824</pub-id><pub-id pub-id-type="pmid">25559567</pub-id></citation></ref>
<ref id="B12"><label>12</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cowell</surname> <given-names>LG</given-names></name> <name><surname>Davila</surname> <given-names>M</given-names></name> <name><surname>Kepler</surname> <given-names>TB</given-names></name> <name><surname>Kelsoe</surname> <given-names>G</given-names></name></person-group>. <article-title>Identification and utilization of arbitrary correlations in models of recombination signal sequences</article-title>. <source>Genome Biol</source> (<year>2002</year>) <volume>3</volume>(<issue>12</issue>):<fpage>RESEARCH0072</fpage>.<pub-id pub-id-type="doi">10.1186/gb-2002-3-12-research0072</pub-id><pub-id pub-id-type="pmid">12537561</pub-id></citation></ref>
<ref id="B13"><label>13</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>AI</given-names></name> <name><surname>Fugmann</surname> <given-names>SD</given-names></name> <name><surname>Cowell</surname> <given-names>LG</given-names></name> <name><surname>Ptaszek</surname> <given-names>LM</given-names></name> <name><surname>Kelsoe</surname> <given-names>G</given-names></name> <name><surname>Schatz</surname> <given-names>DG</given-names></name></person-group>. <article-title>A functional analysis of the spacer of V(D)J recombination signal sequences</article-title>. <source>PLoS Biol</source> (<year>2003</year>) <volume>1</volume>(<issue>1</issue>):<fpage>E1</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pbio.0000001</pub-id><pub-id pub-id-type="pmid">14551903</pub-id></citation></ref>
<ref id="B14"><label>14</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolland</surname> <given-names>DJ</given-names></name> <name><surname>Koohy</surname> <given-names>H</given-names></name> <name><surname>Wood</surname> <given-names>AL</given-names></name> <name><surname>Matheson</surname> <given-names>LS</given-names></name> <name><surname>Krueger</surname> <given-names>F</given-names></name> <name><surname>Stubbington</surname> <given-names>MJ</given-names></name> <etal/></person-group> <article-title>Two mutually exclusive local chromatin states drive efficient V(D)J recombination</article-title>. <source>Cell Rep</source> (<year>2016</year>) <volume>15</volume>(<issue>11</issue>):<fpage>2475</fpage>&#x02013;<lpage>87</lpage>.<pub-id pub-id-type="doi">10.1016/j.celrep.2016.05.020</pub-id><pub-id pub-id-type="pmid">27264181</pub-id></citation></ref>
<ref id="B15"><label>15</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>NM</given-names></name> <name><surname>Loguercio</surname> <given-names>S</given-names></name> <name><surname>Verma-Gaur</surname> <given-names>J</given-names></name> <name><surname>Degner</surname> <given-names>SC</given-names></name> <name><surname>Torkamani</surname> <given-names>A</given-names></name> <name><surname>Su</surname> <given-names>AI</given-names></name> <etal/></person-group> <article-title>Deep sequencing of the murine IgH repertoire reveals complex regulation of nonrandom V gene rearrangement frequencies</article-title>. <source>J Immunol</source> (<year>2013</year>) <volume>191</volume>(<issue>5</issue>):<fpage>2393</fpage>&#x02013;<lpage>402</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1301279</pub-id><pub-id pub-id-type="pmid">23898036</pub-id></citation></ref>
<ref id="B16"><label>16</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levin-Klein</surname> <given-names>R</given-names></name> <name><surname>Fraenkel</surname> <given-names>S</given-names></name> <name><surname>Lichtenstein</surname> <given-names>M</given-names></name> <name><surname>Matheson</surname> <given-names>LS</given-names></name> <name><surname>Bartok</surname> <given-names>O</given-names></name> <name><surname>Nevo</surname> <given-names>Y</given-names></name> <etal/></person-group> <article-title>Clonally stable Vkappa allelic choice instructs Igkappa repertoire</article-title>. <source>Nat Commun</source> (<year>2017</year>) <volume>8</volume>:<fpage>15575</fpage>.<pub-id pub-id-type="doi">10.1038/ncomms15575</pub-id></citation></ref>
<ref id="B17"><label>17</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bettridge</surname> <given-names>J</given-names></name> <name><surname>Na</surname> <given-names>CH</given-names></name> <name><surname>Pandey</surname> <given-names>A</given-names></name> <name><surname>Desiderio</surname> <given-names>S</given-names></name></person-group>. <article-title>H3K4me3 induces allosteric conformational changes in the DNA-binding and catalytic regions of the V(D)J recombinase</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2017</year>) <volume>114</volume>(<issue>8</issue>):<fpage>1904</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1615727114</pub-id><pub-id pub-id-type="pmid">28174273</pub-id></citation></ref>
<ref id="B18"><label>18</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y</given-names></name> <name><surname>Subrahmanyam</surname> <given-names>R</given-names></name> <name><surname>Chakraborty</surname> <given-names>T</given-names></name> <name><surname>Sen</surname> <given-names>R</given-names></name> <name><surname>Desiderio</surname> <given-names>S</given-names></name></person-group>. <article-title>A plant homeodomain in RAG-2 that binds hypermethylated lysine 4 of histone H3 is necessary for efficient antigen-receptor-gene rearrangement</article-title>. <source>Immunity</source> (<year>2007</year>) <volume>27</volume>(<issue>4</issue>):<fpage>561</fpage>&#x02013;<lpage>71</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2007.09.005</pub-id><pub-id pub-id-type="pmid">17936034</pub-id></citation></ref>
<ref id="B19"><label>19</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matthews</surname> <given-names>AG</given-names></name> <name><surname>Kuo</surname> <given-names>AJ</given-names></name> <name><surname>Ramon-Maiques</surname> <given-names>S</given-names></name> <name><surname>Han</surname> <given-names>S</given-names></name> <name><surname>Champagne</surname> <given-names>KS</given-names></name> <name><surname>Ivanov</surname> <given-names>D</given-names></name> <etal/></person-group> <article-title>RAG2 PHD finger couples histone H3 lysine 4 trimethylation with V(D)J recombination</article-title>. <source>Nature</source> (<year>2007</year>) <volume>450</volume>(<issue>7172</issue>):<fpage>1106</fpage>&#x02013;<lpage>10</lpage>.<pub-id pub-id-type="doi">10.1038/nature06431</pub-id><pub-id pub-id-type="pmid">18033247</pub-id></citation></ref>
<ref id="B20"><label>20</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shimazaki</surname> <given-names>N</given-names></name> <name><surname>Tsai</surname> <given-names>AG</given-names></name> <name><surname>Lieber</surname> <given-names>MR</given-names></name></person-group>. <article-title>H3K4me3 stimulates the V(D)J RAG complex for both nicking and hairpinning in trans in addition to tethering in cis: implications for translocations</article-title>. <source>Mol Cell</source> (<year>2009</year>) <volume>34</volume>(<issue>5</issue>):<fpage>535</fpage>&#x02013;<lpage>44</lpage>.<pub-id pub-id-type="doi">10.1016/j.molcel.2009.05.011</pub-id><pub-id pub-id-type="pmid">19524534</pub-id></citation></ref>
<ref id="B21"><label>21</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Batista</surname> <given-names>CR</given-names></name> <name><surname>Li</surname> <given-names>SK</given-names></name> <name><surname>Xu</surname> <given-names>LS</given-names></name> <name><surname>Solomon</surname> <given-names>LA</given-names></name> <name><surname>DeKoter</surname> <given-names>RP</given-names></name></person-group>. <article-title>PU.1 regulates Ig light chain transcription and rearrangement in pre-B cells during B cell development</article-title>. <source>J Immunol</source> (<year>2017</year>) <volume>198</volume>(<issue>4</issue>):<fpage>1565</fpage>&#x02013;<lpage>74</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1601709</pub-id><pub-id pub-id-type="pmid">28062693</pub-id></citation></ref>
<ref id="B22"><label>22</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heizmann</surname> <given-names>B</given-names></name> <name><surname>Kastner</surname> <given-names>P</given-names></name> <name><surname>Chan</surname> <given-names>S</given-names></name></person-group>. <article-title>Ikaros is absolutely required for pre-B cell differentiation by attenuating IL-7 signals</article-title>. <source>J Exp Med</source> (<year>2013</year>) <volume>210</volume>(<issue>13</issue>):<fpage>2823</fpage>&#x02013;<lpage>32</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20131735</pub-id><pub-id pub-id-type="pmid">24297995</pub-id></citation></ref>
<ref id="B23"><label>23</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hodawadekar</surname> <given-names>S</given-names></name> <name><surname>Park</surname> <given-names>K</given-names></name> <name><surname>Farrar</surname> <given-names>MA</given-names></name> <name><surname>Atchison</surname> <given-names>ML</given-names></name></person-group>. <article-title>A developmentally controlled competitive STAT5-PU.1 DNA binding mechanism regulates activity of the Ig kappa E3&#x02019; enhancer</article-title>. <source>J Immunol</source> (<year>2012</year>) <volume>188</volume>(<issue>5</issue>):<fpage>2276</fpage>&#x02013;<lpage>84</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1102239</pub-id></citation></ref>
<ref id="B24"><label>24</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Inlay</surname> <given-names>MA</given-names></name> <name><surname>Tian</surname> <given-names>H</given-names></name> <name><surname>Lin</surname> <given-names>T</given-names></name> <name><surname>Xu</surname> <given-names>Y</given-names></name></person-group>. <article-title>Important roles for E protein binding sites within the immunoglobulin kappa chain intronic enhancer in activating Vkappa Jkappa rearrangement</article-title>. <source>J Exp Med</source> (<year>2004</year>) <volume>200</volume>(<issue>9</issue>):<fpage>1205</fpage>&#x02013;<lpage>11</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20041135</pub-id><pub-id pub-id-type="pmid">15504821</pub-id></citation></ref>
<ref id="B25"><label>25</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>K</given-names></name> <name><surname>Hashimshony</surname> <given-names>T</given-names></name> <name><surname>Sawai</surname> <given-names>CM</given-names></name> <name><surname>Pongubala</surname> <given-names>JM</given-names></name> <name><surname>Skok</surname> <given-names>JA</given-names></name> <name><surname>Aifantis</surname> <given-names>I</given-names></name> <etal/></person-group> <article-title>Regulation of immunoglobulin light-chain recombination by the transcription factor IRF-4 and the attenuation of interleukin-7 signaling</article-title>. <source>Immunity</source> (<year>2008</year>) <volume>28</volume>(<issue>3</issue>):<fpage>335</fpage>&#x02013;<lpage>45</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2007.12.019</pub-id><pub-id pub-id-type="pmid">18280186</pub-id></citation></ref>
<ref id="B26"><label>26</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lazorchak</surname> <given-names>AS</given-names></name> <name><surname>Schlissel</surname> <given-names>MS</given-names></name> <name><surname>Zhuang</surname> <given-names>Y</given-names></name></person-group>. <article-title>E2A and IRF-4/Pip promote chromatin modification and transcription of the immunoglobulin kappa locus in pre-B cells</article-title>. <source>Mol Cell Biol</source> (<year>2006</year>) <volume>26</volume>(<issue>3</issue>):<fpage>810</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1128/MCB.26.3.810-821.2006</pub-id><pub-id pub-id-type="pmid">16428437</pub-id></citation></ref>
<ref id="B27"><label>27</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>S</given-names></name> <name><surname>Turetsky</surname> <given-names>A</given-names></name> <name><surname>Trinh</surname> <given-names>L</given-names></name> <name><surname>Lu</surname> <given-names>R</given-names></name></person-group>. <article-title>IFN regulatory factor 4 and 8 promote Ig light chain kappa locus activation in pre-B cell development</article-title>. <source>J Immunol</source> (<year>2006</year>) <volume>177</volume>(<issue>11</issue>):<fpage>7898</fpage>&#x02013;<lpage>904</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.177.11.7898</pub-id><pub-id pub-id-type="pmid">17114461</pub-id></citation></ref>
<ref id="B28"><label>28</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>S</given-names></name> <name><surname>Wakae</surname> <given-names>K</given-names></name> <name><surname>Anzai</surname> <given-names>Y</given-names></name> <name><surname>Murai</surname> <given-names>K</given-names></name> <name><surname>Tamaki</surname> <given-names>N</given-names></name> <name><surname>Miyazaki</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>E2A and CBP/p300 act in synergy to promote chromatin accessibility of the immunoglobulin kappa locus</article-title>. <source>J Immunol</source> (<year>2012</year>) <volume>188</volume>(<issue>11</issue>):<fpage>5547</fpage>&#x02013;<lpage>60</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1002346</pub-id></citation></ref>
<ref id="B29"><label>29</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sato</surname> <given-names>H</given-names></name> <name><surname>Saito-Ohara</surname> <given-names>F</given-names></name> <name><surname>Inazawa</surname> <given-names>J</given-names></name> <name><surname>Kudo</surname> <given-names>A</given-names></name></person-group>. <article-title>Pax-5 is essential for kappa sterile transcription during Ig kappa chain gene rearrangement</article-title>. <source>J Immunol</source> (<year>2004</year>) <volume>172</volume>(<issue>8</issue>):<fpage>4858</fpage>&#x02013;<lpage>65</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.172.8.4858</pub-id><pub-id pub-id-type="pmid">15067064</pub-id></citation></ref>
<ref id="B30"><label>30</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Degner</surname> <given-names>SC</given-names></name> <name><surname>Verma-Gaur</surname> <given-names>J</given-names></name> <name><surname>Wong</surname> <given-names>TP</given-names></name> <name><surname>Bossen</surname> <given-names>C</given-names></name> <name><surname>Iverson</surname> <given-names>GM</given-names></name> <name><surname>Torkamani</surname> <given-names>A</given-names></name> <etal/></person-group> <article-title>CCCTC-binding factor (CTCF) and cohesin influence the genomic architecture of the Igh locus and antisense transcription in pro-B cells</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2011</year>) <volume>108</volume>(<issue>23</issue>):<fpage>9566</fpage>&#x02013;<lpage>71</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1019391108</pub-id><pub-id pub-id-type="pmid">21606361</pub-id></citation></ref>
<ref id="B31"><label>31</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>C</given-names></name> <name><surname>Yoon</surname> <given-names>HS</given-names></name> <name><surname>Franklin</surname> <given-names>A</given-names></name> <name><surname>Jain</surname> <given-names>S</given-names></name> <name><surname>Ebert</surname> <given-names>A</given-names></name> <name><surname>Cheng</surname> <given-names>HL</given-names></name> <etal/></person-group> <article-title>CTCF-binding elements mediate control of V(D)J recombination</article-title>. <source>Nature</source> (<year>2011</year>) <volume>477</volume>(<issue>7365</issue>):<fpage>424</fpage>&#x02013;<lpage>30</lpage>.<pub-id pub-id-type="doi">10.1038/nature10495</pub-id><pub-id pub-id-type="pmid">21909113</pub-id></citation></ref>
<ref id="B32"><label>32</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ribeiro de Almeida</surname> <given-names>C</given-names></name> <name><surname>Stadhouders</surname> <given-names>R</given-names></name> <name><surname>de Bruijn</surname> <given-names>MJ</given-names></name> <name><surname>Bergen</surname> <given-names>IM</given-names></name> <name><surname>Thongjuea</surname> <given-names>S</given-names></name> <name><surname>Lenhard</surname> <given-names>B</given-names></name> <etal/></person-group> <article-title>The DNA-binding protein CTCF limits proximal Vkappa recombination and restricts kappa enhancer interactions to the immunoglobulin kappa light chain locus</article-title>. <source>Immunity</source> (<year>2011</year>) <volume>35</volume>(<issue>4</issue>):<fpage>501</fpage>&#x02013;<lpage>13</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2011.07.014</pub-id></citation></ref>
<ref id="B33"><label>33</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiang</surname> <given-names>Y</given-names></name> <name><surname>Park</surname> <given-names>SK</given-names></name> <name><surname>Garrard</surname> <given-names>WT</given-names></name></person-group>. <article-title>Vkappa gene repertoire and locus contraction are specified by critical DNase I hypersensitive sites within the Vkappa-Jkappa intervening region</article-title>. <source>J Immunol</source> (<year>2013</year>) <volume>190</volume>(<issue>4</issue>):<fpage>1819</fpage>&#x02013;<lpage>26</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1203127</pub-id></citation></ref>
<ref id="B34"><label>34</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiang</surname> <given-names>Y</given-names></name> <name><surname>Zhou</surname> <given-names>X</given-names></name> <name><surname>Hewitt</surname> <given-names>SL</given-names></name> <name><surname>Skok</surname> <given-names>JA</given-names></name> <name><surname>Garrard</surname> <given-names>WT</given-names></name></person-group>. <article-title>A multifunctional element in the mouse Igkappa locus that specifies repertoire and Ig loci subnuclear location</article-title>. <source>J Immunol</source> (<year>2011</year>) <volume>186</volume>(<issue>9</issue>):<fpage>5356</fpage>&#x02013;<lpage>66</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1003794</pub-id></citation></ref>
<ref id="B35"><label>35</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuxa</surname> <given-names>M</given-names></name> <name><surname>Skok</surname> <given-names>J</given-names></name> <name><surname>Souabni</surname> <given-names>A</given-names></name> <name><surname>Salvagiotto</surname> <given-names>G</given-names></name> <name><surname>Roldan</surname> <given-names>E</given-names></name> <name><surname>Busslinger</surname> <given-names>M</given-names></name></person-group>. <article-title>Pax5 induces V-to-DJ rearrangements and locus contraction of the immunoglobulin heavy-chain gene</article-title>. <source>Genes Dev</source> (<year>2004</year>) <volume>18</volume>(<issue>4</issue>):<fpage>411</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1101/gad.291504</pub-id><pub-id pub-id-type="pmid">15004008</pub-id></citation></ref>
<ref id="B36"><label>36</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>H</given-names></name> <name><surname>Schmidt-Supprian</surname> <given-names>M</given-names></name> <name><surname>Shi</surname> <given-names>Y</given-names></name> <name><surname>Hobeika</surname> <given-names>E</given-names></name> <name><surname>Barteneva</surname> <given-names>N</given-names></name> <name><surname>Jumaa</surname> <given-names>H</given-names></name> <etal/></person-group> <article-title>Yin Yang 1 is a critical regulator of B-cell development</article-title>. <source>Genes Dev</source> (<year>2007</year>) <volume>21</volume>(<issue>10</issue>):<fpage>1179</fpage>&#x02013;<lpage>89</lpage>.<pub-id pub-id-type="doi">10.1101/gad.1529307</pub-id><pub-id pub-id-type="pmid">17504937</pub-id></citation></ref>
<ref id="B37"><label>37</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aoki-Ota</surname> <given-names>M</given-names></name> <name><surname>Torkamani</surname> <given-names>A</given-names></name> <name><surname>Ota</surname> <given-names>T</given-names></name> <name><surname>Schork</surname> <given-names>N</given-names></name> <name><surname>Nemazee</surname> <given-names>D</given-names></name></person-group>. <article-title>Skewed primary Igkappa repertoire and V-J joining in C57BL/6 mice: implications for recombination accessibility and receptor editing</article-title>. <source>J Immunol</source> (<year>2012</year>) <volume>188</volume>(<issue>5</issue>):<fpage>2305</fpage>&#x02013;<lpage>15</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1103484</pub-id></citation></ref>
<ref id="B38"><label>38</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>SG</given-names></name> <name><surname>Ba</surname> <given-names>Z</given-names></name> <name><surname>Du</surname> <given-names>Z</given-names></name> <name><surname>Zhang</surname> <given-names>Y</given-names></name> <name><surname>Hu</surname> <given-names>J</given-names></name> <name><surname>Alt</surname> <given-names>FW</given-names></name></person-group>. <article-title>Highly sensitive and unbiased approach for elucidating antibody repertoires</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2016</year>) <volume>113</volume>(<issue>28</issue>):<fpage>7846</fpage>&#x02013;<lpage>51</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1608649113</pub-id><pub-id pub-id-type="pmid">27354528</pub-id></citation></ref>
<ref id="B39"><label>39</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>YC</given-names></name> <name><surname>Benner</surname> <given-names>C</given-names></name> <name><surname>Mansson</surname> <given-names>R</given-names></name> <name><surname>Heinz</surname> <given-names>S</given-names></name> <name><surname>Miyazaki</surname> <given-names>K</given-names></name> <name><surname>Miyazaki</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>Global changes in the nuclear positioning of genes and intra- and interdomain genomic interactions that orchestrate B cell fate</article-title>. <source>Nat Immunol</source> (<year>2012</year>) <volume>13</volume>(<issue>12</issue>):<fpage>1196</fpage>&#x02013;<lpage>204</lpage>.<pub-id pub-id-type="doi">10.1038/ni.2432</pub-id><pub-id pub-id-type="pmid">23064439</pub-id></citation></ref>
<ref id="B40"><label>40</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stadhouders</surname> <given-names>R</given-names></name> <name><surname>de Bruijn</surname> <given-names>MJ</given-names></name> <name><surname>Rother</surname> <given-names>MB</given-names></name> <name><surname>Yuvaraj</surname> <given-names>S</given-names></name> <name><surname>Ribeiro de Almeida</surname> <given-names>C</given-names></name> <name><surname>Kolovos</surname> <given-names>P</given-names></name> <etal/></person-group> <article-title>Pre-B cell receptor signaling induces immunoglobulin kappa locus accessibility by functional redistribution of enhancer-mediated chromatin interactions</article-title>. <source>PLoS Biol</source> (<year>2014</year>) <volume>12</volume>(<issue>2</issue>):<fpage>e1001791</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pbio.1001791</pub-id></citation></ref>
<ref id="B41"><label>41</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>X</given-names></name> <name><surname>Papasani</surname> <given-names>M</given-names></name> <name><surname>Hao</surname> <given-names>Y</given-names></name> <name><surname>Calamito</surname> <given-names>M</given-names></name> <name><surname>Wei</surname> <given-names>F</given-names></name> <name><surname>Quinn</surname> <given-names>WJ</given-names> <suffix>III</suffix></name> <etal/></person-group> <article-title>YY1 controls Igkappa repertoire and B-cell development, and localizes with condensin on the Igkappa locus</article-title>. <source>EMBO J</source> (<year>2013</year>) <volume>32</volume>(<issue>8</issue>):<fpage>1168</fpage>&#x02013;<lpage>82</lpage>.<pub-id pub-id-type="doi">10.1038/emboj.2013.66</pub-id></citation></ref>
<ref id="B42"><label>42</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ji</surname> <given-names>Y</given-names></name> <name><surname>Resch</surname> <given-names>W</given-names></name> <name><surname>Corbett</surname> <given-names>E</given-names></name> <name><surname>Yamane</surname> <given-names>A</given-names></name> <name><surname>Casellas</surname> <given-names>R</given-names></name> <name><surname>Schatz</surname> <given-names>DG</given-names></name></person-group>. <article-title>The in vivo pattern of binding of RAG1 and RAG2 to antigen receptor loci</article-title>. <source>Cell</source> (<year>2010</year>) <volume>141</volume>(<issue>3</issue>):<fpage>419</fpage>&#x02013;<lpage>31</lpage>.<pub-id pub-id-type="doi">10.1016/j.cell.2010.03.010</pub-id><pub-id pub-id-type="pmid">20398922</pub-id></citation></ref>
<ref id="B43"><label>43</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maman</surname> <given-names>Y</given-names></name> <name><surname>Teng</surname> <given-names>G</given-names></name> <name><surname>Seth</surname> <given-names>R</given-names></name> <name><surname>Kleinstein</surname> <given-names>SH</given-names></name> <name><surname>Schatz</surname> <given-names>DG</given-names></name></person-group>. <article-title>RAG1 targeting in the genome is dominated by chromatin interactions mediated by the non-core regions of RAG1 and RAG2</article-title>. <source>Nucleic Acids Res</source> (<year>2016</year>) <volume>44</volume>(<issue>20</issue>):<fpage>9624</fpage>&#x02013;<lpage>37</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkw633</pub-id><pub-id pub-id-type="pmid">27436288</pub-id></citation></ref>
<ref id="B44"><label>44</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teng</surname> <given-names>G</given-names></name> <name><surname>Maman</surname> <given-names>Y</given-names></name> <name><surname>Resch</surname> <given-names>W</given-names></name> <name><surname>Kim</surname> <given-names>M</given-names></name> <name><surname>Yamane</surname> <given-names>A</given-names></name> <name><surname>Qian</surname> <given-names>J</given-names></name> <etal/></person-group> <article-title>RAG represents a widespread threat to the lymphocyte genome</article-title>. <source>Cell</source> (<year>2015</year>) <volume>162</volume>(<issue>4</issue>):<fpage>751</fpage>&#x02013;<lpage>65</lpage>.<pub-id pub-id-type="doi">10.1016/j.cell.2015.07.009</pub-id><pub-id pub-id-type="pmid">26234156</pub-id></citation></ref>
<ref id="B45"><label>45</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feddersen</surname> <given-names>RM</given-names></name> <name><surname>Van Ness</surname> <given-names>BG</given-names></name></person-group>. <article-title>Corrective recombination of mouse immunoglobulin kappa alleles in Abelson murine leukemia virus-transformed pre-B cells</article-title>. <source>Mol Cell Biol</source> (<year>1990</year>) <volume>10</volume>(<issue>2</issue>):<fpage>569</fpage>&#x02013;<lpage>76</lpage>.<pub-id pub-id-type="doi">10.1128/MCB.10.2.569</pub-id><pub-id pub-id-type="pmid">2153918</pub-id></citation></ref>
<ref id="B46"><label>46</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamagami</surname> <given-names>T</given-names></name> <name><surname>ten Boekel</surname> <given-names>E</given-names></name> <name><surname>Andersson</surname> <given-names>J</given-names></name> <name><surname>Rolink</surname> <given-names>A</given-names></name> <name><surname>Melchers</surname> <given-names>F</given-names></name></person-group>. <article-title>Frequencies of multiple IgL chain gene rearrangements in single normal or kappaL chain-deficient B lineage cells</article-title>. <source>Immunity</source> (<year>1999</year>) <volume>11</volume>(<issue>3</issue>):<fpage>317</fpage>&#x02013;<lpage>27</lpage>.<pub-id pub-id-type="doi">10.1016/S1074-7613(00)80107-7</pub-id><pub-id pub-id-type="pmid">10514010</pub-id></citation></ref>
<ref id="B47"><label>47</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>F</given-names></name> <name><surname>Chen</surname> <given-names>X</given-names></name> <name><surname>Kearney</surname> <given-names>JF</given-names></name></person-group>. <article-title>Development of VH81X transgene-bearing B cells in fetus and adult: sites for expansion and deletion in conventional and CD5/B1 cells</article-title>. <source>Int Immunol</source> (<year>1997</year>) <volume>9</volume>(<issue>4</issue>):<fpage>493</fpage>&#x02013;<lpage>505</lpage>.<pub-id pub-id-type="doi">10.1093/intimm/9.4.493</pub-id><pub-id pub-id-type="pmid">9138009</pub-id></citation></ref>
<ref id="B48"><label>48</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mombaerts</surname> <given-names>P</given-names></name> <name><surname>Iacomini</surname> <given-names>J</given-names></name> <name><surname>Johnson</surname> <given-names>RS</given-names></name> <name><surname>Herrup</surname> <given-names>K</given-names></name> <name><surname>Tonegawa</surname> <given-names>S</given-names></name> <name><surname>Papaioannou</surname> <given-names>VE</given-names></name></person-group>. <article-title>RAG-1-deficient mice have no mature B and T lymphocytes</article-title>. <source>Cell</source> (<year>1992</year>) <volume>68</volume>(<issue>5</issue>):<fpage>869</fpage>&#x02013;<lpage>77</lpage>.<pub-id pub-id-type="doi">10.1016/0092-8674(92)90030-G</pub-id><pub-id pub-id-type="pmid">1547488</pub-id></citation></ref>
<ref id="B49"><label>49</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Langmead</surname> <given-names>B</given-names></name> <name><surname>Trapnell</surname> <given-names>C</given-names></name> <name><surname>Pop</surname> <given-names>M</given-names></name> <name><surname>Salzberg</surname> <given-names>SL</given-names></name></person-group>. <article-title>Ultrafast and memory-efficient alignment of short DNA sequences to the human genome</article-title>. <source>Genome Biol</source> (<year>2009</year>) <volume>10</volume>(<issue>3</issue>):<fpage>R25</fpage>.<pub-id pub-id-type="doi">10.1186/gb-2009-10-3-r25</pub-id><pub-id pub-id-type="pmid">19261174</pub-id></citation></ref>
<ref id="B50"><label>50</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alamyar</surname> <given-names>E</given-names></name> <name><surname>Duroux</surname> <given-names>P</given-names></name> <name><surname>Lefranc</surname> <given-names>MP</given-names></name> <name><surname>Giudicelli</surname> <given-names>V</given-names></name></person-group>. <article-title>IMGT((R)) tools for the nucleotide analysis of immunoglobulin (IG) and T cell receptor (TR) V-(D)-J repertoires, polymorphisms, and IG mutations: IMGT/V-QUEST and IMGT/HighV-QUEST for NGS</article-title>. <source>Methods Mol Biol</source> (<year>2012</year>) <volume>882</volume>:<fpage>569</fpage>&#x02013;<lpage>604</lpage>.<pub-id pub-id-type="doi">10.1007/978-1-61779-842-9_32</pub-id></citation></ref>
<ref id="B51"><label>51</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y</given-names></name> <name><surname>Liu</surname> <given-names>T</given-names></name> <name><surname>Meyer</surname> <given-names>CA</given-names></name> <name><surname>Eeckhoute</surname> <given-names>J</given-names></name> <name><surname>Johnson</surname> <given-names>DS</given-names></name> <name><surname>Bernstein</surname> <given-names>BE</given-names></name> <etal/></person-group> <article-title>Model-based analysis of ChIP-Seq (MACS)</article-title>. <source>Genome Biol</source> (<year>2008</year>) <volume>9</volume>(<issue>9</issue>):<fpage>R137</fpage>.<pub-id pub-id-type="doi">10.1186/gb-2008-9-9-r137</pub-id><pub-id pub-id-type="pmid">18798982</pub-id></citation></ref>
<ref id="B52"><label>52</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dereeper</surname> <given-names>A</given-names></name> <name><surname>Audic</surname> <given-names>S</given-names></name> <name><surname>Claverie</surname> <given-names>JM</given-names></name> <name><surname>Blanc</surname> <given-names>G</given-names></name></person-group>. <article-title>BLAST-EXPLORER helps you building datasets for phylogenetic analysis</article-title>. <source>BMC Evol Biol</source> (<year>2010</year>) <volume>10</volume>:<fpage>8</fpage>.<pub-id pub-id-type="doi">10.1186/1471-2148-10-8</pub-id><pub-id pub-id-type="pmid">20067610</pub-id></citation></ref>
<ref id="B53"><label>53</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dereeper</surname> <given-names>A</given-names></name> <name><surname>Guignon</surname> <given-names>V</given-names></name> <name><surname>Blanc</surname> <given-names>G</given-names></name> <name><surname>Audic</surname> <given-names>S</given-names></name> <name><surname>Buffet</surname> <given-names>S</given-names></name> <name><surname>Chevenet</surname> <given-names>F</given-names></name> <etal/></person-group> <article-title><uri xlink:href="http://Phylogeny.fr">Phylogeny.fr</uri>: robust phylogenetic analysis for the non-specialist</article-title>. <source>Nucleic Acids Res</source> (<year>2008</year>) <volume>36</volume>(<issue>Web Server issue</issue>):<fpage>W465</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkn180</pub-id><pub-id pub-id-type="pmid">18424797</pub-id></citation></ref>
<ref id="B54"><label>54</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edgar</surname> <given-names>RC</given-names></name></person-group>. <article-title>MUSCLE: multiple sequence alignment with high accuracy and high throughput</article-title>. <source>Nucleic Acids Res</source> (<year>2004</year>) <volume>32</volume>(<issue>5</issue>):<fpage>1792</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1093/nar/gkh340</pub-id><pub-id pub-id-type="pmid">15034147</pub-id></citation></ref>
<ref id="B55"><label>55</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anisimova</surname> <given-names>M</given-names></name> <name><surname>Gascuel</surname> <given-names>O</given-names></name></person-group>. <article-title>Approximate likelihood-ratio test for branches: a fast, accurate, and powerful alternative</article-title>. <source>Syst Biol</source> (<year>2006</year>) <volume>55</volume>(<issue>4</issue>):<fpage>539</fpage>&#x02013;<lpage>52</lpage>.<pub-id pub-id-type="doi">10.1080/10635150600755453</pub-id><pub-id pub-id-type="pmid">16785212</pub-id></citation></ref>
<ref id="B56"><label>56</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guindon</surname> <given-names>S</given-names></name> <name><surname>Gascuel</surname> <given-names>O</given-names></name></person-group>. <article-title>A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood</article-title>. <source>Syst Biol</source> (<year>2003</year>) <volume>52</volume>(<issue>5</issue>):<fpage>696</fpage>&#x02013;<lpage>704</lpage>.<pub-id pub-id-type="doi">10.1080/10635150390235520</pub-id><pub-id pub-id-type="pmid">14530136</pub-id></citation></ref>
<ref id="B57"><label>57</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>G</given-names></name> <name><surname>Smith</surname> <given-names>DK</given-names></name> <name><surname>Zhu</surname> <given-names>H</given-names></name> <name><surname>Guan</surname> <given-names>Y</given-names></name> <name><surname>Lam</surname> <given-names>TT-Y</given-names></name> <name><surname>McInerny</surname> <given-names>G</given-names></name></person-group>. <article-title>ggtree: an R package for visualization and annotation of phylogenetic trees with their covariates and other associated data</article-title>. <source>Methods Ecol Evol</source> (<year>2017</year>) <volume>8</volume>(<issue>1</issue>):<fpage>28</fpage>&#x02013;<lpage>36</lpage>.<pub-id pub-id-type="doi">10.1111/2041-210x.12628</pub-id></citation></ref>
<ref id="B58"><label>58</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mammana</surname> <given-names>A</given-names></name> <name><surname>Chung</surname> <given-names>HR</given-names></name></person-group>. <article-title>Chromatin segmentation based on a probabilistic model for read counts explains a large portion of the epigenome</article-title>. <source>Genome Biol</source> (<year>2015</year>) <volume>16</volume>:<fpage>151</fpage>.<pub-id pub-id-type="doi">10.1186/s13059-015-0708-z</pub-id><pub-id pub-id-type="pmid">26206277</pub-id></citation></ref>
<ref id="B59"><label>59</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ernst</surname> <given-names>J</given-names></name> <name><surname>Kellis</surname> <given-names>M</given-names></name></person-group>. <article-title>ChromHMM: automating chromatin-state discovery and characterization</article-title>. <source>Nat Methods</source> (<year>2012</year>) <volume>9</volume>(<issue>3</issue>):<fpage>215</fpage>&#x02013;<lpage>6</lpage>.<pub-id pub-id-type="doi">10.1038/nmeth.1906</pub-id></citation></ref>
<ref id="B60"><label>60</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quinlan</surname> <given-names>AR</given-names></name> <name><surname>Hall</surname> <given-names>IM</given-names></name></person-group>. <article-title>BEDTools: a flexible suite of utilities for comparing genomic features</article-title>. <source>Bioinformatics</source> (<year>2010</year>) <volume>26</volume>(<issue>6</issue>):<fpage>841</fpage>&#x02013;<lpage>2</lpage>.<pub-id pub-id-type="doi">10.1093/bioinformatics/btq033</pub-id><pub-id pub-id-type="pmid">20110278</pub-id></citation></ref>
<ref id="B61"><label>61</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boulesteix</surname> <given-names>AL</given-names></name> <name><surname>Janitza</surname> <given-names>S</given-names></name> <name><surname>Kruppa</surname> <given-names>J</given-names></name> <name><surname>Konig</surname> <given-names>IR</given-names></name></person-group>. <article-title>Overview of random forest methodology and practical guidance with emphasis on computational biology and bioinformatics</article-title>. <source>Wiley Interdiscip Rev Data Min Knowl Discov</source> (<year>2012</year>) <volume>2</volume>(<issue>6</issue>):<fpage>493</fpage>&#x02013;<lpage>507</lpage>.<pub-id pub-id-type="doi">10.1002/widm.1072</pub-id></citation></ref>
<ref id="B62"><label>62</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liaw</surname> <given-names>A</given-names></name> <name><surname>Wiener</surname> <given-names>M</given-names></name></person-group>. <article-title>Classification and regression by randomForest</article-title>. <source>R News</source> (<year>2002</year>) <volume>2/3</volume>:<fpage>18</fpage>&#x02013;<lpage>22</lpage>.</citation></ref>
<ref id="B63"><label>63</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Langmead</surname> <given-names>B</given-names></name> <name><surname>Salzberg</surname> <given-names>SL</given-names></name></person-group>. <article-title>Fast gapped-read alignment with Bowtie 2</article-title>. <source>Nat Methods</source> (<year>2012</year>) <volume>9</volume>(<issue>4</issue>):<fpage>357</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1038/nmeth.1923</pub-id><pub-id pub-id-type="pmid">22388286</pub-id></citation></ref>
<ref id="B64"><label>64</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Revilla-i-Domingo</surname> <given-names>R</given-names></name> <name><surname>Bilic</surname> <given-names>I</given-names></name> <name><surname>Vilagos</surname> <given-names>B</given-names></name> <name><surname>Tagoh</surname> <given-names>H</given-names></name> <name><surname>Ebert</surname> <given-names>A</given-names></name> <name><surname>Tamir</surname> <given-names>IM</given-names></name> <etal/></person-group> <article-title>The B-cell identity factor Pax5 regulates distinct transcriptional programmes in early and late B lymphopoiesis</article-title>. <source>EMBO J</source> (<year>2012</year>) <volume>31</volume>(<issue>14</issue>):<fpage>3130</fpage>&#x02013;<lpage>46</lpage>.<pub-id pub-id-type="doi">10.1038/emboj.2012.155</pub-id><pub-id pub-id-type="pmid">22669466</pub-id></citation></ref>
<ref id="B65"><label>65</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ebert</surname> <given-names>A</given-names></name> <name><surname>McManus</surname> <given-names>S</given-names></name> <name><surname>Tagoh</surname> <given-names>H</given-names></name> <name><surname>Medvedovic</surname> <given-names>J</given-names></name> <name><surname>Salvagiotto</surname> <given-names>G</given-names></name> <name><surname>Novatchkova</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>The distal V(H) gene cluster of the Igh locus contains distinct regulatory elements with Pax5 transcription factor-dependent activity in pro-B cells</article-title>. <source>Immunity</source> (<year>2011</year>) <volume>34</volume>(<issue>2</issue>):<fpage>175</fpage>&#x02013;<lpage>87</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2011.02.005</pub-id><pub-id pub-id-type="pmid">21349430</pub-id></citation></ref>
<ref id="B66"><label>66</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Medvedovic</surname> <given-names>J</given-names></name> <name><surname>Ebert</surname> <given-names>A</given-names></name> <name><surname>Tagoh</surname> <given-names>H</given-names></name> <name><surname>Tamir</surname> <given-names>IM</given-names></name> <name><surname>Schwickert</surname> <given-names>TA</given-names></name> <name><surname>Novatchkova</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>Flexible long-range loops in the VH gene region of the Igh locus facilitate the generation of a diverse antibody repertoire</article-title>. <source>Immunity</source> (<year>2013</year>) <volume>39</volume>(<issue>2</issue>):<fpage>229</fpage>&#x02013;<lpage>44</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2013.08.011</pub-id><pub-id pub-id-type="pmid">23973221</pub-id></citation></ref>
<ref id="B67"><label>67</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mullen</surname> <given-names>AC</given-names></name> <name><surname>Orlando</surname> <given-names>DA</given-names></name> <name><surname>Newman</surname> <given-names>JJ</given-names></name> <name><surname>Loven</surname> <given-names>J</given-names></name> <name><surname>Kumar</surname> <given-names>RM</given-names></name> <name><surname>Bilodeau</surname> <given-names>S</given-names></name> <etal/></person-group> <article-title>Master transcription factors determine cell-type-specific responses to TGF-beta signaling</article-title>. <source>Cell</source> (<year>2011</year>) <volume>147</volume>(<issue>3</issue>):<fpage>565</fpage>&#x02013;<lpage>76</lpage>.<pub-id pub-id-type="doi">10.1016/j.cell.2011.08.050</pub-id></citation></ref>
<ref id="B68"><label>68</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Whyte</surname> <given-names>WA</given-names></name> <name><surname>Orlando</surname> <given-names>DA</given-names></name> <name><surname>Hnisz</surname> <given-names>D</given-names></name> <name><surname>Abraham</surname> <given-names>BJ</given-names></name> <name><surname>Lin</surname> <given-names>CY</given-names></name> <name><surname>Kagey</surname> <given-names>MH</given-names></name> <etal/></person-group> <article-title>Master transcription factors and mediator establish super-enhancers at key cell identity genes</article-title>. <source>Cell</source> (<year>2013</year>) <volume>153</volume>(<issue>2</issue>):<fpage>307</fpage>&#x02013;<lpage>19</lpage>.<pub-id pub-id-type="doi">10.1016/j.cell.2013.03.035</pub-id><pub-id pub-id-type="pmid">23582322</pub-id></citation></ref>
<ref id="B69"><label>69</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vilagos</surname> <given-names>B</given-names></name> <name><surname>Hoffmann</surname> <given-names>M</given-names></name> <name><surname>Souabni</surname> <given-names>A</given-names></name> <name><surname>Sun</surname> <given-names>Q</given-names></name> <name><surname>Werner</surname> <given-names>B</given-names></name> <name><surname>Medvedovic</surname> <given-names>J</given-names></name> <etal/></person-group> <article-title>Essential role of EBF1 in the generation and function of distinct mature B cell types</article-title>. <source>J Exp Med</source> (<year>2012</year>) <volume>209</volume>(<issue>4</issue>):<fpage>775</fpage>&#x02013;<lpage>92</lpage>.<pub-id pub-id-type="doi">10.1084/jem.20112422</pub-id><pub-id pub-id-type="pmid">22473956</pub-id></citation></ref>
<ref id="B70"><label>70</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwickert</surname> <given-names>TA</given-names></name> <name><surname>Tagoh</surname> <given-names>H</given-names></name> <name><surname>Gultekin</surname> <given-names>S</given-names></name> <name><surname>Dakic</surname> <given-names>A</given-names></name> <name><surname>Axelsson</surname> <given-names>E</given-names></name> <name><surname>Minnich</surname> <given-names>M</given-names></name> <etal/></person-group> <article-title>Stage-specific control of early B cell development by the transcription factor Ikaros</article-title>. <source>Nat Immunol</source> (<year>2014</year>) <volume>15</volume>(<issue>3</issue>):<fpage>283</fpage>&#x02013;<lpage>93</lpage>.<pub-id pub-id-type="doi">10.1038/ni.2828</pub-id><pub-id pub-id-type="pmid">24509509</pub-id></citation></ref>
<ref id="B71"><label>71</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>YC</given-names></name> <name><surname>Jhunjhunwala</surname> <given-names>S</given-names></name> <name><surname>Benner</surname> <given-names>C</given-names></name> <name><surname>Heinz</surname> <given-names>S</given-names></name> <name><surname>Welinder</surname> <given-names>E</given-names></name> <name><surname>Mansson</surname> <given-names>R</given-names></name> <etal/></person-group> <article-title>A global network of transcription factors, involving E2A, EBF1 and Foxo1, that orchestrates B cell fate</article-title>. <source>Nat Immunol</source> (<year>2010</year>) <volume>11</volume>(<issue>7</issue>):<fpage>635</fpage>&#x02013;<lpage>43</lpage>.<pub-id pub-id-type="doi">10.1038/ni.1891</pub-id><pub-id pub-id-type="pmid">20543837</pub-id></citation></ref>
<ref id="B72"><label>72</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bossen</surname> <given-names>C</given-names></name> <name><surname>Murre</surname> <given-names>CS</given-names></name> <name><surname>Chang</surname> <given-names>AN</given-names></name> <name><surname>Mansson</surname> <given-names>R</given-names></name> <name><surname>Rodewald</surname> <given-names>HR</given-names></name> <name><surname>Murre</surname> <given-names>C</given-names></name></person-group>. <article-title>The chromatin remodeler Brg1 activates enhancer repertoires to establish B cell identity and modulate cell growth</article-title>. <source>Nat Immunol</source> (<year>2015</year>) <volume>16</volume>(<issue>7</issue>):<fpage>775</fpage>&#x02013;<lpage>84</lpage>.<pub-id pub-id-type="doi">10.1038/ni.3170</pub-id><pub-id pub-id-type="pmid">25985234</pub-id></citation></ref>
<ref id="B73"><label>73</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spanopoulou</surname> <given-names>E</given-names></name> <name><surname>Roman</surname> <given-names>CA</given-names></name> <name><surname>Corcoran</surname> <given-names>LM</given-names></name> <name><surname>Schlissel</surname> <given-names>MS</given-names></name> <name><surname>Silver</surname> <given-names>DP</given-names></name> <name><surname>Nemazee</surname> <given-names>D</given-names></name> <etal/></person-group> <article-title>Functional immunoglobulin transgenes guide ordered B-cell differentiation in Rag-1-deficient mice</article-title>. <source>Genes Dev</source> (<year>1994</year>) <volume>8</volume>(<issue>9</issue>):<fpage>1030</fpage>&#x02013;<lpage>42</lpage>.<pub-id pub-id-type="doi">10.1101/gad.8.9.1030</pub-id><pub-id pub-id-type="pmid">7926785</pub-id></citation></ref>
<ref id="B74"><label>74</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Predeus</surname> <given-names>AV</given-names></name> <name><surname>Gopalakrishnan</surname> <given-names>S</given-names></name> <name><surname>Huang</surname> <given-names>Y</given-names></name> <name><surname>Tang</surname> <given-names>J</given-names></name> <name><surname>Feeney</surname> <given-names>AJ</given-names></name> <name><surname>Oltz</surname> <given-names>EM</given-names></name> <etal/></person-group> <article-title>Targeted chromatin profiling reveals novel enhancers in Ig H and Ig L chain loci</article-title>. <source>J Immunol</source> (<year>2014</year>) <volume>192</volume>(<issue>3</issue>):<fpage>1064</fpage>&#x02013;<lpage>70</lpage>.<pub-id pub-id-type="doi">10.4049/jimmunol.1302800</pub-id><pub-id pub-id-type="pmid">24353267</pub-id></citation></ref>
<ref id="B75"><label>75</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bevington</surname> <given-names>S</given-names></name> <name><surname>Boyes</surname> <given-names>J</given-names></name></person-group>. <article-title>Transcription-coupled eviction of histones H2A/H2B governs V(D)J recombination</article-title>. <source>EMBO J</source> (<year>2013</year>) <volume>32</volume>(<issue>10</issue>):<fpage>1381</fpage>&#x02013;<lpage>92</lpage>.<pub-id pub-id-type="doi">10.1038/emboj.2013.42</pub-id><pub-id pub-id-type="pmid">23463099</pub-id></citation></ref>
<ref id="B76"><label>76</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldmit</surname> <given-names>M</given-names></name> <name><surname>Ji</surname> <given-names>Y</given-names></name> <name><surname>Skok</surname> <given-names>J</given-names></name> <name><surname>Roldan</surname> <given-names>E</given-names></name> <name><surname>Jung</surname> <given-names>S</given-names></name> <name><surname>Cedar</surname> <given-names>H</given-names></name> <etal/></person-group> <article-title>Epigenetic ontogeny of the Igk locus during B cell development</article-title>. <source>Nat Immunol</source> (<year>2005</year>) <volume>6</volume>(<issue>2</issue>):<fpage>198</fpage>&#x02013;<lpage>203</lpage>.<pub-id pub-id-type="doi">10.1038/ni1154</pub-id><pub-id pub-id-type="pmid">15619624</pub-id></citation></ref>
<ref id="B77"><label>77</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kleiman</surname> <given-names>E</given-names></name> <name><surname>Jia</surname> <given-names>H</given-names></name> <name><surname>Loguercio</surname> <given-names>S</given-names></name> <name><surname>Su</surname> <given-names>AI</given-names></name> <name><surname>Feeney</surname> <given-names>AJ</given-names></name></person-group>. <article-title>YY1 plays an essential role at all stages of B-cell differentiation</article-title>. <source>Proc Natl Acad Sci U S A</source> (<year>2016</year>) <volume>113</volume>(<issue>27</issue>):<fpage>E3911</fpage>&#x02013;<lpage>20</lpage>.<pub-id pub-id-type="doi">10.1073/pnas.1606297113</pub-id><pub-id pub-id-type="pmid">27335461</pub-id></citation></ref>
<ref id="B78"><label>78</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>S</given-names></name> <name><surname>Pathak</surname> <given-names>S</given-names></name> <name><surname>Trinh</surname> <given-names>L</given-names></name> <name><surname>Lu</surname> <given-names>R</given-names></name></person-group>. <article-title>Interferon regulatory factors 4 and 8 induce the expression of Ikaros and Aiolos to down-regulate pre-B-cell receptor and promote cell-cycle withdrawal in pre-B-cell development</article-title>. <source>Blood</source> (<year>2008</year>) <volume>111</volume>(<issue>3</issue>):<fpage>1396</fpage>&#x02013;<lpage>403</lpage>.<pub-id pub-id-type="doi">10.1182/blood-2007-08-110106</pub-id><pub-id pub-id-type="pmid">17971486</pub-id></citation></ref>
<ref id="B79"><label>79</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mandal</surname> <given-names>M</given-names></name> <name><surname>Hamel</surname> <given-names>KM</given-names></name> <name><surname>Maienschein-Cline</surname> <given-names>M</given-names></name> <name><surname>Tanaka</surname> <given-names>A</given-names></name> <name><surname>Teng</surname> <given-names>G</given-names></name> <name><surname>Tuteja</surname> <given-names>JH</given-names></name> <etal/></person-group> <article-title>Histone reader BRWD1 targets and restricts recombination to the Igk locus</article-title>. <source>Nat Immunol</source> (<year>2015</year>) <volume>16</volume>(<issue>10</issue>):<fpage>1094</fpage>&#x02013;<lpage>103</lpage>.<pub-id pub-id-type="doi">10.1038/ni.3249</pub-id><pub-id pub-id-type="pmid">26301565</pub-id></citation></ref>
<ref id="B80"><label>80</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>C</given-names></name> <name><surname>Alt</surname> <given-names>FW</given-names></name> <name><surname>Giallourakis</surname> <given-names>C</given-names></name></person-group>. <article-title>PAIRing for distal Igh locus V(D)J recombination</article-title>. <source>Immunity</source> (<year>2011</year>) <volume>34</volume>(<issue>2</issue>):<fpage>139</fpage>&#x02013;<lpage>41</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2011.02.010</pub-id><pub-id pub-id-type="pmid">21349424</pub-id></citation></ref>
<ref id="B81"><label>81</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z</given-names></name> <name><surname>Widlak</surname> <given-names>P</given-names></name> <name><surname>Zou</surname> <given-names>Y</given-names></name> <name><surname>Xiao</surname> <given-names>F</given-names></name> <name><surname>Oh</surname> <given-names>M</given-names></name> <name><surname>Li</surname> <given-names>S</given-names></name> <etal/></person-group> <article-title>A recombination silencer that specifies heterochromatin positioning and ikaros association in the immunoglobulin kappa locus</article-title>. <source>Immunity</source> (<year>2006</year>) <volume>24</volume>(<issue>4</issue>):<fpage>405</fpage>&#x02013;<lpage>15</lpage>.<pub-id pub-id-type="doi">10.1016/j.immuni.2006.02.001</pub-id><pub-id pub-id-type="pmid">16618599</pub-id></citation></ref>
<ref id="B82"><label>82</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>ZM</given-names></name> <name><surname>George-Raizen</surname> <given-names>JB</given-names></name> <name><surname>Li</surname> <given-names>S</given-names></name> <name><surname>Meyers</surname> <given-names>KC</given-names></name> <name><surname>Chang</surname> <given-names>MY</given-names></name> <name><surname>Garrard</surname> <given-names>WT</given-names></name></person-group>. <article-title>Chromatin structural analyses of the mouse Igkappa gene locus reveal new hypersensitive sites specifying a transcriptional silencer and enhancer</article-title>. <source>J Biol Chem</source> (<year>2002</year>) <volume>277</volume>(<issue>36</issue>):<fpage>32640</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1074/jbc.M204065200</pub-id><pub-id pub-id-type="pmid">12080064</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup><uri xlink:href="https://github.com/FelixKrueger/BabrahamLinkON">https://github.com/FelixKrueger/BabrahamLinkON</uri>.</p></fn>
<fn id="fn2"><p><sup>2</sup><uri xlink:href="http://www.bioinformatics.babraham.ac.uk/projects/trim_galore/">http://www.bioinformatics.babraham.ac.uk/projects/trim_galore/</uri>.</p></fn>
<fn id="fn3"><p><sup>3</sup><uri xlink:href="http://www.bioinformatics.babraham.ac.uk/projects/seqmonk/">http://www.bioinformatics.babraham.ac.uk/projects/seqmonk/</uri>.</p></fn>
<fn id="fn4"><p><sup>4</sup><uri xlink:href="https://www.ncbi.nlm.nih.gov/nuccore/NG_005612/">https://www.ncbi.nlm.nih.gov/nuccore/NG_005612/</uri>.</p></fn>
<fn id="fn5"><p><sup>5</sup><uri xlink:href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc&#x0003D;GSE101606">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc&#x0003D;GSE101606</uri>.</p></fn>
</fn-group>
</back>
</article>