<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2017.00222</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Protocols</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>MLPA-Based Analysis of Copy Number Variation in Plant Populations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Samelak-Czajka</surname> <given-names>Anna</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/384218/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Marszalek-Zenczak</surname> <given-names>Malgorzata</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/382878/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Marcinkowska-Swojak</surname> <given-names>Malgorzata</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Kozlowski</surname> <given-names>Piotr</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/233326/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Figlerowicz</surname> <given-names>Marek</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zmienko</surname> <given-names>Agnieszka</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/279724/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute of Computing Science, Faculty of Computing, Poznan University of Technology</institution> <country>Poznan, Poland</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Molecular and Systems Biology, Institute of Bioorganic Chemistry, Polish Academy of Sciences</institution> <country>Poznan, Poland</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Molecular Genetics, Institute of Bioorganic Chemistry, Polish Academy of Sciences</institution> <country>Poznan, Poland</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Scott V. Edwards, Harvard University, USA</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Nathan Lewis Clark, University of Pittsburgh, USA; John Malone, University of Connecticut, USA; Pierre Baduel, John Innes Centre (BBSRC), USA</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Agnieszka Zmienko, <email>akisiel@ibch.poznan.pl</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Evolutionary and Population Genetics, a section of the journal Frontiers in Plant Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>02</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>222</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>02</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Samelak-Czajka, Marszalek-Zenczak, Marcinkowska-Swojak, Kozlowski, Figlerowicz and Zmienko.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Samelak-Czajka, Marszalek-Zenczak, Marcinkowska-Swojak, Kozlowski, Figlerowicz and Zmienko</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Copy number variants (CNVs) are intraspecies duplications/deletions of large DNA segments (>1 kb). A growing number of reports highlight the functional and evolutionary impact of CNV in plants, increasing the need for appropriate tools that enable locus-specific CNV genotyping on a population scale. Multiplex ligation-dependent probe amplification (MLPA) is considered a gold standard in genotyping CNV in humans. Consequently, numerous commercial MLPA assays for CNV-related human diseases have been created. We routinely genotype complex multiallelic CNVs in human and plant genomes using the modified MLPA procedure based on fully synthesized oligonucleotide probes (90&#x2013;200 nt), which greatly simplifies the design process and allows for the development of custom assays. Here, we present a step-by-step protocol for gene-specific MLPA probe design, multiplexed assay setup and data analysis in a copy number genotyping experiment in plants. As a case study, we present the results of a custom assay designed to genotype the copy number status of 12 protein coding genes in a population of 80 <italic>Arabidopsis</italic> accessions. The genes were pre-selected based on whole genome sequencing data and are localized in the genomic regions that display different levels of population-scale variation (non-variable, biallelic, or multiallelic, as well as CNVs overlapping whole genes or their fragments). The presented approach is suitable for population-scale validation of the CNV regions inferred from whole genome sequencing data analysis and for focused analysis of selected genes of interest. It can also be very easily adopted for any plant species, following optimization of the template amount and design of the appropriate control probes, according to the general guidelines presented in this paper.</p>
</abstract>
<kwd-group>
<kwd>structural variation</kwd>
<kwd>MLPA</kwd>
<kwd>1001 Arabidopsis Genomes project</kwd>
<kwd>CNV genotyping</kwd>
<kwd>multiplexing</kwd>
</kwd-group>
<contract-sponsor id="cn001">Narodowe Centrum Nauki<named-content content-type="fundref-id">10.13039/501100004281</named-content></contract-sponsor>
<contract-sponsor id="cn002">Ministerstwo Nauki i Szkolnictwa Wyzszego<named-content content-type="fundref-id">10.13039/501100004569</named-content></contract-sponsor>
<counts>
<fig-count count="6"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="48"/>
<page-count count="16"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>The rise of high-throughput genomics techniques &#x2013; DNA arrays and, more recently, whole-genome sequencing (WGS) &#x2013; has revealed the structural complexity and dynamics of eukaryotic genomes. In particular, the ability to re-sequence and compare hundreds or even thousands of genomes of individuals within one species has paved the way for the investigation of the extent to which individual genomes differ from each other. One type of structural variation that is ubiquitous in the genomes of humans, animals and plants is copy number variation (CNV). This term refers to intraspecies duplications and deletions of large DNA segments, usually >1 kb [although variants >50 bp have been recently included in this spectrum (<xref ref-type="bibr" rid="B2">Alkan et al., 2011</xref>)]. The human genome is the most intensively studied eukaryotic genome in terms of the distribution and functional significance of CNVs and the mechanisms leading to the formation of copy number rearrangements (<xref ref-type="bibr" rid="B45">Zarrei et al., 2015</xref>). However, the number of species for which CNV regions have been inferred on the genome-wide scale is growing rapidly. For plants, this list includes maize, rice, sorghum, <italic>Arabidopsis</italic> (<italic>Arabidopsis thaliana)</italic>, soybean, wheat, and barley (<xref ref-type="bibr" rid="B39">Springer et al., 2009</xref>; <xref ref-type="bibr" rid="B4">Bel&#x00F3; et al., 2010</xref>; <xref ref-type="bibr" rid="B41">Swanson-Wagner et al., 2010</xref>; <xref ref-type="bibr" rid="B8">Cao et al., 2011</xref>; <xref ref-type="bibr" rid="B37">Saintenac et al., 2011</xref>; <xref ref-type="bibr" rid="B46">Zheng et al., 2011</xref>; <xref ref-type="bibr" rid="B33">McHale et al., 2012</xref>; <xref ref-type="bibr" rid="B34">Mu&#x00F1;oz-Amatria&#x00ED;n et al., 2013</xref>; <xref ref-type="bibr" rid="B14">Duitama et al., 2015</xref>; <xref ref-type="bibr" rid="B3">Bai et al., 2016</xref>). As in humans, CNV regions in plants are not uniformly distributed across the chromosomes. Although they are more common in the intergenic regions, they also co-localize with hundreds of protein-coding genes (<xref ref-type="bibr" rid="B41">Swanson-Wagner et al., 2010</xref>; <xref ref-type="bibr" rid="B4">Bel&#x00F3; et al., 2010</xref>; <xref ref-type="bibr" rid="B33">McHale et al., 2012</xref>; <xref ref-type="bibr" rid="B34">Mu&#x00F1;oz-Amatria&#x00ED;n et al., 2013</xref>). The ability to alter the gene structure and copy number makes CNV an important factor that influences gene expression (<xref ref-type="bibr" rid="B47">&#x017B;mie&#x0144;ko et al., 2014</xref>). By the gene dosage effect, CNVs can also affect the interaction of the genes&#x2019; products within protein and metabolic networks (<xref ref-type="bibr" rid="B16">Hanada et al., 2011</xref>; <xref ref-type="bibr" rid="B11">Conant et al., 2014</xref>). Quite often, such variation accounts for adaptive traits or - as shown for humans - can underlie disease (<xref ref-type="bibr" rid="B40">Stankiewicz and Lupski, 2010</xref>; <xref ref-type="bibr" rid="B45">Zarrei et al., 2015</xref>). In plants, a growing number of studies highlight the shaping role of CNVs in genome evolution, phenotypic variation and &#x2013; sometimes rapid - adaptation to environmental challenges (<xref ref-type="bibr" rid="B15">Gaines et al., 2010</xref>; <xref ref-type="bibr" rid="B13">Cook et al., 2012</xref>; <xref ref-type="bibr" rid="B31">Maron et al., 2013</xref>; <xref ref-type="bibr" rid="B10">Chang et al., 2015</xref>; <xref ref-type="bibr" rid="B44">Wang et al., 2015</xref>). Therefore, it is anticipated that the number of genetic studies focused on individual CNVs of interest will grow and that new CNV-associated traits will be revealed.</p>
<p>In-depth analysis of individual CNVs in plants has rarely been conducted (<xref ref-type="bibr" rid="B15">Gaines et al., 2010</xref>; <xref ref-type="bibr" rid="B13">Cook et al., 2012</xref>; <xref ref-type="bibr" rid="B31">Maron et al., 2013</xref>). Likewise, in plants for which the CNV regions were inferred from WGS data, the subsequent validation was not conducted or was limited to the PCR-based detection of CNV deletions (<xref ref-type="bibr" rid="B41">Swanson-Wagner et al., 2010</xref>; <xref ref-type="bibr" rid="B8">Cao et al., 2011</xref>; <xref ref-type="bibr" rid="B42">Tan et al., 2012</xref>; <xref ref-type="bibr" rid="B3">Bai et al., 2016</xref>). Therefore, there is an urgent need to widen the range of experimental studies of CNV in plants to contribute to the creation of high-confidence CNV maps and enhance association studies linking CNVs with phenotypic traits in plant species. In this context, the lack of validated experimental approaches for the analysis of individual CNVs in plants is apparent, as opposed to the well-established methods and standardized protocols available for the human genome.</p>
<p>The range of popular molecular methods used for DNA copy number genotyping in humans is wide (<xref ref-type="bibr" rid="B9">Ceulemans et al., 2012</xref>; <xref ref-type="bibr" rid="B6">Cantsilieris et al., 2013</xref>; <xref ref-type="bibr" rid="B5">Bharuthram et al., 2014</xref>). Among them, multiplex ligation-dependent probe amplification (MLPA), first introduced in 2002 (<xref ref-type="bibr" rid="B38">Schouten et al., 2002</xref>) and later developed by the MRC Holland company, is considered a gold standard in the diagnosis of numerous DNA copy number-related human diseases (<xref ref-type="bibr" rid="B17">H&#x00F6;mig-H&#x00F6;lzel and Savola, 2012</xref>). MLPA is a simple and robust method of relative quantification of DNA sequences on a population scale. The standard multiplex assay utilizes up to 50 probes targeting specific DNA regions (e.g., exons in a gene of interest). Each probe is composed of two half-probes (physically separate DNA fragments, one fully synthetic and one clone-derived) that match the target sequence in directly adjacent positions with their target-specific sequences (TSSs). Successful hybridization of both half-probes to the genomic DNA enables their ligation and linear amplification. The amplification products are then analyzed by capillary electrophoresis. Relative quantification of the signal peaks from fragments of unique size, generated by individual probes in the assay, provides information about the template DNA copy number. MLPA requires little genomic DNA input (<xref ref-type="bibr" rid="B38">Schouten et al., 2002</xref>). Additionally, the genomic sequence targeted by the probes is quite short (50&#x2013;70 nt), which enables use of MLPA for the analysis of regions too small to be detected by the FISH method. MLPA has been shown to be superior to qPCR for gene copy number quantification (<xref ref-type="bibr" rid="B35">Perne et al., 2009</xref>; <xref ref-type="bibr" rid="B7">Cantsilieris et al., 2014</xref>). Additionally, it presents similar performance to droplet digital PCR in accurate quantification of up to eight gene copies, making it suitable for the analysis of multiallelic CNVs, i.e., those that exist in more than two genotypes in a population (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
<p>According to PubMed, the seminal MLPA work (<xref ref-type="bibr" rid="B38">Schouten et al., 2002</xref>) has been cited almost 450 times (&#x223C;220 times within 5 last years). Additionally, &#x223C;2,000 articles in PubMed matched the search keyword &#x201C;Multiplex Ligation-Dependent Probe Amplification&#x201D;. Among these papers, only 16 also matched the search keyword &#x201C;plant&#x201D;. Those that actually described plant applications of MLPA involved alternative applications of this method: the detection of genetically modified organisms (GMO-MLPA) (<xref ref-type="bibr" rid="B36">Rudi et al., 2003</xref>), single nucleotide polymorphism (SNP) genotyping (<xref ref-type="bibr" rid="B43">Thumma et al., 2009</xref>), or gene expression analysis (RT-MLPA) (<xref ref-type="bibr" rid="B24">Li et al., 2009</xref>, <xref ref-type="bibr" rid="B25">2011</xref>, <xref ref-type="bibr" rid="B26">2013</xref>). However, none of these papers presented a primary MLPA application of copy number analysis. Several reasons might account for the fact that the MLPA approach has not been adopted by the plant community. One is much later recognition of the intraspecies variation and CNV prevalence in the plant genomes than in humans. Additionally, the commercial MLPA assays are focused on biomedical studies and cover only humans. Therefore, to assess plant genome variation with MLPA, it is necessary to self-design synthetic probes. It should be noted that, over the years, numerous modifications of the MLPA strategy have been introduced that simplify the probe design procedure (<xref ref-type="bibr" rid="B28">Marcinkowska et al., 2010</xref>; <xref ref-type="bibr" rid="B27">Ling et al., 2015</xref>, and references therein). In the current work, we present the optimized protocol for MLPA-based CNV analysis and provide guidelines for designing and performing MLPA assays in plants. The protocol is based on the MLPA adaptation developed previously by one of us (PK) that involves fully synthetic oligonucleotide probes, 90 to 200 nt in length, and allows for simultaneous genotyping of >30 different positions in the genomic DNA (<xref ref-type="bibr" rid="B23">Kozlowski et al., 2007</xref>). The protocol combines MLPA probe design, synthesis, experimental procedures, data preprocessing and analysis stages into one comprehensive procedure. The lack of MLPA-based genotyping studies in plants highlights the need for such an integrated resource. We also provided the probe design template, developed specifically for the presented MLPA variant. It allows for semi-automatic probe sequence setup, clarifies the idea of probe set composition and shortens the design process by days.</p>
<p>High and low copy level duplications may have different effects on the gene dosage and the phenotype, e.g., by triggering differences in gene expression level or inducing the silencing mechanisms in plants. Therefore, an important aspect of plant CNV genotyping studies is to estimate the actual gene copy numbers in the analyzed lines in order to analyze their influence on the trait of interest (<xref ref-type="bibr" rid="B12">Cook et al., 2014</xref>). To illustrate the performance of the MLPA method for precise DNA copy number genotyping in plant populations, we present exemplar assays for 12 genes with different levels of copy number diversity in a population of 80 <italic>Arabidopsis</italic> ecotypes, including multiallelic CNVs. We also describe the set of experimentally verified normalization control probes and the results of genomic DNA template amount optimization performed for this model species.</p>
<p>An advantage of the presented approach is that the assay - after it has been standardized for the particular organism &#x2013; is always performed in the same conditions, regardless of the probe set composition. It may be utilized for the detailed analysis of a genomic region of interest using a set of MLPA probes scattered along this region or for large-scale validation/genotyping studies of WGS-based predicted CNVs, with 1-2 MLPA probes per inferred CNV.</p>
</sec>
<sec><title>Materials and Equipment</title>
<sec><title>Materials</title>
<list list-type="simple" prefix-word="simple">
<list-item><label>(1)</label><p>High-quality genomic DNA for each analyzed sample, evaluated using a NanoDrop 2000 spectrophotometer (Thermo Scientific) and with standard gel electrophoresis; the working concentration is typically 0.4 to 50 ng/&#x03BC;l, depending on the species (see the following sections).</p>
<list list-type="simple" prefix-word="simple">
<list-item><p>For <italic>Arabidopsis</italic>: We successfully genotyped CNVs using genomic DNA from 3-week-old rosette leaves extracted with a DNeasy Plant Mini Kit (Qiagen).</p></list-item>
</list></list-item>
<list-item><label>(2)</label><p>Self-designed synthetic oligonucleotides (MLPA half-probes; see the following section for the probe design instructions) purchased from Integrated DNA Technologies (or similar provider) as 100 nmol oligo, purified by HPLC (for oligonucleotides up to 100 nt in length) or PAGE (for oligonucleotides over 100 nt in length); the right half-probes should be additionally modified by 5&#x2032; phosphorylation.</p></list-item>
<list-item><label>(3)</label><p>Nuclease-free water (not DEPC-treated) (Ambion, cat. no. AM9938)</p></list-item>
<list-item><label>(4)</label><p>SALSA MLPA EK-1 reagent kit (MRC-Holland, cat. no. EK1-FAM), which includes the following components:</p>
<list list-type="simple" prefix-word="simple">
<list-item><p>SALSA MLPA Buffer</p></list-item>
<list-item><p>SALSA Ligase-65</p></list-item>
<list-item><p>Ligase Buffer A</p></list-item>
<list-item><p>Ligase Buffer B</p></list-item>
<list-item><p>SALSA PCR Primer MIX</p></list-item>
<list-item><p>SALSA Polymerase</p></list-item>
</list></list-item>
<list-item><label>(5)</label><p>Consumables for capillary electrophoresis, depending on the instrument type; here, for the ABI Prism 3130XL Genetic Analyzer:</p>
<list list-type="simple" prefix-word="simple">
<list-item><p>HiDi formamide (Thermo Fisher Scientific, cat. no. 4440753)</p></list-item>
<list-item><p>GeneScan 600 LIZ Size Standard (Thermo Fisher Scientific, cat. no 4366589)</p></list-item>
<list-item><p>POP7 Polymer (Thermo Fisher Scientific, cat. no 4352759).</p></list-item></list>
</list-item></list>
</sec>
<sec><title>Equipment</title>
<list list-type="simple" prefix-word="simple">
<list-item><label>(1)</label><p>0.2 ml PCR strips and suitable caps, e.g., 8-Strip PCR tubes (Starlab, cat. no. I1402-3500) and 8-Strip caps (Starlab, cat. no. I1400-0800).</p></list-item>
<list-item><label>(2)</label><p>Standard and multichannel pipettes.</p></list-item>
<list-item><label>(3)</label><p>Thermocycler with heated lid (e.g., Bio-Rad T100 Thermal Cycler or equivalent).</p></list-item>
<list-item><label>(4)</label><p>Vortex mixer (e.g., ELMI V-3 Sky Line or equivalent).</p></list-item>
<list-item><label>(5)</label><p>Mini laboratory centrifuge with Eppendorf tube adapter and PCR strip adapter (e.g., Labnet Spectrafuge or equivalent).</p></list-item>
<list-item><label>(6)</label><p>Capillary electrophoresis instrument (AppliedBiosystems ABI Prism 3130XL Genetic Analyzer or equivalent) or access to a capillary electrophoresis service provider.</p></list-item>
<list-item><label>(7)</label><p>Software tool for the extraction of the intensity data after size-separation of MLPA reaction products (e.g., GeneMarker by SoftGenetics).</p></list-item></list>
</sec>
</sec>
<sec><title>Stepwise Procedures</title>
<p>The general concept of the MLPA strategy is presented in <bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>. The entire procedure involves three main stages: (A) designing the MLPA probes; (B) performing MLPA assay, which involves half-probes hybridization to DNA template, subsequent ligation and amplification; and (C) data collection and analysis, including the estimation of the copy number genotypes.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p><bold>Overview of the multiplex ligation-dependent probe amplification (MLPA) method.</bold> MLPA is comprised of three main stages: designing the probes, performing the multiplex MLPA assay and data collection and analysis. The three stages are described in the detail in the main text. TSS, target-specific sequence; 5&#x2032;-phos: phosphorylation at the 5&#x2032; end of the oligonucleotide. Note that <italic>Arabidopsis</italic>, as a self-pollinating plant, typically carries pairs of identical alleles. For simplicity, single alleles are depicted.</p></caption>
<graphic xlink:href="fpls-08-00222-g001.tif"/>
</fig>
<sec><title>Stage A: Design the MLPA Probes (Time: Approximately 1 Week + Oligonucleotide Synthesis and Transportation by an External Provider)</title>
<p>The presented MLPA procedure based on fully synthetic oligonucleotide probes allows for simultaneous copy number analysis of &#x223C;30 individual regions in the genomic DNA. Of these, at least 3 to 5 MLPA probes should target the confirmed non-variable control regions, distant from the studied genomic positions. These probes serve as normalization controls in the subsequent analysis of the MLPA data to account for the possible variation of the input DNA template amount and technical issues. The typical targets of the MLPA assays are protein-coding genes, as the changes in their copy number potentially affect the protein level and may contribute to the phenotype. The number of probes designed for each gene and their density in the covered genomic region depend on the user&#x2019;s requirements.</p>
<p>The procedure for individual MLPA probe design has been graphically presented in Supplementary Figure <xref ref-type="supplementary-material" rid="SM2">S1</xref> and is described in detail in the following sections. We used <italic>Arabidopsis</italic> gene <italic>AT1G01040</italic> encoding Dicer-like 1 protein as an example.</p>
<sec><title>Select TSSs for the MLPA Probes</title>
<p><bold>Step 1.</bold> Retrieve the genomic sequence of the gene of interest from the appropriate database, including the exon-intron positions. We recommend localizing the MLPA probes within the exon sequences because they display lower variation than the non-coding regions of genes.</p>
<p>For <italic>Arabidopsis:</italic> Use the gene locus identifier (e.g., <italic>AT1G01040</italic>) to localize that gene in the TAIR10 genomic sequence, available through the <italic>Arabidopsis</italic> genome browser<sup><xref ref-type="fn" rid="fn01">1</xref></sup>, and display its splice variants, when applicable (Protein Coding Gene Models track). In <italic>Arabidopsis</italic>, protein coding genes have five exons on average, each with mean length of &#x223C;240 bp (<xref ref-type="bibr" rid="B22">Koralewski and Krutovsky, 2011</xref>). This length is sufficient for selecting two adjacent TSSs (one for each half-probe). Use the GBrowse navigation tools to zoom in to the selected exon and export its DNA sequence as a FASTA file.</p>
<p><bold>Step 2.</bold> Ensure your sequence does not include any repetitive elements.</p>
<p>For <italic>Arabidopsis</italic>, rice, maize, wheat, and some other crops: Submit the extracted sequence to the CENSOR software tool (<xref ref-type="bibr" rid="B20">Kohany et al., 2006</xref>) that masks the repetitive elements in the query sequence using the collection of repeats for selected animal and plant species. Select a fragment of at least 100 nt that is not interrupted by any masked regions.</p>
<p><bold>Step 3.</bold> If possible, check the selected sequence for the presence of SNPs and small indels.</p>
<p>For <italic>Arabidopsis</italic>: Use the 1001 Genomes Project VCF Subset tool<sup><xref ref-type="fn" rid="fn02">2</xref></sup> to download the subset of VCF files that contain full-genome VCF data for 1135 accessions (as of September 2016) (<xref ref-type="bibr" rid="B1">1001 Genomes Consortium, 2016</xref>). Download SNP information for the region and accessions of interest. Evaluate whether the selected sequence is free of common polymorphisms.</p>
<p><bold>Step 4.</bold> From the selected region, choose two directly adjacent fragments of at least 21 nt (left and right TSS) and adjust their length and position so that the melting temperature (Tm) of each fragment will be as close as possible to 71&#x00B0;C (calculated with the free RaW program available from MRC Holland<sup><xref ref-type="fn" rid="fn03">3</xref></sup> with the following settings: method Go-Oli-Go, salt concentration 0.1 M, oligo concentration 1 &#x03BC;m). Avoid long homopolymer tracts and GC tracts of &#x2265;4 bases.</p>
<p><bold>Step 5.</bold> Join the adjacent left and right TSSs and use the resulting sequence in a homology search against the genomic sequence of the analyzed species to check for its specificity.</p>
<p>For <italic>Arabidopsis</italic>: Perform a BLAST search against <italic>A. thaliana</italic> NCBI reference genome with the following parameters: blastn algorithm, word size 7, match/mismatch scores 2;-3, gap costs 5;2, no sequence masking and filtering, <italic>E</italic>-value threshold 0.001.</p>
<p><bold>Step 6.</bold> Repeat steps 3 to 5 until the pair of adjacent TSSs that satisfies all design criteria is found for a given gene.</p>
</sec>
<sec><title>Design the Half-Probes</title>
<p><bold>Step 7.</bold> Add the respective PCR primer annealing sequence to each TSS and &#x2013; optionally &#x2013; the stuffer sequence, in the following order (see <bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>):</p>
<p>for the left half-probe:</p>
<p>5&#x2032;-left primer annealing sequence &#x2013; stuffer &#x2013; left TSS -3&#x2032;,</p>
<p>where the left primer annealing sequence is GGGTTCCCTAAGGGTTGGA;</p>
<p>for the right half-probe:</p>
<p>5&#x2032;-right TSS &#x2013; stuffer &#x2013; right primer annealing sequence &#x2013; 3&#x2032;,</p>
<p>where the right primer annealing sequence is TCTAGATTGGATCTTGCTGGCGC.</p>
<p>For the stuffer, use the fragment of enterobacteria phage M13 sequence (NCBI/GenBank ID V00604, range: 3-119). This fragment has no significant blastn matches to any eukaryotic genomic sequence deposited in the NCBI/RefSeq Representative Genome Database (accessed July 4th, 2016). It has been successfully applied as a stuffer in our previous MLPA assays performed for <italic>Arabidopsis</italic> and human DNA (<xref ref-type="bibr" rid="B30">Marcinkowska-Swojak et al., 2014</xref>; <xref ref-type="bibr" rid="B19">Klonowska et al., 2015</xref>; <xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
<p><italic>Note:</italic> The addition of the optional stuffer sequence allows the user to adjust the length of the half-probes so that the resulting PCR amplification fragments would be of unique size and differ by 3 nt for probes in the 90-120 nt range and by 4 nt for probes >120 nt long. The length of the two half-probes in the pair should be the same or differ by 1 nt. For example, to obtain the MLPA probe of length 120, the left and right half-probe sequences should each be 60 nt long (and at least 21 nt of each half-probe should constitute TSS).</p>
<p>To facilitate the process of MLPA probe design and combining multiple MLPA probes in one experimental assay, we provided a Microsoft Excel template (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>). This template includes the formulas that automatically adjust the length of the stuffer sequence and add the required adapter sequences to both the left and right half-probes. As a result, the final sequence of the MLPA probe of the desired length is returned. The user can choose the MLPA probe length. Typically, when fewer than the maximal number of MLPA probes are included in the assay, we recommend designing shorter probes to minimize the oligonucleotide synthesis costs. Often, the MLPA assays contain two or more probes targeting adjacent genomic regions. We recommend randomization of these probe MLPA lengths to minimize the influence of the possible biases or artifacts. Likewise, we recommend distributing the control probe lengths to cover the entire range of the MLPA probes in the assay.</p>
<p>For <italic>Arabidopsis</italic>: We provide pre-designed sequences for five control MLPA probes (ctrl1&#x2013;ctrl5) that target genes located on chromosomes 1, 2, 4, and 5. The first gene is <italic>DCL1</italic>, coding for a RNA helicase involved in microRNA processing. The second gene encodes an oxidoreductase belonging to a zinc-binding dehydrogenase family protein. The third non-variable gene is <italic>APG10</italic>, coding for a BBMII isomerase involved in histidine biosynthesis. The fourth gene is <italic>PDF5</italic>, coding for a prefoldin, involved in unfolded protein binding. The fifth gene is <italic>PS2</italic>, coding for a pyrophosphate-specific phosphatase. The lengths of the probes cover the entire range of the MLPA assay (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>). The regions were selected as not copy-number variable in <italic>Arabidopsis</italic> based on WGS data and were experimentally validated in 189 natural accessions (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
</sec>
<sec><title>Order the Oligonucleotide Synthesis</title>
<p>The synthesis of the designed MLPA probes is typically performed by an external service provider, such as Integrated DNA Technologies (IDT).</p>
<p><bold>Step 8.</bold> Order the synthesis of left and right half-probes, each as separate oligonucleotides, at a 100-nmol scale. All right half-probes must be additionally modified at their 5&#x2032; ends (5&#x2032; phosphorylation).</p>
<p><italic>Caution:</italic> 5&#x2032; phosphorylation of the right half-probes is essential for a successful ligation step (described below). The oligonucleotides designed for MLPA assays should be of high purity; therefore, we recommend selecting a PAGE or HPLC purification option, depending on the oligonucleotide length and according to the oligonucleotide manufacturer&#x2019;s recommendations.</p>
<p><bold>Step 9.</bold> Re-dissolve the lyophilized oligonucleotides upon arrival in deionized water to a concentration of 20 &#x03BC;M. Alternatively, the oligonucleotides can be re-dissolved in 10 mM Tris-HCl, pH 8.2.</p>
<p><bold>Step 10.</bold> Store the half-probe stocks at &#x2013;20&#x00B0;C.</p>
</sec>
</sec>
<sec><title>Stage B. Perform MLPA Assay (Time: 2 Days)</title>
<p><italic>Note:</italic> When performing the MLPA assay, keep all reagents, stock solutions and working solutions on ice. Set up the reactions in PCR tubes or strips (recommended) at room temperature, unless indicated otherwise. Depending on the user&#x2019;s experience, we recommend running assays for 8&#x2013;32 samples at once in 1&#x2013;4 PCR strips.</p>
<p><italic>Note:</italic> Whenever applicable, prepare the reagent master mixes for all assayed samples with 10% volume surplus to minimize sample-to-sample variation and save pipetting time. Distribute the master mix to eight tubes of a new PCR strip and then transfer the required amount to all PCR strips containing your samples with a multichannel pipette.</p>
<p><italic>Note:</italic> Perform all incubation steps in a thermocycler, programmed as specified in <bold>Table <xref ref-type="table" rid="T1">1</xref></bold>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Programmed thermocycler conditions for multiplex ligation-dependent probe amplification (MLPA) assay.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Program</th>
<td valign="top" align="left"></td>
<th valign="top" align="left">Action</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="3"><bold>Denaturation (Step 5)</bold></td></tr>
<tr>
<td valign="top" align="left">98&#x00B0;C, 5 min;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Denature samples.</td>
</tr>
<tr>
<td valign="top" align="left">25&#x00B0;C, &#x221E;;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Cool down samples before removing.</td>
</tr>
<tr>
<td valign="top" align="left">Pause</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Proceed to Step 6.</td>
</tr>
<tr>
<td valign="top" align="left" colspan="3"><bold>Hybridization (Steps 9-10)</bold></td></tr>
<tr>
<td valign="top" align="left">95&#x00B0;C, 1 min;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Hybridize half-probes to their genomic targets.</td>
</tr>
<tr>
<td valign="top" align="left">60&#x00B0;C, 16&#x2013;20 h;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left"></td>
</tr>
<tr>
<td valign="top" align="left">54&#x00B0;C, &#x221E;;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Adjust the temperature for the next step.</td>
</tr>
<tr>
<td valign="top" align="left">Pause</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Proceed to Step 11.</td>
</tr>
<tr>
<td valign="top" align="left" colspan="3"><bold>Ligation (Step 14)</bold></td></tr>
<tr>
<td valign="top" align="left">54&#x00B0;C, 15 min;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Ligate adjacently hybridized half-probes.</td>
</tr>
<tr>
<td valign="top" align="left">98&#x00B0;C, 5 min;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Inactivate the enzyme.</td>
</tr>
<tr>
<td valign="top" align="left">20&#x00B0;C, &#x221E;;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Cool down samples before removing.</td>
</tr>
<tr>
<td valign="top" align="left">Pause</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Proceed to Step 15.</td>
</tr>
<tr>
<td valign="top" align="left" colspan="3"><bold>Amplification (Step 18)</bold></td></tr>
<tr>
<td valign="top" align="left">35 cycles of:</td>
<td valign="top" align="left">95&#x00B0;C, 30 s;</td>
<td valign="top" align="left">Amplify the correctly ligated MLPA probes.</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">60&#x00B0;C, 30 s;</td>
<td valign="top" align="left"></td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">72&#x00B0;C, 1 min;</td>
<td valign="top" align="left"></td>
</tr>
<tr>
<td valign="top" align="left">72&#x00B0;C, 20 min;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Perform final extension of PCR products.</td>
</tr>
<tr>
<td valign="top" align="left">4&#x00B0;C, &#x221E;;</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Cool down samples before removing.</td>
</tr>
<tr>
<td valign="top" align="left">End</td>
<td valign="top" align="left"></td>
<td valign="top" align="left">Proceed to Step 19.</td></tr>
</tbody>
</table>
</table-wrap>
<p><italic>Caution:</italic> Do not vortex the tubes containing Ligase-65 or Salsa Polymerase enzymes. Likewise, do not vortex the master mixes after adding any of these enzymes.</p>
<sec><title>Prepare the MLPA Probe Set Mix</title>
<p>The correctly composed assay should include both half-probes (left and right) for each region of interest. Each pair of half-probes should generate a ligation product of unique length in the assay. The concentration of the MLPA probes in the final reaction mixture is very low (see below); therefore, it is convenient to perform a two-step oligonucleotide dilution during the probe set mix preparation as follows.</p>
<p><bold>Step 1.</bold> Melt all half-probe stocks constituting one assay.</p>
<p><bold>Step 2.</bold> Dilute each 20 &#x03BC;M stock with water to a 0.2 &#x03BC;M working solution (200 &#x03BC;l).</p>
<p><bold>Step 3</bold>. Mix 2 &#x03BC;l of each half-probe working solution and fill to 400 &#x03BC;l with water.</p>
<p>The resulting 1 nM MLPA Probe Set Mix will contain all the desired pairs of half-probes in equal concentrations and is directly applicable in the reaction setup.</p>
<p><italic>Note:</italic> MLPA Probe Set Mix can be stored at &#x2013;20&#x00B0;C until later use.</p>
</sec>
<sec><title>Hybridize Half-Probes</title>
<p>For each genomic DNA sample, perform the MLPA assay in a separate tube. We recommend running MLPA assays in multiples of 8 in PCR strips with caps.</p>
<p><italic>Caution:</italic> Replace the strip caps with new ones at each opening during the entire procedure to prevent cross-contamination.</p>
<p><bold>Step 4.</bold> Aliquot 5 &#x03BC;l of genomic DNA (0.4 to 50 ng/&#x03BC;l) to individual strip tubes to obtain a final template amount of 2&#x2013;250 ng per assay, depending on the species.</p>
<p><italic>Note:</italic> We recommend performing template optimization assays for each species.</p>
<p>For <italic>Arabidopsis:</italic> We successfully performed MLPA assays using 2, 5, 10, 15, 30, 60, and 100 ng genomic DNA per assay (see the next section).</p>
<p><bold>Step 5.</bold> Insert the samples into the thermocycler. Heat for 5 mins at 98&#x00B0;C then let the samples cool to 25&#x00B0;C.</p>
<p><bold>Step 6.</bold> Remove the samples from the thermocycler and centrifuge.</p>
<p><bold>Step 7.</bold> Prepare master mix I. Briefly vortex and centrifuge the SALSA MLPA buffer and MLPA Probe Set Mix. Prepare the adequate amount of the master mix I by mixing 1.5 &#x03BC;l of SALSA MLPA buffer and 1.5 &#x03BC;l of 1 nM MLPA Probe Set Mix per sample, with 10% volume surplus. Vortex and centrifuge the tube.</p>
<p><bold>Step 8.</bold> Add 3 &#x03BC;l of the master mix I to each denatured DNA sample and mix briefly by pipetting. Close the strips with the new caps and centrifuge. The reaction volume in each tube should be 8 &#x03BC;l.</p>
<p><bold>Step 9.</bold> Put the samples back into the thermocycler and incubate for 1 min at 95&#x00B0;C, then for 16 to 18 h at 60&#x00B0;C.</p>
<p><bold>Step 10.</bold> Adjust the thermoblock temperature to 54&#x00B0;C before proceeding to the next step.</p>
<p><italic>Caution:</italic> Do NOT remove the samples from the thermocycler!</p>
</sec>
<sec><title>Ligate the Hybridized Half-Probes</title>
<p><bold>Step 11.</bold> Prepare master mix II without enzyme. Briefly vortex and centrifuge Ligase Buffer A and Ligase Buffer B. Mix 3 &#x03BC;l of Ligase Buffer A, 3 &#x03BC;l of Ligase Buffer B, and 25 &#x03BC;l of nuclease-free water per sample, with 10% volume surplus. Vortex and centrifuge the tube.</p>
<p><bold>Step 12.</bold> Centrifuge the tube containing SALSA Ligase-65 enzyme. Add 1 &#x03BC;l of the enzyme per sample with 10% volume surplus to the master mix II. Mix briefly by pipetting. Centrifuge the tube and store on ice until use. Proceed to the next step without delay.</p>
<p><bold>Step 13.</bold> Without removing the strips from the thermocycler, add 32 &#x03BC;l of master mix II to each sample. Mix by pipetting and close the strips with new caps. The reaction volume in each tube should be 40 &#x03BC;l.</p>
<p><bold>Step 14.</bold> Incubate the samples for 15 min at 54&#x00B0;C, followed by heat inactivation of the ligase enzyme (5 min at 98&#x00B0;C). Cool the thermoblock to 20&#x00B0;C and remove the samples.</p>
</sec>
<sec><title>Amplify the Ligated MLPA Probes</title>
<p><bold>Step 15.</bold> Prepare master mix III. Briefly vortex and centrifuge the SALSA PCR primer mix. Mix 2 &#x03BC;l of SALSA PCR primer mix and 7.5 &#x03BC;l of nuclease-free water per sample, with 10% volume surplus. Vortex and centrifuge the tube.</p>
<p><bold>Step 16.</bold> Centrifuge the tube containing SALSA Polymerase enzyme. Heat the tube in hands for approximately 10 s, then add 0.5 &#x03BC;l of the enzyme per sample with 10% volume surplus to master mix III. Mix briefly by pipetting. Centrifuge the tube and store on ice until use.</p>
<p><bold>Step 17.</bold> Add 10 &#x03BC;l of master mix III to each sample and mix by pipetting. Close the strips with new caps and replace in the thermocycler. The final reaction volume in each tube should be 50 &#x03BC;l.</p>
<p><bold>Step 18.</bold> Perform the PCR comprising 35 cycles of: 95&#x00B0;C for 30 s; 60&#x00B0;C for 30 s and 72&#x00B0;C for 1 min, followed by a 20 min final elongation at 72&#x00B0;C. Cool the thermoblock to 4&#x00B0;C.</p>
<p><bold>Step 19.</bold> Store the samples at 4&#x00B0;C, protected from light, until the product size-separation (1&#x2013;3 days).</p>
</sec>
</sec>
<sec><title>Stage C. Collect and Analyze the Data (Time: 1 Day for the Data Collection, Variable for the Analysis)</title>
<sec><title>Size-Separate the PCR Products by Capillary Electrophoresis</title>
<p>The product separation should be performed under denaturing conditions on any standard capillary DNA analyzer. The specific run parameters must be adjusted according to the recommendations of the instrument manufacturer.</p>
<p>We typically use the services of the local Molecular Biology Techniques facility (at the Department of Biology of Adam Mickiewicz University, Poznan, Poland) and separate the samples in ABI Prism 3130XL Genetic Analyzer (Applied Biosystems), using the following procedure.</p>
<p><bold>Step 1.</bold> Each MLPA reaction sample is diluted 20&#x00D7; with nuclease-free water, mixed with 9 &#x03BC;l of HiDi formamide (Thermo Fisher Scientific) containing GeneScan 600 LIZ Size Standard (Thermo Fisher Scientific) and denatured.</p>
<p><bold>Step 2.</bold> Samples are injected at 1.2 kV voltage and separated on ABI Prism 3130XL Genetic Analyzer (Applied Biosystems) at 15 kV, in POP7 separation matrix (Thermo Fisher Scientific).</p>
</sec>
<sec><title>Analyze the Electropherograms</title>
<p>Evaluate the data quality and extract the signal intensity from the electropherograms. Numerous software tools are appropriate for this purpose. Below, we describe the step-by-step analysis performed with GeneMarker (SoftGenetics) (Supplementary Figure <xref ref-type="supplementary-material" rid="SM3">S2</xref>).</p>
<p><italic>Note:</italic> The GeneMarker functions used here are accessible in the limited demo version of the software, freely downloadable from the manufacturer&#x2019;s web site. The details regarding use of these functions are described in the software manual, also available for download.</p>
<p><bold>Step 3.</bold> Load the electropherogram data to GeneMarker.</p>
<p><bold>Step 4.</bold> Analyze the raw data files with the MLPA analysis type option and appropriate DNA standard selected (depending on the capillary electrophoresis conditions). Select the size call method and data normalization approach (Supplementary Figure <xref ref-type="supplementary-material" rid="SM3">S2A</xref>).</p>
<p><italic>Note:</italic> GeneMarker software provides two normalization options (intra-sample &#x201C;Internal Control Probe Normalization&#x201D; and inter-sample &#x201C;Population Normalization&#x201D;) that aim to correct for the variation in signal intensity caused by the differences in the lengths of the probes in the multiplex assay. We typically use the intra-sample normalization against our control probes, although at this step it is not critical, because the range of the probe lengths in our assay (96&#x2013;200 nt) is much smaller than in the case of commercial MLPA assays (130&#x2013;490 nt).</p>
<p><italic>Caution:</italic> Use the same parameter settings for all samples. When applying internal control probe normalization, use the same set of control probes for analysis of all samples in the MLPA assay.</p>
<p><italic>Note:</italic> At the first analysis of a new MLPA assay, run the analysis for a selection of samples using the &#x201C;NONE&#x201D; panel selection. This will allow you to manually create the custom MLPA panel later by indicating the peak positions in your pre-processed samples (see Step 5). If the MLPA panel has already been created, select that panel for the final analysis of all your samples.</p>
<p><bold>Step 5.</bold> Perform this step for the new MLPA assay only. Manually create the probe panel with the Panel Editor (Supplementary Figure <xref ref-type="supplementary-material" rid="SM3">S2B</xref>). Use the pre-processed set of representative MLPA electropherograms (see Step 4) to locate and insert the alleles at the expected positions. Label the alleles with the MLPA probe names. If you want to use the &#x201C;Internal Control Probe Normalization&#x201D; option during the analysis, mark the control probes as 1. Repeat Step 4 to re-run all samples using the newly created panel.</p>
<p><italic>Note:</italic> In our assays, all peak sizes consistently appeared &#x223C;3 bp shorter than the theoretical length of their attributed MLPA probes. This is not an unexpected result because the migration times of the peak maxima depend on many factors, including the amount of the sample injected, the temperature and the dye used. The capillary electrophoresis systems estimate the relative allele size (using internal standard) and do not necessarily report the true fragment size (<xref ref-type="bibr" rid="B32">McCord, 2003</xref>). Therefore, the observed shift is specific to the system and MLPA assay conditions. As long as the peaks are consistently observed at the same positions in all samples under comparison, it does not influence the peak discrimination and subsequent analysis of the MLPA data.</p>
<p><bold>Step 6.</bold> Evaluate the quality of individual electropherograms in accordance with the peak pattern of the size standard, the electrophoresis baseline, signal sloping and overall signal intensity. Samples that show abnormalities should be excluded from the analysis.</p>
<p><bold>Step 7.</bold> Configure the report layout and copy the results to MS Excel or similar program for further analysis (Supplementary Figure <xref ref-type="supplementary-material" rid="SM3">S2C</xref>).</p>
<p><italic>Note:</italic> The processed data can be reported as the fluorescence intensity (peak height) or the peak area values for each allele. The choice of the output typically does not affect the downstream data analysis and we obtained comparable results with both options. We preferably use the fluorescence intensity data.</p>
</sec>
<sec><title>Estimate the DNA Copy Number</title>
<p><bold>Step 8.</bold> Use the normalization controls to perform within-sample normalization of all your sample data before comparison.</p>
<p><italic>For Arabidopsis:</italic> Use at least 3 of the provided control probes (ctrl1&#x2013;ctrl5) for normalization. Divide each intensity value by the average intensity of the control probes, separately for each sample.</p>
<p><bold>Step 9.</bold> For each region analyzed, compare the normalized intensity between the samples. Cluster the samples with the similar intensities and infer the copy numbers from analysis of histograms or two-dimensional plots (see next section). Whenever possible, use the (set of) positive and negative control samples with known copy number status to determine the duplication/deletion intensity thresholds (see the next section for exemplar results).</p>
</sec>
</sec></sec>
<sec><title>Anticipated Results</title>
<sec><title>Exemplar MLPA Assay</title>
<p>Based on the available WGS data from 1001 Arabidopsis Genomes Project (<xref ref-type="bibr" rid="B1">1001 Genomes Consortium, 2016</xref>) and our own analysis of a subset of this data including 80 accessions, originally described in (<xref ref-type="bibr" rid="B8">Cao et al., 2011</xref>), we selected 12 genes that overlapped CNVs with various levels of structural complexity. Genes <italic>AT1G47670</italic> and <italic>AT1G80830</italic> do not present copy number changes. Genes <italic>AT1G32300</italic> and <italic>AT4G19520</italic> are biallelic; more specifically, they display presence-absence variation. The remaining eight genes are multiallelic and present duplications (<italic>AT4G27080, AT5G09590</italic>, and <italic>AT5G61700</italic>) or duplications and deletions (<italic>AT1G27570, AT1G52950, AT3G21960, AT4G27080</italic>, and <italic>AT5G54710</italic>). Additionally, gene <italic>AT5G09590</italic> overlaps CNV only partially, whereas <italic>AT1G52950, AT5G54710</italic>, and <italic>AT1G27570</italic> are members of multigene families and are localized in the regions of high structural diversity (manifested e.g., by the presence of adjacent or overlapping CNVs, presence of nearby transposable element genes or the presence of clusters of highly similar paralogs). To present the performance of the MLPA approach we set up a multiplex assay Ath.test for these genes (<bold>Table <xref ref-type="table" rid="T2">2</xref></bold>). We evaluated the genes&#x2019; copy number status in 80 <italic>Arabidopsis</italic> accessions, characterized in the first stage of 1001 <italic>Arabidopsis</italic> Genomes Project (<xref ref-type="bibr" rid="B8">Cao et al., 2011</xref>). All seeds were obtained from The European <italic>Arabidopsis</italic> Stock Centre<sup><xref ref-type="fn" rid="fn04">4</xref></sup> and grown as described previously (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>The probe composition and gene targets of Ath.test assay.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Probe name</th>
<th valign="top" align="center">Probe length</th>
<th valign="top" align="left">Target genomic site</th>
<th valign="top" align="center">Locus ID</th>
<th valign="top" align="left">Predicted CNV status</th>
<th valign="top" align="center">Source<sup>&#x2217;</sup></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ctrl1</td>
<td valign="top" align="center">96 nt</td>
<td valign="top" align="left">Chr1:25593..25645</td>
<td valign="top" align="center">AT1G01040</td>
<td valign="top" align="left">Non-variable; normalization control</td>
<td valign="top" align="center">a</td>
</tr>
<tr>
<td valign="top" align="left">ctrl2</td>
<td valign="top" align="center">111 nt</td>
<td valign="top" align="left">Chr4:11476533..11476582</td>
<td valign="top" align="center">AT4G21580</td>
<td valign="top" align="left">Non-variable; normalization control</td>
<td valign="top" align="center">a</td>
</tr>
<tr>
<td valign="top" align="left">ctrl3</td>
<td valign="top" align="center">124 nt</td>
<td valign="top" align="left">Chr2:15194440..15194490</td>
<td valign="top" align="center">AT2G36230</td>
<td valign="top" align="left">Non-variable; normalization control</td>
<td valign="top" align="center">a</td>
</tr>
<tr>
<td valign="top" align="left">ctrl4</td>
<td valign="top" align="center">144 nt</td>
<td valign="top" align="left">Chr5:7847361..7847414</td>
<td valign="top" align="center">AT5G23290</td>
<td valign="top" align="left">Non-variable; normalization control</td>
<td valign="top" align="center">a</td>
</tr>
<tr>
<td valign="top" align="left">ctrl5</td>
<td valign="top" align="center">172 nt</td>
<td valign="top" align="left">Chr1:27465468..27465522</td>
<td valign="top" align="center">AT1G73010</td>
<td valign="top" align="left">Non-variable; normalization control</td>
<td valign="top" align="center">a</td>
</tr>
<tr>
<td valign="top" align="left">mlpaA</td>
<td valign="top" align="center">160 nt</td>
<td valign="top" align="left">Chr1:17539289..17539343</td>
<td valign="top" align="center">AT1G47670</td>
<td valign="top" align="left">Non-variable</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaB1; mlpaB2</td>
<td valign="top" align="center">90 nt 148 nt</td>
<td valign="top" align="left">Chr1:30374276..30374321 Chr1:30373647..30373699</td>
<td valign="top" align="center">AT1G80830</td>
<td valign="top" align="left">Non-variable</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaC</td>
<td valign="top" align="center">93 nt</td>
<td valign="top" align="left">Chr1:11651708..11651754</td>
<td valign="top" align="center">AT1G32300</td>
<td valign="top" align="left">Biallelic</td>
<td valign="top" align="center">b</td>
</tr>
<tr>
<td valign="top" align="left">mlpaD1; mlpaD2</td>
<td valign="top" align="center">105 nt 114 nt</td>
<td valign="top" align="left">Chr1:9575624..9575678 Chr1:9577003..9577055</td>
<td valign="top" align="center">AT1G27570</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaE1; mlpaE2</td>
<td valign="top" align="center">136 nt 196 nt</td>
<td valign="top" align="left">Chr1:19726669..19726721 Chr1:19727385..19727439</td>
<td valign="top" align="center">AT1G52950</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaF1; mlpaF2</td>
<td valign="top" align="center">99 nt 120 nt</td>
<td valign="top" align="left">Chr3:7737420..7737467 Chr3:7737872..7737929</td>
<td valign="top" align="center">AT3G21960</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaG1; mlpaG2</td>
<td valign="top" align="center">128 nt 164 nt</td>
<td valign="top" align="left">Chr4:10641616..10641668 Chr4:10644628..10644679</td>
<td valign="top" align="center">AT4G19520</td>
<td valign="top" align="left">Biallelic</td>
<td valign="top" align="center">c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaH</td>
<td valign="top" align="center">180 nt</td>
<td valign="top" align="left">Chr4:13592606..13592658</td>
<td valign="top" align="center">AT4G27080</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaI</td>
<td valign="top" align="center">117 nt</td>
<td valign="top" align="left">Chr4:17705274..17705327</td>
<td valign="top" align="center">AT4G37685</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b</td>
</tr>
<tr>
<td valign="top" align="left">mlpaJ1; mlpaJ2</td>
<td valign="top" align="center">108 nt 156 nt</td>
<td valign="top" align="left">Chr5:2976409..2976464 Chr5:2978013..2978065</td>
<td valign="top" align="center">AT5G09590</td>
<td valign="top" align="left">Multiallelic; part of the gene</td>
<td valign="top" align="center">c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaK1; mlpaK2</td>
<td valign="top" align="center">188 nt 102 nt</td>
<td valign="top" align="left">Chr5:22228424..22228479 Chr5:22229438..22229488</td>
<td valign="top" align="center">AT5G54710</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">b; c</td>
</tr>
<tr>
<td valign="top" align="left">mlpaL</td>
<td valign="top" align="center">132 nt</td>
<td valign="top" align="left">Chr5:24796111..24796161</td>
<td valign="top" align="center">AT5G61700</td>
<td valign="top" align="left">Multiallelic</td>
<td valign="top" align="center">c</td></tr>
</tbody></table>
<table-wrap-foot>
<attrib><italic><sup>&#x2217;</sup>The initial information about the gene CNV status comes from the following resources: a, <xref ref-type="bibr" rid="B48">Zmienko et al. (2016)</xref>; b, Arabidopsis 1001 Genomes Project; c, our unpublished analysis of the WGS data originally presented in <xref ref-type="bibr" rid="B8">Cao et al. (2011)</xref>.</italic></attrib>
</table-wrap-foot>
</table-wrap>
</sec>
<sec><title>Optimization of the Template Amount</title>
<p>The multiplex MLPA-based strategy presented in this paper was originally developed for CNV genotyping of human DNA (<xref ref-type="bibr" rid="B23">Kozlowski et al., 2007</xref>; <xref ref-type="bibr" rid="B28">Marcinkowska et al., 2010</xref>). To adjust it for use with the <italic>Arabidopsis</italic> genome, we aimed to optimize the amount of DNA template. For humans, the typical MLPA assays include 50-250 ng genomic DNA per reaction. In our previous study, we successfully performed MLPA-based copy number analysis using 100 ng <italic>Arabidopsis</italic> genomic DNA (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>). However, because the <italic>Arabidopsis</italic> genome is &#x223C;20 times smaller than the human genome, we expected that the template amount could be substantially reduced without affecting the reaction performance. To evaluate the acceptable range of DNA amount for this species, we used the Col-0 accession, performed serial dilutions of the DNA template and performed MLPA assays for each of the following DNA amounts: 100, 60, 30, 15, 10, 5, and 2 ng, in three replicates. We observed that the intensity data showed little variance across all DNA concentrations tested and the peaks showed very good resolution and similar distribution, regardless of the template amount (<bold>Figures <xref ref-type="fig" rid="F2">2A&#x2013;C</xref></bold>; Supplementary Data Sheet <xref ref-type="supplementary-material" rid="SM5">S1</xref>). The normalized signal intensity data for various template amounts were highly correlated, with the results calculated for 2 ng DNA input showing only slightly lowered correlation than the other amounts (<bold>Figure <xref ref-type="fig" rid="F2">2D</xref></bold>). From this comparison, we concluded that the whole range of tested DNA amounts generates valid data. Below, we used the smallest tested amount of DNA (2 ng) to perform the exemplar Ath.test MLPA assay.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p><bold>Optimization of the DNA template amount for copy number variant (CNV) genotyping in <italic>Arabidopsis</italic>.</bold> The data were obtained from three replicates of MLPA assay per each DNA amount tested: 2, 5, 10, 15, 30, 60, and 100 ng. <bold>(A)</bold> comparison of the peak heights and localization on the electropherograms. <bold>(B,C)</bold> Variance of probe intensity measurements across tested DNA concentrations. Boxplots <bold>(B)</bold> present the distribution of coefficients of variation (CV) calculated separately for each probe and each DNA amount; The linear plot <bold>(C)</bold> presents regression analysis (CV vs input amount). For visibility, mean CV values (all probes, each in three replicates) per input amount are displayed on the plot; <bold>(D)</bold> pairwise correlation of results obtained for all template amounts tested, presented as coefficients of determination (<italic>R</italic><sup>2</sup>) of the MLPA probe intensity data.</p></caption>
<graphic xlink:href="fpls-08-00222-g002.tif"/>
</fig>
</sec>
<sec><title>Gene Copy Number Analysis</title>
<p>We generated MLPA data, processed it in GeneMarker and exported it to a Microsoft Excel worksheet (Supplementary Data Sheet <xref ref-type="supplementary-material" rid="SM5">S1</xref>). Three samples were excluded at this stage due to poor data quality. To enable sample-to-sample comparison, we normalized the data within each sample using the mean signal intensity of the control probes ctrl1&#x2013;ctrl5. The data were then compared and the copy numbers were estimated relative to the Col-0 accession that has the basic copy number of each gene analyzed in this assay (2<italic>n</italic> = 2) and therefore served as the reference sample. To reveal groups of accessions with distinct gene copy numbers, the population data were displayed as dot plots, histograms of the signal intensities or (for genes targeted by two MLPA probes) as 2D plots. We set the duplication/deletion thresholds at &#x003C;0.7 and >1.3 of the relative intensity, respectively, for all genes in the assay. Subsequently, for each gene, the samples passing the threshold values were clustered and the clusters were manually assigned the copy numbers, as demonstrated previously (<xref ref-type="bibr" rid="B30">Marcinkowska-Swojak et al., 2014</xref>; <xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
<sec><title>Non-variable Regions</title>
<p>The probes mlpaA, mlpaB1, and mlpaB2 targeted two genes predicted to have the same copy number in all accessions: <italic>AT1G47670</italic>, coding for lysine histidine transporter-like 8 (mlpaA), and <italic>AT1G80830</italic>, coding for NRAMP1 transporter (mlpaB1 and mlpaB2). For all accessions, the relative signals from these three probes were at the same level as those in Col-0 (mean intensity 1.01, 1.03, and 0.93, respectively, see <bold>Figure <xref ref-type="fig" rid="F3">3A</xref></bold>) and showed very little variance (CV 0.060, 0.089, and 0.064, respectively). Additional evaluation of the mlpaB1 and mlpaB2 probes on a 2D plot revealed that all samples were grouped in one cluster (<bold>Figure <xref ref-type="fig" rid="F3">3B</xref></bold>).</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p><bold>Multiplex ligation-dependent probe amplification assay results for non-variable regions. (A)</bold> relative intensity distribution of the mlpaA probe, targeting the <italic>AT1G47670</italic> gene as well as mlpaB1 and mlpaB2 probes, both targeting the <italic>AT1G80830</italic> gene, among the accessions; <bold>(B)</bold> 2D plot of intensity from probes mlpaB1 and mlpaB2.</p></caption>
<graphic xlink:href="fpls-08-00222-g003.tif"/>
</fig>
</sec>
<sec><title>Biallelic CNVs</title>
<p>We analyzed two genes with presence-absence variation revealed by the WGS data analysis: <italic>AT1G32300</italic> (coding for <sc>D</sc>-arabinono-1,4-lactone oxidase family protein) and <italic>AT4G19520</italic> (coding for TIR-NBS-LRR class disease resistance protein). We designed one probe (mlpaC) for <italic>AT1G32300</italic> exon 1 and two probes, mlpaG1 and mlpaG2, for <italic>AT4G19520</italic> exons 3 and 5, respectively. For <italic>AT1G32300</italic>, we observed a dominant population of samples with mean signal intensity 1.08, indicative of two gene copies per diploid genome. The remaining samples formed a distinct group with mean signal intensity 0.09, indicative of the absence of the analyzed gene in the respective accessions (<bold>Figure <xref ref-type="fig" rid="F4">4A</xref></bold>). In the case of <italic>AT4G19520</italic>, the combined data for the mlpaG1 and mlpaG2 probes revealed the presence of two compact clusters (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>). One cluster included 29 accessions with no difference in copy number relative to Col-0 (mlpaG1 mean intensity 1.03; mlpaG2 mean intensity 1.01). The other cluster included 47 accessions with substantially reduced intensity (mlpaG1 mean intensity 0.14; mlpaG2 mean intensity 0.12), indicative of the deletion.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p><bold>Multiplex ligation-dependent probe amplification assay results for presence-absence CNVs. (A)</bold> relative intensity from probe mlpaC targeting the <italic>AT1G32300</italic> gene, in individual accessions; <bold>(B)</bold> 2D plot of relative signal from mlpaG1 and mlpaG2, both targeting the <italic>AT4G19520</italic> gene. Clusters are colored according to the deduced CNV status.</p></caption>
<graphic xlink:href="fpls-08-00222-g004.tif"/>
</fig>
</sec>
<sec><title>Multiallelic CNVs: One MLPA Probe Per Gene</title>
<p>For three genes that overlap multiallelic CNVs we designed 1 MLPA probe per gene in Ath.test assay (<bold>Figure <xref ref-type="fig" rid="F5">5A</xref></bold>). Gene <italic>AT4G37685</italic> codes for a hypothetical protein and is targeted by the mlpaI probe. Majority of accessions (39) harbor two copies of this gene. Gene deletion was detected in eight accessions and duplication in 30 accessions. Of the latter, 22 accessions had four copies, seven accessions had six copies, and one harbored a very high-level duplication, most likely &#x2265;12 copies.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p><bold>Multiplex ligation-dependent probe amplification results for multiallelic CNVs. (A)</bold> CNV genotyping with one MLPA probe per gene. Histograms present the relative signal distribution from probe mlpaI (targeting the <italic>AT4G37685</italic> gene), probe mlpaL (targeting the <italic>AT5G61700</italic> gene), and probe mlpaH (targeting the <italic>AT4G27080</italic> gene). The histogram bin size is 0.2 in all plots; <bold>(B)</bold> CNV genotyping with two MLPA probes per gene. 2D plots present the relative signal from probes mlpaK1and mlpaK2 (both targeting the <italic>AT5G54710</italic> gene) and from probes mlpaJ1 and mlpaJ2 (both targeting the <italic>AT5G09590</italic> gene). Clusters are colored according to deduced CNV status. The coefficient of determination (<italic>R</italic><sup>2</sup>) is calculated for accessions with assigned copy numbers. <bold>(C)</bold> Genotyping complex multiallelic CNVs. 2D intensity plots present relative signal from probes mlpaF1and mlpaF2 (targeting exon 1 and exon 2 of the <italic>AT3G21960</italic> gene, respectively) and from probes mlpaE1and mlpaE2 (targeting exon 6 and exon 9 of the <italic>AT1G52950</italic> gene, respectively). Clusters are colored according to deduced CNV status. The coefficient of determination (<italic>R</italic><sup>2</sup>) is calculated for subsets of accessions, as detailed in the main text.</p></caption>
<graphic xlink:href="fpls-08-00222-g005.tif"/>
</fig>
<p>Gene <italic>AT5G61700</italic> codes for ATH16, a member of ABC transporter subfamily A and is targeted by probe mlpaL. In most analyzed accessions, the gene exists in two copies per diploid genome. In eight accessions, however, duplications were detected: four copies in three accessions, six copies in two accessions, and &#x2265;10 copies in three accessions. It is worth noting that, in MLPA assays, the signal intensity is non-linearly related to the DNA copy number (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>). This is manifested by reducing the distance between the clusters with different duplication levels for high copy numbers. Consequently, a large number of samples harboring high-level duplications is needed to precisely distinguish the clusters of 8 and more copies from each other.</p>
<p>Gene <italic>AT4G27080</italic> codes for a protein disulfide isomerase that is involved in cell redox homeostasis and is targeted by the mlpaH probe. From the WGS data, we predicted that majority of accessions harbor partial or full duplications of this gene. Likewise, MLPA analysis revealed that only nine accessions harbor two copies of <italic>AT4G27080</italic> gene, while duplications were detected in 68 accessions. Among them, we clearly identified a group of 44 accessions with four copies, but the remaining accessions were less distinctive and formed two heterogeneous groups which we named &#x201C;medium-level duplications&#x201D; (10 accessions) and &#x201C;high-level duplications&#x201D; (14 accessions). For 12 of these &#x201C;high-level duplication&#x201D; accessions, the mlpaH peak intensity counts reached the upper detection limits (see <bold>Notes</bold> section below for additional comments). We concluded that designing two or more MLPA probes targeting this genomic region and repeating the assay with adjusted capillary electrophoresis parameters would be helpful in more accurate distinction of the CNV genotypes or resolution of the structural complexity of the investigated gene.</p>
</sec>
<sec><title>Multiallelic CNVs: Two MLPA Probes Per Gene</title>
<p>For 2 other genes that overlap multiallelic CNVs we designed two MLPA probes per gene (<bold>Figure <xref ref-type="fig" rid="F5">5B</xref></bold>). The <italic>AT5G54710</italic> gene codes for an ankyrin repeat family protein and is positioned between two other ankyrin repeat family protein coding genes, in the region that is highly copy number variable. We used two specific probes (mlpaK1 and mlpaK2), located in the fourth and third exons of <italic>AT5G54710</italic>, respectively, and confirmed that this gene is multiallelic. The high linear correlation of the mlpaK1 and mlpaK2 probe intensities allowed us distinguish several clusters of accessions with distinct copy numbers: 0 copies (2 accessions), 2 copies (54 accessions), 4 copies (8 accessions), 6 copies (6 accessions), and 8 copies (1 accession). We did not assign the integer copy numbers for 6 accessions which displayed uneven duplication level based on the mlpaK1 and mlpaK2 probe signal.</p>
<p>The <italic>AT5G09590</italic> gene, encoding mitochondrial heat shock protein MTHSC70-2, is localized in the breakpoint of a large CNV that encompasses loci <italic>AT5G09590 &#x2013; AT5G09630</italic>. Consequently, <italic>AT5G09590</italic> is only partially duplicated in several accessions. We designed two probes, localized outside of and within the CNV region (mlpaJ1, targeting fourth exon and mlpaJ2, targeting sixth exon, respectively). The results of the MLPA assay clearly revealed that only the 3&#x2032; part of <italic>AT5G09590</italic> (targeted by probe mlpaJ2) is duplicated: 43 accessions harbored four copies, two accessions harbored six copies, and one accession harbored at least 10 copies. The region targeted by probe mlpaJ1 invariantly had two copies in all accessions.</p>
</sec>
<sec><title>Complex Multiallelic CNVs</title>
<p>Some genomic regions, e.g., these that harbor clustered multigene families, may display high structural diversity in the populations. A gene may be fully duplicated/deleted in some accessions while in the other ones only part of this gene may display copy number alteration. Additionally, the duplicated DNA copies within one sample may differ from each other in length and sequence, which may affect the affinity of the MLPA probe to some (but not all) copies. Consequently, the copy number pattern revealed by the MLPA analysis may be complex. Below we present some examples of MLPA analysis in multiallelic CNVs with a complex structure (<bold>Figure <xref ref-type="fig" rid="F5">5C</xref></bold>).</p>
<p>The <italic>AT3G21960</italic> gene is localized in the central part of a &#x223C;50 kb CNV, that encompasses 21 genes, mainly members of the receptor-like protein kinase-related family and genes coding for proteins with unknown domain DUF26. We assayed the <italic>AT3G21960</italic> gene with specific probes targeting exons 1 and 2 (probes mlpaF1 and mlpaF2, respectively). In 30 samples the signals from these probes were highly correlated and formed 4 distinct groups of: 0 copies (1 accession), 2 copies (26 accessions), 4 copies (1 accession) and 6 copies (2 accessions). In 6 accessions, however, only the mlpaF2 probe intensity was elevated (1.83&#x2013;6.54), while mlpaF1 intensity was about 1. On the contrary, the remaining 41 accessions formed a compact cluster, with the mlpaF1 intensity below 0.7 (the value that has been set as the deletion threshold), and the mlpaF2 intensity about 1. A brief evaluation of the <italic>AT3G21960</italic> genomic sequence inferred from WGS data<sup><xref ref-type="fn" rid="fn05">5</xref></sup> (obtained with Pseudogenomes Download Tool) provided evidence that this complex pattern is true, as 519 out of 1135 accessions with available genomic data had 80&#x2013;100% uncalled sites (Ns) in the exon 1 sequence, while only 3 accessions had 80&#x2013;100% uncalled sites in exon 2 sequence.</p>
<p>Complex multiallelic CNVs are often related to the activity of mobile genetic elements, which may trigger partial or full deletion/duplication of the nearby genes. Gene <italic>AT1G52950</italic> codes for a nucleic acid-binding OB fold-like protein and is localized within one CNV region with a nearby transposable element gene <italic>AT1G52960</italic> (the two loci are separated by only 3.6 kb distance). We assayed the copy number status of <italic>AT1G52950</italic> using two probes, mlpaE1 to target exon 6 and mlpaE2 target exon 9. For 69 accessions, we detected compact clusters with distinct copy numbers (0 to 6 copies) and a high correlation between the two measurements (<italic>R</italic><sup>2</sup> = 0.9881). Interestingly, in two cases, the intensity data suggested the existence of one copy and three copies of the <italic>AT1G52950</italic> gene per diploid genome in the surveyed individuals. <italic>Arabidopsis</italic> is a highly self-pollinating species for which most genomic loci are expected to exist in a homozygous state, therefore assaying additional individuals would be necessary to establish the representative gene copy number for these two accessions in a population study. For seven accessions, of which six originated from Southern Tyrol region and 1 was a Spanish relict accession (<xref ref-type="bibr" rid="B1">1001 Genomes Consortium, 2016</xref>), the copy number status indicated by probe mlpaE1 was always higher than the copy number status indicated by probe mlpaE2. This effect may have many reasons, e.g., partial duplication or deletion of a gene of interest, sequence divergence in some duplicated copies that affect the hybridization of one MLPA probe, etc. Unambiguous interpretation of these data would require additional region characterization by sequencing. Nevertheless, the signals from both probes were also well correlated (<italic>R</italic><sup>2</sup> = 0.9856). Finally, one accession displayed an extremely high level of duplications at the mlpaE1 target site while no copy number changes were observed at the mlpaE2 site.</p>
</sec>
<sec><title>Effect of Non-specific Hybridization on MLPA Signal</title>
<p>To present the effect of compromised probe specificity on the MLPA results, we assayed a gene <italic>AT1G27570</italic>, which encodes the phosphatidylinositol 3- and 4-kinase family protein and is localized within the large multiallelic CNV (over 20 kb). We designed two probes, mlpaD1 and mlpaD2, targeting this gene, of which only mlpaD2 was specific to <italic>AT1G27570</italic>. Probe mlpaD1 had an alternative target site (with only two mismatches in the left TSS and one mismatch in the right TSS, distant from the ligation site) in the nearby gene <italic>AT1G27590</italic>, not copy number variable. As a result, the signal from the mlpaD1 probe was elevated by the background signal from the alternative target site. This background signal was stable (due to unchanged copy number of <italic>AT1G27590</italic> gene in all accessions) therefore the high correlation between the data for mlpaD1 and mlpaD2 probes was preserved (<bold>Figure <xref ref-type="fig" rid="F6">6A</xref></bold>). As a rule, we suggest re-designing of the MLPA probes that produce non-specific signal. However, if a set of the control samples that carry confirmed deletion of the gene of interest can be defined, these samples may be used for the data correction. In the present example, we calculated the mean non-specific signal of probe mlpaD1 in the cluster of 15 samples with gene deletions (marked in black color in <bold>Figure <xref ref-type="fig" rid="F6">6A</xref></bold>). This value was then subtracted from the probe mlpaD1 signal in each sample, before estimating the intensity ratio relative to Col-0 accession. The correction improved the relative intensity ratio observed for probe mlpaD1 (<bold>Figure <xref ref-type="fig" rid="F6">6B</xref></bold>). We note here, that the process of data correction had no effect on the overall correlation between the signals from probes mlpaD1 and mlpaD2. This correlation was high (<italic>R</italic><sup>2</sup> = 0.9386), therefore allowing to distinguish the copy number clusters on 2D plots pretty easily both before and after data correction.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p><bold>Effect of non-specific probe hybridization on the MLPA results.</bold> Probes mlpaD1 and mlpaD2 target the <italic>AT1G27570</italic> gene. Additionally, probe mlpaD1 targets also the <italic>AT1G27590</italic> gene. <bold>(A)</bold> 2D intensity plot of relative signal from probes mlpaD1 and mlpaD2; <bold>(B)</bold> 2D intensity plot of relative signal from probe mlpaD1 corrected for the presence of non-specific hybridization signal (see main text for details) and probe mlpaD2. Clusters are colored according to the deduced CNV status and are identical for each sample on both plots. The coefficient of determination (<italic>R</italic><sup>2</sup>) is calculated for all accessions. The deletion of <italic>AT1G27570</italic> gene has been confirmed by PCR in accessions from the &#x201C;0 copies&#x201D; cluster (not shown).</p></caption>
<graphic xlink:href="fpls-08-00222-g006.tif"/>
</fig>
</sec>
</sec></sec>
<sec><title>Notes</title>
<p>Below we included some notes on the limitations of the procedure, common mistakes and possible artifacts related to the presented application.</p>
<sec><title>Probe Design</title>
<p>Oligonucleotide MLPA probes described in this procedure target specific sequences in the genome, typically 45&#x2013;75 bp. Regions located outside of the probe&#x2019;s recognition sequence may have different copy number status. If partial gene duplication/deletion or insertion of duplicated sequence is suspected, additional probes, e.g., covering different exons of the gene should be included in the assay.</p>
<p>Compromised ability of MLPA probe to recognize the target sequence may be the source of false positive results. Sequence changes (SNPs, indels, point mutations) in the target sequence detected by a probe can negatively affect or completely prevent probe binding. The critical positions in the TSS sequence are these constituting the ligation site; the presence of a SNP at or near the ligation site will disrupt the ligation step and result in no signal from the MLPA probe, falsely indicative of deletion of the region in the affected sample (<xref ref-type="bibr" rid="B18">Kim et al., 2016</xref>). Note that the MLPA technique can be also used for detecting small mutations (<xref ref-type="bibr" rid="B29">Marcinkowska-Swojak et al., 2016</xref>), but these applications are not covered in the present protocol.</p>
<p>The accuracy of the results is also strictly dependent on the MLPA probe specificity. If alternative target site exists in the genome (e.g., in a paralogue or a pseudogene), it will generate non-specific signal (see Effect of Non-specific Probe Hybridization on the MLPA Results Section). To this end, for plants with incomplete genome information we strongly advise designing &#x2265;2 MLPA probes per gene, to minimize this risk.</p>
<p>In the case of newly designed MLPA probes we recommend verifying their performance on a (set of) well characterized reference samples. If no product is observed, make sure that the common mistakes interfering with the experimental steps are avoided (see below). If needed, re-design the MLPA probe.</p>
</sec>
<sec><title>Assay Design and Performing</title>
<p>Multiplex ligation-dependent probe amplification results may be compromised by multiple factors that will affect the enzymatic reactions and result in reduced peak signals. These factors include but are not limited to: DNA integrity and contamination, presence of PCR inhibitors in the samples, incomplete DNA denaturation, sample evaporation, suboptimal amount of the sample DNA used. In the Section &#x201C;Stepwise Procedures&#x201D; we included useful tips regarding the sample preparation and assay setup. Additional comments are given below.</p>
<p>If the DNA sample contamination is a suspected problem, perform new DNA extraction. From our experience, we advise using column-based methods, e.g., DNeasy Plant Mini Kit (Qiagen) for DNA extraction (or purification of DNA extracted with other methods) because they produce samples of high purity and comparable amounts.</p>
<p>Use multichannel pipettes to reduce the pipetting time and avoid sample evaporation.</p>
<p>Reduce sample-to-sample variability by simultaneous performing multiple assays, using strips (preferable) or multiwell PCR plates. Use the same MLPA Probe Set Mix preparation for all samples under comparison.</p>
<p>Replacing the strip caps on each opening minimizes the risk of sample cross-contamination.</p>
<p>Follow the capillary electrophoresis protocols (size standard, sample preparation, injection time and voltage) suitable for the instrument used. Decrease injection time if the peaks are out of range. We recommend prior optimization of the DNA template amount in the assay and capillary electrophoresis conditions on a validated reference sample.</p>
<p>Abnormal pictures after capillary electrophoresis may indicate capillary electrophoresis problems but they also may result from the PCR step troubles. See the MLPA troubleshooting wizard by MRC Holland<sup><xref ref-type="fn" rid="fn06">6</xref></sup> for common peak pattern problems and possible solutions.</p>
</sec>
<sec><title>Data Analysis and Copy Number Estimation</title>
<p>It is advisable to manually check the peaks identified by GeneMarker before further data processing. In our assay, we repeatedly observed that the software did not detect the peaks for probe mlpaH in 12 samples and reported &#x201C;0&#x201D; intensity for this probe (Supplementary Figure <xref ref-type="supplementary-material" rid="SM4">S3</xref>). In fact, high intensity peaks from probe mlpaH with their tops flattened (cut) were present in these samples, which indicated that the signal exceeded the capillary electrophoresis system detection limits. We manually corrected the peak localization and used the maximum reported values for copy number calculation, but this likely resulted in underestimation of the gene copy number in these samples in our study (see Section &#x201C;Multiallelic CNVs: One MLPA Probe Per Gene&#x201D;). To accurately quantify the probe signal, repeating the electrophoresis with lower injection time would be necessary. The results from high and low injection time electropherograms may be then merged after internal control probe normalization step, to preserve good resolution of the low intensity peaks.</p>
<p>Multiplex ligation-dependent probe amplification is a relative technique, therefore selecting well validated reference samples with basic copy number of the region of interest (usually two copies) is essential for accurate quantification. However, in case of population scale CNV genotyping of numerous independent genomic regions in a multiplex assay (similar to example provided in this paper) such a reference sample may not exist or remains unknown. Providing that sufficiently large number of samples in the population are genotyped, the presented protocol still allows for inferring the cluster copy numbers without a reference sample, under the assumption that the neighboring clusters of accessions/lines differ by two copies and that the distances between these clusters are &#x223C;equal in the range of 0&#x2013;4 copies (see <xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref> for further discussion on the distances between the clusters in MLPA assays).</p>
</sec>
<sec><title>Validation of the Results</title>
<p>Regardless of the number of probes and samples used, we recommend to verify the positive MLPA results with an independent technique. We advise performing droplet digital PCR (ddPCR) on selected samples, as this approach allows for estimating gene copy numbers at the same or even higher range, as the MLPA procedure described in this protocol (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>). Additionally, ddPCR generates amplicons of &#x223C;60&#x2013;200 bp, therefore allows for genome assaying at similar resolution as MLPA.</p>
</sec>
</sec>
<sec><title>Conclusion</title>
<p>In this work, we described the protocol for the simple MLPA-based CNV genotyping in plants, with particular emphasis on the model plant <italic>Arabidopsis</italic>. We provided a description of the probe design process, experimental setup, and data analysis. We also discussed the results of the exemplar multiplex assay and showed that the MLPA method is very robust and is a rich source of information regarding the CNV in the analyzed samples. The abundant genomic data obtained for a growing number of species as a part of large-scale sequencing projects, highlight CNV as the major contributor to natural diversity at a genotype level (<xref ref-type="bibr" rid="B45">Zarrei et al., 2015</xref>; <xref ref-type="bibr" rid="B1">1001 Genomes Consortium, 2016</xref>; <xref ref-type="bibr" rid="B3">Bai et al., 2016</xref>). Gene duplication has been considered the major factor driving long-term evolution and gene birth by sub- and neofunctionalization of the duplicated copies (<xref ref-type="bibr" rid="B11">Conant et al., 2014</xref>). Some regions in the genome may be more prone to CNV than the others, due to their specific structural features, that will locally induce the mechanisms leading to CNV formation, e.g., non-allelic recombination (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>). The duplication / deletion events may have also consequences on organism&#x2019;s fitness and contribute to the adaptation to environmental challenges, as well as to coevolutionary interactions between host and pathogen or a symbiont (reviewed in: <xref ref-type="bibr" rid="B21">Kondrashov, 2012</xref>, <xref ref-type="bibr" rid="B47">&#x017B;mie&#x0144;ko et al., 2014</xref>). Remarkably, the protein coding genes displaying CNVs are often related to environmental stress response and pathogen resistance (<xref ref-type="bibr" rid="B13">Cook et al., 2012</xref>; <xref ref-type="bibr" rid="B31">Maron et al., 2013</xref>). The creation of high-confidence CNV maps and assessing the gene copy number in large populations will enhance the studies on the evolution of genomes in the context of CNV origin, fixation and the impact on the phenotype. These data can be later combined with the results of the transcriptomic, proteomic, metabolomics, protein interaction, phenotyping, and other studies). We recently used the MLPA method to genotype <italic>MSH2, AT3G18530</italic>, and <italic>AT3G18535</italic> copy number in a set of 189 natural accessions. Based on these results, we were subsequently able to reveal the recurrent nature of <italic>AT3G18530</italic> and <italic>AT3G18535</italic> duplications/deletions and to dissect the structural features that promoted non-allelic homologous recombination, leading to a widespread occurrence of the <italic>AT3G18530</italic> and <italic>AT3G18535</italic> genes deletion in nature (<xref ref-type="bibr" rid="B48">Zmienko et al., 2016</xref>).</p>
<p>This protocol will enable potential users to introduce the MLPA technique in plant genetic and population biology studies. The technique is multiplexable and very well suited for verification of WGS-based analyses or for rapid characterization of copy number status across a region of interest in large populations. Notably, once designed, the individual MLPA probes may be used in various combinations according to one&#x2019;s needs, providing that the lengths of the probes in one assay are unique. We believe that the MLPA protocol presented in the current work will contribute to accelerating the discovery of new associations between CNV and important traits in plants.</p>
</sec>
<sec><title>Author Contributions</title>
<p>AS-C prepared DNA samples, performed MLPA assays, analyzed data, helped prepare figures, and draft the manuscript. MM-Z performed template optimization experiments. MM-S helped design the MLPA probes and set up the assay. PK analyzed the data and helped draft the manuscript. MF contributed to the conception of the work and revised the manuscript. AZ conceived of and designed the study, analyzed data, oversaw the research, prepared figures, and wrote the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was supported by National Centre of Science grants (2011/01/B/NZ2/04816 and 2014/13/B/NZ2/03837). Publication of the results was supported by Polish Ministry of Science and Higher Education, under the KNOW program.</p></fn>
</fn-group>
<ack>
<p>We thank Pawel Wojciechowski and Kaja Milanowska for help with WGS data analysis.</p>
</ack>
<sec sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fpls.2017.00222/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fpls.2017.00222/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.XLSX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Presentation_1.PDF" id="SM2" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Presentation_2.PDF" id="SM3" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Presentation_3.PDF" id="SM4" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_1.XLSX" id="SM5" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><collab>1001 Genomes Consortium</collab> (<year>2016</year>). <article-title>1,135 genomes reveal the global pattern of polymorphism in <italic>Arabidopsis thaliana</italic>.</article-title> <source><italic>Cell</italic></source> <volume>166</volume> <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2016.05.063</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alkan</surname> <given-names>C.</given-names></name> <name><surname>Coe</surname> <given-names>B. P.</given-names></name> <name><surname>Eichler</surname> <given-names>E. E.</given-names></name></person-group> (<year>2011</year>). <article-title>Genome structural variation discovery and genotyping.</article-title> <source><italic>Nat. Rev. Genet.</italic></source> <volume>12</volume> <fpage>363</fpage>&#x2013;<lpage>376</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2958</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bai</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Liao</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>R.</given-names></name> <name><surname>Ge</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>The impact and origin of copy number variations in the <italic>Oryza</italic> species.</article-title> <source><italic>BMC Genomics</italic></source> <volume>17</volume>:<issue>261</issue>. <pub-id pub-id-type="doi">10.1186/s12864-016-2589-2</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bel&#x00F3;</surname> <given-names>A.</given-names></name> <name><surname>Beatty</surname> <given-names>M. K.</given-names></name> <name><surname>Hondred</surname> <given-names>D.</given-names></name> <name><surname>Fengler</surname> <given-names>K. A.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Rafalski</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Allelic genome structural variations in maize detected by array comparative genome hybridization.</article-title> <source><italic>Theor. Appl. Genet.</italic></source> <volume>120</volume> <fpage>355</fpage>&#x2013;<lpage>367</lpage>. <pub-id pub-id-type="doi">10.1007/s00122-009-1128-9</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bharuthram</surname> <given-names>A.</given-names></name> <name><surname>Paximadis</surname> <given-names>M.</given-names></name> <name><surname>Picton</surname> <given-names>A. C. P.</given-names></name> <name><surname>Tiemessen</surname> <given-names>C. T.</given-names></name></person-group> (<year>2014</year>). <article-title>Comparison of a quantitative Real-Time PCR assay and droplet digital PCR for copy number analysis of the CCL4L genes.</article-title> <source><italic>Infect. Genet. Evol.</italic></source> <volume>25</volume> <fpage>28</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1016/j.meegid.2014.03.028</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cantsilieris</surname> <given-names>S.</given-names></name> <name><surname>Baird</surname> <given-names>P. N.</given-names></name> <name><surname>White</surname> <given-names>S. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Molecular methods for genotyping complex copy number polymorphisms.</article-title> <source><italic>Genomics</italic></source> <volume>101</volume> <fpage>86</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2012.10.004</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cantsilieris</surname> <given-names>S.</given-names></name> <name><surname>Western</surname> <given-names>P. S.</given-names></name> <name><surname>Baird</surname> <given-names>P. N.</given-names></name> <name><surname>White</surname> <given-names>S. J.</given-names></name></person-group> (<year>2014</year>). <article-title>Technical considerations for genotyping multi-allelic copy number variation (CNV), in regions of segmental duplication.</article-title> <source><italic>BMC Genomics</italic></source> <volume>15</volume>:<issue>329</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-15-329</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>J.</given-names></name> <name><surname>Schneeberger</surname> <given-names>K.</given-names></name> <name><surname>Ossowski</surname> <given-names>S.</given-names></name> <name><surname>G&#x00FC;nther</surname> <given-names>T.</given-names></name> <name><surname>Bender</surname> <given-names>S.</given-names></name> <name><surname>Fitz</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Whole-genome sequencing of multiple <italic>Arabidopsis thaliana</italic> populations.</article-title> <source><italic>Nat. Genet.</italic></source> <volume>43</volume> <fpage>956</fpage>&#x2013;<lpage>963</lpage>. <pub-id pub-id-type="doi">10.1038/ng.911</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ceulemans</surname> <given-names>S.</given-names></name> <name><surname>van der Ven</surname> <given-names>K.</given-names></name> <name><surname>Del-Favero</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x201C;Targeted screening and validation of copy number variations,&#x201D; in</article-title> <source><italic>Genomic Structural Variants: Methods and Protocols, Methods in Molecular Biology</italic></source> <role>ed.</role> <person-group person-group-type="editor"><name><surname>Feuk</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer Science+Business Media</publisher-name>) <fpage>369</fpage>&#x2013;<lpage>384</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-61779-507-7_18</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>C.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>H.-P.</given-names></name> <name><surname>Ma</surname> <given-names>C.-X.</given-names></name> <name><surname>Sun</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Copy number variation of cytokinin oxidase gene Tackx4 associated with grain weight and chlorophyll content of flag leaf in common wheat.</article-title> <source><italic>PLoS ONE</italic></source> <volume>10</volume>:<issue>e0145970</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0145970</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conant</surname> <given-names>G. C.</given-names></name> <name><surname>Birchler</surname> <given-names>J. A.</given-names></name> <name><surname>Pires</surname> <given-names>J. C.</given-names></name></person-group> (<year>2014</year>). <article-title>Dosage, duplication, and diploidization: clarifying the interplay of multiple models for duplicate gene evolution over time.</article-title> <source><italic>Curr. Opin. Plant Biol.</italic></source> <volume>19</volume> <fpage>91</fpage>&#x2013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1016/j.pbi.2014.05.008</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cook</surname> <given-names>D. E.</given-names></name> <name><surname>Bayless</surname> <given-names>A. M.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Guo</surname> <given-names>X.</given-names></name> <name><surname>Song</surname> <given-names>Q.</given-names></name> <name><surname>Jiang</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Distinct copy number, coding sequence, and locus methylation patterns underlie Rhg1-mediated soybean resistance to soybean cyst nematode.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>165</volume> <fpage>630</fpage>&#x2013;<lpage>647</lpage>. <pub-id pub-id-type="doi">10.1104/pp.114.235952</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cook</surname> <given-names>D. E.</given-names></name> <name><surname>Lee</surname> <given-names>T. G.</given-names></name> <name><surname>Guo</surname> <given-names>X.</given-names></name> <name><surname>Melito</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Bayless</surname> <given-names>A. M.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Copy number variation of multiple genes at Rhg1 mediates nematode resistance in soybean.</article-title> <source><italic>Science</italic></source> <volume>338</volume> <fpage>1206</fpage>&#x2013;<lpage>1209</lpage>. <pub-id pub-id-type="doi">10.1126/science.1228746</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duitama</surname> <given-names>J.</given-names></name> <name><surname>Silva</surname> <given-names>A.</given-names></name> <name><surname>Sanabria</surname> <given-names>Y.</given-names></name> <name><surname>Cruz</surname> <given-names>D. F.</given-names></name> <name><surname>Quintero</surname> <given-names>C.</given-names></name> <name><surname>Ballen</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Whole genome sequencing of elite rice cultivars as a comprehensive information resource for marker assisted selection.</article-title> <source><italic>PLoS ONE</italic></source> <volume>10</volume>:<issue>e0124617</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0124617</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gaines</surname> <given-names>T. A.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Bukun</surname> <given-names>B.</given-names></name> <name><surname>Chisholm</surname> <given-names>S. T.</given-names></name> <name><surname>Shaner</surname> <given-names>D. L.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Gene amplification confers glyphosate resistance in <italic>Amaranthus palmeri</italic>.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>107</volume> <fpage>1029</fpage>&#x2013;<lpage>1034</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0906649107</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanada</surname> <given-names>K.</given-names></name> <name><surname>Sawada</surname> <given-names>Y.</given-names></name> <name><surname>Kuromori</surname> <given-names>T.</given-names></name> <name><surname>Klausnitzer</surname> <given-names>R.</given-names></name> <name><surname>Saito</surname> <given-names>K.</given-names></name> <name><surname>Toyoda</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Functional compensation of primary and secondary metabolites by duplicate genes in <italic>Arabidopsis thaliana</italic>.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>28</volume> <fpage>377</fpage>&#x2013;<lpage>382</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msq204</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>H&#x00F6;mig-H&#x00F6;lzel</surname> <given-names>C.</given-names></name> <name><surname>Savola</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Multiplex ligation-dependent probe amplification (MLPA) in tumor diagnostics and prognostics.</article-title> <source><italic>Diagn. Mol. Pathol.</italic></source> <volume>21</volume> <fpage>189</fpage>&#x2013;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1097/PDM.0b013e3182595516</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>M. J.</given-names></name> <name><surname>Cho</surname> <given-names>S. I.</given-names></name> <name><surname>Chae</surname> <given-names>J. H.</given-names></name> <name><surname>Lim</surname> <given-names>B. C.</given-names></name> <name><surname>Lee</surname> <given-names>J. S.</given-names></name> <name><surname>Lee</surname> <given-names>S. J.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Pitfalls of multiple ligation-dependent probe amplifications in detecting DMD exon deletions or duplications.</article-title> <source><italic>J. Mol. Diagn.</italic></source> <volume>18</volume> <fpage>253</fpage>&#x2013;<lpage>259</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmoldx.2015.11.002</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klonowska</surname> <given-names>K.</given-names></name> <name><surname>Ratajska</surname> <given-names>M.</given-names></name> <name><surname>Czubak</surname> <given-names>K.</given-names></name> <name><surname>Kuzniacka</surname> <given-names>A.</given-names></name> <name><surname>Brozek</surname> <given-names>I.</given-names></name> <name><surname>Koczkowska</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Analysis of large mutations in BARD1 in patients with breast and/or ovarian cancer: the Polish population as an example.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>5</volume>:<issue>10424</issue>. <pub-id pub-id-type="doi">10.1038/srep10424</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kohany</surname> <given-names>O.</given-names></name> <name><surname>Gentles</surname> <given-names>A. J.</given-names></name> <name><surname>Hankus</surname> <given-names>L.</given-names></name> <name><surname>Jurka</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Annotation, submission and screening of repetitive elements in Repbase: RepbaseSubmitter and censor.</article-title> <source><italic>BMC Bioinformatics</italic></source> <volume>7</volume>:<issue>474</issue>. <pub-id pub-id-type="doi">10.1186/1471-2105-7-474</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kondrashov</surname> <given-names>F. A.</given-names></name></person-group> (<year>2012</year>). <article-title>Gene duplication as a mechanism of genomic adaptation to a changing environment.</article-title> <source><italic>Proc. Biol. Sci.</italic></source> <volume>279</volume> <fpage>5048</fpage>&#x2013;<lpage>5057</lpage>. <pub-id pub-id-type="doi">10.1098/rspb.2012.1108</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koralewski</surname> <given-names>T. E.</given-names></name> <name><surname>Krutovsky</surname> <given-names>K. V.</given-names></name></person-group> (<year>2011</year>). <article-title>Evolution of exon-intron structure and alternative splicing.</article-title> <source><italic>PLoS ONE</italic></source> <volume>6</volume>:<issue>e18055</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0018055</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kozlowski</surname> <given-names>P.</given-names></name> <name><surname>Roberts</surname> <given-names>P.</given-names></name> <name><surname>Dabora</surname> <given-names>S.</given-names></name> <name><surname>Franz</surname> <given-names>D.</given-names></name> <name><surname>Bissler</surname> <given-names>J.</given-names></name> <name><surname>Northrup</surname> <given-names>H.</given-names></name><etal/></person-group> (<year>2007</year>). <article-title>Identification of 54 large deletions/duplications in TSC1 and TSC2 using MLPA, and genotype-phenotype correlations.</article-title> <source><italic>Hum. Genet.</italic></source> <volume>121</volume> <fpage>389</fpage>&#x2013;<lpage>400</lpage>. <pub-id pub-id-type="doi">10.1007/s00439-006-0308-9</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Wu</surname> <given-names>H. X.</given-names></name> <name><surname>Dillon</surname> <given-names>S. K.</given-names></name> <name><surname>Southerton</surname> <given-names>S. G.</given-names></name></person-group> (<year>2009</year>). <article-title>Generation and analysis of expressed sequence tags from six developing xylem libraries in <italic>Pinus radiata</italic> D. Don.</article-title> <source><italic>BMC Genomics</italic></source> <volume>10</volume>:<issue>41</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-10-41</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Wu</surname> <given-names>H. X.</given-names></name> <name><surname>Southerton</surname> <given-names>S. G.</given-names></name></person-group> (<year>2011</year>). <article-title>Transcriptome profiling of <italic>Pinus radiata</italic> juvenile wood with contrasting stiffness identifies putative candidate genes involved in microfibril orientation and cell wall mechanics.</article-title> <source><italic>BMC Genomics</italic></source> <volume>12</volume>:<issue>480</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-12-480</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Wu</surname> <given-names>H. X.</given-names></name></person-group> (<year>2013</year>). <article-title>Transcriptome profiling of radiata pine branches reveals new insights into reaction wood formation with implications in plant gravitropism.</article-title> <source><italic>BMC Genomics</italic></source> <volume>14</volume>:<issue>768</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-14-768</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ling</surname> <given-names>X.-Y.</given-names></name> <name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Pan</surname> <given-names>G.</given-names></name> <name><surname>Long</surname> <given-names>H.</given-names></name> <name><surname>Cheng</surname> <given-names>Y.</given-names></name> <name><surname>Xiang</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Preparing long probes by an asymmetric polymerase chain reaction-based approach for multiplex ligation-dependent probe amplification.</article-title> <source><italic>Anal. Biochem.</italic></source> <volume>487</volume> <fpage>8</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1016/j.ab.2015.03.031</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcinkowska</surname> <given-names>M.</given-names></name> <name><surname>Wong</surname> <given-names>K.-K.</given-names></name> <name><surname>Kwiatkowski</surname> <given-names>D. J.</given-names></name> <name><surname>Kozlowski</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Design and generation of MLPA probe sets for combined copy number and small-mutation analysis of human genes: EGFR as an example.</article-title> <source><italic>ScientificWorldJournal.</italic></source> <volume>10</volume> <fpage>2003</fpage>&#x2013;<lpage>2018</lpage>. <pub-id pub-id-type="doi">10.1100/tsw.2010.195</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcinkowska-Swojak</surname> <given-names>M.</given-names></name> <name><surname>Handschuh</surname> <given-names>L.</given-names></name> <name><surname>Wojciechowski</surname> <given-names>P.</given-names></name> <name><surname>Goralski</surname> <given-names>M.</given-names></name> <name><surname>Tomaszewski</surname> <given-names>K.</given-names></name> <name><surname>Kazmierczak</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Simultaneous detection of mutations and copy number variation of NPM1 in the acute myeloid leukemia using multiplex ligation-dependent probe amplification.</article-title> <source><italic>Mutat. Res.</italic></source> <volume>786</volume> <fpage>14</fpage>&#x2013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1016/j.mrfmmm.2016.02.001</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcinkowska-Swojak</surname> <given-names>M.</given-names></name> <name><surname>Klonowska</surname> <given-names>K.</given-names></name> <name><surname>Figlerowicz</surname> <given-names>M.</given-names></name> <name><surname>Kozlowski</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>An MLPA-based approach for high-resolution genotyping of disease-related multi-allelic CNVs.</article-title> <source><italic>Gene</italic></source> <volume>546</volume> <fpage>257</fpage>&#x2013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1016/j.gene.2014.05.072</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maron</surname> <given-names>L. G.</given-names></name> <name><surname>Guimar&#x00E3;es</surname> <given-names>C. T.</given-names></name> <name><surname>Kirst</surname> <given-names>M.</given-names></name> <name><surname>Albert</surname> <given-names>P. S.</given-names></name> <name><surname>Birchler</surname> <given-names>J. A.</given-names></name> <name><surname>Bradbury</surname> <given-names>P. J.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Aluminum tolerance in maize is associated with higher MATE1 gene copy number.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>110</volume> <fpage>5241</fpage>&#x2013;<lpage>5246</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1220766110</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCord</surname> <given-names>B.</given-names></name></person-group> (<year>2003</year>). <source><italic>Troubleshooting Capillary Electrophoresis Systems.</italic></source> Available at: <ext-link ext-link-type="uri" xlink:href="https://pl.promega.com/resources/profiles-in-dna/2003/troubleshooting-capillary-electrophoresis-systems/">https://pl.promega.com/resources/profiles-in-dna/2003/troubleshooting-capillary-electrophoresis-systems/</ext-link></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McHale</surname> <given-names>L. K.</given-names></name> <name><surname>Haun</surname> <given-names>W. J.</given-names></name> <name><surname>Xu</surname> <given-names>W. W.</given-names></name> <name><surname>Bhaskar</surname> <given-names>P. B.</given-names></name> <name><surname>Anderson</surname> <given-names>J. E.</given-names></name> <name><surname>Hyten</surname> <given-names>D. L.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Structural variants in the soybean genome localize to clusters of biotic stress-response genes.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>159</volume> <fpage>1295</fpage>&#x2013;<lpage>1308</lpage>. <pub-id pub-id-type="doi">10.1104/pp.112.194605</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mu&#x00F1;oz-Amatria&#x00ED;n</surname> <given-names>M.</given-names></name> <name><surname>Eichten</surname> <given-names>S. R.</given-names></name> <name><surname>Wicker</surname> <given-names>T.</given-names></name> <name><surname>Richmond</surname> <given-names>T. A.</given-names></name> <name><surname>Mascher</surname> <given-names>M.</given-names></name> <name><surname>Steuernagel</surname> <given-names>B.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Distribution, functional impact, and origin mechanisms of copy number variation in the barley genome.</article-title> <source><italic>Genome Biol.</italic></source> <volume>14</volume>:<issue>R58</issue>. <pub-id pub-id-type="doi">10.1186/gb-2013-14-6-r58</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perne</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Lehmann</surname> <given-names>L.</given-names></name> <name><surname>Groth</surname> <given-names>M.</given-names></name> <name><surname>Stuber</surname> <given-names>F.</given-names></name> <name><surname>Book</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Comparison of multiplex ligation-dependent probe amplification and real-time PCR accuracy for gene copy number quantification using the beta-defensin locus.</article-title> <source><italic>Biotechniques</italic></source> <volume>47</volume> <fpage>1023</fpage>&#x2013;<lpage>1028</lpage>. <pub-id pub-id-type="doi">10.2144/000113300</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rudi</surname> <given-names>K.</given-names></name> <name><surname>Rud</surname> <given-names>I.</given-names></name> <name><surname>Holck</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <article-title>A novel multiplex quantitative DNA array based PCR (MQDA-PCR) for quantification of transgenic maize in food and feed.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>31</volume>:<issue>e62</issue>. <pub-id pub-id-type="doi">10.1007/s00217-009-1155-4</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saintenac</surname> <given-names>C.</given-names></name> <name><surname>Jiang</surname> <given-names>D.</given-names></name> <name><surname>Akhunov</surname> <given-names>E. D.</given-names></name></person-group> (<year>2011</year>). <article-title>Targeted analysis of nucleotide and copy number variation by exon capture in allotetraploid wheat genome.</article-title> <source><italic>Genome Biol.</italic></source> <volume>12</volume>:<issue>R88</issue>. <pub-id pub-id-type="doi">10.1186/gb-2011-12-9-r88</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schouten</surname> <given-names>J. P.</given-names></name> <name><surname>McElgunn</surname> <given-names>C. J.</given-names></name> <name><surname>Waaijer</surname> <given-names>R.</given-names></name> <name><surname>Zwijnenburg</surname> <given-names>D.</given-names></name> <name><surname>Diepvens</surname> <given-names>F.</given-names></name> <name><surname>Pals</surname> <given-names>G.</given-names></name></person-group> (<year>2002</year>). <article-title>Relative quantification of 40 nucleic acid sequences by multiplex ligation-dependent probe amplification.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>30</volume>:<issue>e57</issue>. <pub-id pub-id-type="doi">10.1093/nar/gnf056</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Springer</surname> <given-names>N. M.</given-names></name> <name><surname>Ying</surname> <given-names>K.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name> <name><surname>Ji</surname> <given-names>T.</given-names></name> <name><surname>Yeh</surname> <given-names>C.-T.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name><etal/></person-group> (<year>2009</year>). <article-title>Maize inbreds exhibit high levels of copy number variation (CNV) and presence/absence variation (PAV) in genome content.</article-title> <source><italic>PLoS Genet.</italic></source> <volume>5</volume>:<issue>e1000734</issue>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1000734</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stankiewicz</surname> <given-names>P.</given-names></name> <name><surname>Lupski</surname> <given-names>J. R.</given-names></name></person-group> (<year>2010</year>). <article-title>Structural variation in the human genome and its role in disease.</article-title> <source><italic>Annu. Rev. Med.</italic></source> <volume>61</volume> <fpage>437</fpage>&#x2013;<lpage>455</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-med-100708-204735</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swanson-Wagner</surname> <given-names>R. A.</given-names></name> <name><surname>Eichten</surname> <given-names>S. R.</given-names></name> <name><surname>Kumari</surname> <given-names>S.</given-names></name> <name><surname>Tiffin</surname> <given-names>P.</given-names></name> <name><surname>Stein</surname> <given-names>J. C.</given-names></name> <name><surname>Ware</surname> <given-names>D.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Pervasive gene content variation and copy number variation in maize and its undomesticated progenitor.</article-title> <source><italic>Genome Res.</italic></source> <volume>20</volume> <fpage>1689</fpage>&#x2013;<lpage>1699</lpage>. <pub-id pub-id-type="doi">10.1101/gr.109165.110</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>S.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name> <name><surname>Hou</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Tian</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Variation of presence/absence genes among <italic>Arabidopsis</italic> populations.</article-title> <source><italic>BMC Evol. Biol.</italic></source> <volume>12</volume>:<issue>86</issue>. <pub-id pub-id-type="doi">10.1186/1471-2148-12-86</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thumma</surname> <given-names>B. R.</given-names></name> <name><surname>Matheson</surname> <given-names>B. A.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name> <name><surname>Meeske</surname> <given-names>C.</given-names></name> <name><surname>Meder</surname> <given-names>R.</given-names></name> <name><surname>Downes</surname> <given-names>G. M.</given-names></name><etal/></person-group> (<year>2009</year>). <article-title>Identification of a cis-acting regulatory polymorphism in a eucalypt COBRA-like gene affecting cellulose content.</article-title> <source><italic>Genetics</italic></source> <volume>183</volume> <fpage>1153</fpage>&#x2013;<lpage>1164</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.109.106591</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Xiong</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Jiang</surname> <given-names>L.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Copy number variation at the GL7 locus contributes to grain size diversity in rice.</article-title> <source><italic>Nat. Genet.</italic></source> <volume>47</volume> <fpage>944</fpage>&#x2013;<lpage>948</lpage>. <pub-id pub-id-type="doi">10.1038/ng.3346</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zarrei</surname> <given-names>M.</given-names></name> <name><surname>MacDonald</surname> <given-names>J. R.</given-names></name> <name><surname>Merico</surname> <given-names>D.</given-names></name> <name><surname>Scherer</surname> <given-names>S. W.</given-names></name></person-group> (<year>2015</year>). <article-title>A copy number variation map of the human genome.</article-title> <source><italic>Nat. Rev. Genet.</italic></source> <volume>16</volume> <fpage>172</fpage>&#x2013;<lpage>183</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3871</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>L.-Y.</given-names></name> <name><surname>Guo</surname> <given-names>X.-S.</given-names></name> <name><surname>He</surname> <given-names>B.</given-names></name> <name><surname>Sun</surname> <given-names>L.-J.</given-names></name> <name><surname>Peng</surname> <given-names>Y.</given-names></name> <name><surname>Dong</surname> <given-names>S.-S.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Genome-wide patterns of genetic variation in sweet and grain sorghum (<italic>Sorghum bicolor</italic>).</article-title> <source><italic>Genome Biol.</italic></source> <volume>12</volume>:<issue>R114</issue>. <pub-id pub-id-type="doi">10.1186/gb-2011-12-11-r114</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x017B;mie&#x0144;ko</surname> <given-names>A.</given-names></name> <name><surname>Samelak</surname> <given-names>A.</given-names></name> <name><surname>Koz&#x0142;owski</surname> <given-names>P.</given-names></name> <name><surname>Figlerowicz</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Copy number polymorphism in plant genomes.</article-title> <source><italic>Theor. Appl. Genet.</italic></source> <volume>127</volume> <fpage>1</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1007/s00122-013-2177-7</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zmienko</surname> <given-names>A.</given-names></name> <name><surname>Samelak-Czajka</surname> <given-names>A.</given-names></name> <name><surname>Kozlowski</surname> <given-names>P.</given-names></name> <name><surname>Szymanska</surname> <given-names>M.</given-names></name> <name><surname>Figlerowicz</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title><italic>Arabidopsis thaliana</italic> population analysis reveals high plasticity of the genomic region spanning MSH2 AT3G18530 and AT3G18535 genes and provides evidence for NAHR-driven recurrent CNV events occurring in this location.</article-title> <source><italic>BMC Genomics</italic></source> <volume>17</volume>:<issue>893</issue>. <pub-id pub-id-type="doi">10.1186/s12864-016-3221-1</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="https://gbrowse.arabidopsis.org/cgi-bin/gb2/gbrowse/arabidopsis/">https://gbrowse.arabidopsis.org/cgi-bin/gb2/gbrowse/arabidopsis/</ext-link></p></fn>
<fn id="fn02"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="http://tools.1001genomes.org">http://tools.1001genomes.org</ext-link></p></fn>
<fn id="fn03"><label>3</label><p><ext-link ext-link-type="uri" xlink:href="http://www.mrc-holland.com/">http://www.mrc-holland.com/</ext-link></p></fn>
<fn id="fn04"><label>4</label><p><ext-link ext-link-type="uri" xlink:href="http://arabidopsis.info/">http://arabidopsis.info/</ext-link></p></fn>
<fn id="fn05"><label>5</label><p><ext-link ext-link-type="uri" xlink:href="http://1001genomes.org">http://1001genomes.org</ext-link></p></fn>
<fn id="fn06"><label>6</label><p><ext-link ext-link-type="uri" xlink:href="http://www.mlpa.com/elearning/tswizard/">http://www.mlpa.com/elearning/tswizard/</ext-link></p></fn>
</fn-group>
</back>
</article>