<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fgene.2017.00139</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The Empirical Distribution of Singletons for Geographic Samples of DNA Sequences</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Cubry</surname> <given-names>Philippe</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/430479/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Vigouroux</surname> <given-names>Yves</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/324655/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Fran&#x000E7;ois</surname> <given-names>Olivier</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/65620/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>UMR DIADE, University of Montpellier</institution>, <addr-line>Montpellier</addr-line>, <country>France</country></aff>
<aff id="aff2"><sup>2</sup><institution>TIMC-IMAG UMR 5525, Centre National de la Recherche Scientifique (CNRS), Universit&#x000E9; Grenoble-Alpes</institution>, <addr-line>Grenoble</addr-line>, <country>France</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Samuel A. Cushman, United States Forest Service Rocky Mountain Research Station, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Pablo Orozco-terWengel, Cardiff University, United Kingdom; Ricardo T. Pereyra, University of Gothenburg, Sweden; Rita Rasteiro, University of Bristol, United Kingdom</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Olivier Fran&#x000E7;ois <email>olivier.francois&#x00040;univ-grenoble-alpes.fr</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Evolutionary and Population Genetics, a section of the journal Frontiers in Genetics</p></fn></author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>09</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>139</elocation-id>
<history>
<date date-type="received">
<day>27</day>
<month>03</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>09</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Cubry, Vigouroux and Fran&#x000E7;ois.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Cubry, Vigouroux and Fran&#x000E7;ois</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Rare variants are important for drawing inference about past demographic events in a species history. A singleton is a rare variant for which genetic variation is carried by a unique chromosome in a sample. How singletons are distributed across geographic space provides a local measure of genetic diversity that can be measured at the individual level. Here, we define the empirical distribution of singletons in a sample of chromosomes as the proportion of the total number of singletons that each chromosome carries, and we present a theoretical background for studying this distribution. Next, we use computer simulations to evaluate the potential for the empirical distribution of singletons to provide a description of genetic diversity across geographic space. In a Bayesian framework, we show that the empirical distribution of singletons leads to accurate estimates of the geographic origin of range expansions. We apply the Bayesian approach to estimating the origin of the cultivated plant species <italic>Pennisetum glaucum [L.] R. Br</italic>. (pearl millet) in Africa, and find support for range expansion having started from Northern Mali. Overall, we report that the empirical distribution of singletons is a useful measure to analyze results of sequencing projects based on large scale sampling of individuals across geographic space.</p></abstract>
<kwd-group>
<kwd>genetic diversity</kwd>
<kwd>singletons</kwd>
<kwd>geographic origin</kwd>
<kwd>range expansion</kwd>
<kwd>pearl millet</kwd>
</kwd-group>
<contract-num rid="cn001">ANR-13-BSV7-0017</contract-num>
<contract-num rid="cn001">ANR-11-LABX-0025-01</contract-num>
<contract-sponsor id="cn001">Agence Nationale de la Recherche<named-content content-type="fundref-id">10.13039/501100001665</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="0"/>
<equation-count count="10"/>
<ref-count count="43"/>
<page-count count="10"/>
<word-count count="5480"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>High-throughput sequencing technologies have enabled studies of genomic diversity in model and non-model species at a dramatically increasing rate. Conducted at population and at individual levels, those studies have provided comprehensive surveys of common and rare variation in model species genomes (Weigel and Mott, <xref ref-type="bibr" rid="B42">2009</xref>; 1000 Genomes Project Consortium et al., <xref ref-type="bibr" rid="B2">2010</xref>; International HapMap 3 Consortium, <xref ref-type="bibr" rid="B21">2010</xref>; 1000 Genomes Project Consortium, <xref ref-type="bibr" rid="B1">2015</xref>). For example, the 1000 Genomes Project Consortium (<xref ref-type="bibr" rid="B1">2015</xref>) reported that the majority of variants in human genomes are rare. During the last decade, the role that rare variants play in shaping complex traits has been hotly debated (Pritchard, <xref ref-type="bibr" rid="B33">2001</xref>; Schork et al., <xref ref-type="bibr" rid="B34">2009</xref>; Tennessen et al., <xref ref-type="bibr" rid="B38">2012</xref>), and accurately determining their distribution has become important for medical applications and association studies (Lee et al., <xref ref-type="bibr" rid="B23">2014</xref>; Auer and Lettre, <xref ref-type="bibr" rid="B3">2015</xref>). Beyond humans, rare variation has attracted considerable interest from genome sequencing projects for model organisms, including plants (Zhu et al., <xref ref-type="bibr" rid="B43">2011</xref>; Weigel, <xref ref-type="bibr" rid="B41">2012</xref>; Memon et al., <xref ref-type="bibr" rid="B27">2016</xref>).</p>
<p>Rare variants are also important for drawing inference about past demographic events in a species history (Schraiber and Akey, <xref ref-type="bibr" rid="B35">2015</xref>). Studies of human populations have shown that our species has experienced a complex demographic history, and that a recent period of explosive growth has resulted in an excess of those variants (Coventry et al., <xref ref-type="bibr" rid="B9">2010</xref>; Keinan and Clark, <xref ref-type="bibr" rid="B22">2012</xref>). The analysis of private and rare variation has been used to reveal signals of differential demographic history among populations, and to refine models of human evolution (Marth et al., <xref ref-type="bibr" rid="B25">2004</xref>; Gravel et al., <xref ref-type="bibr" rid="B17">2011</xref>; Mathieson and McVean, <xref ref-type="bibr" rid="B26">2014</xref>). In addition, estimating rare allele frequencies has enabled estimates of gene flow between populations, and has facilitated inference of fine-scale population structure (Slatkin, <xref ref-type="bibr" rid="B36">1985</xref>; Novembre and Slatkin, <xref ref-type="bibr" rid="B28">2009</xref>; O&#x00027;Connor et al., <xref ref-type="bibr" rid="B29">2015</xref>).</p>
<p>In this study, we define the empirical distribution of singletons in a sample of chromosomes as the proportion of the total number of singletons that each chromosome carries, where a singleton is a uniquely represented allele in the sample (Fu and Li, <xref ref-type="bibr" rid="B16">1993</xref>). We provide theoretical and empirical analyses of the distribution of singletons in a sample of chromosomes, and we evaluate the potential for this distribution to provide an accurate description of genetic diversity at the individual level. Using spatial data, we use the distribution of singletons as an individual-based estimate of genetic diversity in geographic space.</p>
<p>The theoretical background for the analysis of the empirical distribution of singletons rely on the distribution of external branch lengths for coalescent genealogies (Blum and Fran&#x000E7;ois, <xref ref-type="bibr" rid="B5">2005</xref>; Caliebe et al., <xref ref-type="bibr" rid="B7">2007</xref>). First, we use coalescent and spatially explicit simulations to evaluate individual contributions to genetic diversity in the sample based on singletons. Then we evaluate the use of the distribution of singletons in an approximate Bayesian Computation (ABC) framework to estimate the geographic origin of range expansions (Beaumont, <xref ref-type="bibr" rid="B4">2010</xref>; Csill&#x000E9;ry et al., <xref ref-type="bibr" rid="B11">2010</xref>). We eventually provide an illustration of our theory by applying the ABC approach to the plant species <italic>Pennisetum glaucum [L.] R. Br</italic>. (pearl millet). Pearl millet is a cereal cultivated in semi-arid regions of Africa and the Indian subcontinent, and it is known to originate in Africa (Clotault et al., <xref ref-type="bibr" rid="B8">2012</xref>). We evaluate the geographic origin of its range expansion by using 146 inbred lines from the whole African range.</p>
</sec>
<sec id="s2">
<title>2. Theory</title>
<p>We consider a sample of <italic>n</italic> chromosomes from a population of <italic>N</italic> haploid organisms. We assume that there are <italic>L</italic> polymorphic loci, and that for each locus, 0 represents the ancestral or reference allele and 1 is the derived allele. A singleton is defined as a derived allele carried by a single chromosome in the sample. The total number of singletons, &#x003BE;<sub>1</sub>, is the number of uniquely represented derived alleles in the sample, and it corresponds to the first component of the site frequency spectrum. We assume that the singletons are distributed over the <italic>n</italic> chromosomes in the sample. More specifically, the number of singletons decomposes as follows</p>
<disp-formula id="E1"><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the number of singletons carried by chromosome <italic>i</italic>. For each <italic>i</italic>, we denote by <italic>p</italic><sub><italic>i</italic></sub> the conditional probability that a singleton is carried by <italic>i</italic>. The <italic>n</italic> values <italic>p</italic><sub>1</sub>, &#x02026;, <italic>p</italic><sub><italic>n</italic></sub> sum up to one, and those values define the <italic>empirical distribution of singletons</italic> in the sample (see below).</p>
<p>Next, we assume that the sample genealogies can be described by coalescent trees (Tavar&#x000E9;, <xref ref-type="bibr" rid="B37">2004</xref>). For a particular locus, a tree is described by <italic>n</italic> tips and <italic>n</italic>&#x02212;1 ancestral nodes. An external branch of the tree connects a tip to an ancestral node. For a given tree, we denote by &#x003C4;<sup>(<italic>i</italic>)</sup> the length of the external branch connecting chromosome <italic>i</italic> to its first ancestor node. The <italic>L</italic> coalescent trees exhibit complex patterns of statistical dependency along the chromosomes due to recombination among loci (Hudson, <xref ref-type="bibr" rid="B19">1990</xref>). Measuring lengths in units of twice the total population size (<italic>N</italic>), and assuming a molecular clock model for mutations, the number of mutations falling on a particular branch of the tree has a Poisson distribution of rate &#x003B8;/2, where &#x003B8; &#x0003D; 2&#x003BC;<italic>N</italic> and &#x003BC; is the per generation mutation rate (Tavar&#x000E9;, <xref ref-type="bibr" rid="B37">2004</xref>). Let &#x02113; be an arbitrary singleton locus. For all <italic>i</italic>, we write</p>
<disp-formula id="E2"><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>&#x02113;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>where <italic>X</italic><sub><italic>i&#x02113;</italic></sub> &#x0003D; 1 if singleton &#x02113; is carried by chromosome <italic>i</italic>, 0 otherwise. In the above formula, the summation runs over all singletons in the sample. Using mathematical properties of conditional distributions for the Poisson process, we have</p>
<disp-formula id="E3"><mml:math id="M4"><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>P</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>&#x02113;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mtext>E&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>. In this formula, the conditional probability that chromosome <italic>i</italic> carries a singleton at locus &#x02113; is given by the ratio of its external branch length to the total length of external branches in the sample genealogy at this locus. The distribution of singletons can be estimated by counting the number of singletons carried by each chromosome and normalizing as follows</p>
<disp-formula id="E4"><mml:math id="M6"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>and the estimate is unbiased</p>
<disp-formula id="E5"><mml:math id="M7"><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>In addition, the number of singletons carried by chromosome <italic>i</italic>, <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, estimates the proportion of genetic diversity carried by chromosome <italic>i</italic></p>
<disp-formula id="E6"><mml:math id="M9"><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02248;</mml:mo><mml:mi>&#x003B8;</mml:mi><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>As a consequence of the theory presented in this section, the individual-based estimates of genetic diversity are unbiased quantities regardless of demographic history, deviations from Hardy-Weinberg equilibrium and linkage disequilibrium. Limitations of the theory include the presence of closely related individuals, which should be removed from the sample prior to analysis. The approach is appropriate for modern sequencing data as soon as a few hundreds of DNA sequences are generated.</p>
<p>The rest of this study will evaluate the use of the empirical distribution of singletons in mapping genetic diversity in geographic space. To provide an elementary example, let us consider a sample of <italic>n</italic> chromosomes from a random mating population of size <italic>N</italic>. Using mathematical results for the neutral coalescent in a random mating population, the expected value of the number of singletons is an unbiased estimator of the genetic diversity in the sample (Fu and Li, <xref ref-type="bibr" rid="B16">1993</xref>)</p>
<disp-formula id="E7"><mml:math id="M10"><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003B8;</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>For the lengths of external branch lengths, we have</p>
<disp-formula id="E8"><mml:math id="M11"><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>/</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>and E[&#x003C4;<sub>1</sub>] &#x0003D; 2 (Blum and Fran&#x000E7;ois, <xref ref-type="bibr" rid="B5">2005</xref>). Here, we expect that each chromosome contributes to genetic diversity equally. The above calculations show that, in a sample of size <italic>n</italic> from a random mating population, the distribution of singletons is uniform over the <italic>n</italic> chromosomes</p>
<disp-formula id="E9"><mml:math id="M12"><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>and we have</p>
<disp-formula id="E10"><mml:math id="M13"><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mtext>E</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mtext>P</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003B8;</mml:mi><mml:mo>/</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>In other words, each individual contributes the same amount of genetic variation to the total sample diversity.</p>
</sec>
<sec id="s3">
<title>3. Simulation methods and data sets</title>
<sec>
<title>3.1. Coalescent simulations of splitting populations</title>
<p>We used the computer program <italic>ms</italic> to perform coalescent simulations for a two-population model (Hudson, <xref ref-type="bibr" rid="B20">2002</xref>). In our simulations, we considered a population split model, in which two populations of sizes <italic>N</italic><sub>1</sub> &#x0003D; 50, 000 and <italic>N</italic><sub>2</sub> &#x0003D; <italic>sN</italic><sub>1</sub> (<italic>s</italic> &#x02208; (0.01;0.5), shrink rate) diverged <italic>t</italic> generations ago (<italic>t</italic> &#x02208; (1, 000;10, 000), split time). Population 1 expanded from an ancestral population of size <italic>N</italic><sub><italic>A</italic></sub> &#x0003D; 5, 000, and the expansion started 10,000 generations ago. Samples of size <italic>n</italic> &#x0003D; 100 were considered and subdivided into subsamples of size 50 from each population. We simulated <italic>L</italic> &#x0003D; 1, 000 unlinked haplotypes using the infinite-site model and an effective mutation rate &#x003B8; &#x02208; (5;10). The <italic>ms</italic> command line was written as follows: ./ms 100 1,000 -t theta -I 2 50 50 -g 1 46.05 -n 2 shrink.rate -eg 0.2 1 0.0 -ej split.time 2 1. The simulated data sets were processed by using the &#x0201C;.geno&#x0201D; format in the <monospace>R</monospace> package LEA (Frichot and Fran&#x000E7;ois, <xref ref-type="bibr" rid="B15">2015</xref>). We summarized the distribution of singletons by computing mean values and standard errors for each subsample. For all simulated samples, we used the <monospace>R</monospace> package <italic>ape</italic> to extract the coalescent trees generated by <italic>ms</italic>, and analyze the distribution of their external branch lengths (Paradis et al., <xref ref-type="bibr" rid="B32">2004</xref>). We used the external branch length distribution to build a theoretical prediction for the distribution of singletons from each tree (see section 2), and summarized the theoretical distributions by computing mean values and standard errors for each subsample. The <italic>L</italic> coalescent simulations were replicated 200 times.</p>
</sec>
<sec>
<title>3.2. Range expansions in Africa</title>
<p>Simulations of range expansions were performed by using the computer program SPLATCHE2 based on an array of 87 by 83 demes modeling the African continent (Currat et al., <xref ref-type="bibr" rid="B13">2004</xref>). The demographic scenarios corresponded to range expansions from a single origin, simulated for a total duration of 1, 600 generations. For each deme, the migration rate was equal to <italic>m</italic> &#x0003D; 0.07, and the growth rate was equal to <italic>r</italic> &#x0003D; 0.1. Additional parameters included an ancestral effective population size of 200 individuals, 200 generations before onset of expansion, and an effective mutation rate of 10<sup>&#x02212;5</sup> per base pair per generation.</p>
<p>Four types of demographic scenarios were considered. Two scenarios considered a &#x0201C;homogeneous&#x0201D; environment, for which the deme carrying capacities were set to a constant value <italic>C</italic> &#x0003D; 100 everywhere in Africa. Two other scenarios considered a heterogeneous environment linked to vegetation. In tropical semi-desert areas, the carrying capacities were set to <italic>C</italic> &#x0003D; 60, and in tropical extreme deserts and rain forests, the carrying capacities were set to <italic>C</italic> &#x0003D; 30. Demographic histories also differed by their geographic source of expansion. Range expansions were started either from an origin in West Africa (Mali, &#x02212;4&#x000B0; E, 13&#x000B0; N) or from an origin in the Sahel area (Chad, 22&#x000B0; E, 20&#x000B0; N).</p>
<p>Ten haploid chromosomes were simulated for 30 population samples through the geographic range considered (300 chromosomes). Genetic variation was surveyed at 30,000 loci, and filtered out for monomorphic loci. From the resulting data sets, we computed the empirical distribution of singletons in each population sample, and compared this measure to expected heterozygosity for each population sample. Data files for running the SPLATCHE2 simulations are provided in Supplementary File <xref ref-type="supplementary-material" rid="SM4">1</xref>. We reproduced the four scenarios by using individual sampling instead of population sampling. Here, individual genotypes were recorded at 300 distinct geographic sites, each obtained from a Gaussian perturbation of population centers with standard error of 2&#x000B0;. The Kriging method was used to interpolate the values of the expected heterozygosity and the empirical distribution of singletons on a geographic map of Africa (Cressie, <xref ref-type="bibr" rid="B10">2015</xref>).</p>
</sec>
<sec>
<title>3.3. Pearl millet data</title>
<p>Whole genome sequencing data were obtained for 146 cultivated accessions of pearl millet (<italic>Pennisetum glaucum [L.] R. Br</italic>.) from the species range in Africa (International Pearl Millet Genome Sequencing Consortium, Varshney et al., <xref ref-type="bibr" rid="B40">2017</xref>). A total of 169,095 SNPs were sampled after filtering out low quality variants, and were used to estimate the distribution of singletons (Supplementary Material <xref ref-type="supplementary-material" rid="SM4">1</xref>).</p>
</sec>
<sec>
<title>3.4. Approximate bayesian computation</title>
<p>We used Approximate Bayesian Computation (ABC) to evaluate the ability of the distribution of singletons to correctly estimate the onset of expansion in a range expanding species, and to estimate a posterior distribution for the location of this origin for cultivated pearl millet. We performed 20,000 range expansion simulations by considering a heterogeneous environment using the computer program SPLATCHE2. The deme carrying capacities were equal to <italic>C</italic> &#x0003D; 100 for tropical semi-desert areas, <italic>C</italic> &#x0003D; 20 for tropical extreme deserts and <italic>C</italic> &#x0003D; 10 for rain forests. Additional parameters included an ancestral effective population size of 200 individuals, 200 generations before onset of expansion, and an effective mutation rate of 10<sup>&#x02212;5</sup> per base pair per generation.</p>
<p>Prior distributions allowed the geographic coordinates of the origin of expansion to vary over the Sahel region. Longitude ranged between &#x02212;16&#x000B0;E and 40&#x000B0;E, and latitude ranged between 5&#x000B0;N and 30&#x000B0;N. Lower prior probabilities were given to extreme latitudes and longitudes as a consequence of unsuitable habitats (water regions). Uninformative prior distributions were considered for the migration rate, the growth rate, the total duration of the demographic phase, the ancestral population size and the time before onset of expansion (Supplementary Table <xref ref-type="supplementary-material" rid="SM5">1</xref>). In simulations, genetic variation was surveyed at 146 geographic sites corresponding to the exact sampling locations of pearl millet accessions. Ten thousands SNPs were simulated for each genotype. When evaluating summary statistics, a fraction of SNPs were removed from the simulated data in order to match with the amount of missing values observed in the original data set.</p>
<p>To define the summary statistics for ABC, we used a histogram for the distribution of singletons in the sample. The 146 accessions were grouped into spatial clusters according to a <italic>k</italic>-means algorithm and individual geographic information (Hartigan and Wong, <xref ref-type="bibr" rid="B18">1979</xref>). The <italic>k</italic>-means algorithm resulted in 14 groups with more than 6 accessions in each group (Figure <xref ref-type="fig" rid="F1">1</xref>). To obtain a histogram, we computed the mean number of singletons in each group, and divided this value by the total number of singletons in the sample (Supplementary Table <xref ref-type="supplementary-material" rid="SM6">2</xref>). Then ABC analysis was performed with the <monospace>R</monospace> package <monospace>abc</monospace> (Blum and Fran&#x000E7;ois, <xref ref-type="bibr" rid="B6">2010</xref>; Csill&#x000E9;ry et al., <xref ref-type="bibr" rid="B12">2012</xref>). Neural network models were used to estimate posterior distributions for the latitude and longitude of the geographic onset of expansion whereas the other parameters were considered as nuisance parameters without any interpretable unit. The tolerance rate was set to 0.05 and 250 neural networks were used in the <monospace>abc</monospace> function.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Geographic distribution of 146 cultivated accessions of pearl millet. Fourteen geographic classes were defined as a result of a <italic>k</italic>-means procedure.</p></caption>
<graphic xlink:href="fgene-08-00139-g0001.tif"/>
</fig>
<p>We first tested the accuracy of our estimates by using simulated data sets as inputs to the inference method. The sampling procedure and the ABC estimation were replicated 100 times, and we evaluated the correlation between coordinates of true origins and their estimated values. Then we considered the pearl millet data, and represented the prior and posterior densities of the geographic onset parameters by using two-dimensional kernel density estimation with 100 grid points in each direction.</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4. Results</title>
<sec>
<title>4.1. Coalescent simulations of splitting populations</title>
<p>To evaluate statistical bias in the estimation of the distribution of singletons, we performed coalescent simulations of samples from two populations with unequal genetic diversity. The two populations diverged from an ancestral population <italic>t</italic> generations ago (<italic>split time</italic>), and at split time, the size of population 2 shrinked to <italic>s</italic> times the size of population 1 (<italic>shrink rate</italic>).</p>
<p>For each simulation, the number of polymorphic loci ranged between 7,883 and 39,761 (average value: 25,265 loci). For a value of the shrink rate <italic>s</italic>&#x02248;1/3, the average proportion of singletons in population 1 was about &#x003C0;<sub>1</sub> &#x0003D; 0.0122, and the average proportion of singletons in population 2 was about &#x003C0;<sub>2</sub> &#x0003D; 0.0078 (&#x003C0;<sub>1</sub>&#x0002B;&#x003C0;<sub>2</sub> &#x0003D; 2/<italic>n</italic>). This result reflected that genetic diversity in population 1 was higher than in population 2. The ratio was about &#x003C0;<sub>1</sub>/&#x003C0;<sub>2</sub> &#x0003D; 1.55 (Figure <xref ref-type="fig" rid="F2">2A</xref>). The individual proportions were concentrated around their mean values with relatively small standard deviations (SD<sub>1</sub> &#x0003D; 0.0010, SD<sub>2</sub> &#x0003D; 0.0008).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Coalescent simulations of two splitting populations (100 chromosomes). <bold>(A)</bold> Empirical distribution of singletons for a value of the shrink rate <italic>s</italic> &#x0003D; 0.33. The dashed lines represent the averaged values for population 1 (expanding) and population 2 (skrinking). <bold>(B)</bold> Predicted and observed (empirical) values of the distribution of singletons for population 1 (left) and population 2 (right).</p></caption>
<graphic xlink:href="fgene-08-00139-g0002.tif"/>
</fig>
<p>The results from 200 replicates provided clear evidence that the empirical distribution of singletons is an unbiased estimate of its theoretical distribution based on coalescent trees (Figure <xref ref-type="fig" rid="F2">2B</xref>). The split time parameter had a weak influence on the distribution of singletons (Pearson correlation test, <italic>P</italic> &#x0003D; 0.64). The ratio &#x003C0;<sub>1</sub>/&#x003C0;<sub>2</sub> reached values between 10 and 40 when the shrink rate was below 10%, and this parameter had a strong influence on the empirical distribution of singletons (Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
</sec>
<sec>
<title>4.2. Range expansions in Africa</title>
<p>For data sets generated under range expansion scenarios, the number of polymorphic loci ranged between 25,453 and 29,321 loci. The number of singletons ranged between 8,835 and 12,653, and the site frequency spectrum showed an excess of rare alleles as expected under explosive population growth. When the onset of expansion was set in Western Africa (cross in Figure <xref ref-type="fig" rid="F3">3</xref>), the maps of the empirical distribution of singletons and expected heterozygosity exhibited similar large-scale geographic patterns (Figure <xref ref-type="fig" rid="F3">3</xref>, Pearson&#x00027;s correlation coefficient 0.78). Because the computation of expected heterozygosities was based on a perfect assignment of samples to their true populations of origin, the interpolated maps corresponding to this measure (Figures <xref ref-type="fig" rid="F3">3B,D</xref>) contained less uncertainty than the maps of singletons (Figures <xref ref-type="fig" rid="F3">3A,C</xref>) that were based on random individual sampling. Considering environmental heterogeneity increased the variability of spatial estimates (Figures <xref ref-type="fig" rid="F3">3C,D</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Individual vs. population sampling after a range expansion simulation scenario (Western origin). <bold>(A,B)</bold> Homogeneous environment. Maps of the empirical distribution of singletons (individual sampling) and expected heterozygosity (true population sampling). <bold>(C,D)</bold> Inhomogeneous environment.</p></caption>
<graphic xlink:href="fgene-08-00139-g0003.tif"/>
</fig>
<p>Next, we compared estimates of heterozygosity for populations to the distribution of singletons in the same populations (Figure <xref ref-type="fig" rid="F4">4</xref>). Differences between maps produced with the empirical distribution of singletons and with expected heterozygosity decreased when the sampled chromosomes were perfectly assigned to their population of origin. The individual and population-based measures provided concordant estimates of genetic diversity in geographic space (Pearson&#x00027;s correlation coefficient 0.51). Similar results were observed when the onset of expansion was set in the Sahel area (20&#x000B0; E, 22&#x000B0; N) and were reported in Figures <xref ref-type="supplementary-material" rid="SM2">S2</xref>, <xref ref-type="supplementary-material" rid="SM3">S3</xref>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Population sampling after a range expansion simulation scenario (Western origin). <bold>(A,B)</bold> Homogeneous environment. Maps of the empirical distribution of singletons (true population sampling) and expected heterozygosity (true population sampling). <bold>(C,D)</bold> Inhomogeneous environment.</p></caption>
<graphic xlink:href="fgene-08-00139-g0004.tif"/>
</fig>
</sec>
<sec>
<title>4.3. Estimates of expansion onsets and application to pearl millet</title>
<p>First, we used the distribution of singletons in ABC to infer origins of range expansion in 100 simulated data sets (Figure <xref ref-type="fig" rid="F5">5</xref>). The results provided evidence of the usefulness of the statistics to identify origins of range expansions. Estimated values for the longitude and latitude of the onset of expansion were highly correlated to the true values for these parameters. Pearson&#x00027;s squared correlation coefficients were equal to <italic>R</italic><sup>2</sup> &#x0003D; 0.950 for the longitude and <italic>R</italic><sup>2</sup> &#x0003D; 0.948 for the latitude (<italic>p</italic>-values &#x0003C; 0.01).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Estimated coordinates of origin against their true values for 100 simulated data sets used as targets for ABC analysis. Pearson&#x00027;s correlation coefficients are reported.</p></caption>
<graphic xlink:href="fgene-08-00139-g0005.tif"/>
</fig>
<p>Next, we used the ABC approach to provide insights on the origin of range expansion of cultivated pearl millet in Africa. A total number of 41,032 singletons were found for 146 individuals, representing 24.27% of all variants. The posterior density for the longitude exhibited a mode around &#x02212;7.52&#x000B0;E (CI:-11.26&#x000B0;E, 0.84&#x000B0;E) (Figure <xref ref-type="fig" rid="F6">6</xref>). For the latitude of origin, the posterior density exhibited a mode around 24.2&#x000B0;N and a large credible interval (CI: 11.03&#x000B0;N, 29.06&#x000B0;N) (Figure <xref ref-type="fig" rid="F6">6</xref>). The most probable location for the origin of expansion of pearl millet in Africa was found near the Mali-Mauritania border (Figure <xref ref-type="fig" rid="F7">7</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Prior and posterior density estimates for the longitude and latitude of the expansion onset for cultivated pearl millet in Africa.</p></caption>
<graphic xlink:href="fgene-08-00139-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Geographic origin of cultivated pearl millet expansion using kernel density estimation.</p></caption>
<graphic xlink:href="fgene-08-00139-g0007.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>How singletons are distributed across geographic space provides a local measure of genetic diversity that can be measured at the individual level. In this study, we developed a theoretical background for the empirical distribution of singletons in a sample of chromosomes. We used simulations to provide evidence that the empirical distribution of singletons measures individual contributions to genetic diversity in the sample. The main advantage of this approach is to provide individual-based (local) estimates of genetic diversity that do not require the definition of populations.</p>
<p>Incorporated in an ABC framework, the empirical distribution of singletons led to accurate estimates of the geographic origin of range expansions in simulations. In ABC, the distribution of singletons was estimated by histograms obtained from clustering algorithms, and the histograms were used as summary statistics for Bayesian inference. Those statistics are appropriate to analyze the results of sequencing projects based on large scale sampling of individuals across geographic space. The method can be viewed as an interesting alternative to phylogenetic approaches when genomic sequences are used.</p>
<p>Potential factors that could bias our estimates of local genetic diversity includes missing data, genotyping errors, related individuals, and the use of a folded site frequency spectrum. Missing values or genotyping errors impacts individual data regardless of geography. By sharing genomic variation locally, related individuals reduce the number of unique variants drastically, and generate bias in global estimates of genetic diversity. Though those errors increase uncertainty in estimates, the biases on geographic estimates remain at small levels. Our ABC analysis took the potential biases into account by simulating the missing data, genotyping errors and the other issues. Alternative methods that could remove the biases would be based on genotype imputation and on the availability of genomic data from a closely related species.</p>
<p>We provided an illustration of the potential of singletons to inform demographic history by studying range expansion of pearl millet in Africa. Pearl millet is a widely grown staple crop in Africa and India, but its precise origin is currently unknown (Tostain, <xref ref-type="bibr" rid="B39">1992</xref>; Oumar et al., <xref ref-type="bibr" rid="B31">2008</xref>; Clotault et al., <xref ref-type="bibr" rid="B8">2012</xref>). When we applied an ABC approach to cultivated pearl millet genomes, we obtained a result supporting the Northern Mali region as the most probable geographic origin of expansion. Although the accuracy of the ABC approach was validated with extensive computer simulations of range expansion, the empirical results pointed out some limitations of our model for the data. The uncertainty around 18&#x000B0; reported for the latitude of origin was high, and improving our estimate would require supplementary information on past environmental conditions, carrying capacities and gene flow between pearl millet and related species. Interestingly, our results rejected an eastern origin for the expansion of the domesticated cereal. This result is consistent with recent archeological studies using both wild and cultivated samples, that pinpointed the Mali-Niger region as the most likely origin of domestication of pearl millet (Manning et al., <xref ref-type="bibr" rid="B24">2011</xref>; Ozainne et al., <xref ref-type="bibr" rid="B30">2014</xref>).</p>
<p>To conclude, singletons are a major component of the site frequency spectrum for many model and non-model species. The density of singletons in genomes has recently proven useful to detect selection in human genomes (Field et al., <xref ref-type="bibr" rid="B14">2016</xref>). Here we showed that the density of singletons in geographic space is useful for providing local estimates of genetic diversity and key insights on the demographic history of a species.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>All authors listed, have made substantial, direct and intellectual contribution to the work, and approved it for publication.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>This work has been partially supported by the Agence Nationale de la Recherche, project AFRICROP, ANR-13-BSV7-0017, and by the LabEx PERSYVAL Lab, ANR-11-LABX-0025-01, funded by the French program Investissement d&#x00027;Avenir.</p>
</ack>
<sec sec-type="supplementary-material" id="s7">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fgene.2017.00139/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fgene.2017.00139/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Image1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S1</label>
<caption><p>Averaged proportion of singletons in population 1, and standard deviations in populations 1 and 2, as functions of the shrink rate.</p></caption></supplementary-material>
<supplementary-material xlink:href="Image2.PNG" id="SM2" mimetype="image/png" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S2</label>
<caption><p>Individual vs. population sampling after a range expansion simulation scenario (Sahel origin). <bold>(A,B)</bold> Homogeneous environment. Maps of the empirical distribution of singletons (individual sampling) and expected heterozygosity (true population sampling). <bold>(C,D)</bold> Inhomogeneous environment.</p></caption></supplementary-material>
<supplementary-material xlink:href="Image3.PNG" id="SM3" mimetype="image/png" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S3</label>
<caption><p>Population sampling after a range expansion simulation scenario (Sahel origin). <bold>(A,B)</bold> Homogeneous environment. Maps of the empirical distribution of singletons (true population sampling) and expected heterozygosity (true population sampling). <bold>(C,D)</bold> Inhomogeneous environment.</p></caption></supplementary-material>
<supplementary-material xlink:href="DataSheet1.XLSX" id="SM4" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table1.DOCX" id="SM5" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table2.DOCX" id="SM6" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table3.DOCX" id="SM7" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><collab>1000 Genomes Project Consortium</collab> <name><surname>Auton</surname> <given-names>A.</given-names></name> <name><surname>Brooks</surname> <given-names>L. D.</given-names></name> <name><surname>Durbin</surname> <given-names>R. M.</given-names></name> <name><surname>Garrison</surname> <given-names>E. P.</given-names></name> <name><surname>Kang</surname> <given-names>H. M.</given-names></name> <name><surname>Korbel</surname> <given-names>J. O.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>A global reference for human genetic variation</article-title>. <source>Nature</source> <volume>526</volume>, <fpage>68</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1038/nature15393</pub-id><pub-id pub-id-type="pmid">26432245</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><collab>1000 Genomes Project Consortium</collab> <name><surname>Abecasis</surname> <given-names>G. R.</given-names></name> <name><surname>Altshuler</surname> <given-names>D.</given-names></name> <name><surname>Auton</surname> <given-names>A.</given-names></name> <name><surname>Brooks</surname> <given-names>L. D.</given-names></name> <name><surname>Durbin</surname> <given-names>R. M.</given-names></name> <name><surname>Gibbs</surname> <given-names>R. A.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>A map of human genome variation from population-scale sequencing</article-title>. <source>Nature</source> <volume>467</volume>, <fpage>1061</fpage>&#x02013;<lpage>1073</lpage>. <pub-id pub-id-type="doi">10.1038/nature09534</pub-id><pub-id pub-id-type="pmid">20981092</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Auer</surname> <given-names>P. L.</given-names></name> <name><surname>Lettre</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Rare variant association studies: considerations, challenges and opportunities</article-title>. <source>Genome Med.</source> <volume>7</volume>:<fpage>16</fpage>. <pub-id pub-id-type="doi">10.1186/s13073-015-0138-2</pub-id><pub-id pub-id-type="pmid">25709717</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beaumont</surname> <given-names>M. A.</given-names></name></person-group> (<year>2010</year>). <article-title>Approximate Bayesian computation in evolution and ecology</article-title>. <source>Ann. Rev. Ecol. Evol. Syst.</source> <volume>41</volume>, <fpage>379</fpage>&#x02013;<lpage>406</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-ecolsys-102209-144621</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blum</surname> <given-names>M. G. B.</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>O.</given-names></name></person-group> (<year>2005</year>). <article-title>Minimal clade size and external branch length under the neutral coalescent</article-title>. <source>Adv. Appl. Probabil.</source> <volume>37</volume>, <fpage>647</fpage>&#x02013;<lpage>662</lpage>. <pub-id pub-id-type="doi">10.1017/S0001867800000409</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blum</surname> <given-names>M. G. B.</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>O.</given-names></name></person-group> (<year>2010</year>). <article-title>Non-linear regression models for approximate Bayesian computation</article-title>. <source>Stat. Comput.</source> <volume>20</volume>, <fpage>63</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-009-9116-0</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caliebe</surname> <given-names>A.</given-names></name> <name><surname>Neininger</surname> <given-names>R.</given-names></name> <name><surname>Krawczak</surname> <given-names>M.</given-names></name> <name><surname>R&#x000F6;sler</surname> <given-names>U.</given-names></name></person-group> (<year>2007</year>). <article-title>On the length distribution of external branches in coalescence trees: genetic diversity within species</article-title>. <source>Theor. Popul. Biol.</source> <volume>72</volume>, <fpage>245</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1016/j.tpb.2007.05.003</pub-id><pub-id pub-id-type="pmid">17643459</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clotault</surname> <given-names>J.</given-names></name> <name><surname>Thuillet</surname> <given-names>A. C.</given-names></name> <name><surname>Buiron</surname> <given-names>M.</given-names></name> <name><surname>De Mita</surname> <given-names>S.</given-names></name> <name><surname>Couderc</surname> <given-names>M.</given-names></name> <name><surname>Haussmann</surname> <given-names>B. I.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Evolutionary history of pearl millet (<italic>Pennisetum glaucum</italic> [L.] R. Br.) and selection on flowering genes since its domestication</article-title>. <source>Mol. Biol. Evol.</source> <volume>29</volume>, <fpage>1199</fpage>&#x02013;<lpage>1212</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msr287</pub-id><pub-id pub-id-type="pmid">22114357</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coventry</surname> <given-names>A.</given-names></name> <name><surname>Bull-Otterson</surname> <given-names>L. M.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Clark</surname> <given-names>A. G.</given-names></name> <name><surname>Maxwell</surname> <given-names>T. J.</given-names></name> <name><surname>Crosby</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Deep resequencing reveals excess rare recent variants consistent with explosive population growth</article-title>. <source>Nat. Commun.</source> <volume>1</volume>:<fpage>131</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms1130</pub-id><pub-id pub-id-type="pmid">21119644</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cressie</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <source>Statistics for Spatial Data</source>. <publisher-loc>New-York, NY</publisher-loc>: <publisher-name>John Wiley and Sons</publisher-name>.</citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Csill&#x000E9;ry</surname> <given-names>K.</given-names></name> <name><surname>Blum</surname> <given-names>M. G.</given-names></name> <name><surname>Gaggiotti</surname> <given-names>O. E.</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>O.</given-names></name></person-group> (<year>2010</year>). <article-title>Approximate Bayesian computation (ABC) in practice</article-title>. <source>Trends Ecol. Evol.</source> <volume>25</volume>, <fpage>410</fpage>&#x02013;<lpage>418</lpage>. <pub-id pub-id-type="doi">10.1016/j.tree.2010.04.001</pub-id><pub-id pub-id-type="pmid">20488578</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Csill&#x000E9;ry</surname> <given-names>K.</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>O.</given-names></name> <name><surname>Blum</surname> <given-names>M. G. B.</given-names></name></person-group> (<year>2012</year>). <article-title>abc: an R package for approximate Bayesian computation (ABC)</article-title>. <source>Methods Ecol. Evol.</source> <volume>3</volume>, <fpage>475</fpage>&#x02013;<lpage>479</lpage>. <pub-id pub-id-type="doi">10.1111/j.2041-210X.2011.00179.x</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Currat</surname> <given-names>M.</given-names></name> <name><surname>Ray</surname> <given-names>N.</given-names></name> <name><surname>Excoffier</surname> <given-names>L.</given-names></name></person-group> (<year>2004</year>). <article-title>SPLATCHE: a program to simulate genetic diversity taking into account environmental heterogeneity</article-title>. <source>Mol. Ecol. Notes</source> <volume>4</volume>, <fpage>139</fpage>&#x02013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1046/j.1471-8286.2003.00582.x</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Field</surname> <given-names>Y.</given-names></name> <name><surname>Boyle</surname> <given-names>E. A.</given-names></name> <name><surname>Telis</surname> <given-names>N.</given-names></name> <name><surname>Gao</surname> <given-names>Z.</given-names></name> <name><surname>Gaulton</surname> <given-names>K. J.</given-names></name> <name><surname>Golan</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Detection of human adaptation during the past 2000 years</article-title>. <source>Science</source> <volume>354</volume>, <fpage>760</fpage>&#x02013;<lpage>764</lpage>. <pub-id pub-id-type="doi">10.1126/science.aag0776</pub-id><pub-id pub-id-type="pmid">27738015</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frichot</surname> <given-names>E.</given-names></name> <name><surname>Fran&#x000E7;ois</surname> <given-names>O.</given-names></name></person-group> (<year>2015</year>). <article-title>LEA: an R package for landscape and ecological association studies</article-title>. <source>Methods Ecol. Evol.</source> <volume>6</volume>, <fpage>925</fpage>&#x02013;<lpage>929</lpage>. <pub-id pub-id-type="doi">10.1111/2041-210X.12382</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>Y. X.</given-names></name> <name><surname>Li</surname> <given-names>W. H.</given-names></name></person-group> (<year>1993</year>). <article-title>Statistical tests of neutrality of mutations</article-title>. <source>Genetics</source> <volume>133</volume>, <fpage>693</fpage>&#x02013;<lpage>709</lpage>. <pub-id pub-id-type="pmid">8454210</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gravel</surname> <given-names>S.</given-names></name> <name><surname>Henn</surname> <given-names>B. M.</given-names></name> <name><surname>Gutenkunst</surname> <given-names>R. N.</given-names></name> <name><surname>Indap</surname> <given-names>A. R.</given-names></name> <name><surname>Marth</surname> <given-names>G. T.</given-names></name> <name><surname>Clark</surname> <given-names>A. G.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Demographic history and rare allele sharing among human populations</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>108</volume>, <fpage>11983</fpage>&#x02013;<lpage>11988</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1019276108</pub-id><pub-id pub-id-type="pmid">21730125</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hartigan</surname> <given-names>J. A.</given-names></name> <name><surname>Wong</surname> <given-names>M. A.</given-names></name></person-group> (<year>1979</year>). <article-title>A <italic>K</italic>-means clustering algorithm</article-title>. <source>Appl. Stat.</source> <volume>28</volume>, <fpage>100</fpage>&#x02013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.2307/2346830</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hudson</surname> <given-names>R. R.</given-names></name></person-group> (<year>1990</year>). <article-title>Gene genealogies and the coalescent process</article-title>. <source>Oxford Surv. Evol. Biol.</source> <volume>7</volume>:<fpage>44</fpage>.</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hudson</surname> <given-names>R. R.</given-names></name></person-group> (<year>2002</year>). <article-title>Generating samples under a Wright-Fisher neutral model of genetic variation</article-title>. <source>Bioinformatics</source> <volume>18</volume>, <fpage>337</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/18.2.337</pub-id><pub-id pub-id-type="pmid">11847089</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><collab>International HapMap 3 Consortium</collab></person-group> (<year>2010</year>). <article-title>Integrating common and rare genetic variation in diverse human populations</article-title>. <source>Nature</source> <volume>467</volume>, <fpage>52</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1038/nature09298</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keinan</surname> <given-names>A.</given-names></name> <name><surname>Clark</surname> <given-names>A. G.</given-names></name></person-group> (<year>2012</year>). <article-title>Recent explosive human population growth has resulted in an excess of rare genetic variants</article-title>. <source>Science</source> <volume>336</volume>, <fpage>740</fpage>&#x02013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1126/science.1217283</pub-id><pub-id pub-id-type="pmid">22582263</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Abecasis</surname> <given-names>G. R.</given-names></name> <name><surname>Boehnke</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>X.</given-names></name></person-group> (<year>2014</year>). <article-title>Rare-variant association analysis: study designs and statistical tests</article-title>. <source>Am. J. Hum. Genet.</source> <volume>95</volume>, <fpage>5</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2014.06.009</pub-id><pub-id pub-id-type="pmid">24995866</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Manning</surname> <given-names>K.</given-names></name> <name><surname>Pelling</surname> <given-names>R.</given-names></name> <name><surname>Higham</surname> <given-names>T.</given-names></name> <name><surname>Schwenniger</surname> <given-names>J. L.</given-names></name> <name><surname>Fuller</surname> <given-names>D. Q.</given-names></name></person-group> (<year>2011</year>). <article-title>4500-year old domesticated pearl millet (<italic>Pennisetum glaucum</italic>) from the Tilemsi Valley, Mali: new insights into an alternative cereal domestication pathway</article-title>. <source>J. Archaeol. Sci.</source> <volume>38</volume>, <fpage>312</fpage>&#x02013;<lpage>322</lpage>. <pub-id pub-id-type="doi">10.1016/j.jas.2010.09.007</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marth</surname> <given-names>G. T.</given-names></name> <name><surname>Czabarka</surname> <given-names>E.</given-names></name> <name><surname>Murvai</surname> <given-names>J.</given-names></name> <name><surname>Sherry</surname> <given-names>S. T.</given-names></name></person-group> (<year>2004</year>). <article-title>The allele frequency spectrum in genome-wide human variation data reveals signals of differential demographic history in three large world populations</article-title>. <source>Genetics</source> <volume>166</volume>, <fpage>351</fpage>&#x02013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.166.1.351</pub-id><pub-id pub-id-type="pmid">15020430</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mathieson</surname> <given-names>I.</given-names></name> <name><surname>McVean</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>Demography and the age of rare variants</article-title>. <source>PLoS Genet.</source> <volume>10</volume>:<fpage>e1004528</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1004528</pub-id><pub-id pub-id-type="pmid">25101869</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Memon</surname> <given-names>S.</given-names></name> <name><surname>Jia</surname> <given-names>X.</given-names></name> <name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>Genomic variations and distinct evolutionary rate of rare alleles in <italic>Arabidopsis thaliana</italic></article-title>. <source>BMC Evol. Biol.</source> <volume>16</volume>:<fpage>25</fpage>. <pub-id pub-id-type="doi">10.1186/s12862-016-0590-7</pub-id><pub-id pub-id-type="pmid">26817829</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Novembre</surname> <given-names>J.</given-names></name> <name><surname>Slatkin</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Likelihood-based inference in isolation-by-distance models using the spatial distribution of low-frequency alleles</article-title>. <source>Evolution</source> <volume>63</volume>:<fpage>2914</fpage>. <pub-id pub-id-type="doi">10.1111/j.1558-5646.2009.00775.x</pub-id><pub-id pub-id-type="pmid">19624728</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Connor</surname> <given-names>T. D.</given-names></name> <name><surname>Fu</surname> <given-names>W.</given-names></name> <name><surname>Turner</surname> <given-names>E.</given-names></name> <name><surname>Mychaleckyj</surname> <given-names>J. C.</given-names></name> <name><surname>Logsdon</surname> <given-names>B.</given-names></name> <name><surname>Auer</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Rare variation facilitates inferences of fine-scale population structure in humans</article-title>. <source>Mol. Biol. Evol.</source> <volume>32</volume>, <fpage>653</fpage>&#x02013;<lpage>660</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msu326</pub-id><pub-id pub-id-type="pmid">25415970</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ozainne</surname> <given-names>S.</given-names></name> <name><surname>Lespez</surname> <given-names>L.</given-names></name> <name><surname>Garnier</surname> <given-names>A.</given-names></name> <name><surname>Ballouche</surname> <given-names>A.</given-names></name> <name><surname>Neumann</surname> <given-names>K.</given-names></name> <name><surname>Pays</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>A question of timing: spatio-temporal structure and mechanisms of early agriculture expansion in West Africa</article-title>. <source>J. Archaeol. Sci.</source> <volume>50</volume>, <fpage>359</fpage>&#x02013;<lpage>368</lpage>. <pub-id pub-id-type="doi">10.1016/j.jas.2014.07.025</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oumar</surname> <given-names>I.</given-names></name> <name><surname>Mariac</surname> <given-names>C.</given-names></name> <name><surname>Pham</surname> <given-names>J.-L.</given-names></name> <name><surname>Vigouroux</surname> <given-names>Y.</given-names></name></person-group> (<year>2008</year>). <article-title>Phylogeny and origin of pearl millet (<italic>Pennisetum glaucum</italic> [L.] R. Br) as revealed by microsatellite loci</article-title>. <source>Theor. Appl. Genet.</source> <volume>117</volume>, <fpage>489</fpage>&#x02013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1007/s00122-008-0793-4</pub-id><pub-id pub-id-type="pmid">18504539</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paradis</surname> <given-names>E.</given-names></name> <name><surname>Claude</surname> <given-names>J.</given-names></name> <name><surname>Strimmer</surname> <given-names>K.</given-names></name></person-group> (<year>2004</year>). <article-title>APE: analyses of phylogenetics and evolution in R language</article-title>. <source>Bioinformatics</source> <volume>20</volume>, <fpage>289</fpage>&#x02013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btg412</pub-id><pub-id pub-id-type="pmid">14734327</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pritchard</surname> <given-names>J. K.</given-names></name></person-group> (<year>2001</year>). <article-title>Are rare variants responsible for susceptibility to complex diseases?</article-title> <source>Am. J. Hum. Genet.</source> <volume>69</volume>, <fpage>124</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1086/321272</pub-id><pub-id pub-id-type="pmid">11404818</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schork</surname> <given-names>N. J.</given-names></name> <name><surname>Murray</surname> <given-names>S. S.</given-names></name> <name><surname>Frazer</surname> <given-names>K. A.</given-names></name> <name><surname>Topol</surname> <given-names>E. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Common vs. rare allele hypotheses for complex diseases</article-title>. <source>Curr. Opin. Genet. Dev.</source> <volume>19</volume>, <fpage>212</fpage>&#x02013;<lpage>219</lpage>. <pub-id pub-id-type="doi">10.1016/j.gde.2009.04.010</pub-id><pub-id pub-id-type="pmid">19481926</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schraiber</surname> <given-names>J. G.</given-names></name> <name><surname>Akey</surname> <given-names>J. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Methods and models for unravelling human evolutionary history</article-title>. <source>Nat. Rev. Genet.</source> <volume>16</volume>, <fpage>727</fpage>&#x02013;<lpage>740</lpage>. <pub-id pub-id-type="doi">10.1038/nrg4005</pub-id><pub-id pub-id-type="pmid">26553329</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Slatkin</surname> <given-names>M.</given-names></name></person-group> (<year>1985</year>). <article-title>Rare alleles as indicators of gene flow</article-title>. <source>Evolution</source> <volume>39</volume>, <fpage>53</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1111/j.1558-5646.1985.tb04079.x</pub-id><pub-id pub-id-type="pmid">28563643</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tavar&#x000E9;</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>Ancestral inference in population genetics</article-title>, in <source>Lectures on Probability Theory and Statistics</source>, Lecture Notes Math. 1837, ed <person-group person-group-type="editor"><name><surname>Picard</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>188</lpage>.</citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tennessen</surname> <given-names>J.</given-names></name> <name><surname>Bigham</surname> <given-names>A.</given-names></name> <name><surname>O&#x00027;Connor</surname> <given-names>T.</given-names></name> <name><surname>Fu</surname> <given-names>W.</given-names></name> <name><surname>Kenny</surname> <given-names>E. E.</given-names></name> <name><surname>Gravel</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Evolution and functional impact of rare coding variation from deep sequencing of human exomes</article-title>. <source>Science</source> <volume>337</volume>, <fpage>64</fpage>&#x02013;<lpage>69</lpage>. <pub-id pub-id-type="doi">10.1126/science.1219240</pub-id><pub-id pub-id-type="pmid">22604720</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tostain</surname> <given-names>S.</given-names></name></person-group> (<year>1992</year>). <article-title>Enzyme diversity in pearl millet (<italic>Pennisetum glaucum</italic> L.)</article-title>. <source>Theor. Appl. Genet.</source> <volume>83</volume>, <fpage>733</fpage>&#x02013;<lpage>742</lpage>. <pub-id pub-id-type="doi">10.1007/BF00226692</pub-id><pub-id pub-id-type="pmid">24202748</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varshney</surname> <given-names>R. K.</given-names></name> <name><surname>Shi</surname> <given-names>C.</given-names></name> <name><surname>Thudi</surname> <given-names>M.</given-names></name> <name><surname>Mariac</surname> <given-names>C.</given-names></name> <name><surname>Wallace</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Pearl millet genome sequence provides a resource to improve agronomic traits in arid environments</article-title>. <source>Nat. Biotechnol.</source> [Epub ahead of print]. Available online at: <ext-link ext-link-type="uri" xlink:href="http://ceg.icrisat.org/ipmgsc/">http://ceg.icrisat.org/ipmgsc/</ext-link> <pub-id pub-id-type="doi">10.1038/nbt.3943</pub-id><pub-id pub-id-type="pmid">28922347</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weigel</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Natural variation in Arabidopsis: from molecular genetics to ecological genomics</article-title>. <source>Plant Physiol.</source> <volume>158</volume>, <fpage>2</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1104/pp.111.189845</pub-id><pub-id pub-id-type="pmid">22147517</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weigel</surname> <given-names>D.</given-names></name> <name><surname>Mott</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>The 1001 genomes project for <italic>Arabidopsis thaliana</italic></article-title>. <source>Genome Biol.</source> <volume>10</volume>:<fpage>107</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2009-10-5-107</pub-id><pub-id pub-id-type="pmid">19519932</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Integrating rare-variant testing, function prediction, and gene network in composite resequencing-based genome-wide association studies (CR-GWAS)</article-title>. <source>Genes Genomes Genet.</source> <volume>1</volume>, <fpage>233</fpage>&#x02013;<lpage>243</lpage>. <pub-id pub-id-type="doi">10.1534/g3.111.000364</pub-id><pub-id pub-id-type="pmid">22384334</pub-id></citation></ref>
</ref-list>
</back>
</article>