<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2018.00570</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>SIPSim: A Modeling Toolkit to Predict Accuracy and Aid Design of DNA-SIP Experiments</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Youngblut</surname> <given-names>Nicholas D.</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/63352/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Barnett</surname> <given-names>Samuel E.</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/456956/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Buckley</surname> <given-names>Daniel H.</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/55415/overview"/>
</contrib>
</contrib-group>
<aff><institution>School of Integrative Plant Science, Cornell University</institution>, <addr-line>Ithaca, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Eoin L. Brodie, Lawrence Berkeley National Laboratory (LBNL), United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Martin Taubert, Friedrich Schiller Universit&#x000E4;t Jena, Germany; Michael Pester, Deutsche Sammlung von Mikroorganismen und Zellkulturen (DSMZ), Germany; Ember Morrissey, West Virginia University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Daniel H. Buckley <email>dbuckley&#x00040;cornell.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Terrestrial Microbiology, a section of the journal Frontiers in Microbiology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>28</day>
<month>03</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>9</volume>
<elocation-id>570</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>09</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>03</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2018 Youngblut, Barnett and Buckley.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Youngblut, Barnett and Buckley</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>DNA Stable isotope probing (DNA-SIP) is a powerful method that links identity to function within microbial communities. The combination of DNA-SIP with multiplexed high throughput DNA sequencing enables simultaneous mapping of <italic>in situ</italic> assimilation dynamics for thousands of microbial taxonomic units. Hence, high throughput sequencing enabled SIP has enormous potential to reveal patterns of carbon and nitrogen exchange within microbial food webs. There are several different methods for analyzing DNA-SIP data and despite the power of SIP experiments, it remains difficult to comprehensively evaluate method accuracy across a wide range of experimental parameters. We have developed a toolset (SIPSim) that simulates DNA-SIP data, and we use this toolset to systematically evaluate different methods for analyzing DNA-SIP data. Specifically, we employ SIPSim to evaluate the effects that key experimental parameters (e.g., level of isotopic enrichment, number of labeled taxa, relative abundance of labeled taxa, community richness, community evenness, and beta-diversity) have on the specificity, sensitivity, and balanced accuracy (defined as the product of specificity and sensitivity) of DNA-SIP analyses. Furthermore, SIPSim can predict analytical accuracy and power as a function of experimental design and community characteristics, and thus should be of great use in the design and interpretation of DNA-SIP experiments.</p></abstract>
<kwd-group>
<kwd>DNA-SIP</kwd>
<kwd>SIP</kwd>
<kwd>method</kwd>
<kwd>microbial</kwd>
<kwd>community</kwd>
<kwd>function</kwd>
<kwd>SIPSim</kwd>
</kwd-group>
<contract-num rid="cn001">DE-SC0010558</contract-num>
<contract-num rid="cn001">DE-SC0004486</contract-num>
<contract-sponsor id="cn001">Department of Energy, Office of Biological &#x00026; Environmental Research Genomic Science Program</contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="0"/>
<equation-count count="8"/>
<ref-count count="36"/>
<page-count count="16"/>
<word-count count="10315"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Stable isotope probing of nucleic acids (DNA-SIP and RNA-SIP) is a powerful culture-independent method for linking microbial metabolic functioning to taxonomic identity (Radajewski et al., <xref ref-type="bibr" rid="B25">2003</xref>). In particular, DNA-SIP has been used in a multitude of environments to identify microbial assimilation of various <sup>13</sup>C- and <sup>15</sup>N-labeled substrates into DNA (Uhl&#x000ED;k et al., <xref ref-type="bibr" rid="B34">2009</xref>). DNA-SIP identifies microbes that assimilate these isotopes into their DNA (&#x0201C;incorporators&#x0201D;) by exploiting the increased buoyant density (BD) of isotopically labeled (&#x0201C;heavy&#x0201D;) DNA relative to unlabeled (&#x0201C;light&#x0201D;) DNA. For example, fully <sup>13</sup>C- and <sup>15</sup>N-labeled DNA will increase in BD by 0.036 and 0.016 g ml<sup>&#x02212;1</sup>, respectively (Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>).</p>
<p>Ideally, isopycnic centrifugation could be used to completely separate labeled and unlabeled DNA fragments based solely on this difference in BD. However, several factors besides BD can influence the position of DNA in isopycnic gradients. For example, G &#x0002B; C content variation within a single genome can produce unlabeled DNA fragments that vary in BD by up to 0.03 g ml<sup>&#x02212;1</sup>, while G &#x0002B; C content variation between microbial genomes can cause the average BD of unlabeled DNA fragments to vary by up to 0.05 g ml<sup>&#x02212;1</sup> (Youngblut and Buckley, <xref ref-type="bibr" rid="B36">2014</xref>). In addition, DNA in SIP experiments will often be partially labeled as a consequence of isotope dilution from unlabeled endogenous substrates. Therefore, it is unlikely that nucleic acid SIP experiments will ever achieve complete separation of labeled and unlabeled DNA.</p>
<p>In the absence of complete separation between labeled and unlabeled DNA, isotope incorporators must be identified using some statistical procedure suitable for comparing the BD distributions of DNA fragments from labeled treatment samples and unlabeled control samples (Pepe-Ranney et al., <xref ref-type="bibr" rid="B22">2016a</xref>). The use of multiplexed high throughput sequencing with DNA-SIP makes it possible to sequence SSU rRNA amplicons across many density gradient fractions and simultaneously determine the BD distributions for thousands of taxa. The problem then becomes one of identifying those taxa that have increased in BD in the isotopically labeled samples relative to the corresponding unlabeled controls.</p>
<p>Different analytical approaches have been applied to DNA-SIP datasets to identify DNA sequences of <sup>13</sup>C-labeled taxa. The earliest and simplest approach to identifying <sup>13</sup>C-labeled DNA sequences (described herein as Heavy-SIP) is to identify SSU rRNA amplicons that occur in &#x0201C;heavy&#x0201D; fractions of CsCl gradients containing <sup>13</sup>C-labeled DNA, but do not occur in either &#x0201C;light&#x0201D; fractions or in &#x0201C;heavy&#x0201D; fractions of unlabeled control gradients (Radajewski et al., <xref ref-type="bibr" rid="B25">2003</xref>; Lueders et al., <xref ref-type="bibr" rid="B17">2004</xref>). More recent approaches include &#x0201C;high resolution stable isotope probing&#x0201D; (HR-SIP) and &#x0201C;quantitative stable isotope probing&#x0201D; (qSIP), which both analyze SSU rRNA amplicons across numerous gradient fractions (Hungate et al., <xref ref-type="bibr" rid="B12">2015</xref>; Pepe-Ranney et al., <xref ref-type="bibr" rid="B22">2016a</xref>,<xref ref-type="bibr" rid="B23">b</xref>). All of these methods differ in the statistical procedures used to detect taxa that incorporate isotopic label. Heavy-SIP often employs either <italic>t</italic>-test, Fisher&#x00027;s exact test, or analogous approaches to compare OTU relative abundance between pairs of fractions (e.g., heavy vs. light). HR-SIP identifies isotopically labeled taxa by evaluating the sequence composition of several high density &#x0201C;heavy&#x0201D; fractions using a differential abundance quantification framework that evaluates sequence count data in isotopically labeled samples relative to their corresponding unlabeled controls. Differential abundance between the &#x0201C;heavy&#x0201D; fractions of labeled and control gradients is tested with DESeq2 (Love et al., <xref ref-type="bibr" rid="B16">2014</xref>), which uses sophisticated statistical methods to reduce technical error and increase analytical power for analysis of microbiome data (McMurdie and Holmes, <xref ref-type="bibr" rid="B18">2014</xref>). In qSIP, SSU rRNA relative abundance values are transformed using qPCR estimates of total SSU rRNA gene copies present within gradient fractions. These normalized data are used to determine the weighted average BD for each taxon in both isotopically labeled samples and corresponding unlabeled controls (Hungate et al., <xref ref-type="bibr" rid="B12">2015</xref>). Incorporators are then determined by using a permutation procedure using 90% confidence intervals to identify those taxa whose BD shifts are unlikely to occur as a result of chance.</p>
<p>While DNA-SIP is a powerful method for the discovery and characterization of microorganisms <italic>in situ</italic>, systematic assessment of the specificity or sensitivity of this method has not been performed. Empirical validations of DNA-SIP methods typically include only one or a few organisms or simple mock communities (Lueders et al., <xref ref-type="bibr" rid="B17">2004</xref>; Buckley et al., <xref ref-type="bibr" rid="B4">2007</xref>; Cupples et al., <xref ref-type="bibr" rid="B7">2007</xref>; Wawrik et al., <xref ref-type="bibr" rid="B35">2009</xref>; Andeer et al., <xref ref-type="bibr" rid="B1">2012</xref>), and such approaches do not adequately replicate the complexity of the DNA fragment BD distributions expected in a typical DNA-SIP experiment (Youngblut and Buckley, <xref ref-type="bibr" rid="B36">2014</xref>). DNA-SIP experiments vary in the diversity of the target community, DNA G &#x0002B; C content distribution, the number of incorporators, incorporator relative abundance, and the atom % excess of labeled DNA. Systematic evaluation of method accuracy should address the effects that all of these variables have on the sensitivity and specificity of detecting isotope incorporators. Since DNA-SIP experiments are costly, technically difficult, and laborious, it is not practical to perform empirical assessment across this full range of variables.</p>
<p>Fortunately, the physics of isopycnic centrifugation have been well characterized mathematically, and the behavior of individual DNA fragments in CsCl gradients is highly reproducible and predictable from first principles (Meselson et al., <xref ref-type="bibr" rid="B19">1957</xref>; Fritsch, <xref ref-type="bibr" rid="B10">1975</xref>; Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>). In addition, genome sequences are available for thousands of diverse microorganisms, and these genomes can be used to simulate DNA fragments representative of community DNA (Youngblut and Buckley, <xref ref-type="bibr" rid="B36">2014</xref>). Hence, we can simulate realistic DNA-SIP data for <italic>in silico</italic> microbial communities that differ in diversity (richness, evenness, and composition), where the relative abundance, genome G &#x0002B; C content, and atom % isotope enrichment are defined for discrete DNA fragments from every genome. We have developed a computational toolset for simulating DNA-SIP data (SIPSim) and used this simulation framework to systematically and objectively evaluate how changes in key SIP experimental parameters are predicted to affect DNA-SIP accuracy.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Theory underlying the simulation framework</title>
<p>DNA stable isotope probing employs isopycnic centrifugation to separate isotopically enriched (&#x0201C;heavy&#x0201D;) DNA molecules from unlabeled (&#x0201C;light&#x0201D;) DNA based on their differences in buoyant density (BD). Isopycnic centrifugation is distinguished from other centrifugation methods in that centrifugation is carried out long enough to both generate a density gradient (typically using CsCl for DNA-SIP) and allow all macromolecules of interest reach sedimentation equilibrium, which is the point at which sedimentation rates equal rates of diffusion (Hearst and Schmid, <xref ref-type="bibr" rid="B11">1973</xref>; Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>). Empirical studies have shown that the average BD (&#x003C1;) of a mixture of DNA molecules is linearly related to the average G &#x0002B; C content for that collection of molecules:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mi>&#x003C1;</mml:mi><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>0.098</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo stretchy='false'>[</mml:mo><mml:mi>G</mml:mi><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>]</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1.66</mml:mn></mml:mrow></mml:math></disp-formula>
<p>where [G &#x0002B; C] is the mole fraction of genome G &#x0002B; C content (Schildkraut et al., <xref ref-type="bibr" rid="B28">1962</xref>; Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>). In addition, empirical studies have also shown that homogeneous mixtures of DNA molecules form a Gaussian distribution in an isopycnic gradient when at sedimentation equilibrium (Meselson et al., <xref ref-type="bibr" rid="B19">1957</xref>; Fritsch, <xref ref-type="bibr" rid="B10">1975</xref>). Therefore, in order to model the BD distribution of a heterogeneous set of genomic DNA fragments, a Gaussian distribution must be estimated for each homogeneous subset of molecules rather than using discrete BD values (as described in Supplementary Material). Based on the work of Meselson et al. (<xref ref-type="bibr" rid="B19">1957</xref>), Fritsch (<xref ref-type="bibr" rid="B10">1975</xref>) derived an equation describing time to reach sedimentation equilibrium, which can be reworked to calculate the standard deviation (&#x003C3;) of the Gaussian distribution:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:mi>&#x003C3;</mml:mi><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mi>L</mml:mi><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1.26</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<disp-formula id="E3"><label>(2.1)</label><mml:math id="M3"><mml:mrow><mml:mi>&#x003B3;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:msup><mml:mi>&#x003C9;</mml:mi><mml:mn>4</mml:mn></mml:msup><mml:msubsup><mml:mi>r</mml:mi><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B2;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where <italic>L</italic> is the effective length of the gradient (cm), <italic>t</italic> is time in seconds, &#x003C9; is the angular velocity (radians sec<sup>&#x02212;1</sup>), <italic>r</italic><sub><italic>p</italic></sub> is the distance of the particle from the axis of rotation (cm), <italic>s</italic> is the sedimentation coefficient of the particle, &#x003B2;&#x000B0; is the coefficient specific to the density gradient medium (e.g., CsCl); <italic>p</italic><sub><italic>p</italic></sub> and <italic>p</italic><sub><italic>m</italic></sub> are the maximum and minimum distances between the gradient and axis of rotation (cm) (Fritsch, <xref ref-type="bibr" rid="B10">1975</xref>). By assuming that sedimentation equilibrium has been reached for all macromolecules of interest, Clay and colleagues derived a simplified equation for determining &#x003C3; from the calculations in Schmid and Hearst (<xref ref-type="bibr" rid="B29">1972</xref>):</p>
<disp-formula id="E4"><label>(3)</label><mml:math id="M4"><mml:mrow><mml:mi>&#x003C3;</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x003C1;</mml:mi><mml:mi>R</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>G</mml:mi><mml:msub><mml:mi>M</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:mrow></mml:math></disp-formula>
<p>where &#x003C1; is the BD of the particle, <italic>R</italic> is the universal gas constant, <italic>T</italic> is the temperature in Kelvins, &#x003B2; is a proportionality constant for aqueous salts of specific densities, <italic>G</italic> is a buoyancy factor as described in Clay et al. (<xref ref-type="bibr" rid="B5">2003</xref>), <italic>M</italic><sub><italic>C</italic></sub> is the molecular weight per base pair of DNA, and <italic>l</italic> is the fragment length (bp). For most DNA-SIP experiments, the assumption of sedimentation equilibrium for all DNA fragments is likely to be unrealistic for relatively short DNA fragments (e.g., &#x0003C;4 kb), given that the time to reach equilibrium is inversely proportional to diffusion and hence rises dramatically with decreasing fragment length (Meselson et al., <xref ref-type="bibr" rid="B19">1957</xref>; Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>; Youngblut and Buckley, <xref ref-type="bibr" rid="B36">2014</xref>). However, the ultracentrifugation durations used in typical DNA-SIP experiments should still generally produce small &#x003C3; values for short DNA fragments according to Equation (2) (Neufeld et al., <xref ref-type="bibr" rid="B20">2007</xref>). Therefore, Equation (3) provides a good approximation for modeling the BD distribution of DNA in density gradients generated in typical DNA-SIP experiments.</p>
<p>The distribution of a heterogeneous mixture of DNA fragments in an isopycnic gradient can thus be modeled by integrating the Gaussian distributions of each homogeneous subset of DNA fragments, where the mean of each Gaussian is determined by Equation (1) and the standard deviation derived from Equation (3). In this way, the BD distribution for a given genome in an isopycnic gradient can be modeled by the following steps: simulate genome fragmentation resulting from DNA extraction, bin gDNA fragments with respect to length and G &#x0002B; C content, model Gaussian distribution for each fragment bin, and then integrate these distributions to describe the cumulative DNA distribution in fractions of the gradient. A strictly Gaussian model predicts that, within a CsCl density gradient, DNA fragments should be undetectable (i.e., probability density &#x0003C; 10<sup>&#x02212;7</sup>) in factions of both high and low BD (Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>). However, empirical observations show DNA to occur throughout CsCl gradients (Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>; Lueders et al., <xref ref-type="bibr" rid="B17">2004</xref>; Leigh et al., <xref ref-type="bibr" rid="B14">2007</xref>). We were able to reconcile the difference between observed and expected DNA distributions as a function of fluid mechanics during gradient reorientation (as described below and in Supplementary Material).</p>
<p>Gradient reorientation impacts the BD distribution of DNA. During isopycnic centrifugation, the buoyant density gradient forms perpendicular to the axis of rotation (Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>), and gradient reorientation during centrifuge deceleration is dramatic, especially for vertical rotors (Flamm et al., <xref ref-type="bibr" rid="B9">1966</xref>). While the distortion of the BD gradient during reorientation has been shown to be minimal in the aggregate (Fisher et al., <xref ref-type="bibr" rid="B8">1964</xref>; Flamm et al., <xref ref-type="bibr" rid="B9">1966</xref>), the inevitable presence of a diffusive boundary layer along the tube wall is sufficient to entrain quantities of DNA, which are small but should be readily detectable by high throughput sequencing methods. The flow field that occurs during gradient reorientation entrains along the tube wall a volume with a dimension proportional to flow velocity, fluid viscosity, and surface topography (Tritton, <xref ref-type="bibr" rid="B33">1977</xref>; Cohen and Dowling, <xref ref-type="bibr" rid="B6">2012</xref>). Following gradient reorientation, DNA from the entrained volume will combine with DNA from the reoriented volume, thereby introducing a small amount of non-BD-equilibrium DNA into each gradient fraction (Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>). The ability of the diffusive boundary to introduce non-BD-equilibrium DNA into gradient fractions can be modeled as a function of rotor geometry and the effect is more pronounced in vertical rotors relative to fixed angle rotors (Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>). Assuming sedimentation equilibrium, BD (&#x003C1;) can be directly related to the distance from axis of rotation (Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>):</p>
<disp-formula id="E5"><label>(4)</label><mml:math id="M5"><mml:mrow><mml:mi>x</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msqrt><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>p</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x003B2;</mml:mi><mml:mo>&#x000B0;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x003C9;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:msubsup><mml:mi>r</mml:mi><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msqrt></mml:mrow></mml:math></disp-formula>
<p>From this calculation, the location of DNA molecules in the centrifuge tube, both during centrifugation and fractionation, can be ascertained by using simple trigonometry along with knowledge of centrifuge tube dimensions and angle to the axis of rotation. A full description of the calculations along with an example can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/nick-youngblut/SIPSim">https://github.com/nick-youngblut/SIPSim</ext-link>. The fraction of a taxon&#x00027;s DNA fragments that are in the boundary layer (<italic>D</italic><sub><italic>ti</italic></sub>) is modeled as:</p>
<disp-formula id="E6"><label>(5)</label><mml:math id="M6"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003B3;</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:math></disp-formula>
<p>where <italic>A</italic><sub><italic>ti</italic></sub> is the pre-fractionation community relative abundance of taxon <italic>t</italic> in gradient <italic>i</italic>, &#x003B3; is a weight parameter determining the contribution of <italic>A</italic><sub><italic>ti</italic></sub> to <italic>A</italic><sub><italic>b</italic></sub>, and &#x003B1; is the baseline fraction DNA in <italic>A</italic><sub><italic>b</italic></sub>.</p>
<p>Assimilation of the commonly used isotopes <sup>13</sup>C and <sup>15</sup>N into genomic DNA produces linear shifts in BD, with a maximum shift of 0.036 and 0.016 g ml<sup>&#x02212;1</sup>, respectively (Birnie and Rickwood, <xref ref-type="bibr" rid="B3">1978</xref>). Thus the shift in BD (&#x003C1;) can be modeled as:</p>
<disp-formula id="E7"><label>(6)</label><mml:math id="M7"><mml:mrow><mml:msub><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mn>13</mml:mn><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>max</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mn>12</mml:mn><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
<p>where <italic>I</italic><sub><italic>i,max</italic></sub> is the maximum possible BD shift if 100% atom excess for isotope <italic>i, A</italic> is the atom % excess of isotope <italic>i</italic>, and &#x003C1;<sub>12<italic>C</italic></sub> is the buoyant density at 0% atom excess. Note that the same formula applies to <sup>15</sup>N.</p>
</sec>
<sec>
<title>SIP data simulation framework overview</title>
<p>Based on the theory described above, our SIP data simulation framework simulates the distribution of gDNA fragments in isopycnic gradients at sedimentation equilibrium. Furthermore, it generates the DNA-SIP datasets obtained from fractionating isopycnic gradient(s) and performing high throughput sequencing on many of the gradient fractions. Our framework also implements all of the DNA-SIP analysis methods assessed in this study (Heavy-SIP, HR-SIP, MW-HR-SIP (see below), qSIP, and &#x00394;BD) and evaluates their accuracy of identifying incorporators or quantifying BD shifts. An overview of our simulation framework is shown in Figure <xref ref-type="fig" rid="F1">1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The SIPSim simulation workflow involves three major stages, which are broken down into multiple steps. Stage 1 involves generating a buoyant density distribution of gDNA fragments for each genome. Stage 2 involves simulating the isopycnic gradients for a particular experimental design. Stage 3 involves generating a DNA-SIP dataset based on the fragment BD value distributions simulated in Stage 1 along with the isopycnic gradient data generated in Stage 2. The output is a table (&#x0201C;DNA-SIP dataset&#x0201D;) of taxon relative abundances in each gradient fraction in each gradient. See section Methods for a more detailed description of the simulation workflow.</p></caption>
<graphic xlink:href="fmicb-09-00570-g0001.tif"/>
</fig>
<p>Our simulation framework is a modular collection of steps that can be grouped in workflow stages that are further broken down into steps (Figure <xref ref-type="fig" rid="F1">1</xref>). The input is a set of reference genomes in fasta format and a text file designating the experimental design, which includes the number of gradients for labeled treatments and unlabeled controls.</p>
<p>Stage 1 involves generating a BD distribution of gDNA fragments for each genome. Step 1a involves simulating the pool of gDNA fragments that is extracted from SIP incubation samples and then loaded into the isopycnic gradients. If amplicon sequence data (e.g., SSU rRNA) is to be generated, amplicons from only the fragments containing the PCR template (&#x0201C;amplicon-fragments&#x0201D;) are sequenced, while shotgun metagenomic sequencing can target all gDNA fragments (&#x0201C;shotgun-fragments&#x0201D;). If &#x02265;1 PCR primer set is provided, amplicon-fragments are generated from genomic regions fully encompassing genome locations that produced amplicons by <italic>in silico</italic> PCR. Alternatively, shotgun-fragments are randomly generated from all possible genomic locations. The fragment size distribution is user-defined (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
<p>As described in Equations (1, 3), the length and G &#x0002B; C content of a DNA fragment can be used to calculate a probability distribution of its location in the gradient, assuming sedimentation equilibrium. Step 1b uses the fragments simulated in Step 1a to generate a 2-dimensional Gaussian kernel density estimation (KDE) for each taxon, which describes the joint probability of obtaining fragments with a certain length and G &#x0002B; C content from that taxon. From this 2D-KDE, a large number of [length, G &#x0002B; C] vectors can be simulated efficiently for more precise estimations of the fragment BD distributions. Fragment BD distributions are calculated for each taxon in Step 1c by sampling [length, G &#x0002B; C] vectors from the 2D-KDE and calculating Gaussian distribution from each, where the mean is based on Eq. 1 and the standard deviation based on Equation (3). The collection of Gaussian distributions for all fragments for each taxon is integrated into a BD distribution for all fragments of a taxon with Monte Carlo error estimation, which involves sampling BD values from the collection of Gaussian distributions and estimating a probability density function (PDF) of the fragment BD distribution as a one-dimensional Gaussian KDE. The result is a list of KDEs, with each describing the probability of detecting the gDNA fragments of a taxon at any point along the isopycnic gradient. These fragment BD distributions are modified in steps 1d and 1e by adding diffusive boundary layer (DBL) effects (see section Theory Underlying the Simulation Framework) and isotope incorporation, respectively. The &#x0201C;smearing&#x0201D; due to DBL effects is modeled as a uniform distribution describing the increased fragment BD uncertainty, and this uncertainty is integrated into the fragment BD distributions by Monte Carlo error estimation as in Step 3b. The BD shift due to isotope incorporation is modeled in a similar manner, except BD uncertainty is a result of inter- and intra-population variation in the amount of isotope incorporated. Variation of isotope incorporation is modeled as a hierarchical set of mixture models (weighted sets of standard distributions; such as two Gaussians), where the parameters for intra-population mixture models that describe the amount of isotope incorporated by each individual are themselves defined by inter-population mixture models that describe how isotope incorporation varies among taxa.</p>
<p>Stage 2 involves simulating the isopycnic gradients for a particular experimental design. Step 2a involves simulating the BD range size of each fraction of each gradient. Sizes are drawn from a user-defined distribution. Step 2b involves simulating the relative abundance distribution of taxa in the gDNA pools loaded into each gradient (&#x0201C;pre-fractionation communities&#x0201D;). The abundance distribution of each pre-fractionation community is user-defined and can vary among gradients. Furthermore, the amount of taxa shared or rank-abundances permutated among communities (i.e., the beta-diversity) is user-defined.</p>
<p>Stage 3 involves generating a DNA-SIP dataset based on the fragment BD distributions simulated in Stage 1 along with the isopycnic gradient data generated in Stage 2. In Step 3a, an OTU (taxon) abundance table is generated by sampling from the fragment BD distributions of each taxon generated in Stage 1, with sampling depth determined by pre-fractionation community abundances simulated in Step 2b. The subsampled fragments are then binned into gradient fractions simulated in Step 2a. The resulting OTU table lists the number of gDNA fragments of each taxon in each gradient fraction in each gradient. If the simulated fragments are amplicons, then PCR amplification efficiency biases are simulated in Step 3b based on the PCR kinetic model described in Suzuki and Giovannoni (<xref ref-type="bibr" rid="B31">1996</xref>). The model assumes that efficiencies decrease as the product concentration increases due to an increased propensity of single stranded products to re-anneal to their homologous complements. Sequence data is simulated in Step 3c by subsampling from the table of fragment counts (the DNA fragment pool), which produces a final table (&#x0201C;DNA-SIP dataset&#x0201D;) of taxon relative abundances in each gradient fraction in each gradient.</p>
</sec>
<sec>
<title>SIP data simulation framework parameters</title>
<p>Unless stated otherwise, we made the following assumptions for all simulations in this study. Community abundance distributions were simulated as lognormal distributions with a mean of 10 and a standard deviation of 2. All taxa were shared among communities, and no rank-abundances were permuted. The total number of fragments in each gradient was 1e<sup>9</sup>. Gradient fragment BD range sizes were sampled from a normal distribution, with a mean of 0.004 and a standard deviation of 0.0015. SSU rRNA amplicon-fragments were simulated using the V4-targeting 16S rRNA primers: 515F and 927R (5&#x02032;-GTGYCAGCMGCMGCGGTRA-3&#x02032;; 5&#x02032;-CCGYC AATTYMTTTRAGTTT-3&#x02032;), as used by Pepe-Ranney et al. (<xref ref-type="bibr" rid="B22">2016a</xref>). The amplicon-fragment size distribution was a left-skewed normal distribution with a mean of &#x0007E;12 kb, which is similar to size distributions produced from common bead beating cell lysis methods (Kauffmann et al., <xref ref-type="bibr" rid="B13">2004</xref>; Roh et al., <xref ref-type="bibr" rid="B27">2006</xref>; Thakuria et al., <xref ref-type="bibr" rid="B32">2008</xref>). A total of 1e<sup>4</sup> amplicon-fragments were simulated per genome, which equated to &#x0003E;100X coverage for the genomic region of interest. Monte Carlo error estimation was conducted with 1e<sup>5</sup> sampling replicates. Ultracentrifugation conditions were set as in Pepe-Ranney et al. (<xref ref-type="bibr" rid="B22">2016a</xref>), with a Beckman TLA-110 rotor spun at 55,000 rpm for 66 h at 20&#x000B0;C and an average density gradient 1.7 g ml<sup>&#x02212;1</sup>. Inter-population variation in isotope incorporation was binary (either 0% or X% atom excess), and intra-population variation was set to zero. Two key parameters were estimated from empirical DNA-SIP data: the bandwidth (smoothing factor) for kernel density estimation, and the gamma parameter in Equation (5). See Table <xref ref-type="supplementary-material" rid="SM1">S1</xref> for a full listing of simulation parameters.</p>
</sec>
<sec>
<title>Implementing DNA-SIP analyses</title>
<p>The HR-SIP method was performed as described in Pepe-Ranney et al. (<xref ref-type="bibr" rid="B22">2016a</xref>,<xref ref-type="bibr" rid="B23">b</xref>). Briefly, we used a &#x0201C;heavy&#x0201D; BD window of 1.71&#x02013;1.75 g ml<sup>&#x02212;1</sup>, a sparsity cutoff of 0.25 (i.e., OTUs must be present in &#x0003E;25% of samples), a log<sub>2</sub> fold change null threshold of 0.25, and a false discovery rate cutoff of 10%. &#x00394;BD was determined as described by Pepe-Ranney et al. (<xref ref-type="bibr" rid="B22">2016a</xref>), with OTU abundances linearly interpolated across 20 evenly spaced values across the gradient BD range. Briefly, &#x00394;BD is calculated as the difference in the center of mass of abundance distributions (i.e., the average BD weighted by relative abundance) between the labeled treatment and unlabeled control.</p>
<p>qSIP was conducted as described in Hungate et al. (<xref ref-type="bibr" rid="B12">2015</xref>), with 90% confidence intervals calculated from 1,000 bootstrap replicates. The variance among qPCR replicates was modeled based on the qPCR data provided in Table <xref ref-type="supplementary-material" rid="SM1">S2</xref> of Hungate et al. (<xref ref-type="bibr" rid="B12">2015</xref>). Specifically, we found the qPCR count variance (&#x003C3;<sup>2</sup>) to increase as a function of the mean (&#x003BC;). The following polynomial regression was found to best describe this relationship and was used for simulating all qPCR count values:</p>
<disp-formula id="E8"><label>(7)</label><mml:math id="M8"><mml:mrow><mml:msup><mml:mi>&#x003C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>5889</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x003BC;</mml:mi><mml:mo>+</mml:mo><mml:mn>0.714</mml:mn><mml:msup><mml:mi>&#x003BC;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:math></disp-formula>
<p>where &#x003BC; was set as the total number of simulated DNA fragments in the gradient fraction (designated in the OTU table from Step 4a).</p>
<p>A range of alternative analytical approaches to Heavy-SIP can be used to detect incorporators. We characterized four additional approaches to identifying labeled OTUs. Method 1 identifies as labeled any OTU that occurs in &#x0201C;heavy&#x0201D; fractions of the labeled gradient. Method 2 identifies as labeled any taxa present in the &#x0201C;heavy&#x0201D; fractions of the labeled treatment and absent from the &#x0201C;heavy&#x0201D; fractions of the control gradient. Method 3 identifies as labeled any taxa present in the &#x0201C;heavy&#x0201D; fractions of the labeled treatment and absent in the &#x0201C;light&#x0201D; fractions of the labeled treatment. Method 4 identifies as labeled any taxa present in the &#x0201C;heavy&#x0201D; fractions of the labeled treatment and absent from both the &#x0201C;heavy&#x0201D; fractions of the control and the &#x0201C;light&#x0201D; fractions of the labeled treatment. Of these four approaches, Method 1 provided the highest accuracy (Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>) and so this is the method that we used to represent &#x0201C;Heavy-SIP.&#x0201D;</p>
<p>We hypothesized that HR-SIP sensitivity could be improved by altering the &#x0201C;heavy&#x0201D; BD window (1.71&#x02013;1.75 g ml<sup>&#x02212;1</sup>) in which sequence composition is compared between treatment and control. We used SIPSim to evaluate a range of analytical approaches (data not shown) and found that the analysis of multiple windows (hereby called &#x0201C;MW-HR-SIP&#x0201D;) resulted in a significant improvement in sensitivity relative to HR-SIP. MW-HR-SIP evaluates sequence composition within BD windows of: 1.70&#x02013;1.73, 1.72&#x02013;1.75, 1.74&#x02013;1.77 g ml<sup>&#x02212;1</sup> (Figure <xref ref-type="supplementary-material" rid="SM1">S3</xref>) while adjusting for multiple comparisons.</p>
</sec>
<sec>
<title>Datasets</title>
<p>The genome dataset used to simulate genomic DNA fragments was obtained from Genbank (Benson et al., <xref ref-type="bibr" rid="B2">2008</xref>). From a list of all bacterial genomes designated as &#x0201C;complete,&#x0201D; one representative was chosen per species in order to reduce the bias toward highly represented species. We found the dataset to contain a rather high proportion (&#x0007E;12%) of low G &#x0002B; C organisms (&#x0003C;30% G &#x0002B; C); most of which were obligate endosymbionts. We randomly sampled a subset of these low G &#x0002B; C genomes in order to reduce the proportion of low G &#x0002B; C organisms to just 1% of the genome dataset. The resulting dataset consisted of 1,147 bacterial genomes.</p>
<p>In order to simulate empirical data from Lueders et al. (<xref ref-type="bibr" rid="B17">2004</xref>), the genome sequences of <italic>Methanosarcina barkeri</italic> MS and <italic>Methylobacterium extorquens</italic> AM1 were downloaded from Genbank. Amplicon-fragments were simulated with the primers Ar109f (5&#x02032;-ACKGCTCAGTAACACGT-3&#x02032;), Ar915r (5&#x02032;-GTGCTCCCCCGCCAATTCCT-3&#x02032;), Ba519f (5&#x02032;-CAGCMGCCGCGGTAANWC-3&#x02032;), and Ba907r (5&#x02032;-CCGTCAATTCMTTTRAGTT-3&#x02032;). Atom % excess was assumed to be 100%, and isopycnic centrifugation conditions were simulated as specified in Lueders et al. (<xref ref-type="bibr" rid="B17">2004</xref>).</p>
<p>For model evaluation (see Supporting Results), we downloaded the genomes <italic>Clostridium ljungdahlii</italic> DSM 13528, <italic>Escherichia coli</italic> 1303, and <italic>Streptomyces pratensis</italic> <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="ATCC_33331">ATCC 33331</ext-link> from Genbank.</p>
<p>We compared the properties of DNA fragment BD distributions between simulated DNA-SIP data and empirical DNA-SIP data obtained from a soil community. The DNA-SIP dataset from Youngblut and colleagues consisted of SSU rRNA MiSeq sequences (V4 region) of &#x0007E;24 fractions per gradient from 6 gradients of unlabeled controls. This dataset was generated using agricultural soils (15 g per sample) amended with a complex substrate mixture (3.6 mg C per g soil), and incubated aerobically at 50% water holding capacity and at room temperature. DNA was extracted following destructive sampling of replicates at days 1, 3, 6, 14, 30, and 48. DNA was subject to CsCl centrifugation (TLA110 Beckman rotor, 55,000 rpm, 66 h, 1.69 g ml<sup>&#x02212;1</sup> average gradient density) and fractionated (100 &#x003BC;l fractions) using methods which have previously been described in detail (Pepe-Ranney et al., <xref ref-type="bibr" rid="B23">2016b</xref>). Since the use of these samples in the present context is limited to the analysis of DNA buoyant density distribution properties in the gradient, only control samples containing unlabeled DNA were examined. These data were subsampled to obtain a total richness equal to the 1,147 OTUs in our reference genome dataset. The sequence data is available from the NCBI under BioProject <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="PRJNA382302">PRJNA382302</ext-link>.</p>
</sec>
<sec>
<title>Software implementation</title>
<p>The SIP simulation framework was mostly written in Python v2.7.11, with some accompanying code written in C&#x0002B;&#x0002B; v4.9.2 and R v3.2.3 (R Core Team, <xref ref-type="bibr" rid="B26">2016</xref>). MFEprimer v2.0 was used to perform <italic>in silico</italic> PCR (Qu et al., <xref ref-type="bibr" rid="B24">2009</xref>). The software, along with documentation and examples, can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/nick-youngblut/SIPSim">https://github.com/nick-youngblut/SIPSim</ext-link>. All genomes were downloaded from Genbank with the R package <italic>genomes</italic> v2.12.0 (Stubben, <xref ref-type="bibr" rid="B30">2014</xref>), and all data analysis was conducted in R with the following packages: ggplot2 v2.1.0, dplyr v0.4.3, tidyr v 0.4.1, and cowplot v0.6.2.</p>
<p>Further methodological details are provided in the Supplementary Material.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Model validation and parameter estimation</title>
<p>The SIPSim model starts with a set of user-designated genomes and user-designated experimental parameters (e.g., number of gradient fractions, desired community characteristics, desired isotopic labeling characteristics) as described (see section Methods and Supplementary Material). Briefly, the genomes are fragmented as would occur during DNA extraction, isotopic labeling is applied to some number of genomes as specified by the user, the BD distributions are determined for each DNA fragment and fragment collections are then binned into gradient fractions, fragments are sampled from each fraction as would occur during amplification and DNA sequencing of SSU rRNA genes, and then the relative abundance is calculated for each OTU (Figure <xref ref-type="fig" rid="F1">1</xref>, see also section Methods). The model produces results that are highly similar to those observed in empirical experiments, including the ability to detect DNA fragments throughout the density gradient (Figure <xref ref-type="fig" rid="F2">2</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>SIPSim output provides data that approximates results obtained from DNA-SIP experiments. The CsCl gradient BD distributions of diverse amplicon fragments (<italic>n</italic> &#x0003D; 1,147 taxa) are depicted such that the distribution of each taxon is represented by a different color. All taxa in the control had 0% atom excess <sup>13</sup>C, while 10% of taxa in the treatment were randomly assigned 100% atom excess <sup>13</sup>C. Most unlabeled amplicon fragments occur within the range of 1.69&#x02013;1.72 g ml<sup>&#x02212;1</sup>, while <sup>13</sup>C-labeled taxa are shifted into higher BD fractions (pre-sequencing, top panels). During the process of high-throughput DNA sequencing amplicon fragments are randomly sampled from each fraction, and this sampling effect alters the shape of the fragment distributions observed in DNA-SIP experiments (post-sequencing, middle panel) relative to the actual distribution of DNA in the gradient (top panels). Typically, data from DNA-SIP experiments are transformed into relative abundance values (post-sequencing, bottom panel) prior to analysis. Identification of taxa that have incorporated isotope requires comparison of amplicon fragment relative abundance distributions in treatment relative to control gradients. The dashed vertical line is provided as a point of reference and designates the theoretical buoyant density of an unlabeled DNA fragment with 50% G &#x0002B; C (as modeled in Equation 1).</p></caption>
<graphic xlink:href="fmicb-09-00570-g0002.tif"/>
</fig>
<p>The development of the simulation model was guided by established centrifugal theory and by comparison of simulated results to empirical data (as in section Methods and Supplementary Material). First, we performed a simple evaluation of model performance by recreating results from a prior DNA-SIP experiment with <italic>Methanosarcina barkeri</italic> MS and <italic>Methylobacterium extorquens</italic> AM1 (Lueders et al., <xref ref-type="bibr" rid="B17">2004</xref>) (Figure <xref ref-type="supplementary-material" rid="SM1">S4</xref>). Simulated DNA distributions (both in terms of total DNA and SSU rRNA gene amplicon copies) significantly and strongly correlated with the empirical data for both taxa (<italic>p</italic> &#x0003C; 0.003 for all comparisons; see Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>). In addition, the simulated SSU rRNA gene amplicon-fragment BD distributions were shifted 0.007 g ml<sup>&#x02212;1</sup> toward the middle of the BD gradient relative to the shotgun-fragments (&#x0201C;total DNA&#x0201D;), a phenomenon also observed in the empirical data. This central tendency for SSU rRNA amplicon-fragments reflects G &#x0002B; C conservation of the <italic>rrn</italic> operon, as previously described (Youngblut and Buckley, <xref ref-type="bibr" rid="B36">2014</xref>).</p>
<p>Next, we evaluated SIPSim results by comparing to empirical data. For this purpose SIPSim output was compared to results obtained with unlabeled DNA from soil (see section Datasets). The empirical data was derived from an experiment in which an unlabeled nutrient mixture was added to soil and DNA was extracted at 1, 3, 6, 14, 30, and 48 days. These six DNA samples were equilibrated in CsCl gradients, fractionated by BD, and SSU rRNA gene amplicons were sequenced for &#x0007E;24 fractions from each gradient. The simulation run included 1,147 microbial genomes (see section Methods), hence the soil data was resampled to 1,147 OTUs in order to standardize the richness of the simulated and empirical data. The empirical results reveal several interesting features about the distribution of DNA in CsCl gradients. First, we observed that variance in DNA fragment BD is positively correlated with its relative abundance and that most taxa with relative abundances &#x0003E;0.1% are detected in all CsCl gradient fractions (Figure <xref ref-type="fig" rid="F3">3</xref>). Second, we observed that taxonomic similarity is auto-correlated with fraction BD, so that fractions of similar BD have similar nucleic acid composition (Figure <xref ref-type="fig" rid="F3">3</xref>). Finally, we observed that changes in community composition caused dramatic shifts in the sequence composition of &#x0201C;heavy fractions&#x0201D; even in the absence of isotopic labeling (Figure <xref ref-type="fig" rid="F3">3</xref>). We applied three different analytical approaches to assess similarity between empirical and simulated data: correlation in Shannon Diversity with respect to fraction BD, correlation in Jaccard dissimilarity with respect to fraction BD, and correlation in OTU BD range with respect to taxon relative abundance. We found that variance between simulated and empirical results was significantly less than the variance observed between replicate empirical samples (Figure <xref ref-type="supplementary-material" rid="SM1">S5</xref>). Furthermore, we used these comparisons between simulated and empirical results (Figure <xref ref-type="supplementary-material" rid="SM1">S5</xref>) to optimize parameter values for use in the SIPSim model (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>, as described in Supplementary Material).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Empirical DNA-SIP data shows that unlabeled DNA is found widely within the gradient and that changes in beta-diversity can alter the composition of &#x0201C;heavy&#x0201D; fractions in the absence of isotopically labeled substrates. The DNA is from soil communities incubated for 1, 3, 6, 14, 30, or 48 days following the addition of an unlabeled nutrient mixture. SSU rRNA genes were amplified and sequenced from approximately 24 fractions from each gradient, these amplicons were used to identify the BD variance of amplicon fragments derived from discrete OTUs. The DNA concentration of each gradient fraction was measured using Picogreen assay <bold>(A)</bold>. These values are normalized to the maximum concentration within each gradient. The amplicon diversity within each gradient fraction was measured using the Shannon Index, showing that the diversity of heavy fractions differs between samples even in the absence of isotopic labeling <bold>(B)</bold>. The correlograms <bold>(C)</bold> reveal autocorrelation (measured with Mantel tests) between taxonomic similarity and fraction BD within each gradient. The variance in OTU BD is positively correlated with OTU pre-fractionation relative abundance, with highly abundant OTUs found throughout the CsCl gradient <bold>(D)</bold>. To improve clarity, single OTUs in <bold>(D)</bold> were binned into hexagons, with darker shading indicating more OTUs.</p></caption>
<graphic xlink:href="fmicb-09-00570-g0003.tif"/>
</fig>
</sec>
<sec>
<title>The influence of isotope incorporation on DNA-SIP accuracy</title>
<p>We hypothesized that both the number of taxa that incorporate isotope and the atom % excess isotope incorporation per taxon would substantially affect the accuracy of DNA-SIP methods. To test these predictions, we simulated DNA-SIP datasets for both <sup>13</sup>C-labeled samples and unlabeled controls (3 replicates of each), while varying both the number of incorporators (1, 5, 10, 25, or 50% of taxa) and the atom % excess isotope incorporation for each taxon (0, 15, 25, 50, 75, or 100 atom % excess <sup>13</sup>C). Taxa in the control were always set to 0% atom excess isotope incorporation. Each simulation was replicated 10 times, with differing taxa randomly designated as incorporators in each replicate. We evaluated 7 methods used to analyze DNA-SIP data including: HR-SIP, qSIP, 4 different approaches to &#x0201C;Heavy-SIP,&#x0201D; and MW-HR-SIP. Heavy-SIP includes a family of approaches (see section Methods) in which incorporators are identified on the basis of presence-absence in heavy and/or light fractions (Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>).</p>
<p>The model predicts that both the number of incorporators and the amount of isotope incorporated affected accuracy (Figure <xref ref-type="fig" rid="F4">4</xref>). However, the predicted effect of these parameters on specificity and sensitivity varied depending on the analytical method (Figure <xref ref-type="fig" rid="F4">4</xref>). Specificity is the proportion of true negatives observed out of all true negatives expected, and so specificity declines in direct relation to an increase in the number of false positives. For example, a specificity of 0.8 would generate 200 false positives in a sample of 1,000 unlabeled taxa. Specificity, as measured across a wide range in parameters, was predicted to be highest for MW-HR-SIP (1.00 &#x000B1; 0; ave. &#x000B1; s.d.) and HR-SIP (1.00 &#x000B1; 0), substantially lower for qSIP (0.88 &#x000B1; 0.06), and very low for Heavy-SIP (0.28 &#x000B1; 0.16) (Figure <xref ref-type="fig" rid="F4">4</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>SIPSim predicts that DNA-SIP methods vary in accuracy depending on the <sup>13</sup>C atom % excess of DNA and the number of taxa that incorporate isotope. Points and bars represent means and standard deviations, respectively (<italic>n</italic> &#x0003D; 10 simulations). Specificity indicates the fraction of true negatives that are identified correctly. Sensitivity indicates the fraction of labeled taxa (true positives) identified correctly. Balanced accuracy is the product of specificity and sensitivity. The <italic>x</italic>-axis indicates the amount of <sup>13</sup>C isotope present in taxa that are labeled, and different colors are used to indicate the percentage of taxa that have incorporated <sup>13</sup>C as indicated by the legend. Heavy-SIP identifies incorporators solely based on OTU presence in &#x0201C;heavy&#x0201D; gradient fractions, and this approach is shown because it has better balanced accuracy than any of the other Heavy-SIP approaches that were analyzed (see Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>).</p></caption>
<graphic xlink:href="fmicb-09-00570-g0004.tif"/>
</fig>
<p>Sensitivity is the fraction of true positives observed out of all true positives expected. For example, a sensitivity of 0.7 means that a method failed to detect 30% of the incorporators present. Both qSIP and Heavy-SIP are predicted to have relatively high sensitivity (median values of 0.91 and 0.93, respectively), and the sensitivity of these methods was largely insensitive to the atom % excess of DNA and the number of incorporators (Figure <xref ref-type="fig" rid="F4">4</xref>). In contrast, the sensitivities of both HR-SIP and MW-HR-SIP were predicted to be highly responsive to the atom % excess of DNA, and the number of incorporators (Figure <xref ref-type="fig" rid="F4">4</xref>). For these methods, sensitivity is predicted to decline in proportion to the atom % excess <sup>13</sup>C label in DNA.</p>
<p>Balanced accuracy is calculated as the mean of specificity and sensitivity. The model predicts a tradeoff in balanced accuracy in relation to the atom % excess <sup>13</sup>C of DNA. MW-HR-SIP had the highest predicted accuracy of any of the 7 methods tested when % atom excess <sup>13</sup>C exceeded 50%, but qSIP has higher accuracy at lower levels of isotope incorporation (Figure <xref ref-type="fig" rid="F4">4</xref>). This tradeoff in balanced accuracy resulted from a difference in the tolerance for false positives. For example, MW-HR-SIP produced nearly zero false positives but as a result of its high specificity, it lost sensitivity at lower levels of isotope incorporation. In contrast, qSIP detected labeled taxa across a wider range of isotope incorporation, but it did so at the cost of a large number of false positives.</p>
</sec>
<sec>
<title>The influence of experimental parameters on DNA-SIP accuracy</title>
<p>We also evaluated the effects of sequencing effort and fraction size on DNA-SIP accuracy because SIP experiments often vary in the number of fractions analyzed per gradient and the number of sequences analyzed per fraction. The model shows that sequencing effort has different consequences for method specificity and sensitivity. We found that the specificities of HR-SIP and MW-HR-SIP are predicted to be independent of sequencing effort, while those of Heavy-SIP and qSIP actually declined with sequencing effort (Figure <xref ref-type="fig" rid="F5">5</xref>). The predicted decrease in specificity for Heavy-SIP and qSIP indicates that the rate of false discovery for these methods increases in proportion to the number of sequences analyzed, while the rate of false discovery in HR-SIP and MW-HR-SIP is unaffected by the number of sequences analyzed. In contrast, we found that method sensitivity improved with sequencing effort for all analytical methods (Figure <xref ref-type="fig" rid="F5">5</xref>). This predicted increase in sensitivity is caused by an increase in statistical power caused by sampling more sequences from each gradient fraction. We further show that improvements in sensitivity are predicted to be greatest for taxa present at low relative abundance (Figure <xref ref-type="fig" rid="F6">6</xref>). For example, with MW-HR-SIP the sensitivity of detection for a 50% atom <sup>13</sup>C enriched OTU present at 0.001 relative abundance is predicted to be nearly zero if less than 1,000 sequences are analyzed per gradient fraction, but sensitivity improves dramatically as sequencing depth increases (Figure <xref ref-type="fig" rid="F6">6</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>SIPSim predicts that DNA-SIP methods vary in accuracy depending on the number of sequences analyzed per gradient fraction. Points and bars represent means and standard deviations, respectively (<italic>n</italic> &#x0003D; 10 simulations). Specificity indicates the fraction of true negatives that are identified correctly. Sensitivity indicates the fraction of labeled taxa (true positives) identified correctly. Balanced accuracy is the product of specificity and sensitivity. The <italic>x</italic>-axis indicates the amount of <sup>13</sup>C isotope present in taxa that are labeled, and different colors are used to indicate the average number of sequences determined per gradient fraction as described by the legend.</p></caption>
<graphic xlink:href="fmicb-09-00570-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>SIPSim predicts that the sensitivity of detecting OTUs that have low relative abundance and low atom % <sup>13</sup>C enrichment increases with added sequencing effort per gradient fraction. Each panel provides results from simulations conducted at different levels of atom % <sup>13</sup>C enrichment as indicated (0, 15, 25, 50, 75, and 100 atom % <sup>13</sup>C), incorporators were identified using MW-HR-SIP, and sensitivity was assessed by binning OTUs into 10 different abundance classes. Sensitivity indicates the fraction of labeled taxa (true positives) identified correctly. The <italic>x</italic>-axis indicates the mean relative abundance of the taxa being evaluated, and different colors are used to indicate the average number of sequences determined per gradient fraction as described by the legend. Points and bars represent means and standard deviations, respectively (<italic>n</italic> &#x0003D; 10 simulations).</p></caption>
<graphic xlink:href="fmicb-09-00570-g0006.tif"/>
</fig>
<p>SIP methods also will often vary in the number of fractions collected per gradient, and so we evaluated the effect of fraction size on DNA-SIP accuracy. The model predicts that the use of smaller fractions (i.e., collecting more fractions per gradient) tends to improve sensitivity and overall accuracy for HR-SIP, MW-HR-SIP, and qSIP, though the effects are modest (Figure <xref ref-type="supplementary-material" rid="SM1">S7</xref>). In contrast, Heavy-SIP is predicted to improve in accuracy when larger fractions are used, though this effect is also somewhat modest (Figure <xref ref-type="supplementary-material" rid="SM1">S7</xref>).</p>
</sec>
<sec>
<title>The influence of community variation on DNA-SIP accuracy</title>
<p>All DNA-SIP analyses rely upon comparisons made between isotopically enriched experimental treatments and their corresponding unlabeled controls. In real SIP experiments, the composition of replicate post incubation communities are likely to vary somewhat as a result of sample heterogeneity and incubation effects. However, the simulations described above assume random sampling from identical pre-fractionation (post-incubation) community structures. We hypothesized that an increase in variation in community composition between treatment and control samples would decrease the accuracy of DNA-SIP analyses. To test this hypothesis, we generated simulations in which isotope incorporation was held constant (50 atom % excess <sup>13</sup>C; 10% of OTUs are incorporators) but beta-diversity was varied among 3 replicate treatment and 3 replicate control samples. We varied beta-diversity in two ways: (i) using permutation to vary the rank abundance of a fixed proportion of community members and (ii) varying the proportion of taxa shared between communities. For each simulation scenario, we calculated the mean Bray-Curtis distance among communities in order to provide a real-world metric for gauging the potential accuracy of actual DNA-SIP experiments.</p>
<p>The model predicts that increased beta-diversity among samples impacts the accuracy of DNA-SIP methods (Figure <xref ref-type="fig" rid="F7">7</xref>). The model predicts that accuracy is impacted more by the number of taxa shared between samples than by differences in taxon abundance (Figure <xref ref-type="supplementary-material" rid="SM1">S8</xref>). The sensitivity of all methods is predicted to decline as beta-diversity increases, falling from approximately 0.9 for both qSIP and MW-HR-SIP when samples shared 100% of their OTUs to 0.64 and 0.7 for qSIP and MW-HR-SIP, respectively, when samples shared only 80% of their OTUs (Figure <xref ref-type="supplementary-material" rid="SM1">S8</xref>). Increasing the beta-diversity between samples was predicted to have little effect on the specificity of qSIP but diminished the specificity of HR-SIP and MW-HR-SIP (Figure <xref ref-type="fig" rid="F7">7</xref>). MW-HR-SIP was predicted to have greater balanced accuracy than qSIP as the Bray-Curtis distance between treatment and control samples increased from of 0.0 to 0.4, but at higher levels of distance the two methods performed with similar accuracy (Figure <xref ref-type="fig" rid="F7">7</xref>). Regardless, it is clear that sample-to-sample variation has an overall negative impact on DNA-SIP accuracy, and this result emphasizes the importance of minimizing experimental variation between unlabeled controls and labeled treatments in SIP experiments.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>SIPSim predicts that DNA-SIP methods differ in their sensitivity to community dissimilarity between replicate samples. Beta diversity, expressed as Bray-Curtis dissimilarity, was varied between simulated replicates (3 replicates each for <sup>12</sup>C-control and <sup>13</sup>C-treatment gradients) to determine the effect that community dissimilarity between replicates has on method accuracy. Variation in beta diversity was simulated by systematically varying two parameters: the percent of taxa shared between replicate samples (80, 85, 90, 95, or 100%) and the percent of taxa whose rank abundances that were permuted (0, 5, 10, 15, or 20%), with 10 simulation replicates for each parameter set. The blue lines are LOESS curves fit to accuracy values for all simulations (<italic>n</italic> &#x0003D; 250), and the gray regions represent 99% confidence intervals. For all simulations, 10% of the community were incorporators (50% atom excess <sup>13</sup>C).</p></caption>
<graphic xlink:href="fmicb-09-00570-g0007.tif"/>
</fig>
</sec>
<sec>
<title>Using DNA-SIP data to quantify atom % excess</title>
<p>So far, we have focused on the accuracy of DNA-SIP methods with respect to the identification of taxa that incorporate isotope into their DNA. However, changes in DNA BD can also be used to quantify the isotope enrichment of DNA from particular taxa. Two approaches have been used to evaluate isotope enrichment from DNA-SIP data: qSIP and &#x00394;BD, with the latter being a complementary analysis to HR-SIP (Pepe-Ranney et al., <xref ref-type="bibr" rid="B22">2016a</xref>). Both &#x00394;BD and qSIP derive quantitative estimates from measuring taxon BD shifts (and thus atom % excess) in the labeled treatment gradient(s) vs. their unlabeled counterparts. The &#x00394;BD method attempts to measure the extent of the BD shift directly from the compositional sequence data, while qSIP utilizes relative abundances transformed by qPCR counts of total SSU rRNA copies. Therefore, &#x00394;BD accuracy likely suffers from compositional effects inherent to HTS datasets, while qSIP accuracy is dependent on qPCR accuracy and variation.</p>
<p>We assessed the quantification accuracy of both methods using the simulations described previously, where either the amount of isotope incorporation or sample beta-diversity was varied. The model predicts that &#x00394;BD produced estimates of isotope incorporation that are closer on average to the true value compared to qSIP, but &#x00394;BD values had much higher variance than qSIP estimates (Figure <xref ref-type="fig" rid="F8">8</xref>). Furthermore, the predicted variance in &#x00394;BD atom % excess <sup>13</sup>C estimates increased substantially with even moderate increases in beta-diversity between samples, while the qSIP estimations were largely invariant across the simulation parameter space (Figure <xref ref-type="supplementary-material" rid="SM1">S9</xref>A). Both methods are predicted to consistently misestimate true <sup>13</sup>C atom % excess, though the effect was greater for qSIP, with qSIP underestimating <sup>13</sup>C atom % excess by 30.2&#x02013;39.2% for fully labeled DNA (Figure <xref ref-type="fig" rid="F8">8B</xref>).</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>SIPSim predicts that &#x00394;BD and qSIP vary in their accuracy at estimating <sup>13</sup>C atom % excess of labeled DNA fragments. The accuracy of both methods declines as the amount of <sup>13</sup>C in DNA increases <bold>(A)</bold>, but accuracy is not affected by the percent of taxa that are labeled; values indicate the mean and standard deviation (<italic>n</italic> &#x0003D; 10 simulations). Probability density plots indicate that estimates of <sup>13</sup>C atom % excess made using &#x00394;<italic>BD</italic> have greater variance than those made using <italic>qSIP</italic>, but both estimates systematically underestimate levels of isotope incorporation <bold>(B)</bold>. Each vertical pair of panels indicates the probability density for estimates made across different levels of isotope incorporation (15, 25, 50, and 100 atom % excess), and the dashed line indicates the actual level of isotopic enrichment. For the calculation of probability density, 10% of taxa were labeled using the level of enrichment indicated in each panel.</p></caption>
<graphic xlink:href="fmicb-09-00570-g0008.tif"/>
</fig>
<p>We further investigated several factors to determine whether they impact the estimation of <sup>13</sup>C atom % excess from DNA-SIP data. The model predicts that the size of gradient fractions (Figure <xref ref-type="supplementary-material" rid="SM1">S10</xref>) and the depth of sequencing per fraction (Figure <xref ref-type="supplementary-material" rid="SM1">S11</xref>) had little impact on estimation of <sup>13</sup>C atom % excess for OTUs. Furthermore, the model predicts that combining techniques, by using MW-HR-SIP to first identify labeled taxa and then using qSIP to calculate the <sup>13</sup>C atom % excess of OTUs, reduces the variance of <sup>13</sup>C atom % excess estimates but does not correct for the systematic underestimation of <sup>13</sup>C atom % excess (Figures <xref ref-type="supplementary-material" rid="SM1">S10</xref>, <xref ref-type="supplementary-material" rid="SM1">S11</xref>). Finally, we attempted to determine why qSIP systematically underestimates the <sup>13</sup>C atom % excess of OTUs. We find that the degree to which qSIP underestimates <sup>13</sup>C atom % excess is predicted to increase in proportion to the actual <sup>13</sup>C atom % excess of each OTU (Figure <xref ref-type="supplementary-material" rid="SM1">S12</xref>). This outcome could be explained by qSIP&#x00027;s use of weighted averaging in estimating <sup>13</sup>C atom % excess. The use of weighted averages to estimate the BD of each OTU assumes a roughly Gaussian BD distribution, which we show not to be the case (Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>). The BD distributions of &#x0201C;heavy&#x0201D; DNA fragments will be left skewed in a CsCl gradient and the weighted average of a left skewed distribution will cause systematic underestimation of <sup>13</sup>C atom % excess with the degree of underestimation increasing in proportion to the BD of the DNA, as observed.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>Our simulation framework (SIPSim) provides a tractable platform for evaluating the accuracy of DNA-SIP methods and for developing new methods to analyze DNA-SIP data. Given the laborious nature of DNA-SIP experiments, it is impractical to use empirical analyses with mock communities to evaluate the range of parameter values that can be investigated readily through simulation (e.g., we simulated &#x0003E;1,000 SIP experiments in this effort). In addition, both the physics of density gradient centrifugation and the physical properties of genomic DNA are well established, making the simulation of DNA-SIP data both tractable and reliable. Without rigorous assessment of DNA-SIP methods, it is difficult to determine the likelihood of false negatives (Type II error) and false positives (Type I error) across the wide range of experimental conditions in which DNA-SIP has been employed in the literature. Issues of Type I and Type II statistical error are compounded by the nature of high-throughput sequencing data, where it is necessary to make many thousands of comparisons to identify OTUs that change in response to treatment. This multiple comparison problem has major implications for statistical power and the likelihood of false detection (Paulson et al., <xref ref-type="bibr" rid="B21">2013</xref>). We have used SIPSim to test the effects of multiple parameters on the accuracy of current methods for analyzing DNA-SIP data.</p>
<p>Different approaches for detecting isotope incorporators result in substantial differences in sensitivity and specificity. The model predicts that both qSIP and MW-HR-SIP are superior to several different &#x0201C;Heavy-SIP&#x0201D; approaches (Figures <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="fig" rid="F7">7</xref> and Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>). The qSIP method is predicted to have high sensitivity but low specificity, resulting in a large number of false positives (8 &#x000B1; 0.3 to 15 &#x000B1; 0.7% of the unlabeled taxa which were evaluated were misidentified as labeled; Figures <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="fig" rid="F7">7</xref>). In contrast, MW-HR-SIP is predicted to have high specificity and negligible false positives (Figures <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="fig" rid="F7">7</xref>), but had lower sensitivity (more false negatives). This tradeoff between specificity and sensitivity can be contextualized by considering a community that contains 1,100 taxa, 55 of which are isotopically labeled. If these 55 taxa are labeled at 50% atom excess <sup>13</sup>C, both methods do a good job of detecting labeled taxa (true positives: MW-HR-SIP, 51 &#x000B1; 2; qSIP, 50 &#x000B1; 2), but qSIP detects many false positives (false positives: MW-HR-SIP, 0 &#x000B1; 1; qSIP, 126 &#x000B1; 8). If these 55 taxa are instead labeled at 25% atom excess <sup>13</sup>C then MW-HR-SIP detects fewer labeled taxa (true positives: MW-HR-SIP, 33 &#x000B1; 3; qSIP, 50 &#x000B1; 2), but qSIP continues to detect many false positives (false positives: MW-HR-SIP, 1 &#x000B1; 0; qSIP, 122 &#x000B1; 8). In these examples, &#x0003E;97% of the taxa identified by MW-HR-SIP are truly labeled, while only about 29% of those identified by qSIP are actually labeled (note that this example contextualizes the number of false positives relative to the sum of true and false positives, while specificity is formally defined as the number of true negatives observed relative to true negatives expected). It is possible that the low specificity of qSIP could be caused by the fact that this method employs 90% confidence intervals to identify as <sup>13</sup>C-labeled those taxa that have a large BD increase in response to <sup>13</sup>C-labeling. It is possible that modification of qSIP to employ 99% confidence intervals could result in an improvement of specificity, however, such a change would also certainly diminish sensitivity and so the overall impact on balanced accuracy is difficult to predict at this time. Further, improvement of DNA-SIP analyses should be facilitated by use of the SIPSim framework.</p>
<p>In regards to methods used to quantify the atom % excess of individual taxa from DNA-SIP data, we found that the utility of qSIP or &#x00394;BD is predicted to vary depending on the hypothesis being evaluated. &#x00394;BD produced more accurate estimates of mean <sup>13</sup>C atom % excess than qSIP (Figure <xref ref-type="fig" rid="F8">8</xref> and Figure <xref ref-type="supplementary-material" rid="SM1">S8</xref>), and so this approach may be suitable when seeking to make relative comparisons in the degree of labeling between large groups of taxa (as described in Pepe-Ranney et al., <xref ref-type="bibr" rid="B23">2016b</xref>). However, the high variability of this approach causes &#x00394;BD to be unreliable in determining differences in atom % excess <sup>13</sup>C at the scale of individual OTUs. Alternatively, qSIP is predicted to produce much more stable estimates of atom % excess <sup>13</sup>C among individual taxa, but the method is predicted to produce systematic underestimates of isotope incorporation. We hypothesize that qSIP is underestimating atom % excess <sup>13</sup>C because it uses weighted averaging to calculate the BD of each OTU. We expect that a statistical approach less sensitive to the violations of normality that occur in CsCl gradients may improve the ability of qSIP to accurately estimate atom % excess <sup>13</sup>C values.</p>
<p>The SIPSim framework makes it possible to evaluate hypothetical outcomes of DNA-SIP experiments before they are performed and to evaluate the accuracy of DNA-SIP data analysis methods. For brevity, we have only focused on a few key variables that could affect the accuracy of DNA-SIP methods. However, SIPSim can also be used to assess the accuracy of DNA-SIP methods across a range of possible real-world scenarios. For instance, spatial or population-level heterogeneity could result in taxa that are not homogeneously labeled (Lennon and Jones, <xref ref-type="bibr" rid="B15">2011</xref>). Such systematic heterogeneity in labeling would manifest as &#x0201C;split&#x0201D; (bimodal or multimodal) distributions of DNA fragments in an isopycnic gradient. It would be challenging to evaluate such scenarios empirically, but SIPSim can be readily used to evaluate a range of such scenarios. SIPSim also provides a toolkit for developing and improving analytical methods used in DNA-SIP experiments.</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>Conclusion</title>
<p>With our newly developed simulation toolset, we determined that MW-HR-SIP is predicted to have the lowest false positive rate of all methods tested for analyzing DNA-SIP data. The use of MW-HR-SIP resulted in a negligible number of false positives and its ability to detect true positives varied in relation to OTU isotopic enrichment and relative abundance. Sensitivity is predicted to improve with increases in sequencing effort. Generally, SIPSim predicts that the specificities of all DNA-SIP methods decline with increased beta-diversity among replicate samples. Thus, given that accuracy is predicted to decline most rapidly between a mean Bray-Curtis distance of 0 and 0.2 for all methods evaluated (Figure <xref ref-type="fig" rid="F7">7</xref>), we recommend that researchers strive for mean Bray-Curtis distances of &#x0003C;0.2 among replicate samples used in SIP experiments (i.e., between treatments and their corresponding controls).</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>NY developed the code for SIPSim, performed and analyzed all simulations, developed the figures and wrote the manuscript. SB developed the approach to modeling diffusive boundary layers in equilibrium density gradients. DB conceived the simulation approach, supervised and directed the research, and edited the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>We thank Chuck Pepe-Ranney for helpful discussions on the modeling approach used in this work.</p>
</ack>
<sec sec-type="supplementary-material" id="s7">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmicb.2018.00570/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmicb.2018.00570/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Presentation1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andeer</surname> <given-names>P.</given-names></name> <name><surname>Strand</surname> <given-names>S. E.</given-names></name> <name><surname>Stahl</surname> <given-names>D. A.</given-names></name></person-group> (<year>2012</year>). <article-title>High-sensitivity stable-isotope probing by a quantitative terminal restriction fragment length polymorphism protocol</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>78</volume>, <fpage>163</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.05973-11</pub-id><pub-id pub-id-type="pmid">22038597</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benson</surname> <given-names>D. A.</given-names></name> <name><surname>Karsch-Mizrachi</surname> <given-names>I.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name> <name><surname>Ostell</surname> <given-names>J.</given-names></name> <name><surname>Wheeler</surname> <given-names>D. L.</given-names></name></person-group> (<year>2008</year>). <article-title>GenBank</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>D25</fpage>&#x02013;<lpage>D30</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkp1024</pub-id><pub-id pub-id-type="pmid">18073190</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Birnie</surname> <given-names>G. D.</given-names></name> <name><surname>Rickwood</surname> <given-names>D.</given-names></name></person-group> (<year>1978</year>). <source>Centrifugal Separations in Molecular and Cell Biology.</source> <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Butterworths</publisher-name>.</citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buckley</surname> <given-names>D. H.</given-names></name> <name><surname>Huangyutitham</surname> <given-names>V.</given-names></name> <name><surname>Hsu</surname> <given-names>S.-F.</given-names></name> <name><surname>Nelson</surname> <given-names>T. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Stable isotope probing with <sup>15</sup>N achieved by disentangling the effects of genome G&#x0002B;C content and isotope enrichment on DNA density</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>73</volume>, <fpage>3189</fpage>&#x02013;<lpage>3195</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.02609-06</pub-id><pub-id pub-id-type="pmid">17369331</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clay</surname> <given-names>O.</given-names></name> <name><surname>Douady</surname> <given-names>C. J.</given-names></name> <name><surname>Carels</surname> <given-names>N.</given-names></name> <name><surname>Hughes</surname> <given-names>S.</given-names></name> <name><surname>Bucciarelli</surname> <given-names>G.</given-names></name> <name><surname>Bernardi</surname> <given-names>G.</given-names></name></person-group> (<year>2003</year>). <article-title>Using analytical ultracentrifugation to study compositional variation in vertebrate genomes</article-title>. <source>Eur. Biophys. J.</source> <volume>32</volume>, <fpage>418</fpage>&#x02013;<lpage>426</lpage>. <pub-id pub-id-type="doi">10.1007/s00249-003-0294-y</pub-id><pub-id pub-id-type="pmid">12684711</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Cohen</surname> <given-names>I. M.</given-names></name> <name><surname>Dowling</surname> <given-names>D. R.</given-names></name></person-group> (eds.). (<year>2012</year>). <article-title>Boundary layers and related topics</article-title>, in <source>Fluid Mechanics, 5th Edn</source>., ed <person-group person-group-type="editor"><name><surname>Kundu Pijush</surname> <given-names>K.</given-names></name></person-group> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>361</fpage>&#x02013;<lpage>419</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cupples</surname> <given-names>A. M.</given-names></name> <name><surname>Shaffer</surname> <given-names>E. A.</given-names></name> <name><surname>Chee-Sanford</surname> <given-names>J. C.</given-names></name> <name><surname>Sims</surname> <given-names>G. K.</given-names></name></person-group> (<year>2007</year>). <article-title>DNA buoyant density shifts during <sup>15</sup>N-DNA stable isotope probing</article-title>. <source>Microbiol. Res.</source> <volume>162</volume>, <fpage>328</fpage>&#x02013;<lpage>334</lpage>. <pub-id pub-id-type="doi">10.1016/j.micres.2006.01.016</pub-id><pub-id pub-id-type="pmid">16563712</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>W. D.</given-names></name> <name><surname>Cline</surname> <given-names>G. B.</given-names></name> <name><surname>Anderson</surname> <given-names>N. G.</given-names></name></person-group> (<year>1964</year>). <article-title>Density gradient centrifugation in angle-head rotors</article-title>. <source>Anal. Biochem.</source> <volume>9</volume>, <fpage>477</fpage>&#x02013;<lpage>482</lpage>. <pub-id pub-id-type="pmid">14239485</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Flamm</surname> <given-names>W. G.</given-names></name> <name><surname>Bond</surname> <given-names>H. E.</given-names></name> <name><surname>Burr</surname> <given-names>H. E.</given-names></name></person-group> (<year>1966</year>). <article-title>Density-Gradient centrifugation of DNA in a fixed-angle rotor</article-title>. <source>Biochim. Biophys. Acta</source> <volume>129</volume>, <fpage>310</fpage>&#x02013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.1016/0005-2787(66)90373-X</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fritsch</surname> <given-names>A.</given-names></name></person-group> (<year>1975</year>). <source>Preparative Density Gradient Centrifugations</source>. <publisher-loc>Geneva, SA</publisher-loc>: <publisher-name>Beckman Instrument International</publisher-name></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hearst</surname> <given-names>J. E.</given-names></name> <name><surname>Schmid</surname> <given-names>C. W.</given-names></name></person-group> (<year>1973</year>). <article-title>Density gradient sedimentation equilibrium</article-title>, in <source>Methods in Enzymology</source>, ed <person-group person-group-type="editor"><name><surname>Hirs</surname> <given-names>S. N. T.</given-names></name></person-group> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>111</fpage>&#x02013;<lpage>127</lpage>.</citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hungate</surname> <given-names>B. A.</given-names></name> <name><surname>Mau</surname> <given-names>R. L.</given-names></name> <name><surname>Schwartz</surname> <given-names>E.</given-names></name> <name><surname>Caporaso</surname> <given-names>J. G.</given-names></name> <name><surname>Dijkstra</surname> <given-names>P.</given-names></name> <name><surname>van Gestel</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Quantitative microbial ecology through stable isotope probing</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>81</volume>, <fpage>7570</fpage>&#x02013;<lpage>7581</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.02280-15</pub-id><pub-id pub-id-type="pmid">26296731</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kauffmann</surname> <given-names>I. M.</given-names></name> <name><surname>Schmitt</surname> <given-names>J.</given-names></name> <name><surname>Schmid</surname> <given-names>R. D.</given-names></name></person-group> (<year>2004</year>). <article-title>DNA isolation from soil samples for cloning in different hosts</article-title>. <source>Appl. Microbiol. Biotechnol.</source> <volume>64</volume>, <fpage>665</fpage>&#x02013;<lpage>670</lpage>. <pub-id pub-id-type="doi">10.1007/s00253-003-1528-8</pub-id><pub-id pub-id-type="pmid">14758515</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leigh</surname> <given-names>M. B.</given-names></name> <name><surname>Pellizari</surname> <given-names>V. H.</given-names></name> <name><surname>Uhl&#x000ED;k</surname> <given-names>O.</given-names></name> <name><surname>Sutka</surname> <given-names>R.</given-names></name> <name><surname>Rodrigues</surname> <given-names>J.</given-names></name> <name><surname>Ostrom</surname> <given-names>N. E.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Biphenyl-utilizing bacteria and their functional genes in a pine root zone contaminated with polychlorinated biphenyls (PCBs)</article-title>. <source>ISME J.</source> <volume>1</volume>, <fpage>134</fpage>&#x02013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1038/ismej.2007.26</pub-id><pub-id pub-id-type="pmid">18043623</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lennon</surname> <given-names>J. T.</given-names></name> <name><surname>Jones</surname> <given-names>S. E.</given-names></name></person-group> (<year>2011</year>). <article-title>Microbial seed banks: the ecological and evolutionary implications of dormancy</article-title>. <source>Nat. Rev. Microbiol.</source> <volume>9</volume>, <fpage>119</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro2504</pub-id><pub-id pub-id-type="pmid">21233850</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Love</surname> <given-names>M. I.</given-names></name> <name><surname>Huber</surname> <given-names>W.</given-names></name> <name><surname>Anders</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2</article-title>. <source>Genome Biol.</source> <volume>15</volume>:<fpage>550</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-014-0550-8</pub-id><pub-id pub-id-type="pmid">25516281</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lueders</surname> <given-names>T.</given-names></name> <name><surname>Manefield</surname> <given-names>M.</given-names></name> <name><surname>Friedrich</surname> <given-names>M. W.</given-names></name></person-group> (<year>2004</year>). <article-title>Enhanced sensitivity of DNA- and rRNA-based stable isotope probing by fractionation and quantitative analysis of isopycnic centrifugation gradients</article-title>. <source>Environ. Microbiol.</source> <volume>6</volume>, <fpage>73</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1046/j.1462-2920.2003.00536.x</pub-id><pub-id pub-id-type="pmid">14686943</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McMurdie</surname> <given-names>P. J.</given-names></name> <name><surname>Holmes</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Waste not, want not: why rarefying microbiome data is inadmissible</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003531</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003531</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meselson</surname> <given-names>M.</given-names></name> <name><surname>Stahl</surname> <given-names>F. W.</given-names></name> <name><surname>Vinograd</surname> <given-names>J.</given-names></name></person-group> (<year>1957</year>). <article-title>Equilibrium sedimentation of macromolecules in density gradients</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>43</volume>, <fpage>581</fpage>&#x02013;<lpage>588</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.43.7.581</pub-id><pub-id pub-id-type="pmid">16590059</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neufeld</surname> <given-names>J. D.</given-names></name> <name><surname>Vohra</surname> <given-names>J.</given-names></name> <name><surname>Dumont</surname> <given-names>M. G.</given-names></name> <name><surname>Lueders</surname> <given-names>T.</given-names></name> <name><surname>Manefield</surname> <given-names>M.</given-names></name> <name><surname>Friedrich</surname> <given-names>M. W.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>DNA stable-isotope probing</article-title>. <source>Nat. Protoc.</source> <volume>2</volume>, <fpage>860</fpage>&#x02013;<lpage>866</lpage>. <pub-id pub-id-type="doi">10.1038/nprot.2007.109</pub-id><pub-id pub-id-type="pmid">17446886</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paulson</surname> <given-names>J. N.</given-names></name> <name><surname>Stine</surname> <given-names>O. C.</given-names></name> <name><surname>Bravo</surname> <given-names>H. C.</given-names></name> <name><surname>Pop</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Differential abundance analysis for microbial marker-gene surveys</article-title>. <source>Nat. Methods</source> <volume>10</volume>, <fpage>1200</fpage>&#x02013;<lpage>1202</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2658</pub-id><pub-id pub-id-type="pmid">24076764</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pepe-Ranney</surname> <given-names>C.</given-names></name> <name><surname>Campbell</surname> <given-names>A. N.</given-names></name> <name><surname>Koechli</surname> <given-names>C. N.</given-names></name> <name><surname>Berthrong</surname> <given-names>S.</given-names></name> <name><surname>Buckley</surname> <given-names>D. H.</given-names></name></person-group> (<year>2016a</year>). <article-title>Unearthing the ecology of soil microorganisms using a high resolution DNA-SIP approach to explore cellulose and xylose metabolism in soil</article-title>. <source>Front. Microbiol.</source> <volume>7</volume>:<fpage>703</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2016.00703</pub-id><pub-id pub-id-type="pmid">27242725</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pepe-Ranney</surname> <given-names>C.</given-names></name> <name><surname>Koechli</surname> <given-names>C.</given-names></name> <name><surname>Potrafka</surname> <given-names>R.</given-names></name> <name><surname>Andam</surname> <given-names>C.</given-names></name> <name><surname>Eggleston</surname> <given-names>E.</given-names></name> <name><surname>Garcia-Pichel</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2016b</year>). <article-title>Non-cyanobacterial diazotrophs mediate dinitrogen fixation in biological soil crusts during early crust formation</article-title>. <source>ISME J.</source> <volume>10</volume>, <fpage>287</fpage>&#x02013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1038/ismej.2015.106</pub-id><pub-id pub-id-type="pmid">26114889</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qu</surname> <given-names>W.</given-names></name> <name><surname>Shen</surname> <given-names>Z.</given-names></name> <name><surname>Zhao</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>MFEprimer: multiple factor evaluation of the specificity of PCR primers</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>276</fpage>&#x02013;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btn614</pub-id><pub-id pub-id-type="pmid">19038987</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radajewski</surname> <given-names>S.</given-names></name> <name><surname>McDonald</surname> <given-names>I. R.</given-names></name> <name><surname>Murrell</surname> <given-names>J. C.</given-names></name></person-group> (<year>2003</year>). <article-title>Stable-isotope probing of nucleic acids: a window to the function of uncultured microorganisms</article-title>. <source>Curr. Opin. Biotechnol.</source> <volume>14</volume>, <fpage>296</fpage>&#x02013;<lpage>302</lpage>. <pub-id pub-id-type="doi">10.1016/S.0958-1669(03)00064-8</pub-id><pub-id pub-id-type="pmid">12849783</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><collab>R Core Team</collab></person-group> (<year>2016</year>). <source>R: A Language and Environment for Statistical Computing R Foundation for Statistical Computing</source>, <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Core Team</publisher-name>.</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roh</surname> <given-names>C.</given-names></name> <name><surname>Villatte</surname> <given-names>F.</given-names></name> <name><surname>Kim</surname> <given-names>B.-G.</given-names></name> <name><surname>Schmid</surname> <given-names>R. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Comparative study of methods for extraction and purification of environmental DNA from soil and sludge samples</article-title>. <source>Appl. Biochem. Biotechnol.</source> <volume>134</volume>, <fpage>97</fpage>&#x02013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1385/ABAB:134:2:97</pub-id><pub-id pub-id-type="pmid">16943632</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schildkraut</surname> <given-names>C. L.</given-names></name> <name><surname>Marmur</surname> <given-names>J.</given-names></name> <name><surname>Doty</surname> <given-names>P.</given-names></name></person-group> (<year>1962</year>). <article-title>Determination of the base composition of deoxyribonucleic acid from its buoyant density in CsCl</article-title>. <source>J. Mol. Biol.</source> <volume>4</volume>, <fpage>430</fpage>&#x02013;<lpage>443</lpage>. <pub-id pub-id-type="pmid">14498379</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmid</surname> <given-names>C. W.</given-names></name> <name><surname>Hearst</surname> <given-names>J. E.</given-names></name></person-group> (<year>1972</year>). <article-title>Sedimentation equilibrium of DNA samples heterogeneous in density</article-title>. <source>Biopolymers</source> <volume>11</volume>, <fpage>1913</fpage>&#x02013;<lpage>1918</lpage>. <pub-id pub-id-type="doi">10.1002/bip.1972.360110911</pub-id><pub-id pub-id-type="pmid">5072737</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Stubben</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <source>Genomes: Genome Sequencing Project Metadata. R Package Version 3.6.0</source>. <pub-id pub-id-type="doi">10.18129/B9.bioc.genomes</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suzuki</surname> <given-names>M. T.</given-names></name> <name><surname>Giovannoni</surname> <given-names>S. J.</given-names></name></person-group> (<year>1996</year>). <article-title>Bias caused by template annealing in the amplification of mixtures of 16S rRNA genes by PCR</article-title>. <source>Appl. Environ. Microbiol</source>. <volume>62</volume>, <fpage>625</fpage>&#x02013;<lpage>630</lpage>. <pub-id pub-id-type="pmid">8593063</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thakuria</surname> <given-names>D.</given-names></name> <name><surname>Schmidt</surname> <given-names>O.</given-names></name> <name><surname>Mac Si&#x000FA;rt&#x000E1;in</surname> <given-names>M.</given-names></name> <name><surname>Egan</surname> <given-names>D.</given-names></name> <name><surname>Doohan</surname> <given-names>F. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Importance of DNA quality in comparative soil microbial community structure analyses</article-title>. <source>Soil Biol. Biochem.</source> <volume>40</volume>, <fpage>1390</fpage>&#x02013;<lpage>1403</lpage>. <pub-id pub-id-type="doi">10.1016/j.soilbio.2007.12.027</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tritton</surname> <given-names>D. J.</given-names></name></person-group> (<year>1977</year>). <article-title>Boundary layers and related topics</article-title>, in <source>Physical Fluid Dynamics</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Van Nostrand Reinhold Company, Ltd.</publisher-name>), <fpage>101</fpage>&#x02013;<lpage>118</lpage>.</citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Uhl&#x000ED;k</surname> <given-names>O.</given-names></name> <name><surname>Jecn&#x000E1;</surname> <given-names>K.</given-names></name> <name><surname>Leigh</surname> <given-names>M. B.</given-names></name> <name><surname>Mackov&#x000E1;</surname> <given-names>M.</given-names></name> <name><surname>Macek</surname> <given-names>T.</given-names></name></person-group> (<year>2009</year>). <article-title>DNA-based stable isotope probing: a link between community structure and function</article-title>. <source>Sci. Total Environ.</source> <volume>407</volume>, <fpage>3611</fpage>&#x02013;<lpage>3619</lpage>. <pub-id pub-id-type="doi">10.1016/j.scitotenv.2008.05.012</pub-id><pub-id pub-id-type="pmid">18573518</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wawrik</surname> <given-names>B.</given-names></name> <name><surname>Callaghan</surname> <given-names>A. V.</given-names></name> <name><surname>Bronk</surname> <given-names>D. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Use of inorganic and organic nitrogen by <italic>Synechococcus</italic> spp. and diatoms on the west florida shelf as measured using stable isotope probing</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>75</volume>, <fpage>6662</fpage>&#x02013;<lpage>6670</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.01002-09</pub-id><pub-id pub-id-type="pmid">19734334</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Youngblut</surname> <given-names>N. D.</given-names></name> <name><surname>Buckley</surname> <given-names>D. H.</given-names></name></person-group> (<year>2014</year>). <article-title>Intra-genomic variation in G &#x0002B; C content and its implications for DNA stable isotope probing</article-title>. <source>Environ. Microbiol. Rep.</source> <volume>6</volume>, <fpage>767</fpage>&#x02013;<lpage>775</lpage>. <pub-id pub-id-type="doi">10.1111/1758-2229.12201</pub-id><pub-id pub-id-type="pmid">25139123</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This material is based upon work supported by the Department of Energy, Office of Biological and Environmental Research Genomic Science Program under Award Numbers DE-SC0010558 and DE-SC0004486.</p>
</fn>
</fn-group>
</back>
</article>