<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">855076</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.855076</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Cell Type Diversity Statistic: An Entropy-Based Metric to Compare Overall Cell Type Composition Across Samples</article-title>
<alt-title alt-title-type="left-running-head">Karagiannis et al.</alt-title>
<alt-title alt-title-type="right-running-head">Cell Type Diversity Statistic</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Karagiannis</surname>
<given-names>Tanya T</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/650365/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Monti</surname>
<given-names>Stefano</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/61455/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sebastiani</surname>
<given-names>Paola</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/38742/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Institute for Clinical Research and Health Policy Studies</institution>, <institution>Tufts Medical Center</institution>, <addr-line>Boston</addr-line>, <addr-line>MA</addr-line>, <country>United States</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Bioinformatics Program</institution>, <institution>Boston University</institution>, <addr-line>Boston</addr-line>, <addr-line>MA</addr-line>, <country>United States</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Division of Computational Biomedicine</institution>, <institution>Boston University School of Medicine</institution>, <addr-line>Boston</addr-line>, <addr-line>MA</addr-line>, <country>United States</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Department of Biostatistics</institution>, <institution>Boston University School of Public Health</institution>, <addr-line>Boston</addr-line>, <addr-line>MA</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/552766/overview">Tao Huang</ext-link>, Shanghai Institute of Nutrition and Health, (CAS), China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1645243/overview">Cong Liang</ext-link>, Tianjin University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1651250/overview">Alsu Missarova</ext-link>, European Bioinformatics Institute (EMBL-EBI), United Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Tanya T Karagiannis, <email>tkaragiannis@tuftsmedicalcenter.org</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>855076</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Karagiannis, Monti and Sebastiani.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Karagiannis, Monti and Sebastiani</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Changes of cell type composition across samples can carry biological significance and provide insight into disease and other conditions. Single cell transcriptomics has made it possible to study cell type composition at a fine resolution. Most single cell studies investigate compositional changes between samples for each cell type independently, not accounting for the fixed number of cells per sample in sequencing data. Here, we provide a metric of the distribution of cell type proportions in a sample that can be used to compare the overall distribution of cell types across multiple samples and biological conditions. This is the first method to measure overall cell type composition at the single cell level. We use the method to assess compositional changes in peripheral blood mononuclear cells (PBMCs) related to aging and extreme old age using multiple single cell datasets from individuals of four age groups across the human lifespan.</p>
</abstract>
<kwd-group>
<kwd>single cell transcriptomic analysis</kwd>
<kwd>cell type composition</kwd>
<kwd>sample level analysis</kwd>
<kwd>sample-to-sample comparison</kwd>
<kwd>diversity statistics</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Tissues are composed of heterogenous cell types that demonstrate differences in biological function (<xref ref-type="bibr" rid="B22">Raj and van Oudenaarden, 2008</xref>; <xref ref-type="bibr" rid="B9">Choi and Kim, 2019</xref>). Gene expression profiling methods such as single cell RNA-sequencing (scRNA-seq) have made it possible to profile the genome-wide gene expression levels for each single cell of a sample, to account for cell-to-cell variability (<xref ref-type="bibr" rid="B8">Chen et al., 2019</xref>; <xref ref-type="bibr" rid="B27">Tanay and Regev, 2017</xref>; <xref ref-type="bibr" rid="B9">Choi and Kim, 2019</xref>), and to identify and characterize cell types in a given tissue (<xref ref-type="bibr" rid="B15">Jaitin et al., 2014</xref>; <xref ref-type="bibr" rid="B18">Macosko et al., 2015</xref>; <xref ref-type="bibr" rid="B33">Zheng et al., 2017</xref>). ScRNA-seq has been extensively applied in multiple research areas to study cell types and states, as well as cell types compositional changes, across diseases and conditions (<xref ref-type="bibr" rid="B25">Shalek et al., 2014</xref>; <xref ref-type="bibr" rid="B4">Baron et al., 2016</xref>; <xref ref-type="bibr" rid="B20">Muraro et al., 2016</xref>; <xref ref-type="bibr" rid="B30">Villani et al., 2017</xref>; <xref ref-type="bibr" rid="B5">Butler et al., 2018</xref>; <xref ref-type="bibr" rid="B23">Schaum et al., 2018</xref>; <xref ref-type="bibr" rid="B19">Mathys et al., 2019</xref>; <xref ref-type="bibr" rid="B29">Velmeshev et al., 2019</xref>).</p>
<p>Most methods to analyze cell type composition at a single cell level model each cell type independently from other cell types (<xref ref-type="bibr" rid="B13">Haber et al., 2017</xref>; <xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>; <xref ref-type="bibr" rid="B14">Hashimoto et al., 2019</xref>; <xref ref-type="bibr" rid="B32">Wilk et al., 2020</xref>; <xref ref-type="bibr" rid="B34">Zheng et al., 2020</xref>; <xref ref-type="bibr" rid="B35">Zhu et al., 2020</xref>). For example, changes of peripheral blood mononuclear cells (PBMCs) composition observed between supercentenarians and younger age controls in <xref ref-type="bibr" rid="B14">Hashimoto et al., 2019</xref> were assessed for each cell type <italic>independently</italic> using a Wilcoxon rank sum test. Other studies have taken a similar approach when assessing compositional changes between groups of samples at the single cell level (<xref ref-type="bibr" rid="B13">Haber et al., 2017</xref>; <xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>; <xref ref-type="bibr" rid="B14">Hashimoto et al., 2019</xref>; <xref ref-type="bibr" rid="B32">Wilk et al., 2020</xref>; <xref ref-type="bibr" rid="B34">Zheng et al., 2020</xref>; <xref ref-type="bibr" rid="B35">Zhu et al., 2020</xref>). However, high throughput sequencing data are in fact compositional (<xref ref-type="bibr" rid="B12">Gloor et al., 2016</xref>, <xref ref-type="bibr" rid="B11">2017</xref>; <xref ref-type="bibr" rid="B16">Lin and Peddada, 2020</xref>). The approach we propose rests on the observation that a sample in scRNA-seq data is composed of cell abundances across cell types that are in constrained proportions, given the total number of cells in the sample (<xref ref-type="bibr" rid="B12">Gloor et al., 2016</xref>; <xref ref-type="bibr" rid="B11">Gloor et al., 2017</xref>; <xref ref-type="bibr" rid="B16">Lin and Peddada, 2020</xref>). In other words, the proportion of cell types within a sample are in fact dependent on each other: if the proportion of one type increases, then others need to decrease (<xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>). It is thus necessary to account for this dependency when assessing overall cell type compositional changes across samples. In addition, there is no method that provides a numerical summary of a sample overall cell type composition that can be used to compare samples in different conditions (<xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>).</p>
<p>Here, we introduce a statistic to summarize the distribution of the proportions of cell types in a sample. Using three single cell transcriptomic datasets of PBMCs comprising four age groups, we show the utility of this statistic to describe changes in PBMCs composition in aging and extreme old age.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>Materials and Methods</title>
<p>
<bold>Cell type diversity statistic.</bold> The statistic makes three assumptions: 1) To make different samples of cells comparable, cell abundances must be normalized based on the total number of cells in a sample; 2) After conditioning on the total number of cells in a sample (<xref ref-type="bibr" rid="B11">Gloor et al., 2017</xref>), the cell type composition data is a simplex (<xref ref-type="bibr" rid="B2">Aitchison, 1982</xref>), and when the proportion of one cell type changes, the proportion of the other cell types must change as well to maintain the total fixed; and 3) To make the statistic comparable across different cell type resolutions, the statistic must be normalized. Formally, we denote by <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> the proportion of cell type <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> in a sample s with <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> cells, so that <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1.</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</p>
<p>The statistic is adapted from alpha diversity measures applied in ecology and microbiome studies (<xref ref-type="bibr" rid="B31">Whittaker, 1972</xref>; <xref ref-type="bibr" rid="B21">Olde Loohuis et al., 2018</xref>; <xref ref-type="bibr" rid="B7">Calle, 2019</xref>). We measure the overall cell type composition of a sample by the adjusted entropy<disp-formula id="equ1">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>In the formula, <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:mtext>log</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the maximum value of <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> that is reached when <inline-formula id="inf7">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>k</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> for all indexes <inline-formula id="inf8">
<mml:math id="m9">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> so that the distribution is uniform. The minimum value of <inline-formula id="inf9">
<mml:math id="m10">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is 0, which corresponds to a mass-point distribution with <inline-formula id="inf10">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for all indexes <inline-formula id="inf11">
<mml:math id="m12">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> but one. The adjusted entropy <inline-formula id="inf12">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> therefore ranges between <inline-formula id="inf13">
<mml:math id="m14">
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. A sample with more uniformity in cell type proportions, and hence more variability, will result in a greater cell type diversity statistic and <inline-formula id="inf14">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> in a sample with equal proportions of all cell types. A sample with cell type proportions that are skewed towards specific cell types, and less variability, will have a lower statistic and <inline-formula id="inf15">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> when all cells are of one type.</p>
<p>
<bold>Data.</bold> To demonstrate the utility of the cell type diversity statistic, we analyzed three single cell transcriptomic datasets of PBMCs representing regular aging and extreme old age. One dataset comprised samples of 7 centenarians from the New England Centenarian Study (NECS) (<xref ref-type="bibr" rid="B24">Sebastiani and Perls, 2012</xref>) and 2 younger age controls. We downloaded a publicly available scRNA-seq dataset of PBMCs from 45 younger age controls (<xref ref-type="bibr" rid="B28">van der Wijst et al., 2018</xref>), which we will refer to as NATGEN, and a publicly available scRNA-seq dataset of PBMCs from 5 younger age controls and 7 supercentenarians, which we will refer to as PNAS (<xref ref-type="bibr" rid="B14">Hashimoto et al., 2019</xref>). We integrated these datasets and stratified the samples into four age groups of the human lifespan: 12 subjects of younger age (20&#x2013;39), 26 subjects of middle age (40&#x2013;59), 14 subjects of older age (60&#x2013;89), and 14 subjects of extreme longevity (100&#x2013;119). Data processing steps and identification of the 12 cell types are described in the Supplement.</p>
<p>
<bold>Application of cell type diversity statistic.</bold> We integrated the datasets to generate a matrix of cell type abundances across samples from all three datasets. We calculated the cell type proportions for each sample such that the sum of the cell type proportions for a particular sample equals to 1. We applied the cell type diversity statistic to different cell type resolutions: 1) based on the proportions of lymphocytes and myeloid cells; and 2) based on the proportions of the 12 lymphocyte and myeloid subpopulations that were detected in the data. For both resolutions, we measured the cell type diversity statistic per sample and compared the differences of the statistics between the four age groups using ANOVA and pairwise T-tests with significance level 0.05.</p>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>Results and Discussion</title>
<p>We applied the cell type diversity statistic to the cell type proportions from the three scRNA-seq datasets of younger age individuals and centenarians to assess overall compositional changes across four age groups: younger age (20&#x2013;39), middle age (40&#x2013;59), older age (60&#x2013;89), and extreme old age (100&#x2013;119&#xa0;years of age). We first calculated the cell type proportions for each sample across the four age groups (<xref ref-type="fig" rid="F1">Figure 1A</xref>, <xref ref-type="sec" rid="s11">Supplementary Table S1</xref>) and we observed a shift in the distribution of cell proportions from lymphocyte and myeloid cell types from younger ages to centenarians (<xref ref-type="fig" rid="F1">Figure 1A</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Cell type diversity statistic to summarize PBMCs composition across age groups. <bold>(A)</bold>. Proportions of 12 cell types discovered in scRNA-seq of PBMCs from different age groups. Each bar represents the proportions of lymphocyte (blue-green gradient) and myeloid (red-yellow gradient) cell types (<italic>y</italic>-axis) in a sample. <bold>(B)</bold>. Each boxplot represents the distribution of the diversity statistic of the proportions of lymphocyte and myeloid cells in younger, middle, older, and extreme old age individuals (<italic>x</italic>-axis). The differences of the statistics across age groups were statistically significant (F-test <italic>p</italic>-value &#x3d; 0.0001873) <bold>(C)</bold>. Each boxplot represents the distribution of the diversity statistic of the proportions of the 12 cell types grouped by younger, middle, older, and extreme old age (<italic>x</italic>-axis). The differences of the statistics across age groups were statistically significant (F-test <italic>p</italic>-value &#x3d; 0.0001875). The diversity statistic was significantly higher, in the extreme old age group compared to each younger age control group: younger age group (t-test <italic>p</italic>-value &#x3d; 0.00115), middle age group (t-test <italic>p</italic>-value &#x3d; 0.00016), and older age group (t-test <italic>p</italic>-value &#x3d; 0.00363).</p>
</caption>
<graphic xlink:href="fgene-13-855076-g001.tif"/>
</fig>
<p>We then calculated the cell type diversity statistic to measure the variability of the proportion of lymphocyte and myeloid cells in each sample (<xref ref-type="sec" rid="s11">Supplementary Table S2</xref>). Comparing the cell type diversity statistics across the four age groups, we found a significant difference in the distribution of the statistics across the four age groups (F-test <italic>p</italic>-value &#x3d; 0.0001873) (<xref ref-type="fig" rid="F1">Figure 1B</xref>). The increased value of the cell type diversity statistic in the extreme old age group is consistent with the shift in abundances from lymphocytes to myeloid cells, which is an expected change in the immune system with aging (<xref ref-type="bibr" rid="B10">Geiger et al., 2013</xref>). We also applied the cell type diversity statistic to measure the variability of the proportions of 12 lymphocyte and myeloid subpopulations in each sample (<xref ref-type="sec" rid="s11">Supplementary Table S3</xref>). We again found a significant difference in the distribution of the statistic in the four age groups (F-test <italic>p</italic>-value &#x3d; 0.0001875) (<xref ref-type="fig" rid="F1">Figure 1C</xref>). Specifically, centenarians had significantly increased cell type diversity statistics compared to each younger age control group: younger age group (t-test <italic>p</italic>-value &#x3d; 0.00115), middle age group (t-test <italic>p</italic>-value &#x3d; 0.00016), and older age group (t-test <italic>p</italic>-value &#x3d; 0.00363) (<xref ref-type="fig" rid="F1">Figure 1C</xref>). The pattern of the cell type diversity with age groups suggests that centenarians have a more uniform distribution of cell types compared to individuals of younger ages even at a finer resolution of cell types.</p>
<p>The analyses illustrate how the cell type diversity statistic can be used in combination with visualizations of cell type proportions to provide a numerical summary of the distribution of cell types in different conditions. We showed an application of this metric in the context of aging to summarize changes of the distribution of cell types across different age groups, at different resolutions. The metric showed a significant change of the distribution of 12 cell types in extreme old age compared to younger age groups, as well as a significant change of the proportion of lymphocytes and myeloid cells that are biologically relevant to aging (<xref ref-type="bibr" rid="B10">Geiger et al., 2013</xref>). Although in our analysis the distribution of the cell type diversity statistics did not change with different cell type resolutions, in other applications the statistic could change since the distribution of the proportions of subpopulations of cells can be very different.</p>
<p>One major challenge in the analysis of single cell transcriptomics data is in the identification and annotation of cell types. There are varying methods to identify cell types (<xref ref-type="bibr" rid="B3">Andrews et al., 2021</xref>; <xref ref-type="bibr" rid="B1">Adil et al., 2021</xref>; <xref ref-type="bibr" rid="B26">Shekhar and Menon, 2019</xref>; <xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>) and the resolution of cell type for analysis should be selected based on the biological question of interest (<xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>). Another challenge of this type of analyses is accounting for cell types that are not detectable under specific conditions. Other metrics are needed to account for cell types that are not detected in all conditions.</p>
<p>The cell type diversity statistic is applied as a global summary of cell type composition, and additional analyses are required to quantify individual cell type changes and to adjust this analysis for additional covariates. The recent method scCoda uses a Bayesian Dirichlet regression model to examine individuals cell type changes and accounts for the constrained proportions in single cell composition data is particularly promising (<xref ref-type="bibr" rid="B6">B&#xfc;ttner et al., 2021</xref>).</p>
<p>Entropy as a metric to study composition level data has been applied in many fields including analyses of microbiome data (<xref ref-type="bibr" rid="B31">Whittaker, 1972</xref>; <xref ref-type="bibr" rid="B21">Olde Loohuis et al., 2018</xref>; <xref ref-type="bibr" rid="B7">Calle, 2019</xref>). The importance in applying this metric to single cell transcriptomics is that it accounts for the constrained proportions of cell types in each sample, and ignoring these constraints can results in inconsistencies when assessing compositional changes (<xref ref-type="bibr" rid="B12">Gloor et al., 2016</xref>; <xref ref-type="bibr" rid="B11">Gloor et al., 2017</xref>; <xref ref-type="bibr" rid="B7">Calle, 2019</xref>; <xref ref-type="bibr" rid="B17">Luecken and Theis, 2019</xref>).</p>
</sec>
<sec sec-type="conclusion" id="s4">
<title>Conclusion</title>
<p>We present the cell type diversity statistic, an entropy-based measure to assess and summarize the overall cell type composition of samples in single cell gene expression data. The diversity statistic allows for the investigation of global cell type compositional changes applicable to studying disease and other conditions at the single cell level. We demonstrate the utility of this method by its application to single cell datasets of aging and extreme old age, and show that it can reveal novel changes in composition in aging at different resolutions.</p>
</sec>
</body>
<back>
<sec id="s12">
<title>Code Availability Statement</title>
<p>The cell type diversity statistic is available as a function in R at <ext-link ext-link-type="uri" xlink:href="https://github.com/tanya-karagiannis/Cell-Type-Diversity-Statistic">https://github.com/tanya-karagiannis/Cell-Type-Diversity-Statistic</ext-link>. The function can be applied to a matrix of cell type proportions per sample, a Seurat object, and a Single Cell Experiment object.</p>
</sec>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: The data that support these findings are publicly available and were accessed from several repositories. NATGEN single cell expression data and subject level data were publicly available as referenced in (<xref ref-type="bibr" rid="B28">van der Wijst et al., 2018</xref>): <ext-link ext-link-type="uri" xlink:href="https://molgenis58.target.rug.nl/scrna-seq/">https://molgenis58.target.rug.nl/scrna-seq/</ext-link>. PNAS single cell expression data and subject level data was available as referenced in (<xref ref-type="bibr" rid="B14">Hashimoto et al., 2019</xref>): <ext-link ext-link-type="uri" xlink:href="http://gerg.gsc.riken.jp/SC2018/">http://gerg.gsc.riken.jp/SC2018/</ext-link>. NECS will be available from Synapse (URL <ext-link ext-link-type="uri" xlink:href="https://adknowledgeportal.synapse.org/Explore/Projects/DetailsPage?Grant%20Number=UH2AG064704">https://adknowledgeportal.synapse.org/Explore/Projects/DetailsPage?Grant%20Number&#x003D;UH2AG064704</ext-link>).</p>
</sec>
<sec id="s6">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by Boston University IRB. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>TK, PS, and SM conceived of the presented method for single cell transcriptomics data. TK implemented the method and wrote the paper with feedback from all authors. All authors contributed to the final version of the manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by NIH-NIA UH2AG064704.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s11">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.855076/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2022.855076/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.docx" id="SM1" mimetype="application/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="DataSheet2.XLSX" id="SM2" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adil</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Jan</surname>
<given-names>A. T.</given-names>
</name>
<name>
<surname>Asger</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Single-Cell Transcriptomics: Current Methods and Challenges in Data Acquisition and Analysis</article-title>. <source>Front. Neurosci.</source> <volume>15</volume>. <pub-id pub-id-type="doi">10.3389/fnins.2021.591122</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aitchison</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1982</year>). <article-title>The Statistical Analysis of Compositional Data</article-title>. <source>J. R. Stat. Soc. Ser. B (Methodological)</source> <volume>44</volume>, <fpage>139</fpage>&#x2013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1111/j.2517-6161.1982.tb01195.x</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Andrews</surname>
<given-names>T. S.</given-names>
</name>
<name>
<surname>Kiselev</surname>
<given-names>V. Y.</given-names>
</name>
<name>
<surname>McCarthy</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hemberg</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Tutorial: Guidelines for the Computational Analysis of Single-Cell RNA Sequencing Data</article-title>. <source>Nat. Protoc.</source> <volume>16</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1038/s41596-020-00409-w</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baron</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Veres</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wolock</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Faust</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Gaujoux</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vetere</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell Population Structure</article-title>. <source>Cel Syst.</source> <volume>3</volume>, <fpage>346</fpage>&#x2013;<lpage>360</lpage>. <comment>e4</comment>. <pub-id pub-id-type="doi">10.1016/j.cels.2016.08.011</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Butler</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hoffman</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Smibert</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Papalexi</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Integrating Single-Cell Transcriptomic Data across Different Conditions, Technologies, and Species</article-title>. <source>Nat. Biotechnol.</source> <volume>36</volume>, <fpage>411</fpage>&#x2013;<lpage>420</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.4096</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>B&#xfc;ttner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ostner</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>M&#xfc;ller</surname>
<given-names>C. L.</given-names>
</name>
<name>
<surname>Theis</surname>
<given-names>F. J.</given-names>
</name>
<name>
<surname>Schubert</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>scCODA Is a Bayesian Model for Compositional Single-Cell Data Analysis</article-title>. <source>Nat. Commun.</source> <volume>12</volume>, <fpage>6876</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-27150-6</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Calle</surname>
<given-names>M. L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Statistical Analysis of Metagenomics Data</article-title>. <source>Genomics Inform.</source> <volume>17</volume>, <fpage>e6</fpage>. <pub-id pub-id-type="doi">10.5808/GI.2019.17.1.e6</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ning</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Single-Cell RNA-Seq Technologies and Related Computational Data Analysis</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <fpage>317</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00317</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Choi</surname>
<given-names>Y. H.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>J. K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Dissecting Cellular Heterogeneity Using Single-Cell RNA Sequencing</article-title>. <source>Mol. Cell</source> <volume>42</volume>, <fpage>189</fpage>&#x2013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.14348/molcells.2019.2446</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Geiger</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>de Haan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Florian</surname>
<given-names>M. C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>The Ageing Haematopoietic Stem Cell Compartment</article-title>. <source>Nat. Rev. Immunol.</source> <volume>13</volume>, <fpage>376</fpage>&#x2013;<lpage>389</lpage>. <pub-id pub-id-type="doi">10.1038/nri3433</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gloor</surname>
<given-names>G. B.</given-names>
</name>
<name>
<surname>Macklaim</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Pawlowsky-Glahn</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Egozcue</surname>
<given-names>J. J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Microbiome Datasets Are Compositional: And This Is Not Optional</article-title>. <source>Front. Microbiol.</source> <volume>8</volume>, <fpage>2224</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2017.02224</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gloor</surname>
<given-names>G. B.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Pawlowsky-Glahn</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Egozcue</surname>
<given-names>J. J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>It&#x27;s All Relative: Analyzing Microbiome Data as Compositions</article-title>. <source>Ann. Epidemiol.</source> <volume>26</volume>, <fpage>322</fpage>&#x2013;<lpage>329</lpage>. <pub-id pub-id-type="doi">10.1016/j.annepidem.2016.03.003</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haber</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Biton</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rogel</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Herbst</surname>
<given-names>R. H.</given-names>
</name>
<name>
<surname>Shekhar</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Smillie</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>A Single-Cell Survey of the Small Intestinal Epithelium</article-title>. <source>Nature</source> <volume>551</volume>, <fpage>333</fpage>&#x2013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1038/nature24489</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hashimoto</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kouno</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ikawa</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hayatsu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Miyajima</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yabukami</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Single-cell Transcriptomics Reveals Expansion of Cytotoxic CD4 T Cells in Supercentenarians</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>116</volume>, <fpage>24242</fpage>&#x2013;<lpage>24251</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1907883116</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jaitin</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Kenigsberg</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Keren-Shaul</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Elefant</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Paul</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zaretsky</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Massively Parallel Single-Cell RNA-Seq for Marker-free Decomposition of Tissues into Cell Types</article-title>. <source>Science</source> <volume>343</volume>, <fpage>776</fpage>&#x2013;<lpage>779</lpage>. <pub-id pub-id-type="doi">10.1126/science.1247651</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Peddada</surname>
<given-names>S. D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Analysis of Microbial Compositions: a Review of Normalization and Differential Abundance Analysis</article-title>. <source>npj Biofilms Microbiomes</source> <volume>6</volume>, <fpage>60</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1038/s41522-020-00160-w</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luecken</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Theis</surname>
<given-names>F. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Current Best Practices in Single-Cell RNA-Seq Analysis: a Tutorial</article-title>. <source>Mol. Syst. Biol.</source> <volume>15</volume>, <fpage>e8746</fpage>. <pub-id pub-id-type="doi">10.15252/msb.20188746</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Macosko</surname>
<given-names>E. Z.</given-names>
</name>
<name>
<surname>Basu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nemesh</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shekhar</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Highly Parallel Genome-wide Expression Profiling of Individual Cells Using Nanoliter Droplets</article-title>. <source>Cell</source> <volume>161</volume>, <fpage>1202</fpage>&#x2013;<lpage>1214</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.05.002</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mathys</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Davila-Velderrain</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mohammadi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>J. Z.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Single-cell Transcriptomic Analysis of Alzheimer&#x27;s Disease</article-title>. <source>Nature</source> <volume>570</volume>, <fpage>332</fpage>&#x2013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-1195-2</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muraro</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Dharmadhikari</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gr&#xfc;n</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Groen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Dielen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Jansen</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>A Single-Cell Transcriptome Atlas of the Human Pancreas</article-title>. <source>Cel Syst.</source> <volume>3</volume>, <fpage>385</fpage>&#x2013;<lpage>394</lpage>. <comment>e3</comment>. <pub-id pub-id-type="doi">10.1016/j.cels.2016.09.002</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olde Loohuis</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Mangul</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ori</surname>
<given-names>A. P. S.</given-names>
</name>
<name>
<surname>Jospin</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Koslicki</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H. T.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Transcriptome Analysis in Whole Blood Reveals Increased Microbial Diversity in Schizophrenia</article-title>. <source>Transl Psychiatry</source> <volume>8</volume>, <fpage>96</fpage>&#x2013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1038/s41398-018-0107-9</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Raj</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>van Oudenaarden</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Nature, Nurture, or Chance: Stochastic Gene Expression and its Consequences</article-title>. <source>Cell</source> <volume>135</volume>, <fpage>216</fpage>&#x2013;<lpage>226</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2008.09.050</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schaum</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Karkanias</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Neff</surname>
<given-names>N. F.</given-names>
</name>
<name>
<surname>May</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Quake</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Wyss-Coray</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Single-cell Transcriptomics of 20 Mouse Organs Creates a Tabula Muris</article-title>. <source>Nature</source> <volume>562</volume>, <fpage>367</fpage>&#x2013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0590-4</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sebastiani</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Perls</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The Genetics of Extreme Longevity: Lessons from the New England Centenarian Study</article-title>. <source>Front. Genet.</source> <volume>3</volume>. <pub-id pub-id-type="doi">10.3389/fgene.2012.00277</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shalek</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Shuga</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Trombetta</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Gennert</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Single-cell RNA-Seq Reveals Dynamic Paracrine Control of Cellular Variation</article-title>. <source>Nature</source> <volume>510</volume>, <fpage>363</fpage>&#x2013;<lpage>369</lpage>. <pub-id pub-id-type="doi">10.1038/nature13437</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Shekhar</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Menon</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Identification of Cell Types from Single-Cell Transcriptomic Data</article-title>,&#x201d; in <source>Computational Methods for Single-Cell Data Analysis, Methods in Molecular Biology</source>. Editor <person-group person-group-type="editor">
<name>
<surname>Yuan</surname>
<given-names>G.-C.</given-names>
</name>
</person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>45</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-9057-3_4</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tanay</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Regev</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Scaling Single-Cell Genomics from Phenomenology to Mechanism</article-title>. <source>Nature</source> <volume>541</volume>, <fpage>331</fpage>&#x2013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1038/nature21350</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>van der Wijst</surname>
<given-names>M. G. P.</given-names>
</name>
<name>
<surname>Brugge</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>de Vries</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Deelen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Swertz</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Franke</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Single-cell RNA Sequencing Identifies Celltype-specific Cis-eQTLs and Co-expression QTLs</article-title>. <source>Nat. Genet.</source> <volume>50</volume>, <fpage>493</fpage>&#x2013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-018-0089-9</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Velmeshev</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Schirmer</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jung</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Haeussler</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Perez</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mayer</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Single-cell Genomics Identifies Cell Type-specific Molecular Changes in Autism</article-title>. <source>Science</source> <volume>364</volume>, <fpage>685</fpage>&#x2013;<lpage>689</lpage>. <pub-id pub-id-type="doi">10.1126/science.aav8130</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Villani</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Reynolds</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sarkizova</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Shekhar</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Fletcher</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Single-cell RNA-Seq Reveals New Types of Human Blood Dendritic Cells, Monocytes, and Progenitors</article-title>. <source>Science</source> <volume>356</volume>. <pub-id pub-id-type="doi">10.1126/science.aah4573</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whittaker</surname>
<given-names>R. H.</given-names>
</name>
</person-group> (<year>1972</year>). <article-title>Evolution and Measurement of Species Diversity</article-title>. <source>Taxon</source> <volume>21</volume>, <fpage>213</fpage>&#x2013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.2307/1218190</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wilk</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Rustagi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>N. Q.</given-names>
</name>
<name>
<surname>Roque</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mart&#xed;nez-Col&#xf3;n</surname>
<given-names>G. J.</given-names>
</name>
<name>
<surname>McKechnie</surname>
<given-names>J. L.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A Single-Cell Atlas of the Peripheral Immune Response in Patients with Severe COVID-19</article-title>. <source>Nat. Med.</source> <volume>26</volume>, <fpage>1070</fpage>&#x2013;<lpage>1076</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-020-0944-y</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>G. X.</given-names>
</name>
<name>
<surname>Terry</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Belgrader</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ryvkin</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bent</surname>
<given-names>Z. W.</given-names>
</name>
<name>
<surname>Wilson</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Massively Parallel Digital Transcriptional Profiling of Single Cells</article-title>. <source>Nat. Commun.</source> <volume>8</volume>, <fpage>14049</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms14049</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wen</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A Human Circulating Immune Cell Landscape in Aging and COVID-19</article-title>. <source>Protein Cell</source> <volume>11</volume>, <fpage>740</fpage>&#x2013;<lpage>770</lpage>. <pub-id pub-id-type="doi">10.1007/s13238-020-00762-2</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhuang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Single-Cell Sequencing of Peripheral Mononuclear Cells Reveals Distinct Immune Response Landscapes of COVID-19 and Influenza Patients</article-title>. <source>Immunity</source> <volume>53</volume>, <fpage>685</fpage>&#x2013;<lpage>696</lpage>. <comment>e3</comment>. <pub-id pub-id-type="doi">10.1016/j.immuni.2020.07.009</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>