<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="methods-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Immunol.</journal-id>
<journal-title>Frontiers in Immunology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Immunol.</abbrev-journal-title>
<issn pub-type="epub">1664-3224</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fimmu.2023.1223471</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Immunology</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>sc-ImmuCC: hierarchical annotation for immune cell types in single-cell RNA-seq</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Jiang</surname>
<given-names>Ying</given-names>
</name>
<xref ref-type="author-notes" rid="fn003">
<sup>&#x2020;</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Ziyi</given-names>
</name>
<xref ref-type="author-notes" rid="fn003">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/245508"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Han</surname>
<given-names>Na</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/848885"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shang</surname>
<given-names>Jingzhe</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2148582"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wu</surname>
<given-names>Aiping</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/546382"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>State Key Laboratory of Common Mechanism Research for Major Diseases, Suzhou Institute of Systems Medicine, Chinese Academy of Medical Sciences &amp; Peking Union Medical College</institution>, <addr-line>Suzhou, Jiangsu</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Thomas Hartung, Johns Hopkins University, United States</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Aar&#xf3;n V&#xe1;zquez Jim&#xe9;nez, National Institute of Genomic Medicine (INMEGEN), Mexico; Tian Tian, Children&#x2019;s Hospital of Philadelphia, United States</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Aiping Wu, <email xlink:href="mailto:wap@ism.cams.cn">wap@ism.cams.cn</email>
</p>
</fn>
<fn fn-type="equal" id="fn003">
<p>&#x2020;These authors have contributed equally to this work and share first authorship</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1223471</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Jiang, Chen, Han, Shang and Wu</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Jiang, Chen, Han, Shang and Wu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Accurately identifying immune cell types in single-cell RNA-sequencing (scRNA-Seq) data is critical to uncovering immune responses in health or disease conditions. However, the high heterogeneity and sparsity of scRNA-Seq data, as well as the similarity in gene expression among immune cell types, poses a great challenge for accurate identification of immune cell types in scRNA-Seq data. Here, we developed a tool named sc-ImmuCC for hierarchical annotation of immune cell types from scRNA-Seq data, based on the optimized gene sets and ssGSEA algorithm. sc-ImmuCC simulates the natural differentiation of immune cells, and the hierarchical annotation includes three layers, which can annotate nine major immune cell types and 29 cell subtypes. The test results showed its stable performance and strong consistency among different tissue datasets with average accuracy of 71-90%. In addition, the optimized gene sets and hierarchical annotation strategy could be applied to other methods to improve their annotation accuracy and the spectrum of annotated cell types and subtypes. We also applied sc-ImmuCC to a dataset composed of COVID-19, influenza, and healthy donors, and found that the proportion of monocytes in patients with COVID-19 and influenza was significantly higher than that in healthy people. The easy-to-use sc-ImmuCC tool provides a good way to comprehensively annotate immune cell types from scRNA-Seq data, and will also help study the immune mechanism underlying physiological and pathological conditions.</p>
</abstract>
<kwd-group>
<kwd>immune cell identification</kwd>
<kwd>scRNA-seq</kwd>
<kwd>hierarchical annotation</kwd>
<kwd>immune cell signature sets</kwd>
<kwd>ssGSEA</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="38"/>
<page-count count="12"/>
<word-count count="5989"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Systems Immunology</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<title>Introduction</title>
<p>The infiltration and quantity of immune cells in tissues are closely related to the occurrence, treatment response, and development of diseases (<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B3">3</xref>), such as cancer and inflammatory disease (<xref ref-type="bibr" rid="B4">4</xref>). Quantitatively identifying immune cells in tissues can provide new insights and methods for disease treatment and prevention. Traditional experimental methods, such as flow cytometry (<xref ref-type="bibr" rid="B5">5</xref>), affinity purification (<xref ref-type="bibr" rid="B6">6</xref>), and immunohistochemistry (<xref ref-type="bibr" rid="B7">7</xref>), are capable of qualitatively and quantitatively measuring immune cells. However, these methods are of limited use when the markers for some cell types are not clear, and they impractical in the case of large-scale samples due to their time consumption (<xref ref-type="bibr" rid="B8">8</xref>).</p>
<p>Computational strategies have gradually been developed to obtain the constitution of immune cells directly from tissue omics data (<xref ref-type="bibr" rid="B9">9</xref>, <xref ref-type="bibr" rid="B10">10</xref>). Computational methods have several advantages over traditional experimental methods, including high throughput, automation, and the recognition of unknown cell types (<xref ref-type="bibr" rid="B11">11</xref>). Quantification of immune cells from omics data is an important strategy to infer the number and proportion of different immune cell subsets from high-throughput sequencing data. Previously, we have developed ImmuCC and seq-ImmuCC to infer the relative compositions of immune cell types in mouse tissues from microarray mRNA expression or bulk RNA-seq data (<xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B13">13</xref>). These methods have been widely used to quantify the immune cell compositions of mouse tissues. However, these bulk expression profile-based methods for quantifying immune cells lack the ability to analyze individual cells and therefore cannot help effectively analyze immune cell heterogeneity.</p>
<p>Single-cell RNA sequencing (scRNA-seq) technology has emerged as a powerful technique for studying the heterogeneity and complexity of RNA transcripts within individual cells (<xref ref-type="bibr" rid="B14">14</xref>&#x2013;<xref ref-type="bibr" rid="B16">16</xref>), and for identifying the composition of cell types and functions within different tissues, organs and organisms (<xref ref-type="bibr" rid="B17">17</xref>). The advancements in scRNA-Seq technology now allow us to obtain single cell gene expression and quantitative composition data of various immune cells in tissues. To identify immune cells in scRNA-Seq data, clustering followed by manual annotation is commonly used (<xref ref-type="bibr" rid="B18">18</xref>, <xref ref-type="bibr" rid="B19">19</xref>). However, this strategy is usually labor-intensive and subjective with respect to the selection of signature genes, which causes poor reproducibility of cell annotation by different researchers. Moreover, some rare cell types cannot be directly clustered, and some abundant cell types with different anchor genes selected during integration can lead to undetermined annotation results.</p>
<p>In recent years, an increasing number of automated cell type annotation methods have been developed (<xref ref-type="bibr" rid="B20">20</xref>). These annotation methods can be categorized into three types: (1) marker gene-based methods, such as SCINA (<xref ref-type="bibr" rid="B21">21</xref>), scCATCH (<xref ref-type="bibr" rid="B22">22</xref>) and SCSA (<xref ref-type="bibr" rid="B23">23</xref>), which use prior knowledge for cell type annotation; (2) similarity-based methods, such as SingleR (<xref ref-type="bibr" rid="B24">24</xref>),which are based on correlations between query cells and predefined reference cell types, they assign the label of the type with maximum correlation; (3) supervised classification-based methods, scPred (<xref ref-type="bibr" rid="B25">25</xref>) using a combination of unbiased feature selection from a reduced-dimension space, and machine-learning probability-based prediction method to annotate cell types. Garnett (<xref ref-type="bibr" rid="B26">26</xref>) constructs a reference cell type hierarchy and uses elastic net regression for cell type prediction. Although these tools have provided powerful annotation performance for single cell sequencing data, their major limitation is the lack of specialized annotation tools for immune cells. Achieving consistent annotations for immune cell types is a challenge, as the same immune cell is often not in the same layer label in different tools. For instance, T cells belong to the major cell type, whereas CD4 T cells and CD8 T cells are subtypes of the former, the output of most tools is often a mixture of major cell types and subtypes. Moreover, the similar gene expression profiles of immune cell types cause difficulty in distinguishing them accurately. As a result, it is difficult to obtain accurate annotations and uniform output labels for immune cell types from scRNA-Seq data.</p>
<p>To address this issue, we have developed a hierarchical immune cell annotation method, namely sc-ImmuCC, based on the hierarchical lineage differentiation of immune cells. The core concept of sc-ImmuCC is to define the signature gene sets for differential immune cell types and to recognize them based on the enrichment scores of the signature genes. In sc-ImmuCC, the major types of immune cells are first annotated, followed by annotation of the subtypes of each cell type individually, which can reduce the interference between similar cell types, such as T cells and NK cells, and improve the accuracy of subtype annotation by avoiding cluttered annotation labels. The hierarchical annotation strategy of sc-ImmuCC can not only effectively distinguish immune cells with defined signature genes, but also provides an open framework for integrating more knowledge for future annotation.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<title>Materials and methods</title>
<sec id="s2_1">
<title>Summary of sc-ImmuCC</title>
<p>Briefly, sc-ImmuCC is a hierarchical method for annotating immune cell types based on signature gene sets and ssGSEA [the single-sample GSEA, an extension of Gene Set Enrichment Analysis (GSEA)]. Our objective is to collect representative and discriminative signature genes to the extent possible, and the genes are retained for later calculation. There are three main steps in the sc-ImmuCC model. The first step is signature gene selection. For the first layer of the immune cell types, canonical cell markers are mainly selected. For the second and third layers, not only the canonical cell markers, but also some functional feature genes that are highly expressed in RNA-Seq or scRNA-Seq data and some signature genes from the disease data sources in Ingenuity Pathway Analysis (IPA) are also included. The second step is calculation of the enrichment scores. According to the defined gene sets, the ssGSEA algorithm is used to calculate the immune cell enrichment scores of each cell hierarchically. Finally, according to the enrichment score, the cell types are annotated, and the largest score value is selected and converted into a cell type label.</p>
</sec>
<sec id="s2_2">
<title>Hierarchical immune cell types</title>
<p>According to the differentiation lineage of the immune cells, we divided the annotation process of immune cells into three layers (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1A</bold>
</xref>). The first layer consists of nine major immune cell types: T cells, B cells, monocytes, macrophages, dendritic cells (DC), natural killer cells (NK), innate lymphoid cells (ILC), mast cells, and neutrophils. The second layer of cells is the subtype of the first layer of cells, mainly including the ILC subtypes: ILC1, ILC2 and ILC3; B cell subtypes: na&#xef;ve B cells, memory B cells and plasma cells; T cell subtypes: CD4 T cells and CD8 T cells; NK subtypes: NK<sup>_bright</sup> and NK<sup>_dim</sup>; DC subtypes: plasmacytoid DC (pDC) and conventional DC (cDC); monocyte subtypes: classic monocytes and non-classical monocytes; and macrophage subtypes: M1 macrophage and M2 macrophage. There are a total of 16 cell subtypes in the second layer. The third layer is a more specific subtype classification for the CD4 T cells and CD8 T cells in the second layer. The CD4 T cells include CD4_naive, CD4_central_memory, CD4_effector_memory, regulatory T cell (Treg), T follicular helper cell (Tfh), T helper (Th) cell 1, Th2 and Th17 subtypes, and the CD8 T cells include CD8_naive, CD8_central_memory, CD8_effector_memory, cytotoxic cells, and exhausted cells subtypes. The subsequent annotation process is performed in each layer separately.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Overview of sc-ImmuCC. <bold>(A)</bold> Immune cell types and layers in sc-ImmuCC. <bold>(B)</bold> Signature gene sets for each cell type were collected and screened separately based on the cell types of different layers (left), and the process of sc-ImmuCC hierarchically annotating immune cells (right).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1223471-g001.tif"/>
</fig>
</sec>
<sec id="s2_3">
<title>Definition of signature gene sets</title>
<p>The signature genes for each immune cell type were obtained by integrating the cell markers reported in the literature. Canonical experimentally validated signature genes were used as the marker genes for the first layer of cell types. Except to classical signature genes collected from the literature, signature genes derived from scRNA-seq, bulk RNA-seq data, and IPA were used as sources for the signature genes of the second and third layer cells (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1B</bold>
</xref>). Genes with low expression values were removed based on their average expression among multiple datasets. After filtering the low expressed marker genes, 11 gene sets were used to distinguish different immune cell types (<xref ref-type="supplementary-material" rid="SM2">
<bold>Tables S6-16</bold>
</xref>), including the signature gene sets used to distinguish immune and non-immune cells, immune cells of nine major types, namely B cells, T cells, macrophages, monocytes, neutrophils, mast cells, innate lymphoid cells, dendritic cells and natural killer cells. The details of signature gene sets (<xref ref-type="supplementary-material" rid="SM2">
<bold>Tables S6-16</bold>
</xref>) were shown in <xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary Data</bold>
</xref>.</p>
</sec>
<sec id="s2_4">
<title>Dataset and preprocessing</title>
<p>All single cell gene expression datasets used in this study were obtained from public accessions, downloaded from the NCBI GEO, 10x Genomics, and EMBL-EBI ArrayExpress databases (<xref ref-type="supplementary-material" rid="SM1">
<bold>Table S1</bold>
</xref>). We used the original cell type annotation provided by each publication as ground truth, and systematically standardized the original labels. The specific original labels and their corresponding standardized labels can be found in <xref ref-type="supplementary-material" rid="SM1">
<bold>Table S3</bold>
</xref> (see <xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary Data</bold>
</xref>). For cells with ambiguous original annotations, such as &#x201c;Mono/Mac&#x201d;, &#x201c;NK/T&#x201d;, and &#x201c;T/B doublets&#x201d;, we discard them and retain only the cell type with a single annotation label.</p>
<p>To test the effectiveness and robustness of the method, 14 independent scRNA-Seq datasets were used, covering a wide range of immune cell types and tissue sources. For the PBMC dataset, the five common immune cell types, namely T cells, B cells, dendritic cells, monocytes, and natural killer cells, were retained, and some rare cell types such as platelets were removed from the testing dataset. Considering the imbalance in the E-MTAB-11536 dataset, we randomly sampled it for annotating the cell types in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref>. For cell types with more than 50,000 cells, we randomly selected 12,000 cells for testing. For cell types with more than 20,000 and less than 50,000 cells, we randomly selected 8,000 cells. For cell types with less than 10,000 cells, we chose all of them to test (<xref ref-type="supplementary-material" rid="SM1">
<bold>Table S4</bold>
</xref>).</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Evaluation of sc-ImmuCC on some annotated datasets. <bold>(A)</bold> UMAP plot the PBMC dataset with the original cell label (left) and the cell label annotated with sc-ImmuCC (right). <bold>(B)</bold> Pheatmap comparing the cell types from the original publication (rows) to those inferred by sc-ImmuCC (columns) for E-MTAB-11536. The color represents the recall and precision score (as a percentage) of each original cell type predicted by sc-ImmuCC. <bold>(C)</bold> Sankey plot for the sc-ImmuCC annotations on Layer 2. Cell type annotations by sc-ImmuCC (right) with the original cell type annotations in the dataset (left). <bold>(D)</bold> Sankey plot for sc-ImmuCC annotations on Layer 3.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1223471-g002.tif"/>
</fig>
<p>All the data used in the tests without any filtering, correction, or normalization. For the datasets used in <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3</bold>
</xref> and <xref ref-type="fig" rid="f4">
<bold>4</bold>
</xref>, non-immune cells were removed in advance if they were included, without any additional processing. In the testing <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3</bold>
</xref> and <xref ref-type="fig" rid="f4">
<bold>4</bold>
</xref>, the dataset GSE131907 underwent pre-filtering to exclude non-immune cells, and the original data was directly used for testing in <xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S5B</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Performance of sc-ImmuCC in comparison with other methods. Each dataset was generated by randomly selecting 3000 cells per cell type. If the number of this type was less than 3000, we included all of them. Each dataset was randomly sampled five times; each tool was tested separately; and the results were averaged over five times. <bold>(A)</bold> Boxplots show the accuracy on layer 1 in different cell types. <bold>(B)</bold> Boxplots show the recall and precision score on layer 1 in different datasets. <bold>(C)</bold> Boxplots show the overall accuracy on layer 2 and layer 3 of sc-ImmuCC, SingleR and ImmClassifier. <bold>(D)</bold> Boxplots show the recall and precision on layer 2 and layer 3 of sc-ImmuCC, SingleR and ImmClassifier. t.test was conducted for sc-ImmuCC with other tools. *p &lt; 0.05, **p &lt; 0.01, ***p &lt; 0.001, ****p &lt; 0.0001. ns, not significant.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1223471-g003.tif"/>
</fig>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Performance of Garnett and SCINA with optimized gene sets and hierarchical annotations. <bold>(A)</bold> Overall accuracy comparison using SCINA&#x2019;s original gene set (three cell types) and optimized gene set (nine cell types) in layer 1. <bold>(B)</bold> Boxplots show the F1-score of SCINA with different gene sets in layer 1. <bold>(C)</bold> Sankey plot of SCINA in layer 2. <bold>(D)</bold> Overall accuracy comparison using Garnett_pbmc classifier and Garnett train with optimized gene set in layer 1. <bold>(E)</bold> Boxplots show the F1-score of Gernett with different classifier in layer 1. <bold>(F)</bold> Sankey plot of Garnett using the hierarchical train and one-step train in layer 2. t.test was conducted for Optimized gene set with original gene sets. **p &lt; 0.01, ***p &lt; 0.001, ****p &lt; 0.0001. ns, not significant.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1223471-g004.tif"/>
</fig>
</sec>
<sec id="s2_5">
<title>Reference cell type annotation tools</title>
<p>We selected five existing methods for performance comparison with sc-ImmuCC: SingleR, ImmClassifier, Garnett, SCINA and scCATCH. SingleR is a similarity-based method; ImmClassifier and Garnett are based on machine learning methods, and are both based on hierarchy. SCINA and scCATCH are marker gene-based methods. The former is based on a custom gene set, and the latter is based on the marker gene database.</p>
<p>To perform the tool comparison, ImmClassifier was run with default parameters. We ran SingleR with default parameters and ran Garnett with the hsPBMC pretrained classifier. The SCINA R package using precompiled immune cell signatures from the RCC patients, and we labeled cell types which not included in the precompiled gene set as &#x201c;unknown&#x201d;. scCATCH was used with default parameters, and we choose the &#x201c;Bone&#x201d;, &#x201c;Lung&#x201d; and &#x201c;Blood&#x201d; tissue sources.</p>
<p>To balance the datasets and randomness of the methods, we adopted a random sampling method for each of the four datasets. We randomly selected 3,000 cells per cell type (&gt;3,000 cell) each time; if there were less than 3,000 cells per cell type, all of them were included each time (<xref ref-type="supplementary-material" rid="SM1">
<bold>Table S5</bold>
</xref>). Five random samplings were carried out in total, and the final results were the average of the five.</p>
</sec>
<sec id="s2_6">
<title>Statistical analysis and visualization</title>
<p>The basic statistical analyses presented in <xref ref-type="fig" rid="f4">
<bold>Figures&#xa0;4</bold>
</xref> and <xref ref-type="fig" rid="f5">
<bold>5</bold>
</xref> were performed with the &#x201c;ggpubr&#x201d; package. UMAP plots were obtained by &#x201c;Seurat&#x201d; package (<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B28">28</xref>). For the datasets GSE144744 and GSE131907, the top 2,000 most variable genes were used for PCA. The top 15 principal components were used to generate the UMAP and tSNE plots. Sankey plots were generated using &#x201c;networkD3&#x201d; package (<xref ref-type="bibr" rid="B29">29</xref>). Boxplots were generated using &#x201c;ggplot2&#x201d; package (<xref ref-type="bibr" rid="B30">30</xref>). Pheatmaps are drawn with the &#x201c;pheatmap&#x201d; package.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Applications of sc-ImmuCC. <bold>(A)</bold> tSNE plot colored by annotated cell types. <bold>(B)</bold> Stacked diagram of cell composition for each group. <bold>(C)</bold> Boxplots show the cell proportion of each group in layer 1. T tests were conducted for each cell type between the disease and HD groups. *P &lt; 0.05, **P &lt; 0.01, and ***P &lt; 0.001. <bold>(D)</bold> Fraction of cell types in COVID_Mild and COVID_Server in layer 1. <bold>(E)</bold> Cell proportion between the influenza and HD groups in layer 2. <bold>(F)</bold> The dot plot showing the marker expression profiles of different groups under sc-ImmuCC annotation. <bold>(G)</bold> CD8 T cell subtypes proportion between the disease and HD groups.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1223471-g005.tif"/>
</fig>
<p>The Seurat R package is integrated into sc-ImmuCC as a visualization tool to help users gain a more intuitive understanding of the annotation results and gene expression in the data. Seurat was not used as a data preprocessing tool. No clustering was performed using Seurat. Seurat was employed for principal component analysis of the data and added the annotated cell types to the Seurat object. The annotation results were visualized in UMAP and tSNE plots based on PCA and annotated cell types.</p>
</sec>
<sec id="s2_7">
<title>Input and output of sc-ImmuCC</title>
<p>The input for sc-ImmuCC is a single-cell count matrix with cells unique barcodes as column names and gene names as row names. The input matrix does not require any filtering, correction, or normalization. However, users have the option to perform these operations if desired. From our testing, we found that normalization does not affect the annotation results.</p>
<p>The output of sc-ImmuCC is a CSV file containing the annotation results for each cell type at the first, second, and third layer, along with corresponding tSNE, UMAP, DotPlot, and Heatmap plots. For second and third-layer annotations, if the cell number for a specific cell type is less than 50, graphical representation is omitted, and only the CSV annotation file is provided.</p>
</sec>
<sec id="s2_8">
<title>Codes and availability</title>
<p>Codes and scripts of the sc-ImmuCC method were written in R version 4.1.1 and Bioconductor (<xref ref-type="bibr" rid="B31">31</xref>) version 3.14, installation instruction, usage and example codes can be found at <uri xlink:href="https://github.com/wuaipinglab/scImmuCC">https://github.com/wuaipinglab/scImmuCC</uri>. The ssGSEA enrichment scores were calculated with the &#x201c;GSVA&#x201d; package (<xref ref-type="bibr" rid="B31">31</xref>). The relevant images are drawn by the &#x201c;Seurat&#x201d; package.</p>
</sec>
</sec>
<sec id="s3" sec-type="results">
<title>Results</title>
<sec id="s3_1">
<title>Overview of the sc-ImmuCC method</title>
<p>The stepwise differentiation of immune cells in nature provides a good reference framework to identify immune cell types from single-cell sequencing data hierarchically. By simulating as the hierarchical differentiation of immune cells, here we propose sc-ImmuCC, a tool to annotate the immune cell types within scRNA-seq data. The immune cell types included in sc-ImmuCC can be hierarchically classified into three layers according to the differentiation lineage trees, with nine major cell types in the first layer, 16 cell subtypes in the second layer, and 13 T cell subtypes in the third layer (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1A</bold>
</xref>).</p>
<p>Cell differentiation is primarily driven by differential gene expression, leading to the formation of a diverse range of cell types. To identify specific cell types, it is crucial to identify the signature genes that are highly expressed in each cell type. sc-ImmuCC aims to identify immune cell types in scRNA-Seq data by screening the signature gene sets of different cell types, calculating the enrichment scores hierarchically through the ssGSEA algorithm, and providing annotations based on the scores. Three key steps are included in the sc-ImmuCC model (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1B</bold>
</xref>, see Methods for more details. First, we conducted expression screening of all gene sets and select genes with higher average expression across multiple datasets as signature gene sets (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1B, C</bold>
</xref>). Second, we hierarchically identified nine major immune cell types and their subtypes (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1A</bold>
</xref>). Third, to evaluate the robustness and applicability of the method, we applied it to different datasets, including PBMC datasets, enriched immune cell datasets, and tumor tissue datasets.</p>
</sec>
<sec id="s3_2">
<title>Performance of sc-ImmuCC on immune cell datasets</title>
<p>The performance of sc-ImmuCC in the first layer was evaluated on two different datasets. On a testing PBMC dataset, sc-ImmuCC achieved an overall accuracy of 85%, and the accuracy of T cells, B cells, DCs, NKs and monocytes were 86%, 99%, 75%, 92%, and 76%, respectively (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2A</bold>
</xref>). Some labelled NK cells were annotated as T cells in our model for their highly expressed CD3D, CD3E and GNLY (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2A</bold>
</xref>), which suggests that our method may can correct some mis-annotations. Similarly, in the T cell-dominated cluster, a small fraction of monocytes was relabeled as B cells by sc-ImmuCC due to their highly expressed CD79A and CD79B but not LYZ (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2A</bold>
</xref>).</p>
<p>To further assess the performance of sc-ImmuCC in annotating complex datasets, an immune-enriched dataset consisting of 16 tissues was selected. We compared the annotation results of sc-ImmunCC with the manually curated annotations of the original paper. In the first layer, sc-ImmuCC achieved an overall accuracy of 88% (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref>). The F1-score of T cells, B cells, DCs, NK cells, monocytes, macrophages and mast were 91%, 94%, 71%, 91%, 81%, 90%, and 99%. However, there were relatively high rates of misclassification in the innate lymphoid cells, with 31% of those cells labelled as T cells and 53.7% as monocytes (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref>). Upon examining the expression levels of some known marker genes associated with innate lymphoid cells in this cell cluster, it was observed that the expression of these genes was very low in most cells (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2B</bold>
</xref>). Hence, we performed a secondary evaluation of sc-ImmuCC&#x2019;s performance in annotating ILC cells using the GSE146771 dataset. The results demonstrated an annotation accuracy of over 50% for ILC cells in this dataset (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2C</bold>
</xref>). The distinction between the GSE146771 dataset and the dataset depicted in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref> lies in the higher expression levels of signature genes associated with ILC cells in the former (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2D</bold>
</xref>). This suggests that sc-ImmuCC did not accurately identify the innate lymphoid cells in E_MTAB_11536 dataset due to their low expression of the signature genes.</p>
<p>At the second layer, the sc-ImmuCC demonstrates a high level of annotation accuracy for T cells, B cells, DCs, NKs, monocytes, macrophages and ILCs subtypes (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2C</bold>
</xref>), with an overall accuracy rate of 78-90%, except for macrophages which had an accuracy rate of 70% (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2C</bold>
</xref>). As expected, the performance of sc-ImmuCC decreased at the third layer for the CD4 and CD8 T cells. Though most subtypes of the CD8 T cells could be accurately annotated, the overall accuracy for the CD4 T cell subtypes decreased significantly (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2D</bold>
</xref>). This may suggest that further optimization of gene sets for CD4 T cell and CD8 T cell subtypes to improve discrimination is needed, as many signature genes do not exist in the scRNA-Seq data. Besides, the original literature annotation for subtypes may not be accurate, as the signature gene expression of CD4 T cell subtypes is not clearly evident based on the original annotations. For example, IL17A, IL17F, RORA, and RORC were not significantly expressed in Th17 (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2D</bold>
</xref>).</p>
</sec>
<sec id="s3_3">
<title>Comparison of sc-ImmuCC with other methods</title>
<p>To further assess the performance of sc-ImmuCC, we compared it with five representative methods for single cell annotation: ImmClassifier (<xref ref-type="bibr" rid="B32">32</xref>), Garnett, SingleR, SCINA, and scCATCH. The reference methods were chosen as the representative ones from hierarchical-based, correlation-based, and gene-based approaches. All methods were used on four independent scRNA-Seq datasets covering various tissue sources, including PBMC, metastatic lung adenocarcinoma, and non-small-cell lung cancer. All methods were compared using four metrics, namely overall accuracy, precision, recall, and F1-score.</p>
<p>Overall, almost all methods performed well at the first layer, sc-ImmuCC achieved an average overall annotation correctness of 71%, 80%, 85%, and 79% on the four datasets, respectively. It is worth noting that, although our tool achieved only 71% accuracy on the pbmc_68k dataset, all tools generated annotations for this dataset that deviated from the original ones. The best-performing method on this dataset, SingleR, achieved 73% accuracy (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S3A</bold>
</xref>). Apart from a slightly lower overall accuracy observed for DC cells, the accuracy of sc-ImmuCC for other cell types exceeded 85% (<xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3A</bold>
</xref>). sc-ImmuCC demonstrated achieves comparable performance to other tools on the four datasets (<xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3B</bold>
</xref>). Except for the pbmc_68k dataset, the F1-score on tumor datasets and cross-tissue datasets outperformed other tools (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S3B</bold>
</xref>).</p>
<p>At the second and third layers, we compared sc-ImmuCC with the ImmClassifier and SingleR methods. sc-ImmuCC demonstrated comparable accuracy for T cells, B cells and DC cells as SingleR and ImmClassifier (<xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3A, B</bold>
</xref>). sc-ImmuCC exhibited significantly superior annotation performance for some other cell types, such as dendritic cells, monocytes, and macrophages (<xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3C, D</bold>
</xref> and <xref ref-type="supplementary-material" rid="SM1">
<bold>S3C, D</bold>
</xref>). Due to the hierarchical annotation strategy, sc-ImmuCC reduced the interference of some similar signature gene expressions between cell types from different branches. At the third layer, sc-ImmuCC outperformed the other two methods in identifying the subtypes of CD8 T cells, but its performance for the CD4 T cells was equally limited as of ImmClassifier and SingleR (<xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3C, D</bold>
</xref> and <xref ref-type="supplementary-material" rid="SM1">
<bold>S3E</bold>
</xref>).</p>
</sec>
<sec id="s3_4">
<title>Optimized gene sets and hierarchical annotation facilitate other tools</title>
<p>To assess the impact of gene set optimization and hierarchical annotation on other tools, we applied these two strategies to two methods, SCINA and Garnett. SCINA is a marker-based annotation method that uses precompiled immune cell signatures from RCC patients, including classical monocytes, CD19_B and NK_dim cell types, and other cell types not included may be annotated as &#x201c;unknown&#x201d;. With our defined first layer signature gene sets, which included nine cell types, SCINA achieved a significant improvement in the overall annotation accuracy. For the lung adenocarcinoma and colorectal cancer datasets, the accuracy improved by 6% and 15%, respectively, and our method exhibited a comparable integrated performance to that of the original gene set in the cross-tissue datasets (<xref ref-type="fig" rid="f4">
<bold>Figures&#xa0;4A, B</bold>
</xref>). Subsequently, we used the second layer signature gene set on SCINA and found that almost all subtypes could be identified except for the M1 macrophage and natural killer cell subtypes (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4C</bold>
</xref>). Particularly, the precision for the subtypes of B cells and monocytes exceeded 75%. The recall for the B cell subtypes is also exceeded 91% (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S4A</bold>
</xref>).</p>
<p>Similar results were also observed when applying the optimized gene set and hierarchical annotation to Garnett. By using our first layer gene sets with Garnett, the accuracy improved by 8%-26% (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4D</bold>
</xref>). Not only did the number of annotation types increase, but also the overall performance has also improved (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4E</bold>
</xref>). For the second layer annotation, we first integrated the optimized gene sets to create a comprehensive gene set for one-step training. Additionally, we trained classifier with each gene set separately to achieve a hierarchical annotation. By comparing the annotated results with those from the original literature, we found that the hierarchically trained classifier had better annotations for most cell subtypes. In contrast, the one-step-trained classifier assigned most cells as first layer immune cell types and failed to obtain finer annotations (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4F</bold>
</xref>). Except for NK<sup>_bright</sup>, the recall for most subtypes exceeded 90%, and the precision rate for the DC and monocyte subtypes was more than is over 80% (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S4B, C</bold>
</xref>).</p>
<p>These results demonstrate that the strategies of optimized gene sets and hierarchical annotation can not only identify more immune cell types, but also generate more accurate annotation results. Overall, our results highlight the potential of these strategies for improving the accuracy and completeness of cell type annotation in scRNA-seq data, even when the strategies are applied to other annotation methods.</p>
</sec>
<sec id="s3_5">
<title>Immune-cell profiling in COVID-19 and influenza using sc-ImmuCC</title>
<p>We employed sc-ImmuCC to annotate COVID-19 data (<xref ref-type="bibr" rid="B33">33</xref>) to gain insights into cellular composition and immune responses across various pathologies. A previous study examined the single-cell transcriptome of PBMCs from COVID-19 and influenza patients and presented a comprehensive overview of their immunophenotypes at the single-cell level (<xref ref-type="bibr" rid="B34">34</xref>). However, by focusing on some subset of immune cell types, the study did not fully labelled immune cell subsets. To address this, we re-annotated this dataset and conducted an in-depth analysis of immune cell composition across different pathological states.</p>
<p>We identified five major cell types, namely T cells, B cells, dendritic cells, monocytes, and NK cells, along with their subtypes (<xref ref-type="fig" rid="f5">
<bold>Figures&#xa0;5A</bold>
</xref> and <xref ref-type="supplementary-material" rid="SM1">
<bold>S5A</bold>
</xref>). Due to the small number of DC cells, we only annotated them at the first layer. The relative proportions of immune cells in PBMCs from the disease groups were altered compared to those in healthy donors (<xref ref-type="fig" rid="f5">
<bold>Figures&#xa0;5B, C</bold>
</xref>). Interestingly, severe COVID-19 and influenza showed some similarity in terms of a significant increase in monocyte proportion compared to healthy donors (<xref ref-type="fig" rid="f5">
<bold>Figures&#xa0;5C</bold>
</xref> and <xref ref-type="supplementary-material" rid="SM1">
<bold>S5B</bold>
</xref>), which is consistent with the literature, with subtle differences in DCs. In severe COVID-19, the proportion of monocytes was significantly increased, whereas the proportions of T cells and NK cells were decreased (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5C</bold>
</xref>). The proportion of monocytes was also significantly higher in severe COVID-19 patients compared to mild COVID-19 patients (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5D</bold>
</xref>). In influenza, the proportion of CD8 T cells was significantly increased, and CD4 T cells were significantly decreased (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5E</bold>
</xref>).</p>
<p>The expression profiles of immune cell markers in COVID-19, influenza patients, and healthy donors showed that common monocyte markers such as LYZ, VCAN and FCN1 were more significantly expressed in influenza patients, which may reflect stronger pro-inflammatory signals in influenza patients. Some inflammation-related genes such as <italic>NFKB1, NFKB2</italic>, and <italic>TGFB1</italic> were specifically upregulated in COVID-19, whereas genes such as <italic>STAT3, STAT1, TLR4, NLRP3</italic>, and <italic>PYCARD</italic> were specifically upregulated in influenza (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5F</bold>
</xref>). In addition, our results provide annotation of NK cell subtypes that were not previously identified in the literature.</p>
<p>Finally, we observed differences in the composition of CD8 T cell subtypes between healthy and disease states (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S5C</bold>
</xref>). In the GSE149689 dataset, the proportion of CD8 central memory cells was slightly lower in the COVID-19 and influenza groups compared to the healthy group, whereas CD8 cytotoxic cells showed the opposite trend, although not significant (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5G</bold>
</xref>). Nevertheless, sc-ImmuCC allowed for the annotation of subtypes at every hierarchical layer and was effective in annotating diverse source datasets. Applying it to a tumor dataset demonstrated its ability to accurately distinguish between immune cells and non-immune cells (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S5D</bold>
</xref>).</p>
</sec>
</sec>
<sec id="s4" sec-type="discussion">
<title>Discussion</title>
<p>Owing to the bias and uncertainty of the scRNA-Seq technology in detecting the biological characteristics of immune cells, accurately identifying immune cells and improving the accuracy of recognition is critically important. In this study, we designed a method called sc-ImmuCC to identify immune cell types from scRNA-Seq data. The performance of sc-ImmuCC was validated on various independent datasets with good robustness and accuracy. Compared with similar existing tools, sc-ImmuCC provides a hierarchical annotation for immune cells, which has several benefits. First, the concept of hierarchy is based on simulating the natural differentiation of immune cells. In the human body, hematopoietic stem cells (HSCs) continuously replenish all types of blood cells through a series of lineage-restricted steps. The immune cells are classified by lineage stratification, which allows for a more comprehensive annotation of the corresponding cell types. Second, hierarchical annotation avoids interference from similar gene expression profiles by not comparing them between different branches. Finally, hierarchical annotation is an open computational framework that can be continuously integrated and extended to new immune cell types in the future.</p>
<p>Moreover, optimized gene sets and hierarchical strategies can also improve the performance of other similar methods. For example, although Garnett is hierarchical, its trained classifiers do not provide proper output labels for some cell types, which makes it difficult for users to obtain a simple and clear result. Using the hsPBMC pretrained classifier will output T cells, CD4 T cells, and CD8 T cells, which are not at the same layer, whereas using the hsLung pretrained classifier cannot distinguish dendritic cells, monocytes, and macrophages. For users who want to understand the specific composition of these cells in the data in detail, this may be inconvenient. However, after using the optimized gene set to self-train with Garnett, we can not only distinguish dendritic cells, monocytes, and macrophages well in the first layer, but also train each subtype of cells separately and then perform hierarchical annotation with the dataset, thus greatly improving the annotation performance of each subtype of cells. Hierarchical annotation reduces errors in identifying cell subtypes, whereas gene set selection simplifies complex gene data to highlight the relevant genes. However, gene set selection may miss some relevant genes due to differences in regulation and tissue specificity. To improve accuracy, multiple tools should be used in hierarchical annotation and diverse strategies should be applied in gene set selection, such as utilizing diverse gene expression datasets, biological pathway databases, and functional enrichment analysis tools to cover a wide range of signature genes comprehensively.</p>
<p>Using the sc-ImmuCC method, we can effectively distinguish immune cells and non-immune cells in healthy and disease data, and quantify the composition of major immune cells across different tissues or organs in human scRNA-Seq data. An example is to use sc-ImmuCC to annotate immune cells in COVID-19 patients. With a large number of clinical and laboratory studies published on COVID-19 (<xref ref-type="bibr" rid="B35">35</xref>, <xref ref-type="bibr" rid="B36">36</xref>), &#x201c;omics&#x201d; approaches have played an important role in the study of COVID-19 and generated massive amounts of data at an unprecedented rate (<xref ref-type="bibr" rid="B37">37</xref>). Accurately identifying immune cell types in COVID-19 patients and dissecting the immune response in COVID-19 may aid the development of vaccines and antiviral drugs (<xref ref-type="bibr" rid="B38">38</xref>). sc-ImmuCC can accurately identify immune cell types hierarchically, thus providing a more precise understanding of cell subtypes and enabling more accurate comparisons of the changes in the cell composition and gene expression of interest across different diseases. By identifying immune cells in different pathological states, valuable immune cell background information can be extracted, thus providing more reliable evidence for the diagnosis and treatment of diseases.</p>
<p>Although using sc-ImmuCC to annotate the subtypes of cells at the third layer is not yet perfect, we hope that this study can promote the development of algorithms that can achieve this goal. In the future, we will continue to improve the annotation for the CD4 and CD8 T cell subtypes, such as optimizing gene sets, assigning weights to genes in the corresponding cell type based on their contribution values in the gene expression matrix, and integrating other methods. For the selection of gene sets, we will consider selecting genes according to the source of the data, and testing whether there are tissue-specific differences in the gene sets. Another limitation of sc-ImmuCC is that cannot detect new cell types. It is an annotation method based on a given specific cell type gene set and assigns cells to the cell type with the highest enrichment score calculated. This limitation needs to be addressed in future work, and for unidentifiable types, they can be improved by being identified as &#x201c;unknown&#x201d;. Currently, our method cannot fully annotate all immune cell types, such as &#x3b3;&#x3b4; T cells, which are known to play a crucial role in tumor defense, were not included in our signature gene set. In addition, some cell types with a small proportion, such as eosinophils and basophils, were also not included due to too few datasets to test. Consequently, sc-ImmuCC has limitations when studying these less common immune cell types due to the absence of appropriate test datasets for certain uncommon cell types. In the context of certain allergic diseases or parasitic infection sequencing data, the cell types annotated by our method may not possess sufficient accuracy and comprehensiveness for practical utilization. Therefore, further refinement of our method by including more cell types in the signature gene sets and finding more available test datasets will be an important future direction. Most of the previous research including our own method, can only identify differentiated terminal cells but cannot annotate immune cells undergoing continuous differentiation. In the future, it may be possible to construct a reference tree of immune cell evolution and project cells directly onto the tree to obtain their specific location in the differentiation path, to better understand the distribution and function of immune cells for a comprehensive and accurate study of immune cells.</p>
<p>Overall, the performance of sc-ImmuCC mainly depends on the given signature gene sets currently. The public datasets of immune cells were collected from different sources and tissues. If we further divide each subtype according to different tissue sources, perhaps there can be more accurate annotation of immune cells in scRNA-Seq data. The study of precise identification of immune cell types holds scientific significance and clinical application prospects, and further research will promote immunology&#x2019;s development.</p>
</sec>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="s10">
<bold>Supplementary Material</bold>
</xref>. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>YJ, ZC and AW conceived and designed the study. YJ collected the data and preprocessed it. YJ, ZC and AW analyzed the data and results. NH and JS contributed to the discussion and analysis of the studies. YJ, ZC and AW wrote the paper. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by the Non-profit Central Research Institute Fund of Chinese Academy of Medical Sciences (grant number 2021-PT180-001); the CAMS Innovation Fund for Medical Sciences (CIFMS) (grant number 2021-I2M-1-061); the National Key Plan for Scientific Research and Development of China (grant number 2021YFC2301305); the National Natural Science Foundation of China (grant number 92169106); the National Natural Science Foundation of China (32200544); and the CAMS Innovation Fund for Medical Sciences (grant number 2022-I2M-2-004).</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>The authors are grateful to the members of the Wu lab for helpful discussions.</p>
</ack>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fimmu.2023.1223471/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fimmu.2023.1223471/full#supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document"/>
<supplementary-material xlink:href="DataSheet_2.xlsx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ino</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Yamazaki-Itoh</surname> <given-names>R</given-names>
</name>
<name>
<surname>Shimada</surname> <given-names>K</given-names>
</name>
<name>
<surname>Iwasaki</surname> <given-names>M</given-names>
</name>
<name>
<surname>Kosuge</surname> <given-names>T</given-names>
</name>
<name>
<surname>Kanai</surname> <given-names>Y</given-names>
</name>
<etal/>
</person-group>. <article-title>Immune cell infiltration as an indicator of the immune microenvironment of pancreatic cancer</article-title>. <source>Br J Cancer</source> (<year>2013</year>) <volume>108</volume>:<page-range>914&#x2013;23</page-range>. doi: <pub-id pub-id-type="doi">10.1038/bjc.2013.32</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Man</surname> <given-names>YG</given-names>
</name>
<name>
<surname>Stojadinovic</surname> <given-names>A</given-names>
</name>
<name>
<surname>Mason</surname> <given-names>J</given-names>
</name>
<name>
<surname>Avital</surname> <given-names>I</given-names>
</name>
<name>
<surname>Bilchik</surname> <given-names>A</given-names>
</name>
<name>
<surname>Bruecher</surname> <given-names>B</given-names>
</name>
<etal/>
</person-group>. <article-title>Tumor-infiltrating immune cells promoting tumor invasion and metastasis: existing theories</article-title>. <source>J Cancer</source> (<year>2013</year>) <volume>4</volume>:<fpage>84</fpage>&#x2013;<lpage>95</lpage>. doi: <pub-id pub-id-type="doi">10.7150/jca.5482</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guzik</surname> <given-names>TJ</given-names>
</name>
<name>
<surname>Skiba</surname> <given-names>DS</given-names>
</name>
<name>
<surname>Touyz</surname> <given-names>RM</given-names>
</name>
<name>
<surname>Harrison</surname> <given-names>DG</given-names>
</name>
</person-group>. <article-title>The role of infiltrating immune cells in dysfunctional adipose tissue</article-title>. <source>Cardiovasc Res</source> (<year>2017</year>) <volume>113</volume>:<page-range>1009&#x2013;23</page-range>. doi: <pub-id pub-id-type="doi">10.1093/cvr/cvx108</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grivennikov</surname> <given-names>SI</given-names>
</name>
<name>
<surname>Greten</surname> <given-names>FR</given-names>
</name>
<name>
<surname>Karin</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>Immunity, inflammation, and cancer</article-title>. <source>Cell</source> (<year>2010</year>) <volume>140</volume>:<page-range>883&#x2013;99</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2010.01.025</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>YR</given-names>
</name>
<name>
<surname>O'koren</surname> <given-names>EG</given-names>
</name>
<name>
<surname>Hotten</surname> <given-names>DF</given-names>
</name>
<name>
<surname>Kan</surname> <given-names>MJ</given-names>
</name>
<name>
<surname>Kopin</surname> <given-names>D</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>ER</given-names>
</name>
<etal/>
</person-group>. <article-title>A protocol for the comprehensive flow cytometric analysis of immune cells in normal and inflamed murine non-lymphoid tissues</article-title>. <source>PloS One</source> (<year>2016</year>) <volume>11</volume>:<elocation-id>e0150606</elocation-id>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0150606</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Watkins</surname> <given-names>SK</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Watkins</surname> <given-names>KE</given-names>
</name>
<name>
<surname>Hurwitz</surname> <given-names>AA</given-names>
</name>
</person-group>. <article-title>Isolation of immune cells from primary tumors</article-title>. <source>J Vis Exp</source> (<year>2012</year>) (<issue>64</issue>):<elocation-id>e3952</elocation-id>. doi: <pub-id pub-id-type="doi">10.3791/3952</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Basa</surname> <given-names>RC</given-names>
</name>
<name>
<surname>Davies</surname> <given-names>V</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X</given-names>
</name>
<name>
<surname>Murali</surname> <given-names>B</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>J</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>B</given-names>
</name>
<etal/>
</person-group>. <article-title>Decreased anti-tumor cytotoxic immunity among microsatellite-stable colon cancers from African americans</article-title>. <source>PloS One</source> (<year>2016</year>) <volume>11</volume>:<elocation-id>e0156660</elocation-id>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0156660</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Progress and challenge for computational quantification of tissue immune cells</article-title>. <source>Briefings Bioinf</source> (<year>2021</year>) <volume>22</volume>(<issue>5</issue>):<fpage>bbaa358</fpage>. doi: <pub-id pub-id-type="doi">10.1093/bib/bbaa358</pub-id>
</citation>
</ref>
<ref id="B9">
<label>9</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Finotello</surname> <given-names>F</given-names>
</name>
<name>
<surname>Trajanoski</surname> <given-names>Z</given-names>
</name>
</person-group>. <article-title>Quantifying tumor-infiltrating immune cells from transcriptomics data</article-title>. <source>Cancer Immunol Immunother</source> (<year>2018</year>) <volume>67</volume>:<page-range>1031&#x2013;40</page-range>. doi: <pub-id pub-id-type="doi">10.1007/s00262-018-2150-z</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Newman</surname> <given-names>AM</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>CL</given-names>
</name>
<name>
<surname>Green</surname> <given-names>MR</given-names>
</name>
<name>
<surname>Gentles</surname> <given-names>AJ</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>W</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Y</given-names>
</name>
<etal/>
</person-group>. <article-title>Robust enumeration of cell subsets from tissue expression profiles</article-title>. <source>Nat Methods</source> (<year>2015</year>) <volume>12</volume>:<page-range>453&#x2013;7</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nmeth.3337</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brbic</surname> <given-names>M</given-names>
</name>
<name>
<surname>Zitnik</surname> <given-names>M</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>S</given-names>
</name>
<name>
<surname>Pisco</surname> <given-names>AO</given-names>
</name>
<name>
<surname>Altman</surname> <given-names>RB</given-names>
</name>
<name>
<surname>Darmanis</surname> <given-names>S</given-names>
</name>
<etal/>
</person-group>. <article-title>MARS: discovering novel cell types across heterogeneous single-cell experiments</article-title>. <source>Nat Methods</source> (<year>2020</year>) <volume>17</volume>:<page-range>1200&#x2013;6</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41592-020-00979-3</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>A</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>T</given-names>
</name>
<name>
<surname>Qin</surname> <given-names>FX</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Inference of immune cell composition on the expression profiles of mouse tissue</article-title>. <source>Sci Rep</source> (<year>2017</year>) <volume>7</volume>:<fpage>40508</fpage>. doi: <pub-id pub-id-type="doi">10.1038/srep40508</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Quan</surname> <given-names>L</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>A</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>X</given-names>
</name>
<etal/>
</person-group>. <article-title>Seq-ImmuCC: cell-centric view of tissue transcriptome measuring cellular compositions of immune microenvironment from mouse RNA-seq data</article-title>. <source>Front Immunol</source> (<year>2018</year>) <volume>9</volume>:<elocation-id>1286</elocation-id>. doi: <pub-id pub-id-type="doi">10.3389/fimmu.2018.01286</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname> <given-names>F</given-names>
</name>
<name>
<surname>Barbacioru</surname> <given-names>C</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Nordman</surname> <given-names>E</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>C</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>N</given-names>
</name>
<etal/>
</person-group>. <article-title>mRNA-seq whole-transcriptome analysis of a single cell</article-title>. <source>Nat Methods</source> (<year>2009</year>) <volume>6</volume>:<page-range>377&#x2013;82</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nmeth.1315</pub-id>
</citation>
</ref>
<ref id="B15">
<label>15</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Svensson</surname> <given-names>V</given-names>
</name>
<name>
<surname>Vento-Tormo</surname> <given-names>R</given-names>
</name>
<name>
<surname>Teichmann</surname> <given-names>SA</given-names>
</name>
</person-group>. <article-title>Exponential scaling of single-cell RNA-seq in the past decade</article-title>. <source>Nat Protoc</source> (<year>2018</year>) <volume>13</volume>:<fpage>599</fpage>&#x2013;<lpage>604</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nprot.2017.149</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hwang</surname> <given-names>B</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>JH</given-names>
</name>
<name>
<surname>Bang</surname> <given-names>D</given-names>
</name>
</person-group>. <article-title>Single-cell RNA sequencing technologies and bioinformatics pipelines</article-title>. <source>Exp Mol Med</source> (<year>2018</year>) <volume>50</volume>:<fpage>1</fpage>&#x2013;<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s12276-018-0071-8</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jovic</surname> <given-names>D</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>H</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>L</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>F</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y</given-names>
</name>
</person-group>. <article-title>Single-cell RNA sequencing technologies and applications: a brief overview</article-title>. <source>Clin Transl Med</source> (<year>2022</year>) <volume>12</volume>:<fpage>e694</fpage>. doi: <pub-id pub-id-type="doi">10.1002/ctm2.694</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kiselev</surname> <given-names>VY</given-names>
</name>
<name>
<surname>Andrews</surname> <given-names>TS</given-names>
</name>
<name>
<surname>Hemberg</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>Challenges in unsupervised clustering of single-cell RNA-seq data</article-title>. <source>Nat Rev Genet</source> (<year>2019</year>) <volume>20</volume>:<page-range>273&#x2013;82</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41576-018-0088-9</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Clarke</surname> <given-names>ZA</given-names>
</name>
<name>
<surname>Andrews</surname> <given-names>TS</given-names>
</name>
<name>
<surname>Atif</surname> <given-names>J</given-names>
</name>
<name>
<surname>Pouyabahar</surname> <given-names>D</given-names>
</name>
<name>
<surname>Innes</surname> <given-names>BT</given-names>
</name>
<name>
<surname>Macparland</surname> <given-names>SA</given-names>
</name>
<etal/>
</person-group>. <article-title>Tutorial: guidelines for annotating single-cell transcriptomic maps using automated and manual methods</article-title>. <source>Nat Protoc</source> (<year>2021</year>) <volume>16</volume>:<page-range>2749&#x2013;64</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41596-021-00534-0</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>B</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Mora</surname> <given-names>A</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X</given-names>
</name>
</person-group>. <article-title>Automatic cell type identification methods for single-cell RNA sequencing</article-title>. <source>Comput Struct Biotechnol J</source> (<year>2021</year>) <volume>19</volume>:<page-range>5874&#x2013;87</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.csbj.2021.10.027</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>D</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>X</given-names>
</name>
<name>
<surname>Choi</surname> <given-names>JH</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>S</given-names>
</name>
<etal/>
</person-group>. <article-title>SCINA: a semi-supervised subtyping algorithm of single cells and bulk samples</article-title>. <source>Genes (Basel)</source> (<year>2019</year>) <volume>10</volume>(<issue>7</issue>):<fpage>53</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/genes10070531</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shao</surname> <given-names>X</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>J</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>X</given-names>
</name>
<name>
<surname>Xue</surname> <given-names>R</given-names>
</name>
<name>
<surname>Ai</surname> <given-names>N</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>X</given-names>
</name>
</person-group>. <article-title>scCATCH: automatic annotation on cell types of clusters from single-cell RNA sequencing data</article-title>. <source>iScience</source> (<year>2020</year>) <volume>23</volume>:<fpage>100882</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.isci.2020.100882</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>SCSA: a cell type annotation tool for single-cell RNA-seq data</article-title>. <source>Front Genet</source> (<year>2020</year>) <volume>11</volume>:<elocation-id>490</elocation-id>. doi: <pub-id pub-id-type="doi">10.3389/fgene.2020.00490</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aran</surname> <given-names>D</given-names>
</name>
<name>
<surname>Looney</surname> <given-names>AP</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>L</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>E</given-names>
</name>
<name>
<surname>Fong</surname> <given-names>V</given-names>
</name>
<name>
<surname>Hsu</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>Reference-based analysis of lung single-cell sequencing reveals a transitional profibrotic macrophage</article-title>. <source>Nat Immunol</source> (<year>2019</year>) <volume>20</volume>:<page-range>163&#x2013;72</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41590-018-0276-y</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alquicira-Hernandez</surname> <given-names>J</given-names>
</name>
<name>
<surname>Sathe</surname> <given-names>A</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>HP</given-names>
</name>
<name>
<surname>Nguyen</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Powell</surname> <given-names>JE</given-names>
</name>
</person-group>. <article-title>scPred: accurate supervised method for cell-type classification from single-cell RNA-seq data</article-title>. <source>Genome Biol</source> (<year>2019</year>) <volume>20</volume>:<fpage>264</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13059-019-1862-5</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pliner</surname> <given-names>HA</given-names>
</name>
<name>
<surname>Shendure</surname> <given-names>J</given-names>
</name>
<name>
<surname>Trapnell</surname> <given-names>C</given-names>
</name>
</person-group>. <article-title>Supervised classification enables rapid annotation of cell atlases</article-title>. <source>Nat Methods</source> (<year>2019</year>) <volume>16</volume>:<page-range>983&#x2013;6</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41592-019-0535-3</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stuart</surname> <given-names>T</given-names>
</name>
<name>
<surname>Butler</surname> <given-names>A</given-names>
</name>
<name>
<surname>Hoffman</surname> <given-names>P</given-names>
</name>
<name>
<surname>Hafemeister</surname> <given-names>C</given-names>
</name>
<name>
<surname>Papalexi</surname> <given-names>E</given-names>
</name>
<name>
<surname>Mauck</surname> <given-names>WM</given-names>
<suffix>3rd</suffix>
</name>
<etal/>
</person-group>. <article-title>Comprehensive integration of single-cell data</article-title>. <source>Cell</source> (<year>2019</year>) <volume>177</volume>:<fpage>1888</fpage>&#x2013;<lpage>1902.e1821</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2019.05.031</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hao</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Hao</surname> <given-names>S</given-names>
</name>
<name>
<surname>Andersen-Nissen</surname> <given-names>E</given-names>
</name>
<name>
<surname>Mauck</surname> <given-names>WM</given-names>
<suffix>3rd</suffix>
</name>
<name>
<surname>Zheng</surname> <given-names>S</given-names>
</name>
<name>
<surname>Butler</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>Integrated analysis of multimodal single-cell data</article-title>. <source>Cell</source> (<year>2021</year>) <volume>184</volume>:<fpage>3573</fpage>&#x2013;<lpage>3587.e3529</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2021.04.048</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gandrud</surname> <given-names>C</given-names>
</name>
<name>
<surname>Allaire</surname> <given-names>JJ</given-names>
</name>
<name>
<surname>Russell</surname> <given-names>K</given-names>
</name>
<name>
<surname>Lewis</surname> <given-names>BW</given-names>
</name>
<name>
<surname>Kuo</surname> <given-names>K</given-names>
</name>
<name>
<surname>Sese</surname> <given-names>C</given-names>
</name>
<etal/>
</person-group>. <article-title>networkD3: D3 JavaScript network graphs from R</article-title>. (<year>2017</year>). Available at: <uri xlink:href="https://cran.r-project.org/web/packages/networkD3/networkD3.pdf">https://cran.r-project.org/web/packages/networkD3/networkD3.pdf</uri>.</citation>
</ref>
<ref id="B30">
<label>30</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname> <given-names>H</given-names>
</name>
</person-group>. <article-title>Ggplot2: elegant graphics for data analysis</article-title>. <source>ggplot2: Elegant Graphics Data Anal</source> (<year>2009</year>). doi: <pub-id pub-id-type="doi">10.1007/978-0-387-98141-3</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gentleman</surname> <given-names>RC</given-names>
</name>
<name>
<surname>Carey</surname> <given-names>VJ</given-names>
</name>
<name>
<surname>Bates</surname> <given-names>DJ</given-names>
</name>
<name>
<surname>Bolstad</surname> <given-names>BM</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>Bioconductor: open software development for computational biology and bioinformatics." genome biology, 5(10), R80</article-title>. <source>Genome Biol</source> (<year>2004</year>) <volume>5</volume>:<fpage>article R80</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/gb-2004-5-10-r80</pub-id>
</citation>
</ref>
<ref id="B32">
<label>32</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>X</given-names>
</name>
<name>
<surname>Gosline</surname> <given-names>SJC</given-names>
</name>
<name>
<surname>Pflieger</surname> <given-names>LT</given-names>
</name>
<name>
<surname>Wallet</surname> <given-names>P</given-names>
</name>
<name>
<surname>Iyer</surname> <given-names>A</given-names>
</name>
<name>
<surname>Guinney</surname> <given-names>J</given-names>
</name>
<etal/>
</person-group>. <article-title>Knowledge-based classification of fine-grained immune cell types in single-cell RNA-seq data</article-title>. <source>Brief Bioinform</source> (<year>2021</year>) <volume>22</volume>(<issue>5</issue>):<fpage>bbab039</fpage>. doi: <pub-id pub-id-type="doi">10.1093/bib/bbab039</pub-id>
</citation>
</ref>
<ref id="B33">
<label>33</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>F</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>S</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>B</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>YM</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W</given-names>
</name>
<name>
<surname>Song</surname> <given-names>ZG</given-names>
</name>
<etal/>
</person-group>. <article-title>A new coronavirus associated with human respiratory disease in China</article-title>. <source>Nature</source> (<year>2020</year>) <volume>579</volume>:<page-range>265&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41586-020-2008-3</pub-id>
</citation>
</ref>
<ref id="B34">
<label>34</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>JS</given-names>
</name>
<name>
<surname>Park</surname> <given-names>S</given-names>
</name>
<name>
<surname>Jeong</surname> <given-names>HW</given-names>
</name>
<name>
<surname>Ahn</surname> <given-names>JY</given-names>
</name>
<name>
<surname>Shin</surname> <given-names>EC</given-names>
</name>
</person-group>. <article-title>Immunophenotyping of COVID-19 and influenza highlights the role of type I interferons in development of severe COVID-19</article-title>. <source>Sci Immunol</source> (<year>2020</year>) <volume>5</volume>:<elocation-id>eabd1554</elocation-id>. doi: <pub-id pub-id-type="doi">10.1126/sciimmunol.abd1554</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Speranza</surname> <given-names>E</given-names>
</name>
<name>
<surname>Williamson</surname> <given-names>BN</given-names>
</name>
<name>
<surname>Feldmann</surname> <given-names>F</given-names>
</name>
<name>
<surname>Sturdevant</surname> <given-names>GL</given-names>
</name>
<name>
<surname>P&#xe9;rez</surname> <given-names>L--</given-names>
</name>
<name>
<surname>Meade-White</surname> <given-names>K</given-names>
</name>
<etal/>
</person-group>. <article-title>Single-cell RNA sequencing reveals SARS-CoV-2 infection dynamics in lungs of African green monkeys</article-title>. <source>Sci Trans Med</source> <volume>13</volume>(<issue>578</issue>):<elocation-id>eabe8146</elocation-id>. doi: <pub-id pub-id-type="doi">10.1126/scitranslmed.abe8146</pub-id>
</citation>
</ref>
<ref id="B36">
<label>36</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moss</surname> <given-names>P</given-names>
</name>
</person-group>. <article-title>The T cell immune response against SARS-CoV-2</article-title>. <source>Nat Immunol</source> (<year>2022</year>) <volume>23</volume>:<page-range>186&#x2013;93</page-range>. doi: <pub-id pub-id-type="doi">10.1038/s41590-021-01122-w</pub-id>
</citation>
</ref>
<ref id="B37">
<label>37</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>W</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>J</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W</given-names>
</name>
<name>
<surname>Pei</surname> <given-names>G</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Q</given-names>
</name>
<etal/>
</person-group>. <article-title>Delineating COVID-19 immunological features using single-cell RNA sequencing</article-title>. <source>Innovation (Camb)</source> (<year>2022</year>) <volume>3</volume>:<fpage>100289</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.xinn.2022.100289</pub-id>
</citation>
</ref>
<ref id="B38">
<label>38</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Carpp</surname> <given-names>LN</given-names>
</name>
<name>
<surname>Miller</surname> <given-names>HER</given-names>
</name>
<name>
<surname>Zager</surname> <given-names>M</given-names>
</name>
<name>
<surname>Newell</surname> <given-names>EW</given-names>
</name>
<name>
<surname>Gottardo</surname> <given-names>R</given-names>
</name>
</person-group>. <article-title>Single-cell immunology of SARS-CoV-2 infection</article-title>. <source>Nat Biotechnol</source> (<year>2022</year>) <volume>40</volume>:<fpage>30</fpage>&#x2013;<lpage>41</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41587-021-01131-y</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>