<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioinform.</journal-id>
<journal-title>Frontiers in Bioinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioinform.</abbrev-journal-title>
<issn pub-type="epub">2673-7647</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1191961</article-id>
<article-id pub-id-type="doi">10.3389/fbinf.2023.1191961</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioinformatics</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Omics data integration in computational biology viewed through the prism of machine learning paradigms</article-title>
<alt-title alt-title-type="left-running-head">Fouch&#xe9; and Zinovyev</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fbinf.2023.1191961">10.3389/fbinf.2023.1191961</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Fouch&#xe9;</surname>
<given-names>Aziz</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2244598/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zinovyev</surname>
<given-names>Andrei</given-names>
</name>
<xref ref-type="aff" rid="aff5">
<sup>5</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/106132/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Institut Curie</institution>, <institution>PSL Research University</institution>, <addr-line>Paris</addr-line>, <country>France</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Institut National de la Sant&#x00E9; et de la Recherche M&#x00E9;dicale</institution>, <addr-line>Paris</addr-line>, <country>France</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>CBIO-Centre for Computational Biology</institution>, <institution>ParisTech</institution>, <institution>PSL Research University</institution>, <addr-line>Paris</addr-line>, <country>France</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Ecole Normale Sup&#xe9;rieure Paris-Saclay</institution>, <addr-line>Cachan</addr-line>, <country>France</country>
</aff>
<aff id="aff5">
<sup>5</sup>
<institution>In Silico R&#x26;D</institution>, <institution>Evotec</institution>, <addr-line>Toulouse</addr-line>, <country>France</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1920825/overview">Elias S. Manolakos</ext-link>, National and Kapodistrian University of Athens, Greece</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2096950/overview">Shengquan Chen</ext-link>, Nankai University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1456798/overview">Somnath Tagore</ext-link>, Columbia University, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Aziz Fouch&#xe9;, <email>aziz.fouche@curie.fr</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>3</volume>
<elocation-id>1191961</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>03</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Fouch&#xe9; and Zinovyev.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Fouch&#xe9; and Zinovyev</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Important quantities of biological data can today be acquired to characterize cell types and states, from various sources and using a wide diversity of methods, providing scientists with more and more information to answer challenging biological questions. Unfortunately, working with this amount of data comes at the price of ever-increasing data complexity. This is caused by the multiplication of data types and batch effects, which hinders the joint usage of all available data within common analyses. Data integration describes a set of tasks geared towards embedding several datasets of different origins or modalities into a joint representation that can then be used to carry out downstream analyses. In the last decade, dozens of methods have been proposed to tackle the different facets of the data integration problem, relying on various paradigms. This review introduces the most common data types encountered in computational biology and provides systematic definitions of the data integration problems. We then present how machine learning innovations were leveraged to build effective data integration algorithms, that are widely used today by computational biologists. We discuss the current state of data integration and important pitfalls to consider when working with data integration tools. We eventually detail a set of challenges the field will have to overcome in the coming years.</p>
</abstract>
<kwd-group>
<kwd>single-cell</kwd>
<kwd>data integration</kwd>
<kwd>machine learning</kwd>
<kwd>batch effect</kwd>
<kwd>multi-omics</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Integrative Bioinformatics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>This last decade has witnessed a sharp increase in the amount and complexity of data produced for cellular biology, thanks to an ever-growing number of bulk and single-cell profiling assays. These technologies allowed scientists to study heterogeneous cell populations through many biological feature spaces (or <italic>modalities</italic>) such as mRNA expression (<xref ref-type="bibr" rid="B51">Klein et al., 2015</xref>; <xref ref-type="bibr" rid="B65">Macosko et al., 2015</xref>), DNA methylation (<xref ref-type="bibr" rid="B41">Guo et al., 2013</xref>) and chromatin accessibility (<xref ref-type="bibr" rid="B12">Buenrostro et al., 2015a</xref>; <xref ref-type="bibr" rid="B13">Buenrostro et al., 2015b</xref>), and protein abundance (<xref ref-type="bibr" rid="B2">Aebersold and Mann, 2003</xref>; <xref ref-type="bibr" rid="B94">Westermeier and Marouga, 2005</xref>; <xref ref-type="bibr" rid="B84">Tibes et al., 2006</xref>). These assays can be carried out either in bulk, which yields for each sample a single averaged molecular profile, or at the single-cell level, which provides an exquisite insight into cell states and types present in the cell population. In particular, carrying out biological assays at the single-cell level snapshots cells at various points of a dynamical process, which can then be leveraged for various applications such as lineage tracing (<xref ref-type="bibr" rid="B73">Schiebinger et al., 2019</xref>), transcriptional dynamics (<xref ref-type="bibr" rid="B54">La Manno et al., 2018</xref>), inference of transcriptional trajectories (<xref ref-type="bibr" rid="B22">Chen H. et al., 2019</xref>) and many more.</p>
<p>In addition, during the last few years, there have been several joint assays proposed to profile single cells through several modalities simultaneously, such as scM&#x26;T-seq for transcriptome and methylome (<xref ref-type="bibr" rid="B4">Angermueller et al., 2016</xref>), sc-GEM for genotype, transcriptome and methylome (<xref ref-type="bibr" rid="B25">Cheow et al., 2016</xref>), CITE-seq for transcriptome and surface proteins (<xref ref-type="bibr" rid="B77">Stoeckius et al., 2017</xref>), or SNARE-seq for transcriptome and chromatin accessibility (<xref ref-type="bibr" rid="B23">Chen S. et al., 2019</xref>). It is also worth mentioning spatial transcriptomics, which yields measurements from a small number of cells in each well while also providing positional information of cells within the biological tissue (<xref ref-type="bibr" rid="B75">St&#xe5;hl et al., 2016</xref>). Finally, important phenotypical information can be obtained from microscopic imaging data, such as whole slide imaging (<xref ref-type="bibr" rid="B69">Pantanowitz et al., 2011</xref>).</p>
<p>Hand-to-hand with the surge of biological modalities, there has been an explosion in the number of available datasets helped by various scientific initiatives to make biological data more easily available (<xref ref-type="bibr" rid="B26">Conesa and Beck, 2019</xref>); among these initiatives, one can mention atlases of entire organisms such as the Tabula Muris (<xref ref-type="bibr" rid="B72">Schaum et al., 2018</xref>) and Human (<xref ref-type="bibr" rid="B49">Tabula Sapiens Consortium et al., 2022</xref>) Consortia. We would also like to talk about disease-based atlas such as The Cancer Genome Atlas (TGCA) database (<xref ref-type="bibr" rid="B91">Weinstein et al., 2013</xref>), and the IMMUcan database (<xref ref-type="bibr" rid="B15">Camps et al., 2023</xref>) which provides an exquisite insight into the nature of tumor microenvironment. When tackling difficult biological questions, using data gathered across different sources or modalities is enticing. On the one hand, combining data from different sources helps to provide a comprehensive view of the biological object of interest. For example, it can facilitate the discovery of rare but relevant cell types or states, or help quantify the relative abundance of cell types across a collection of biological samples. On the other hand, having different modalities at their disposal allows scientists to link them together, possibly leading to exciting mechanistic discoveries. Finally, there can be an emergent property where analyzing a biological object through several modalities simultaneously could yield superior information compared to analyzing each modality individually.</p>
<p>Unfortunately, there are several obstacles to overcome before data from several sources and modalities can be used within an analysis pipeline. First, the multiplicity of sources comes at the price of all sorts of batch effects, as datasets can come from different replicas, technologies, individuals, or even species. Then, combining datasets containing measurements from different modalities is a major computational challenge, especially when samples are not linked across datasets, as there is no trivial common space to embed samples together. Therefore, there is a real need for methods and tools that would be able to tie together biological datasets across datasets (or <italic>batches</italic>) and modalities. In this review, we investigate this question through the prism of machine learning paradigms, and present how a few of these concepts are today widely used within popular, state-of-the-art data integration methods.</p>
</sec>
<sec id="s2">
<title>2 Data integration links biological datasets across batches or modalities</title>
<p>Data integration describes a set of problems that represent different facets of the question of tying together biological datasets across batches and modalities: <italic>vertical</italic>, <italic>horizontal</italic>, <italic>diagonal</italic> and <italic>mosaic</italic> integration (<xref ref-type="bibr" rid="B6">Argelaguet et al., 2021</xref>), which indicate the nature of anchors that exist between datasets (<xref ref-type="fig" rid="F1">Figure 1A</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Data integration describes a set of problems aiming to tie together data across different origins or modalities. <bold>(A)</bold> A biological object can be profiled through multiple batches (columns) and modalities (rows), and not all batches necessarily contain measurements for all modalities. <bold>(B)</bold> Vertical integration (VI) consists in using cells or samples as anchors to deduce links between features across modalities. <bold>(C)</bold> Horizontal integration (HI) consists in using overlapping features as anchors to jointly analyze data coming from different sources. <bold>(D)</bold> Diagonal integration (DI) consists in embedding together several batches with non-overlapping modalities. Mosaic integration (MI) is the problem of missing modalities inference.</p>
</caption>
<graphic xlink:href="fbinf-03-1191961-g001.tif"/>
</fig>
<p>In vertical integration (VI), each dataset contains a set of measurements carried out on the same set of samples (separate bulk experiments with matched samples in different modalities or single-cells measured through joint assays) (<xref ref-type="fig" rid="F1">Figure 1B</xref>). VI identifies links between biological features, such as scRNA-seq transcript counts and scATAC-seq peaks, which can help formulate mechanistic hypotheses across modalities. VI methods usually rely on dimensionality reduction, matrix factorization, or modeling. Some can be endowed with additional biological knowledge, such as pathway data and functional interaction between features across modalities.</p>
<p>Horizontal integration (HI) describes the complementary task where several datasets have been acquired in the same biological modality, allowing multiple batches to be expressed within a common features space (<xref ref-type="fig" rid="F1">Figure 1C</xref>). HI&#x2019;s primary use is to correct batch effects between datasets that can be explained by experimenter variation, different sequencing technologies, or inter-individual biological specificities (e.g., species, sex, or ethnicity). HI has been a very popular research topic for the last few years, and many HI tools have been proposed to this day. They can rely on a large variety of computational paradigms such as nearest neighbors, clustering, deep neural networks, matrix factorization, manifold alignment, and many more. Some tools may require additional priors, such as selecting a reference dataset or having access to cell types as labels.</p>
<p>When no trivial anchoring exists between datasets, diagonal (DI) or mosaic integration (MI) formalisms must be used. DI describes the framework where each dataset is measured in a different biological modality, while MI allows pairs of datasets to be measured in overlapping modalities (<xref ref-type="fig" rid="F1">Figure 1D</xref>). DI and MI are the most challenging facets of data integration and are subject to active research. Methods proposed to perform DI and MI usually rely on advanced machine learning paradigms capable of high levels of abstraction, such as deep neural networks, manifold alignment, or transport theory. Some tools operate in a completely unsupervised fashion, while others require additional information to help them bridge the gap between modalities.</p>
<p>Data integration of biological data is tightly related to several machine learning topics such as domain adaptation (<xref ref-type="bibr" rid="B68">Pan et al., 2010</xref>; <xref ref-type="bibr" rid="B99">You et al., 2019</xref>; <xref ref-type="bibr" rid="B35">Farahani et al., 2021</xref>), data fusion (<xref ref-type="bibr" rid="B21">Castanedo, 2013</xref>; <xref ref-type="bibr" rid="B37">Gao et al., 2020</xref>) and manifold alignment (<xref ref-type="bibr" rid="B89">Wang et al., 2011</xref>). Therefore, it is unsurprising to observe strategies leveraging similar machine learning paradigms such as supervised dimensionality reduction, matrix factorization, nearest neighbors, optimal transport, or deep autoencoders. Interestingly, new methods in all these domains go hand-to-hand with advances in machine learning, with many recent methods featuring advanced machine learning concepts. This is arguably a natural evolution as data complexity and quantity increase, which motivates the need for more powerful models capable of increased levels of abstraction.</p>
</sec>
<sec id="s3">
<title>3 Horizontal integration (HI) links batches anchored by their common modality</title>
<p>Horizontal integration (HI) describes the situation where several batches are all gathered in a common modality with overlapping feature spaces. It is worth noting that depending on the tool, there may only suffice that each pair of datasets contains an overlapping feature space (e.g., dataset A containing features <inline-formula id="inf1">
<mml:math id="m1">
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>, dataset B containing features <inline-formula id="inf2">
<mml:math id="m2">
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> and dataset C containing features <inline-formula id="inf3">
<mml:math id="m3">
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>). HI is a convenient framework in which cells can directly be compared across different batches due to their feature space overlap, which allows the use of natural concepts such as distances, neighborhoods, or similarity measures. Many tools have been proposed to tackle HI, and we gathered a non-exhaustive list of them in (<xref ref-type="table" rid="T1">Table 1</xref>). As we can see, these methods use various strategies to identify similar cells across batches and embed cells into a joint space. Some require additional information, such as reference datasets or cell labels. The remainder of this section is devoted to describing the main computational principles and machine learning paradigms HI methods rely on and providing some rationale and guidelines about each of them.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>A non-exhaustive list of horizontal integration (HI) tools aiming to jointly embed single-cell datasets measured in the same modality into a common space. BA, Bayesian; NN, Nearest Neighbors; DAE, Deep Autoencoders; DR, Dimensionality Reduction; IC, Iterative Clustering; MF, Matrix Factorization; MA, Manifold Alignment; RE, Regression; FR, Framework.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Tool</th>
<th align="center">Strategy</th>
<th align="center">Input</th>
<th align="center">Output</th>
<th align="center">Year</th>
<th align="center">References</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">ComBAT</td>
<td align="center">BA</td>
<td align="center">RNA-seq</td>
<td align="center">Gene space</td>
<td align="center">2007</td>
<td align="center">
<xref ref-type="bibr" rid="B48">Johnson et al. (2007)</xref>
</td>
</tr>
<tr>
<td align="center">MNN</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Gene space</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B42">Haghverdi et al. (2018)</xref>
</td>
</tr>
<tr>
<td align="center">scmap</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Clustering</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B50">Kiselev et al. (2018)</xref>
</td>
</tr>
<tr>
<td align="center">scvi</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, spatial</td>
<td align="center">Embedding</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B61">Lopez et al. (2018)</xref>
</td>
</tr>
<tr>
<td align="center">ingest</td>
<td align="center">DR</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B95">Wolf et al. (2018)</xref>
</td>
</tr>
<tr>
<td align="center">CONOS</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Graph</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B9">Barkas et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">Scanorama</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B44">Hie et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">scAlign</td>
<td align="center">DAE</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B47">Johansen and Quon (2019)</xref>
</td>
</tr>
<tr>
<td align="center">Harmony</td>
<td align="center">CL</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B52">Korsunsky et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">Seurat v3</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Gene space</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B78">Stuart et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">LIGER</td>
<td align="center">MF</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B93">Welch et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">DESC</td>
<td align="center">DAE</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B57">Li et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">BBKNN</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Graph</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B70">Pola&#x144;ski et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">SpaGE</td>
<td align="center">NN</td>
<td align="center">RNA-seq, spatial</td>
<td align="center">Embedding</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B1">Abdelaal et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">Tangram</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, spatial</td>
<td align="center">Embedding</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B10">Biancalani et al. (2021)</xref>
</td>
</tr>
<tr>
<td align="center">Canek</td>
<td align="center">NN</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B62">Loza et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">CAPITAL</td>
<td align="center">MA</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B79">Sugihara et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">SCISSOR</td>
<td align="center">RE</td>
<td align="center">RNA-seq</td>
<td align="center">Graph</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B80">Sun et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">Transmorph</td>
<td align="center">FR</td>
<td align="center">RNA-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B36">Fouch&#xe9; et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">DAPCA</td>
<td align="center">MF</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2023</td>
<td align="center">
<xref ref-type="bibr" rid="B67">Mirkes et al. (2023)</xref>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Many HI methods rely on manifold alignment strategies to integrate batches together (<xref ref-type="fig" rid="F2">Figure 2A</xref>), allowing them to consider the whole data structure instead of matching individual cells. Perhaps the oldest and most natural manifold alignment technique is Procrustes analysis (<xref ref-type="bibr" rid="B40">Gower, 1975</xref>), named after the mythical greek thug who cut or stretched his victims so that they fit the length of their bed. This is an old and intuitive machine learning paradigm mostly used for shape alignment that aims at projecting query datasets onto a reference one while only allowing simple transformations (rotation, rescaling, and shifting). Procrustes-based methods are not often used to integrate single-cell data, although some attempts can be found in the literature (<xref ref-type="bibr" rid="B34">Eto et al., 2018</xref>). First introduced to infer cell differentiation trajectories (<xref ref-type="bibr" rid="B73">Schiebinger et al., 2019</xref>), discrete optimal transport (OT) theory and its extensions (Gromov-Wasserstein, partial OT, unbalanced OT) is the most popular paradigm used for manifold alignment-based HI. It aims to align cells as discrete probability distributions represented as weighted point clouds in a metric space based on pairwise cell-cell cost matrices between batches that are often distance matrices. OT and its extensions have been successfully applied to horizontal and diagonal data integration (<xref ref-type="bibr" rid="B19">Cao et al., 2022b</xref>; <xref ref-type="bibr" rid="B28">Demetci et al., 2022</xref>). Manifold alignment-based HI is a powerful paradigm, but it can sometimes struggle to solve complex alignment tasks (for instance, when the structure of a dataset presents ambiguous symmetries or when some batches contain specific cell types that must not be aligned).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Horizontal integration describes the problem of embedding together datasets measured along the same biological modality. Different types of popular machine learning approaches are commonly used to match similar cells across batches. <bold>(A)</bold> Manifold alignment techniques find the projection that create the optimal overlap between two point clouds. <bold>(B)</bold> Nearest neighbors techniques identifies similar cells across datasets based on a similarity measure. <bold>(C)</bold> Deep autoencoders (AEC) learn a joint latent representation of the data in which batch effects are regressed out.</p>
</caption>
<graphic xlink:href="fbinf-03-1191961-g002.tif"/>
</fig>
<p>Another class of HI methods seeks similar cells across batches, operating at the single-cell level rather than at a global level (<xref ref-type="fig" rid="F2">Figure 2B</xref>). Some are based on the nearest neighbors approach like mutual nearest neighbors (MNN) (<xref ref-type="bibr" rid="B42">Haghverdi et al., 2018</xref>), CONOS (<xref ref-type="bibr" rid="B9">Barkas et al., 2019</xref>), Scanorama (<xref ref-type="bibr" rid="B44">Hie et al., 2019</xref>), Seurat (<xref ref-type="bibr" rid="B71">Satija et al., 2015</xref>; <xref ref-type="bibr" rid="B14">Butler et al., 2018</xref>; <xref ref-type="bibr" rid="B78">Stuart et al., 2019</xref>; <xref ref-type="bibr" rid="B43">Hao et al., 2021</xref>) that include different integration schemes such as CCA and robust PCA (RPCA), or BBKNN (<xref ref-type="bibr" rid="B70">Pola&#x144;ski et al., 2020</xref>). All nearest neighbors-based methods rely on the hypothesis that batch effects are almost orthogonal to biological effects, which would allow identifying similar cells across batches through simple orthogonal projection. They then apply various strategies to end up with a joint representation of cells like correction vectors or joint graph construction. These methods tend to work best when facing slight to moderate batch effects and generally fail when batch effects are far from being orthogonal to relevant biological signals. They tend to scale well to large datasets thanks to various optimizations during nearest neighbors computation like nearest neighbors descent (<xref ref-type="bibr" rid="B30">Dong et al., 2011</xref>). Another metric-based approach is described in Harmony (<xref ref-type="bibr" rid="B52">Korsunsky et al., 2019</xref>), which is probably the most used tool in practice for HI of single-cell data. It uses an iterative algorithm of successive biased clustering across batches and correction. First, cells are clustered across datasets with such a bias that penalizes clusters of cells with a homogeneous batch of origin. Then, cells of a given cluster are pooled towards each other. An optimality criterion is tested at each iteration to assess whether batch mixing is sufficient, using a local purity metric called Local Inverse Simpson&#x2019;s Index (LISI). Due to its simplicity and availability with both Python and R packages, Harmony is widely used today and still achieves respectable results in benchmarks (<xref ref-type="bibr" rid="B3">Anaissi et al., 2022</xref>) despite being limited when facing strong batch effects (<xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>).</p>
<p>Deep autoencoders (DAEs) (and more recently variational autoencoders) have been popular tools in single-cell for a few years already and excel at performing a variety of complex preprocessing tasks, such as dimensionality reduction (<xref ref-type="bibr" rid="B90">Wang and Gu, 2018</xref>), or denoising and correcting dropouts (<xref ref-type="bibr" rid="B33">Eraslan et al., 2019</xref>), as well as acting as generative models (<xref ref-type="bibr" rid="B87">Trong et al., 2020</xref>). DAEs are neural networks that leverage a bottleneck structure to learn a compressed data representation in a low dimensional space, which can then be exploited for various tasks (<xref ref-type="fig" rid="F2">Figure 2C</xref>). DAE is a powerful framework to carry out horizontal data integration with tools such as scvi (<xref ref-type="bibr" rid="B61">Lopez et al., 2018</xref>), scAlign (<xref ref-type="bibr" rid="B47">Johansen and Quon, 2019</xref>) or DESC (<xref ref-type="bibr" rid="B57">Li et al., 2020</xref>). In particular, scANVI, part of the scvi framework, is the top performer tool in the (<xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>) atlas-scale benchmark. DAEs generally have high computational capabilities thanks to the fact to be able to exploit GPU acceleration during training. The main downside of DAEs is the large amounts of data necessary for their training and their lack of interpretability, though there are efforts to improve on the latter point (<xref ref-type="bibr" rid="B81">Svensson et al., 2020</xref>; <xref ref-type="bibr" rid="B86">Treppner et al., 2022</xref>).</p>
<p>In an attempt to organize these methods into a common framework, we introduced Transmorph (<xref ref-type="bibr" rid="B36">Fouch&#xe9; et al., 2022</xref>), an open-source computational framework that allows the user to assemble custom HI pipelines from basic algorithmic blocks. This framework focuses on methods that combine a matching step, identifying similar cells across batches, and an embedding step, where these correspondences are used to generate a joint representation of all datasets. Transmorph also gives access to pre-build HI pipelines, HI quality assessment routines, benchmarking datasets and easy access to other state-of-the-art HI tools such as Harmony (<xref ref-type="bibr" rid="B52">Korsunsky et al., 2019</xref>) and scvi (<xref ref-type="bibr" rid="B61">Lopez et al., 2018</xref>). We hope to see more initiatives deployed in the next years in this sense to provide frameworks that can help organize the field of HI methods.</p>
<p>Despite the myriad approaches proposed to tackle HI, it remains challenging today to correct strong batch effects. For instance, (<xref ref-type="bibr" rid="B85">Tran et al., 2020</xref>; <xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>), showed that if several methods can satisfyingly remove moderate batch effects, integrating datasets across species remains difficult for unsupervised methods which do not require cell labeling information. Also, many methods rely on finding first an overlapping feature space between all datasets, which can be an obstacle when building large atlases combining many batches of varying quality, where the number of common features can shrink drastically. Finally, the problem of selecting appropriate metrics to assess data integration quality is still difficult. Most benchmarks use a mixture of metrics to measure different aspects of the data integration task such as batch mixture, label clustering or topology preservation, depending on the information available:<list list-type="simple">
<list-item>
<p>&#x2022; Batch mixture metrics such as batch-LISI are commonly used to measure how much the data integration procedure brought cells from different datasets close to one another. These metrics are popular because they do not require additional information, such as cell types or states, and can be used as unsupervised tools. Unfortunately, a good integration does not necessarily imply good batch mixture metrics, as two datasets without overlapping cell types should not be mixed after integration; similarly, projecting all datasets together onto a single point would result in perfect batch mixing, but all the biological information would be lost. For these reasons, even though batch mixture metrics are quite informative and widely used, most benchmarks also include other integration metrics to compensate for these limitations.</p>
</list-item>
<list-item>
<p>&#x2022; Label clustering metrics, such as normalized mutual information or adjusted Rand index, provide an additional axis to measure data integration quality by assessing if cells of similar type cluster together after integration. Label clustering metrics are usually quite good for controlling the data integration quality if cell types can be identified confidently. The main downside of these metrics is the necessity to have high-confidence cell labels available before integration, which is often not the case (especially as one of the purposes of data integration is to be carried out before clustering and cell type inference).</p>
</list-item>
<list-item>
<p>&#x2022; Finally, topology preservation metrics assess how data integration has preserved relations between the different cells and penalize cases where cells that were close before integration have been brought far apart by the algorithm (meaning cells that were initially similar but are dissimilar after integration). Topology can be biology-driven by observing the conservation of signals related to specific cell processes, such as cell cycle or other transcriptomic trajectories, or data-driven with algorithms as simple as comparing the <italic>k</italic>-nearest neighbors of a cell before and after integration and penalizing the differences.</p>
</list-item>
</list>
</p>
<p>Evaluating the quality of a HI can be daunting, as shown by the large variety of metrics that have been developed for it. In practice, we often use a batch mixture metric such as LISI, complemented by a secondary metric that can be either a label clustering metric if high-confidence labels are available and a topology preservation metric otherwise.</p>
</sec>
<sec id="s4">
<title>4 Vertical integration (VI) connects modalities measured in the same cells</title>
<p>Vertical integration (VI) uses several datasets containing individual measurements from the same cells obtained from joint single-cell assays measured through different biological features (e.g., gene expression and chromatin accessibility) to infer relations between the different modalities (<xref ref-type="table" rid="T2">Table 2</xref>). VI is usually declined into two variants, namely, <italic>local</italic> VI and <italic>global</italic> VI. Local VI identifies links between individual features (such as genes and methylated promoters), and can be used to formulate hypotheses of direct or indirect biological interactions between the omics layers (e.g., gene expression and accessibility of a chromatin region), with methods like LMM (<xref ref-type="bibr" rid="B88">Van Der Wijst et al., 2018</xref>) or Spearman&#x2019;s rank correlation coefficient (<xref ref-type="bibr" rid="B27">Cuomo et al., 2020</xref>). On the other hand, global VI links features across different modalities via global factors that can be related to biological processes (e.g., identifying a group of genes and chromatin regions to correspond to proliferation activity).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>A non-exhaustive list of global vertical integration (VI) tools that can be used to learn relations between features across modalities from joint single-cell assays. FC, Feature Correlation; MD, Matrix Decomposition; NN, Nearest Neighbors; DAE, Deep Autoencoders; TM, Topic Modelling.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Tool</th>
<th align="center">Strategy</th>
<th align="center">Input</th>
<th align="center">Year</th>
<th align="center">References</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">CCA</td>
<td align="center">FC</td>
<td align="center">Any</td>
<td align="center">1936</td>
<td align="center">
<xref ref-type="bibr" rid="B45">Hotelling (1992)</xref>
</td>
</tr>
<tr>
<td align="center">RGCCA</td>
<td align="center">FC</td>
<td align="center">Any</td>
<td align="center">2011</td>
<td align="center">
<xref ref-type="bibr" rid="B83">Tenenhaus and Tenenhaus (2011)</xref>
</td>
</tr>
<tr>
<td align="center">JIVE</td>
<td align="center">MD</td>
<td align="center">Any</td>
<td align="center">2013</td>
<td align="center">
<xref ref-type="bibr" rid="B60">Lock et al. (2013)</xref>
</td>
</tr>
<tr>
<td align="center">SGCCA</td>
<td align="center">FC</td>
<td align="center">Any</td>
<td align="center">2014</td>
<td align="center">
<xref ref-type="bibr" rid="B82">Tenenhaus et al. (2014)</xref>
</td>
</tr>
<tr>
<td align="center">MOFA</td>
<td align="center">MD</td>
<td align="center">Any</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B7">Argelaguet et al. (2018)</xref>; <xref ref-type="bibr" rid="B5">Argelaguet et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">DIABLO</td>
<td align="center">FC</td>
<td align="center">Any</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B74">Singh et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">scAI</td>
<td align="center">MD</td>
<td align="center">RNA-seq, epigenomic</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B46">Jin et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">Seurat v4</td>
<td align="center">NN</td>
<td align="center">Any</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B43">Hao et al. (2021)</xref>
</td>
</tr>
<tr>
<td align="center">scMM</td>
<td align="center">DAE</td>
<td align="center">Any</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B66">Minoura et al. (2021)</xref>
</td>
</tr>
<tr>
<td align="center">SMILE</td>
<td align="center">DAE</td>
<td align="center">Any</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B97">Xu et al. (2022b)</xref>
</td>
</tr>
<tr>
<td align="center">MIRA</td>
<td align="center">TM</td>
<td align="center">RNA-seq, chromatin state</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B64">Lynch et al. (2022)</xref>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>A family of global VI tools are based on a methodology inspired by canonical correlation analysis (CCA) (<xref ref-type="bibr" rid="B45">Hotelling, 1992</xref>), which use joint feature measurements across datasets to identify correlated features across modalities (<xref ref-type="fig" rid="F3">Figure 3A</xref>). RGCCA (<xref ref-type="bibr" rid="B83">Tenenhaus and Tenenhaus, 2011</xref>) extended this framework to simultaneously allow the analysis of more than 2 datasets. These concepts have been refined in (<xref ref-type="bibr" rid="B82">Tenenhaus et al., 2014</xref>) and DIABLO (<xref ref-type="bibr" rid="B74">Singh et al., 2019</xref>) to achieve better feature selection.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Two main strategies are used for vertical integration of joint assays. <bold>(A)</bold> Local strategies link features across modalities via pairwise correspondence. <bold>(B)</bold> Global strategies link features across modalities via common biological factors.</p>
</caption>
<graphic xlink:href="fbinf-03-1191961-g003.tif"/>
</fig>
<p>On the other hand, other popular global VI tools are based on matrix decomposition algorithms (<xref ref-type="fig" rid="F3">Figure 3B</xref>) (<xref ref-type="bibr" rid="B60">Lock et al., 2013</xref>; <xref ref-type="bibr" rid="B7">Argelaguet et al., 2018</xref>; <xref ref-type="bibr" rid="B5">2020</xref>; <xref ref-type="bibr" rid="B46">Jin et al., 2020</xref>). These tools generally aim to decompose each data matrix into a component explained by global factors, a component containing dataset-specific and modality-specific factors, and a noise term. They mostly differ by their exact decomposition model and specific strategies used to infer its parameters.</p>
<p>If deep autoencoders did wonders for HI, they were also successfully applied to VI problems (<xref ref-type="bibr" rid="B66">Minoura et al., 2021</xref>) by using two distinct encoders and decoders using a shared latent space into which both modalities are projected. This strategy notably allows the network to &#x201c;translate&#x201d; a modality into another. We can also mention the recent MIRA method (<xref ref-type="bibr" rid="B64">Lynch et al., 2022</xref>), which leverages a variational autoencoder approach to learn gene expression and chromatin accessibility shared topics.</p>
<p>Overall, the VI framework has allowed the growth of methods taking advantage of the powerful sample anchoring across datasets, with many approaches proposed inspired by statistics and machine learning. A few important benchmarks have been carried out to assess the quality of VI tools, notably (<xref ref-type="bibr" rid="B16">Cantini et al., 2021</xref>) which focuses on joint dimensionality reduction (jDR) methods. Due to the difficulty of setting up joint assays and the inability of these methods to function without matched cells, there is a crucial need for diagonal integration (DI) tools that aim to integrate datasets across batches and modalities.</p>
</sec>
<sec id="s5">
<title>5 Diagonal and mosaic integration jointly embed non- or partially-anchored datasets</title>
<p>Diagonal integration (DI) and mosaic integration (MI) are two data integration frameworks for single-cell data that do not require datasets to be acquired through matched biological assays (<xref ref-type="table" rid="T3">Table 3</xref>). In this paragraph, we use DI indistinguishably from MI. The goal is to leverage datasets structure and possibly external information, such as genomic locations, pathways, or partial sample or modality overlap to infer complete bonds between cells across modalities without relying on explicit sample anchoring (<xref ref-type="fig" rid="F4">Figures 4A, B</xref>). DI generally aims to build a joint embedding of datasets into a common latent space, while MI focuses on inferring missing modalities from partially anchored datasets. Let us focus on the two main families of methods that exist for tackling DI: manifold alignment and deep autoencoders. These two machine learning paradigms can handle high levels of abstraction, which seems required to tackle DI in the general case.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>A non-exhaustive list of diagonal (DI) and mosaic integration (MI) tools that integrate single-cell datasets gathered across different biological samples and modalities. MA, Manifold Alignment; MF, Matrix Factorization; MMD, Maximum Mean Discrepancy; NN, Nearest Neighbors; DAE, Deep Autoencoders; OT, Optimal Transport; GW, Gromov-Wasserstein; LI, Linear Inference.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Tool</th>
<th align="center">Strategy</th>
<th align="center">Input</th>
<th align="center">Output</th>
<th align="center">Year</th>
<th align="center">References</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">MATCHER</td>
<td align="center">MA</td>
<td align="center">RNA-seq, epigenetic</td>
<td align="center">Gen. Model</td>
<td align="center">2017</td>
<td align="center">
<xref ref-type="bibr" rid="B92">Welch et al. (2017)</xref>
</td>
</tr>
<tr>
<td align="center">CoupledNMF</td>
<td align="center">MF</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Clustering</td>
<td align="center">2018</td>
<td align="center">
<xref ref-type="bibr" rid="B32">Duren et al. (2018)</xref>
</td>
</tr>
<tr>
<td align="center">MMD-MA</td>
<td align="center">MMD</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B59">Liu et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">LIGER</td>
<td align="center">MF</td>
<td align="center">RNA-seq, ATAC-seq, scMethyl</td>
<td align="center">Embedding</td>
<td align="center">2019</td>
<td align="center">
<xref ref-type="bibr" rid="B93">Welch et al. (2019)</xref>
</td>
</tr>
<tr>
<td align="center">UnionCom</td>
<td align="center">MA</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B17">Cao et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">bindSC</td>
<td align="center">NN</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B31">Dou et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">SCIM</td>
<td align="center">DAE</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2020</td>
<td align="center">
<xref ref-type="bibr" rid="B76">Stark et al. (2020)</xref>
</td>
</tr>
<tr>
<td align="center">MultiVI</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Embedding</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B8">Ashuach et al. (2021)</xref>
</td>
</tr>
<tr>
<td align="center">COBOLT</td>
<td align="center">DAE</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2021</td>
<td align="center">
<xref ref-type="bibr" rid="B39">Gong et al. (2021)</xref>
</td>
</tr>
<tr>
<td align="center">Pamona</td>
<td align="center">OT</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B19">Cao et al. (2022b)</xref>
</td>
</tr>
<tr>
<td align="center">Polarbear</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B100">Zhang et al. (2022a)</xref>
</td>
</tr>
<tr>
<td align="center">GLUE</td>
<td align="center">DAE</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B20">Cao and Gao (2022)</xref>
</td>
</tr>
<tr>
<td align="center">SCOT</td>
<td align="center">GW</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B28">Demetci et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">scJoint</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B58">Lin et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">sciCAN</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B96">Xu et al. (2022a)</xref>
</td>
</tr>
<tr>
<td align="center">scDART</td>
<td align="center">DAE</td>
<td align="center">RNA-seq, ATAC-seq</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B101">Zhang et al. (2022b)</xref>
</td>
</tr>
<tr>
<td align="center">StabMap</td>
<td align="center">LI</td>
<td align="center">Any</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B38">Ghazanfar et al. (2022)</xref>
</td>
</tr>
<tr>
<td align="center">UINMF</td>
<td align="center">MF</td>
<td align="center">RNA-seq, ATAC-seq, spatial</td>
<td align="center">Embedding</td>
<td align="center">2022</td>
<td align="center">
<xref ref-type="bibr" rid="B53">Kriebel and Welch (2022)</xref>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Several strategies can be carried out to tackle the diagonal integration computational challenge <bold>(A)</bold> A biological object (e.g., a population of cells). can be profiled using different assays, without obvious means to link both representations. <bold>(B)</bold> Knowledge of interaction between features across modalities can be obtained from vertical integration of external datasets generated using joint assays. This information can then be leveraged to compare cells between batches even if they are not expressed in the same modality, which allows to use horizontal integration tools. <bold>(C)</bold> Datasets can be independently encoded into abstractions that can then be matched in an unsupervised fashion to build a joint representation of datasets. <bold>(D)</bold> Datasets can be jointly encoded into a unique abstraction, for instance through a learning process using a deep autoencoder framework, that can then be used as a joint embedding of datasets.</p>
</caption>
<graphic xlink:href="fbinf-03-1191961-g004.tif"/>
</fig>
<p>Manifold alignment methods (<xref ref-type="bibr" rid="B92">Welch et al., 2017</xref>; <xref ref-type="bibr" rid="B59">Liu et al., 2019</xref>; <xref ref-type="bibr" rid="B17">Cao et al., 2020</xref>; <xref ref-type="bibr" rid="B19">Cao et al., 2022b</xref>; <xref ref-type="bibr" rid="B28">Demetci et al., 2022</xref>) for DI operate similarly as in the HI case and work under the assumption stating that smooth point clouds alignment corresponds to meaningful biological correspondence (<xref ref-type="fig" rid="F4">Figure 4C</xref>). This allows them to work in an unsupervised fashion without requiring additional knowledge other than data matrices. Despite working accurately in some cases, it has been shown this hypothesis is far from being universal (<xref ref-type="bibr" rid="B98">Xu and McCord, 2022</xref>). In this article, the authors show that under some simple data tweaking, such as missing cell types or different sample sizes, manifold alignment DI methods can generate erroneous embeddings featuring clusters with mixed cell types. This is concerning, as validating DI is a challenging task, given that it is rarely the case to have reliable cell type labels across modalities at disposal. Therefore, we suggest that these unsupervised manifold alignment methods must be used carefully and only when integration quality control is feasible. In other cases, it is preferable to choose another DI method that allows the user to provide additional information that helps bridge the gap across modalities.</p>
<p>As for HI and VI, deep autoencoders are powerful tools for solving DI tasks, with several advantages. First, they can take advantage of GPU acceleration built in deep learning libraries to greatly speed up the training process, and naturally scale to very large datasets. The second benefit of using these neural networks is that they offer the possibility to train a separate encoder and decoder for each biological modality, which helps capture modality-specific factors compared to manifold alignment algorithms where all omics layers are treated similarly. These separate encoders generally share a joint latent space (<xref ref-type="fig" rid="F4">Figure 4D</xref>), with some form of penalty to force latent representations to overlap. They also present an algorithmic structure that facilitates the introduction of external biological guidance, like in the GLUE tool (<xref ref-type="bibr" rid="B20">Cao and Gao, 2022</xref>), which uses a guidance graph as prior knowledge about functional relationships between features across modalities. We would also like to mention in this category the Polarbear tool (<xref ref-type="bibr" rid="B100">Zhang R. et al., 2022</xref>), which leverages deep autoencoders to notably translate single-cell data between RNA-seq and ATAC-seq.</p>
<p>To the best of our knowledge, there do not exist at the time of writing a large-scale, independent benchmark of DI methods like for HI (<xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>). This is arguably difficult to set up due to the number of single-cell modalities available today, given the fact that, in addition, not all methods can deal with all modalities. Some may also require specific prior knowledge, and output type may vary. Furthermore, there is a lack of reliable metrics for assessing the quality of DI methods and real-life benchmarking datasets. A first breakthrough is to note in this direction, with a NIPS single-cell analysis competition organized recently which gave access to a public multimodal dataset containing single-cell gene expression, protein expression, and chromatin accessibility using CITE-seq and Multiome (<xref ref-type="bibr" rid="B55">Lance et al., 2022</xref>). With the democratization of such datasets, benchmarking DI methods will become more accessible, which will help standardize the field and identify the best-performing methods for each scenario.</p>
<p>To finish, there is a growing interest in integrating single-cell data with other related data modalities, such as whole slide images or spatial transcriptomics. There is a particular interest in deconvoluting spatial transcriptomic spots by integrating them with a single-cell RNA-seq dataset obtained from a similar same tissue. This is a current challenge, and several methods have been proposed for this task, notably benchmarked in (<xref ref-type="bibr" rid="B56">Li et al., 2022</xref>).</p>
<p>Overall, DI is arguably the most challenging data integration problem, and solving it is still a very active research area. This very convenient data integration paradigm is extremely versatile, as it theoretically does not need any anchoring (cells or features) between the different datasets. In practice, if many DI tools indeed work in a completely unsupervised way leveraging data topology such as MMD-MA (<xref ref-type="bibr" rid="B59">Liu et al., 2019</xref>), Pamona (<xref ref-type="bibr" rid="B20">Cao and Gao, 2022</xref>) or SCOT (<xref ref-type="bibr" rid="B28">Demetci et al., 2022</xref>), others require additional information to bridge the gap between modalities like GLUE (<xref ref-type="bibr" rid="B18">Cao et al., 2022a</xref>) or MultiVI (<xref ref-type="bibr" rid="B8">Ashuach et al., 2021</xref>) which can take a covariate design matrix as an optional parameter. For the moment, it appears that these biased methods offer more control on the results, as data topology can be misleading in practice and yield aberrant results (<xref ref-type="bibr" rid="B98">Xu and McCord, 2022</xref>). Therefore, using DI tools that can be enriched with biological context seems to be the best choice in the applications where such context can be obtained in a reliable way, typically when integrating datasets where strong covariates exist between modalities.</p>
</sec>
<sec sec-type="discussion" id="s6">
<title>6 Discussion</title>
<p>Data integration consists of distinct challenges depending on the anchoring that exists between datasets, and each facet of DI requires distinct tools that leverage various algorithmic strategies. For instance, metric-based methods excel at solving HI tasks, whereas linear matrix analysis methods excel at solving VI tasks. Machine learning paradigms with high abstraction levels, such as manifold alignment methods and deep neural networks, are excellent assets for dealing with DI and MI problems, the latter also performing well at HI and VI tasks. Overall, VI methods are pretty good at solving the task, HI methods are capable of dealing with small to moderate batch effects but still struggle to mitigate significant batch effects such as inter-species data, and DI/MI problems are arguably still unsolved in the general case.</p>
<p>We talked about the Transmorph framework that articulates computational blocks to conceive HI pipelines, but this is not the only framework that exists which is related to data integration. We can cite MUON (<xref ref-type="bibr" rid="B11">Bredikhin et al., 2022</xref>), which facilitates the handling of data consisting of different modalities, Polyphony (<xref ref-type="bibr" rid="B24">Cheng et al., 2022</xref>), which carries out transfer learning across datasets by leveraging data integration algorithms, or SinCast (<xref ref-type="bibr" rid="B29">Deng et al., 2022</xref>) which is specialized in cell type inference by mapping a query onto an atlas.</p>
<p>It is essential to note that there are important pitfalls to data integration that must not be overlooked. The primary issue that can be encountered is named <italic>overcorrection</italic> and describes an undesirable event where a data integration method incorrectly aligns cells that do not share the same biological type or state. This typically happens when batch effects are too strong, when a dataset contains specific cell types, when cell type distribution is highly imbalanced, or when there is little anchoring between batches. Overcorrection can be difficult to detect when there is no easy access to cell labels and is a critical issue that hinder every subsequent analysis step. Indeed, it can lead to cells belonging to the same cluster without sharing critical biological properties such as cell type or states. Other issues are worth noting even though they are not exclusive to the data integration task, such as the difficulty in differentiating between true zeros and missing values in RNA-seq datasets or the fact that different modalities are often expressed using different data types (e.g., binary or integer data) which may be difficult to handle jointly within mathematical frameworks. Finally, data integration tools based on abstract machine learning paradigms such as deep autoencoders often comes at the cost of a decrease in model interpretability which is an important downside for any health-related application. However, many efforts are made to overcome this issue (<xref ref-type="bibr" rid="B81">Svensson et al., 2020</xref>; <xref ref-type="bibr" rid="B86">Treppner et al., 2022</xref>) and we expect to see many more in the years to come.</p>
<p>There is always an urgent need for large-scale, independent benchmarks like the HI benchmark proposed in (<xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>), or the VI benchmark carried out in (<xref ref-type="bibr" rid="B16">Cantini et al., 2021</xref>). To the best of our knowledge, there is still a lack of large-scale independent DI and MI benchmarks. Two things are necessary to carry out such benchmarks: high-quality datasets and reliable metrics. A list of potential datasets can be found in (<xref ref-type="bibr" rid="B6">Argelaguet et al., 2021</xref>). There is no clear consensus about which quality assessment metric to use, and most benchmarks like (<xref ref-type="bibr" rid="B63">Luecken et al., 2022</xref>) opt for a mixture of metrics that cover several aspects of data integration: conservation of biological variance (CBV) metrics which measure how close similar cells (type or state) are after integration, and removal of batch effects (RBE) metrics. Some CBV metrics are label-based, such as normalized mutual information (NMI), adjusted Rand index (ARI), average silhouette width (ASW), class local inverse Simpson&#x2019;s index (cLISI), isolated label F1 (ILF) and isolated label silhouette (ILS), others are label-free and generally assess the conservation of biological processes such as cell cycle, highly variable genes, and transcriptomic trajectories. RBE metrics include batch-PC regression, batch-ASW, graph connectivity, iLISI, and kBet. We often observe a tradeoff between CBV and RBE, which can lead to different methods choice depending on the application, whether it is preferable to have good dataset mixing or conservation of subtle biological signals.</p>
<p>To conclude, years of algorithmic and computational advances made it possible to solve most HI and VI problems with satisfying performance, with only the most complicated instances still being problematic (e.g., HI of many batches with strong batch effects). Solving DI and MI is the next computational challenge. The most promising approaches that have been developed to tackle it are based on deep learning models, particularly deep autoencoders. It has been shown that purely unsupervised DI may not be a well-posed problem and could suffer fundamental flaws (<xref ref-type="bibr" rid="B98">Xu and McCord, 2022</xref>), which greatly incentivizes using knowledge-driven tools that allow the user to include external information to enhance models with functional information that link features across modalities. Finally, apart from developing new tools, there is also an urgent need to enrich the data integration ecosystem with organizing frameworks, standardized benchmarks, datasets, and quality assessment metrics.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Author contributions</title>
<p>AF and AZ wrote the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the French government under the management of Agence Nationale de la Recherche as part of the &#x201c;Investissements d&#x2019;avenir&#x201d; program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute). These funding sources had no role in the design, execution, and interpretation of the results of this study.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>AZ was employed by the Evotec company.</p>
<p>The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abdelaal</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Mourragui</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mahfouz</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Reinders</surname>
<given-names>M. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Spage: Spatial gene enhancement using scrna-seq</article-title>. <source>Nucleic acids Res.</source> <volume>48</volume>, <fpage>e107</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa740</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aebersold</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mann</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Mass spectrometry-based proteomics</article-title>. <source>Nature</source> <volume>422</volume>, <fpage>198</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1038/nature01511</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Anaissi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zandavi</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Suleiman</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Alyassine</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Braytee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vafaee</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A benchmark of pre-processing effect on single cell RNA sequencing integration methods. Preprint</article-title>, In <comment>Review</comment>. <pub-id pub-id-type="doi">10.21203/rs.3.rs-2249309/v1</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Angermueller</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>H. J.</given-names>
</name>
<name>
<surname>Macaulay</surname>
<given-names>I. C.</given-names>
</name>
<name>
<surname>Teng</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>T. X.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Parallel single-cell sequencing links transcriptional and epigenetic heterogeneity</article-title>. <source>Nat. methods</source> <volume>13</volume>, <fpage>229</fpage>&#x2013;<lpage>232</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.3728</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argelaguet</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Arnol</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bredikhin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Deloro</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Velten</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Marioni</surname>
<given-names>J. C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>MOFA&#x2b;: A statistical framework for comprehensive integration of multi-modal single-cell data</article-title>. <source>Genome Biol.</source> <volume>21</volume>, <fpage>111</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-02015-1</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argelaguet</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Cuomo</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Stegle</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Marioni</surname>
<given-names>J. C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Computational principles and challenges in single-cell data integration</article-title>. <source>Nat. Biotechnol.</source> <volume>39</volume>, <fpage>1202</fpage>&#x2013;<lpage>1215</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-021-00895-7</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argelaguet</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Velten</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Arnol</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Dietrich</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zenz</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Marioni</surname>
<given-names>J. C.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Multi-omics factor analysis&#x2014;A framework for unsupervised integration of multi-omics data sets</article-title>. <source>Mol. Syst. Biol.</source> <volume>14</volume>, <fpage>e8124</fpage>. <pub-id pub-id-type="doi">10.15252/msb.20178124</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ashuach</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gabitto</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Yosef</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MultiVI: Deep generative model for the integration of multi-modal data</article-title>. <source>Bioinformatics</source> <volume>2021</volume>. <pub-id pub-id-type="doi">10.1101/2021.08.20.457057</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barkas</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Petukhov</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Nikolaeva</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lozinsky</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Demharter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Khodosevich</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Joint analysis of heterogeneous single-cell rna-seq dataset collections</article-title>. <source>Nat. methods</source> <volume>16</volume>, <fpage>695</fpage>&#x2013;<lpage>698</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-019-0466-z</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Biancalani</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Scalia</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Buffoni</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Avasthi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Sanger</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Deep learning and alignment of spatially resolved single-cell transcriptomes with tangram</article-title>. <source>Nat. methods</source> <volume>18</volume>, <fpage>1352</fpage>&#x2013;<lpage>1362</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-021-01264-7</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bredikhin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kats</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Stegle</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>MUON: Multimodal omics analysis framework</article-title>. <source>Genome Biol.</source> <volume>23</volume>, <fpage>42</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-021-02577-8</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Buenrostro</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H. Y.</given-names>
</name>
<name>
<surname>Greenleaf</surname>
<given-names>W. J.</given-names>
</name>
</person-group> (<year>2015a</year>). <article-title>Atac-seq: A method for assaying chromatin accessibility genome-wide</article-title>. <source>Curr. Protoc. Mol. Biol.</source> <volume>109</volume>, <fpage>21</fpage>. <pub-id pub-id-type="doi">10.1002/0471142727.mb2129s109</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Buenrostro</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Litzenburger</surname>
<given-names>U. M.</given-names>
</name>
<name>
<surname>Ruff</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gonzales</surname>
<given-names>M. L.</given-names>
</name>
<name>
<surname>Snyder</surname>
<given-names>M. P.</given-names>
</name>
<etal/>
</person-group> (<year>2015b</year>). <article-title>Single-cell chromatin accessibility reveals principles of regulatory variation</article-title>. <source>Nature</source> <volume>523</volume>, <fpage>486</fpage>&#x2013;<lpage>490</lpage>. <pub-id pub-id-type="doi">10.1038/nature14590</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Butler</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hoffman</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Smibert</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Papalexi</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Integrating single-cell transcriptomic data across different conditions, technologies, and species</article-title>. <source>Nat. Biotechnol.</source> <volume>36</volume>, <fpage>411</fpage>&#x2013;<lpage>420</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.4096</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Camps</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>No&#xeb;l</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Liechti</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Massenet-Regad</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Rigade</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>G&#xf6;tz</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2023</year>). <article-title>Meta-analysis of human cancer single-cell rna-seq datasets using the immucan database</article-title>. <source>Cancer Res.</source> <volume>83</volume>, <fpage>363</fpage>&#x2013;<lpage>373</lpage>. <pub-id pub-id-type="doi">10.1158/0008-5472.can-22-0074</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cantini</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zakeri</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hernandez</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Naldi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Thieffry</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Remy</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Benchmarking joint multi-omics dimensionality reduction approaches for the study of cancer</article-title>. <source>Nat. Commun.</source> <volume>12</volume>, <fpage>124</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-20430-7</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Unsupervised topological alignment for single-cell multi-omics integration</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>i48</fpage>&#x2013;<lpage>i56</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa443</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2022a</year>). <article-title>A unified computational framework for single-cell data integration with optimal transport</article-title>. <source>Nat. Commun.</source> <volume>13</volume>, <fpage>7419</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-35094-8</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2022b</year>). <article-title>Manifold alignment for heterogeneous single-cell multi-omics data integration using pamona</article-title>. <source>Bioinformatics</source> <volume>38</volume>, <fpage>211</fpage>&#x2013;<lpage>219</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab594</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>Z.-J.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Multi-omics single-cell data integration and regulatory inference with graph-linked embedding</article-title>. <source>Nat. Biotechnol.</source> <volume>40</volume>, <fpage>1458</fpage>&#x2013;<lpage>1466</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-022-01284-4</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Castanedo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A review of data fusion techniques</article-title>. <source>Sci. world J.</source> <volume>2013</volume>, <fpage>1</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1155/2013/704504</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Albergante</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Hsu</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Lareau</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Lo Bosco</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). <article-title>Single-cell trajectories reconstruction, exploration and mapping of omics data with stream</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>1903</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-09670-4</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lake</surname>
<given-names>B. B.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019b</year>). <article-title>High-throughput sequencing of the transcriptome and chromatin accessibility in the same cell</article-title>. <source>Nat. Biotechnol.</source> <volume>37</volume>, <fpage>1452</fpage>&#x2013;<lpage>1457</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-019-0290-0</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Keller</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Qu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Gehlenborg</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Polyphony: An interactive transfer learning framework for single-cell data analysis</article-title>. <source>IEEE Trans. Vis. Comput. Graph</source> <volume>29</volume>, <fpage>591</fpage>. <pub-id pub-id-type="doi">10.1109/TVCG.2022.3209408</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheow</surname>
<given-names>L. F.</given-names>
</name>
<name>
<surname>Courtois</surname>
<given-names>E. T.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Viswanathan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>R. Z.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Single-cell multimodal profiling reveals cellular epigenetic heterogeneity</article-title>. <source>Nat. methods</source> <volume>13</volume>, <fpage>833</fpage>&#x2013;<lpage>836</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.3961</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Conesa</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Beck</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Making multi-omics data accessible to researchers</article-title>. <source>Sci. data</source> <volume>6</volume>, <fpage>251</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-019-0258-4</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cuomo</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Seaton</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>McCarthy</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Martinez</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bonder</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Garcia-Bernardo</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Single-cell rna-sequencing of differentiating ips cells reveals dynamic genetic effects on gene expression</article-title>. <source>Nat. Commun.</source> <volume>11</volume>, <fpage>810</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-14457-z</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Demetci</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Santorella</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sandstede</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Noble</surname>
<given-names>W. S.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Scot: Single-cell multi-omics alignment with optimal transport</article-title>. <source>J. Comput. Biol.</source> <volume>29</volume>, <fpage>3</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2021.0446</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>L&#xea; Cao</surname>
<given-names>K.-A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Sincast: A computational framework to predict cell identities in single-cell transcriptomes using bulk atlases as references</article-title>. <source>Briefings Bioinforma.</source> <volume>23</volume>, <fpage>bbac088</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbac088</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Moses</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Efficient k-nearest neighbor graph construction for generic similarity measures</article-title>,&#x201d; in <conf-name>Proceedings of the 20th international conference on World wide web</conf-name>, <conf-loc>Hyderabad, India</conf-loc>, <conf-date>March 28 - April 01, 2011</conf-date>, <fpage>577</fpage>&#x2013;<lpage>586</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mohanty</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Unbiased integration of single cell multi-omics data</article-title>. <comment>BioRxiv.</comment>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duren</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zamanighomi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Satpathy</surname>
<given-names>A. T.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H. Y.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Integrative analysis of single-cell genomics data by coupled nonnegative matrix factorizations</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>115</volume>, <fpage>7723</fpage>&#x2013;<lpage>7728</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1805681115</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eraslan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Simon</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Mircea</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mueller</surname>
<given-names>N. S.</given-names>
</name>
<name>
<surname>Theis</surname>
<given-names>F. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Single-cell rna-seq denoising using a deep count autoencoder</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>390</fpage>&#x2013;<lpage>414</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-018-07931-2</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Eto</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hirota</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Seno</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Matsuda</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Asymmetric integration of single-cell transcriptomic data using latent dirichlet allocation and procrustes analysis</article-title>,&#x201d; in <conf-name>2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) (IEEE)</conf-name>, <conf-loc>Madrid, Spain</conf-loc>, <conf-date>Dec. 3 2018 to Dec. 6 2018</conf-date>, <fpage>2129</fpage>. <comment>&#x2013;2135</comment>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Farahani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Voghoei</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Rasheed</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Arabnia</surname>
<given-names>H. R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A brief review of domain adaptation</article-title>. <source>Adv. data Sci. Inf. Eng.</source> <volume>2021</volume>, <fpage>877</fpage>&#x2013;<lpage>894</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-71704-9_65</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fouch&#xe9;</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chadoutaud</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Delattre</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Zinovyev</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>transmorph: a unifying computational framework for single-cell data integration</article-title>. <comment>bioRxiv</comment>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A survey on deep learning for multimodal data fusion</article-title>. <source>Neural Comput.</source> <volume>32</volume>, <fpage>829</fpage>&#x2013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1162/neco_a_01273</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ghazanfar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guibentif</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Marioni</surname>
<given-names>J. C.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Stabmap: Mosaic single cell data integration using non-overlapping features</article-title>. <comment>bioRxiv</comment>, <fpage>2022</fpage>&#x2013;<lpage>2102</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gong</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Purdom</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Cobolt: Integrative analysis of multimodal single-cell sequencing data</article-title>. <source>Genome Biol.</source> <volume>22</volume>, <fpage>351</fpage>&#x2013;<lpage>421</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-021-02556-z</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gower</surname>
<given-names>J. C.</given-names>
</name>
</person-group> (<year>1975</year>). <article-title>Generalized procrustes analysis</article-title>. <source>Psychometrika</source> <volume>40</volume>, <fpage>33</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1007/bf02291478</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Single-cell methylome landscapes of mouse embryonic stem cells and early embryos analyzed using reduced representation bisulfite sequencing</article-title>. <source>Genome Res.</source> <volume>23</volume>, <fpage>2126</fpage>&#x2013;<lpage>2135</lpage>. <pub-id pub-id-type="doi">10.1101/gr.161679.113</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haghverdi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lun</surname>
<given-names>A. T. L.</given-names>
</name>
<name>
<surname>Morgan</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Marioni</surname>
<given-names>J. C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors</article-title>. <source>Nat. Biotechnol.</source> <volume>36</volume>, <fpage>421</fpage>&#x2013;<lpage>427</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.4091</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Andersen-Nissen</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mauck</surname>
<given-names>W. M.</given-names>
<suffix>III</suffix>
</name>
<name>
<surname>Zheng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Butler</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Integrated analysis of multimodal single-cell data</article-title>. <source>Cell</source> <volume>184</volume>, <fpage>3573</fpage>&#x2013;<lpage>3587.e29</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2021.04.048</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hie</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bryson</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Berger</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Efficient integration of heterogeneous single-cell transcriptomes using scanorama</article-title>. <source>Nat. Biotechnol.</source> <volume>37</volume>, <fpage>685</fpage>&#x2013;<lpage>691</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-019-0113-3</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hotelling</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1992</year>). &#x201c;<article-title>Relations between two sets of variates</article-title>,&#x201d; in <source>Breakthroughs in statistics: Methodology and distribution</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Kotz</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>N. L.</given-names>
</name>
</person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>). <pub-id pub-id-type="doi">10.1007/978-1-4612-4380-9_14</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nie</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>scAI: an unsupervised approach for the integrative analysis of parallel single-cell transcriptomic and epigenomic profiles</article-title>. <source>Genome Biol.</source> <volume>21</volume>, <fpage>25</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-1932-8</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johansen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Quon</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>scAlign: a tool for alignment, integration, and rare cell identification from scRNA-seq data</article-title>. <source>Genome Biol.</source> <volume>20</volume>, <fpage>166</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1766-4</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>W. E.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rabinovic</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Adjusting batch effects in microarray expression data using empirical bayes methods</article-title>. <source>Biostatistics</source> <volume>8</volume>, <fpage>118</fpage>&#x2013;<lpage>127</lpage>. <pub-id pub-id-type="doi">10.1093/biostatistics/kxj037</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<collab>Tabula Sapiens Consortium,</collab> <person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Karkanias</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Krasnow</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Pisco</surname>
<given-names>A. O.</given-names>
</name>
<name>
<surname>Quake</surname>
<given-names>S. R.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>The tabula sapiens: A multiple-organ, single-cell transcriptomic atlas of humans</article-title>. <source>Science</source> <volume>376</volume>, <fpage>eabl4896</fpage>. <pub-id pub-id-type="doi">10.1126/science.abl4896</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kiselev</surname>
<given-names>V. Y.</given-names>
</name>
<name>
<surname>Yiu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hemberg</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>scmap: projection of single-cell RNA-seq data across data sets</article-title>. <source>Nat. Methods</source> <volume>15</volume>, <fpage>359</fpage>&#x2013;<lpage>362</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.4644</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Klein</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Mazutis</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Akartuna</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Tallapragada</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Veres</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>V.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells</article-title>. <source>Cell</source> <volume>161</volume>, <fpage>1187</fpage>&#x2013;<lpage>1201</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.04.044</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Korsunsky</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Millard</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Slowikowski</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Fast, sensitive and accurate integration of single-cell data with Harmony</article-title>. <source>Nat. Methods</source> <volume>16</volume>, <fpage>1289</fpage>&#x2013;<lpage>1296</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-019-0619-0</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kriebel</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>Welch</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Uinmf performs mosaic integration of single-cell multi-omic datasets using nonnegative matrix factorization</article-title>. <source>Nat. Commun.</source> <volume>13</volume>, <fpage>780</fpage>&#x2013;<lpage>817</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-28431-4</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>La Manno</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Soldatov</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zeisel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Braun</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Hochgerner</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Petukhov</surname>
<given-names>V.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Rna velocity of single cells</article-title>. <source>Nature</source> <volume>560</volume>, <fpage>494</fpage>&#x2013;<lpage>498</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0414-6</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lance</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Luecken</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Burkhardt</surname>
<given-names>D. B.</given-names>
</name>
<name>
<surname>Cannoodt</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rautenstrauch</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Laddach</surname>
<given-names>A. C.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Multimodal single cell data integration challenge: Results and lessons learned</article-title>. <comment>bioRxiv</comment>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Benchmarking spatial and single-cell transcriptomics integration methods for transcript distribution prediction and cell type deconvolution</article-title>. <source>Nat. Methods</source> <volume>19</volume>, <fpage>662</fpage>&#x2013;<lpage>670</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-022-01480-9</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lyu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Stambolian</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Deep learning enables accurate clustering with batch effect removal in single-cell rna-seq analysis</article-title>. <source>Nat. Commun.</source> <volume>11</volume>, <fpage>2338</fpage>&#x2013;<lpage>2414</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-15851-3</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>T.-Y.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>W. H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y. R.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Scjoint integrates atlas-scale single-cell rna-seq and atac-seq data with transfer learning</article-title>. <source>Nat. Biotechnol.</source> <volume>40</volume>, <fpage>703</fpage>&#x2013;<lpage>710</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-021-01161-6</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vert</surname>
<given-names>J.-P.</given-names>
</name>
<name>
<surname>Noble</surname>
<given-names>W. S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Jointly embedding multiple single-cell omics measurements</article-title>. <source>Algorithms Bioinform</source> <volume>143</volume>, <fpage>10</fpage>. <pub-id pub-id-type="doi">10.4230/LIPIcs.WABI.2019.10</pub-id>
</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lock</surname>
<given-names>E. F.</given-names>
</name>
<name>
<surname>Hoadley</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Marron</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Nobel</surname>
<given-names>A. B.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Joint and individual variation explained (jive) for integrated analysis of multiple data types</article-title>. <source>Ann. Appl. Stat.</source> <volume>7</volume>, <fpage>523</fpage>&#x2013;<lpage>542</lpage>. <pub-id pub-id-type="doi">10.1214/12-AOAS597</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lopez</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Regier</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cole</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Yosef</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep generative modeling for single-cell transcriptomics</article-title>. <source>Nat. Methods</source> <volume>15</volume>, <fpage>1053</fpage>&#x2013;<lpage>1058</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-018-0229-2</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Loza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Teraguchi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Standley</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Diez</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Unbiased integration of single cell transcriptome replicates</article-title>. <source>NAR Genomics Bioinforma.</source> <volume>4</volume>, <fpage>lqac022</fpage>. <comment>lqac022</comment>. <pub-id pub-id-type="doi">10.1093/nargab/lqac022</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luecken</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>B&#xfc;ttner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chaichoompu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Danese</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Interlandi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mueller</surname>
<given-names>M. F.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Benchmarking atlas-level data integration in single-cell genomics</article-title>. <source>Nat. Methods</source> <volume>19</volume>, <fpage>41</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-021-01336-8</pub-id>
</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lynch</surname>
<given-names>A. W.</given-names>
</name>
<name>
<surname>Theodoris</surname>
<given-names>C. V.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>H. W.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X. S.</given-names>
</name>
<name>
<surname>Meyer</surname>
<given-names>C. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>MIRA: Joint regulatory modeling of multimodal expression and chromatin accessibility in single cells</article-title>. <source>Nat. Methods</source> <volume>19</volume>, <fpage>1097</fpage>&#x2013;<lpage>1108</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-022-01595-z</pub-id>
</citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Macosko</surname>
<given-names>E. Z.</given-names>
</name>
<name>
<surname>Basu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nemesh</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shekhar</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets</article-title>. <source>Cell</source> <volume>161</volume>, <fpage>1202</fpage>&#x2013;<lpage>1214</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.05.002</pub-id>
</citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Minoura</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Abe</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Nam</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Nishikawa</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shimamura</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A mixture-of-experts deep generative model for integrated analysis of single-cell multiomics data</article-title>. <source>Cell Rep. methods</source> <volume>1</volume>, <fpage>100071</fpage>. <pub-id pub-id-type="doi">10.1016/j.crmeth.2021.100071</pub-id>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mirkes</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Bac</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fouch&#xe9;</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Stasenko</surname>
<given-names>S. V.</given-names>
</name>
<name>
<surname>Zinovyev</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gorban</surname>
<given-names>A. N.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Domain adaptation principal component analysis: Base linear method for learning with out-of-distribution data</article-title>. <source>Entropy</source> <volume>25</volume>, <fpage>33</fpage>. <pub-id pub-id-type="doi">10.3390/e25010033</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Tsang</surname>
<given-names>I. W.</given-names>
</name>
<name>
<surname>Kwok</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Domain adaptation via transfer component analysis</article-title>. <source>IEEE Trans. neural Netw.</source> <volume>22</volume>, <fpage>199</fpage>&#x2013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1109/tnn.2010.2091281</pub-id>
</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pantanowitz</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Valenstein</surname>
<given-names>P. N.</given-names>
</name>
<name>
<surname>Evans</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Kaplan</surname>
<given-names>K. J.</given-names>
</name>
<name>
<surname>Pfeifer</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Wilbur</surname>
<given-names>D. C.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Review of the current state of whole slide imaging in pathology</article-title>. <source>J. pathology Inf.</source> <volume>2</volume>, <fpage>36</fpage>. <pub-id pub-id-type="doi">10.4103/2153-3539.83746</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pola&#x144;ski</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Miao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Meyer</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Teichmann</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>J.-E.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>BBKNN: Fast batch alignment of single cell transcriptomes</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>964</fpage>&#x2013;<lpage>965</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz625</pub-id>
</citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Satija</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Farrell</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Gennert</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Schier</surname>
<given-names>A. F.</given-names>
</name>
<name>
<surname>Regev</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Spatial reconstruction of single-cell gene expression data</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>495</fpage>&#x2013;<lpage>502</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3192</pub-id>
</citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schaum</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Karkanias</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Neff</surname>
<given-names>N. F.</given-names>
</name>
<name>
<surname>May</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Quake</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Wyss-Coray</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Single-cell transcriptomics of 20 mouse organs creates a tabula muris: The tabula muris consortium</article-title>. <source>Nature</source> <volume>562</volume>, <fpage>367</fpage>&#x2013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0590-4</pub-id>
</citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schiebinger</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tabaka</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cleary</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Subramanian</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Solomon</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Optimal-transport analysis of single-cell gene expression identifies developmental trajectories in reprogramming</article-title>. <source>Cell</source> <volume>176</volume>, <fpage>928</fpage>&#x2013;<lpage>943.e22</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.01.006</pub-id>
</citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shannon</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Gautier</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Rohart</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Vacher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tebbutt</surname>
<given-names>S. J.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Diablo: An integrative approach for identifying key molecular drivers from multi-omics assays</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>3055</fpage>&#x2013;<lpage>3062</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty1054</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>St&#xe5;hl</surname>
<given-names>P. L.</given-names>
</name>
<name>
<surname>Salm&#xe9;n</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Vickovic</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lundmark</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Navarro</surname>
<given-names>J. F.</given-names>
</name>
<name>
<surname>Magnusson</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Visualization and analysis of gene expression in tissue sections by spatial transcriptomics</article-title>. <source>Science</source> <volume>353</volume>, <fpage>78</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaf2403</pub-id>
</citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stark</surname>
<given-names>S. G.</given-names>
</name>
<name>
<surname>Ficek</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Locatello</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bonilla</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chevrier</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Singer</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Scim: Universal single-cell matching with unpaired feature sets</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>i919</fpage>&#x2013;<lpage>i927</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa843</pub-id>
</citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoeckius</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hafemeister</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Stephenson</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Houck-Loomis</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Chattopadhyay</surname>
<given-names>P. K.</given-names>
</name>
<name>
<surname>Swerdlow</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Simultaneous epitope and transcriptome measurement in single cells</article-title>. <source>Nat. methods</source> <volume>14</volume>, <fpage>865</fpage>&#x2013;<lpage>868</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.4380</pub-id>
</citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stuart</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Butler</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hoffman</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hafemeister</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Papalexi</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mauck</surname>
<given-names>W. M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Comprehensive integration of single-cell data</article-title>. <source>Cell</source> <volume>177</volume>, <fpage>1888</fpage>&#x2013;<lpage>1902.e21</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.05.031</pub-id>
</citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sugihara</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kato</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mori</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kawahara</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Alignment of single-cell trajectory trees with CAPITAL</article-title>. <source>Nat. Commun.</source> <volume>13</volume>, <fpage>5972</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-33681-3</pub-id>
</citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Moran</surname>
<given-names>A. E.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>L.-Y.</given-names>
</name>
<name>
<surname>Qian</surname>
<given-names>D. Z.</given-names>
</name>
<name>
<surname>Schedin</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Identifying phenotype-associated subpopulations by integrating bulk and single-cell sequencing data</article-title>. <source>Nat. Biotechnol.</source> <volume>40</volume>, <fpage>527</fpage>&#x2013;<lpage>538</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-021-01091-3</pub-id>
</citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Svensson</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Gayoso</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yosef</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Pachter</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Interpretable factor models of single-cell rna-seq via variational autoencoders</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>3418</fpage>&#x2013;<lpage>3421</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa169</pub-id>
</citation>
</ref>
<ref id="B82">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tenenhaus</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Philippe</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Guillemot</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Le Cao</surname>
<given-names>K.-A.</given-names>
</name>
<name>
<surname>Grill</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Frouin</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Variable selection for generalized canonical correlation analysis</article-title>. <source>Biostatistics</source> <volume>15</volume>, <fpage>569</fpage>&#x2013;<lpage>583</lpage>. <pub-id pub-id-type="doi">10.1093/biostatistics/kxu001</pub-id>
</citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tenenhaus</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tenenhaus</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Regularized generalized canonical correlation analysis</article-title>. <source>Psychometrika</source> <volume>76</volume>, <fpage>257</fpage>&#x2013;<lpage>284</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-011-9206-8</pub-id>
</citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tibes</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hennessy</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Andreeff</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mills</surname>
<given-names>G. B.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>Reverse phase protein array: Validation of a novel proteomic technology and utility for analysis of primary leukemia specimens and hematopoietic stem cells</article-title>. <source>Mol. cancer Ther.</source> <volume>5</volume>, <fpage>2512</fpage>&#x2013;<lpage>2521</lpage>. <pub-id pub-id-type="doi">10.1158/1535-7163.mct-06-0334</pub-id>
</citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tran</surname>
<given-names>H. T. N.</given-names>
</name>
<name>
<surname>Ang</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Chevrier</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>N. Y. S.</given-names>
</name>
<name>
<surname>Goh</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A benchmark of batch-effect correction methods for single-cell rna sequencing data</article-title>. <source>Genome Biol.</source> <volume>21</volume>, <fpage>12</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1850-9</pub-id>
</citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Treppner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Binder</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hess</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Interpretable generative deep learning: An illustration with single cell gene expression data</article-title>. <source>Hum. Genet.</source> <volume>141</volume>, <fpage>1481</fpage>&#x2013;<lpage>1498</lpage>. <pub-id pub-id-type="doi">10.1007/s00439-021-02417-6</pub-id>
</citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Trong</surname>
<given-names>T. N.</given-names>
</name>
<name>
<surname>Mehtonen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gonz&#xe1;lez</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Kramer</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Hautam&#xe4;ki</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hein&#xe4;niemi</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Semisupervised generative autoencoder for single-cell data</article-title>. <source>J. Comput. Biol.</source> <volume>27</volume>, <fpage>1190</fpage>&#x2013;<lpage>1203</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2019.0337</pub-id>
</citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Van Der Wijst</surname>
<given-names>M. G.</given-names>
</name>
<name>
<surname>Brugge</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>De Vries</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Deelen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Swertz</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Study</surname>
<given-names>L. C.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Single-cell rna sequencing identifies celltype-specific cis-eqtls and co-expression qtls</article-title>. <source>Nat. Genet.</source> <volume>50</volume>, <fpage>493</fpage>&#x2013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-018-0089-9</pub-id>
</citation>
</ref>
<ref id="B89">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Krafft</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Mahadevan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Manifold alignment</article-title>,&#x201d; in <source>Manifold Learning: Theory and Applications</source>, <fpage>95</fpage>&#x2013;<lpage>120</lpage>.</citation>
</ref>
<ref id="B90">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Vasc: Dimension reduction and visualization of single-cell rna-seq data by deep variational autoencoder</article-title>. <source>Genomics, proteomics Bioinforma.</source> <volume>16</volume>, <fpage>320</fpage>&#x2013;<lpage>331</lpage>. <pub-id pub-id-type="doi">10.1016/j.gpb.2018.08.003</pub-id>
</citation>
</ref>
<ref id="B91">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weinstein</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Collisson</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mills</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Mills Shaw</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ozenberger</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ellrott</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>The cancer genome atlas pan-cancer analysis project</article-title>. <source>Nat. Genet.</source> <volume>45</volume>, <fpage>1113</fpage>&#x2013;<lpage>1120</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2764</pub-id>
</citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Welch</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Hartemink</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Prins</surname>
<given-names>J. F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>MATCHER: Manifold alignment reveals correspondence between single cell transcriptome and epigenome dynamics</article-title>. <source>Genome Biol.</source> <volume>18</volume>, <fpage>138</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-017-1269-0</pub-id>
</citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Welch</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Kozareva</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ferreira</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vanderburg</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Macosko</surname>
<given-names>E. Z.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Single-cell multi-omic integration compares and contrasts features of brain cell identity</article-title>. <source>Cell</source> <volume>177</volume>, <fpage>1873</fpage>&#x2013;<lpage>1887.e17</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.05.006</pub-id>
</citation>
</ref>
<ref id="B94">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Westermeier</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Marouga</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Protein detection methods in proteomics research</article-title>. <source>Biosci. Rep.</source> <volume>25</volume>, <fpage>19</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1007/s10540-005-2845-1</pub-id>
</citation>
</ref>
<ref id="B95">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolf</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Angerer</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Theis</surname>
<given-names>F. J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Scanpy: Large-scale single-cell gene expression data analysis</article-title>. <source>Genome Biol.</source> <volume>19</volume>, <fpage>15</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-017-1382-0</pub-id>
</citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Begoli</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>McCord</surname>
<given-names>R. P.</given-names>
</name>
</person-group> (<year>2022a</year>). <article-title>sciCAN: single-cell chromatin accessibility and gene expression data integration via cycle-consistent adversarial network</article-title>. <source>npj Syst. Biol. Appl.</source> <volume>8</volume>, <fpage>33</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1038/s41540-022-00245-6</pub-id>
</citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Das</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>McCord</surname>
<given-names>R. P.</given-names>
</name>
</person-group> (<year>2022b</year>). <article-title>SMILE: Mutual information learning for integration of single-cell omics data</article-title>. <source>Bioinformatics</source> <volume>38</volume>, <fpage>476</fpage>&#x2013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab706</pub-id>
</citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>McCord</surname>
<given-names>R. P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Diagonal integration of multimodal single-cell data: Potential pitfalls and paths forward</article-title>. <source>Nat. Commun.</source> <volume>13</volume>, <fpage>3505</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-31104-x</pub-id>
</citation>
</ref>
<ref id="B99">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>You</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>M. I.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Universal domain adaptation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>, <conf-loc>Long Beach, CA, USA</conf-loc>, <conf-date>June 15 2019 to June 20 2019</conf-date>, <fpage>2720</fpage>.</citation>
</ref>
<ref id="B100">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Meng-Papaxanthos</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Vert</surname>
<given-names>J.-p.</given-names>
</name>
<name>
<surname>Noble</surname>
<given-names>W. S.</given-names>
</name>
</person-group> (<year>2022a</year>). <article-title>Multimodal single-cell translation and alignment with semi-supervised learning</article-title>. <source>J. Comput. Biol.</source> <volume>29</volume>, <fpage>1198</fpage>&#x2013;<lpage>1212</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2022.0264</pub-id>
</citation>
</ref>
<ref id="B101">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2022b</year>). <article-title>scDART: integrating unmatched scRNA-seq and scATAC-seq data and learning cross-modality relationship simultaneously</article-title>. <source>Genome Biol.</source> <volume>23</volume>, <fpage>139</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-022-02706-x</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>