<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioinform.</journal-id>
<journal-title>Frontiers in Bioinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioinform.</abbrev-journal-title>
<issn pub-type="epub">2673-7647</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">846922</article-id>
<article-id pub-id-type="doi">10.3389/fbinf.2022.846922</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioinformatics</subject>
<subj-group>
<subject>Technology and Code</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>MIntO: A Modular and Scalable Pipeline For Microbiome Metagenomic and Metatranscriptomic Data Integration</article-title>
<alt-title alt-title-type="left-running-head">Saenz et al.</alt-title>
<alt-title alt-title-type="right-running-head">MIntO: Microbiome Integrated Meta-Omics</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Saenz</surname>
<given-names>Carmen</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1608316/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Nigro</surname>
<given-names>Eleonora</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1620422/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gunalan</surname>
<given-names>Vithiagaran</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1651290/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Arumugam</surname>
<given-names>Manimozhiyan</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/294404/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Novo Nordisk Foundation Center for Basic Metabolic Research</institution>, <institution>Faculty of Health and Medical Sciences</institution>, <institution>University of Copenhagen</institution>, <addr-line>Copenhagen</addr-line>, <country>Denmark</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/38684/overview">Joao Carlos Setubal</ext-link>, University of S&#xe3;o Paulo, Brazil</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1150544/overview">Bin Hu</ext-link>, Los Alamos National Laboratory (DOE), United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/626971/overview">Martin H&#xf6;lzer</ext-link>, Robert Koch Institute (RKI), Germany</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1630094/overview">Fabio Sanchez</ext-link>, University of S&#xe3;o Paulo, Brazil</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Manimozhiyan Arumugam, <email>arumugam@sund.ku.dk</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Genomic Analysis, a section of the journal Frontiers in Bioinformatics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>2</volume>
<elocation-id>846922</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Saenz, Nigro, Gunalan and Arumugam.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Saenz, Nigro, Gunalan and Arumugam</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Omics technologies have revolutionized microbiome research allowing the characterization of complex microbial communities in different biomes without requiring their cultivation. As a consequence, there has been a great increase in the generation of omics data from metagenomes and metatranscriptomes. However, pre-processing and analysis of these data have been limited by the availability of computational resources, bioinformatics expertise and standardized computational workflows to obtain consistent results that are comparable across different studies. Here, we introduce MIntO (Microbiome Integrated meta-Omics), a highly versatile pipeline that integrates metagenomic and metatranscriptomic data in a scalable way. The distinctive feature of this pipeline is the computation of gene expression profile through integrating metagenomic and metatranscriptomic data taking into account the community turnover and gene expression variations to disentangle the mechanisms that shape the metatranscriptome across time and between conditions. The modular design of MIntO enables users to run the pipeline using three available modes based on the input data and the experimental design, including <italic>de novo</italic> assembly leading to metagenome-assembled genomes. The integrated pipeline will be relevant to provide unique biochemical insights into microbial ecology by linking functions to retrieved genomes and to examine gene expression variation. Functional characterization of community members will be crucial to increase our knowledge of the microbiome&#x2019;s contribution to human health and environment. MIntO v1.0.1 is available at <ext-link ext-link-type="uri" xlink:href="https://github.com/arumugamlab/MIntO">https://github.com/arumugamlab/MIntO</ext-link>.</p>
</abstract>
<kwd-group>
<kwd>omics integration</kwd>
<kwd>metagenomic</kwd>
<kwd>metatranscriptomic</kwd>
<kwd>pipeline</kwd>
<kwd>gene expression</kwd>
<kwd>community turnover</kwd>
<kwd>microbial ecology</kwd>
<kwd>microbiome</kwd>
</kwd-group>
<contract-sponsor id="cn001">Novo Nordisk Foundation Center for Basic Metabolic Research<named-content content-type="fundref-id">10.13039/501100011747</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Horizon 2020<named-content content-type="fundref-id">10.13039/501100007601</named-content>
</contract-sponsor>
<contract-sponsor id="cn003">Danmarks Frie Forskningsfond<named-content content-type="fundref-id">10.13039/501100011958</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The human microbiome is a complex congregation of microbes comprising trillions of microbial cells present in our bodies (<xref ref-type="bibr" rid="B6">Bashan et al., 2016</xref>). Microbe-microbe and microbe-host interactions confer a variety of physiological benefits to the hosts and impact their susceptibility to disease. For instance, the microbial niche can provide metabolic functions different from the host genome, most of which are encoded by genes that have not yet been discovered (<xref ref-type="bibr" rid="B42">Nicholson et al., 2012</xref>; <xref ref-type="bibr" rid="B13">Donia and Fischbach, 2015</xref>).</p>
<p>Studying these microbial communities is a challenging task, which has recently been made easier by high-throughput sequencing approaches which generate omics data such as metagenomes and metatranscriptomes. These omics methods have revolutionized microbiome research by allowing the characterization of complex microbial communities in different biomes without requiring their cultivation. Metagenomic data enables the genomic and taxonomic characterization of microbial community composition and, depending on the sequencing strategy employed, can allow the recovery of Metagenome-Assembled Genomes (MAGs) (<xref ref-type="bibr" rid="B2">Almeida et al., 2019</xref>; <xref ref-type="bibr" rid="B57">Stewart et al., 2019</xref>; <xref ref-type="bibr" rid="B51">Saheb Kashaf et al., 2022</xref>). However, it can only unravel the functional potential in a sample (<xref ref-type="bibr" rid="B49">Quince et al., 2017</xref>). In contrast, metatranscriptomic data identifies the pool of genes that are transcribed under a specific condition, which gives a more accurate picture of the processes and molecular activity occurring in the microbial community (<xref ref-type="bibr" rid="B53">Satinsky et al., 2014</xref>; <xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>). Hence, by analyzing both metagenomes and metatranscriptomes, we can have deeper insights into the functional potential as well as the actual activity of microbial communities (<xref ref-type="bibr" rid="B70">Wang et al., 2020</xref>; <xref ref-type="bibr" rid="B63">Tl&#xe1;skal et al., 2021</xref>).</p>
<p>In recent years, the application of high-throughput sequencing approaches in microbiome research has greatly increased together with the generation of large amounts of data (<xref ref-type="bibr" rid="B48">Qin et al., 2010</xref>; <xref ref-type="bibr" rid="B20">Human Microbiome Project Consortium, 2012</xref>; <xref ref-type="bibr" rid="B47">Pasolli et al., 2019</xref>). As a consequence, the pre-processing and analysis of such data have been limited by the availability of computational resources and bioinformatics expertise. In addition, there is a lack of standardized protocols to handle and analyze multi-omics data sets in a more consistent manner, making the comparisons between different studies and findings more challenging. Standardizing the way omics data are handled ensures a degree of consistency of the results across different studies. Furthermore, making the workflows semi-automatic will allow the analysis of complex microbial communities by users with limited bioinformatic skills.</p>
<p>Standard metagenomic and metatranscriptomic approaches entail 1) read curation, 2) <italic>de novo</italic> assembly and/or co-assembly, 3) binning, 4) gene prediction, 5) annotation of predicted genes at taxonomic and functional level and 6) quantification of gene abundances and transcripts. However, most of the computational pipelines developed so far can only analyze metagenomic or metatranscriptomic data individually and only few, reported in <xref ref-type="table" rid="T1">Table 1</xref>, can handle both meta-omics data (<xref ref-type="bibr" rid="B26">Kim et al., 2016</xref>; <xref ref-type="bibr" rid="B41">Narayanasamy et al., 2016</xref>; <xref ref-type="bibr" rid="B59">Tamames and Puente-S&#xe1;nchez, 2018</xref>; <xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>; <xref ref-type="bibr" rid="B55">Sequeira et al., 2019</xref>; <xref ref-type="bibr" rid="B64">Van Damme et al., 2021</xref>). Furthermore, only one of them (<xref ref-type="bibr" rid="B64">Van Damme et al., 2021</xref>) can combine two sequencing technologies (Nanopore or long-sequences and Illumina or short-sequences) to recover MAGs.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Features of pipelines that handle metagenomic and metatranscriptomic data in comparison to MIntO: Steps, capacities and approaches.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">FMAP <xref ref-type="bibr" rid="B26">Kim et al. (2016)</xref>; <xref ref-type="bibr" rid="B52">Salazar et al. (2019)</xref>
</th>
<th align="center">IMP <xref ref-type="bibr" rid="B41">Narayanasamy et al. (2016)</xref>
</th>
<th align="center">MOSCA <xref ref-type="bibr" rid="B55">Sequeira et al. (2019)</xref>
</th>
<th align="center">SqueezeMeta <xref ref-type="bibr" rid="B59">Tamames and Puente-S&#xe1;nchez, (2018)</xref>
</th>
<th align="center">MUFFIN <xref ref-type="bibr" rid="B64">Van Damme et al. (2021)</xref>
</th>
<th align="center">MIntO (2021)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">data source</td>
<td align="left">short reads</td>
<td align="left">paired-end short reads</td>
<td align="left">paired-end short reads</td>
<td align="left">paired-end short reads</td>
<td align="left">paired-end Illumina reads (short reads) and Nanopore-based reads (long reads)</td>
<td align="left">paired-end Illumina reads (short reads) and Nanopore-based reads (long reads)</td>
</tr>
<tr>
<td align="left">quality and read length control</td>
<td align="left">only quality control</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">host genome removal</td>
<td align="left">only human genome removal</td>
<td align="left">Yes</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">rRNA removal</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">taxonomy assignment</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">
<italic>de novo</italic> assembly/co-assembly</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">combining short and long reads</td>
<td align="left">optionally, include long reads</td>
</tr>
<tr>
<td align="left">binning</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">gene prediction</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">function annotation</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">alignment to reference database/genomes</td>
<td align="left">alignment to reference database</td>
<td align="left">Yes</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">alignment to retrieved MAGs</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">normalization</td>
<td align="left">RPKM</td>
<td align="left">RPKM</td>
<td align="left">TMM, RLE</td>
<td align="left">RPKM</td>
<td align="left">TPM</td>
<td align="left">TPM, Marker genes</td>
</tr>
<tr>
<td align="left">visualization</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">No</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">local installation</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">gene expression computation</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">Yes</td>
</tr>
<tr>
<td align="left">differential analysis/Downstream analysis</td>
<td align="left">differentially-abundant genes analysis</td>
<td align="left">No</td>
<td align="left">differential gene expression analysis</td>
<td align="left">No</td>
<td align="left">No</td>
<td align="left">No</td>
</tr>
<tr>
<td align="left">Software dependencies installed by the user before using the pipeline</td>
<td align="left">Perl, R, Statistics::R, DIAMOND or USEARCH, Bio::DB::Taxonomy, XML::LibXML</td>
<td align="left">Python3, pip, impy, Conda, Docker/Singularity</td>
<td align="left">MOSGUITO, and Conda, or Docker/Singularity</td>
<td align="left">Conda</td>
<td align="left">Nextflow and Conda or Docker/Singularity</td>
<td align="left">FetchMGs, Conda</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Overall, the pipelines shown in <xref ref-type="table" rid="T1">Table 1</xref> integrate metagenomic and metatranscriptomic data by comparing the abundances of genes and their respective transcripts. To the best of our knowledge, none of these (<xref ref-type="table" rid="T1">Table 1</xref>) considers the community composition and gene expression alterations as the underlying processes that shape the community transcript levels (<xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>) when integrating metagenomic and metatranscriptomic data. However, perturbations of the transcript levels can be a consequence of two factors: the variation in the expression of genes encoded by the organisms in the community, and/or by changes in the abundance of these members and their related genes in a process known as community turnover (<xref ref-type="bibr" rid="B53">Satinsky et al., 2014</xref>; <xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>). Hence, the integration of abundances of genes and the respective transcripts represents the gene expression profiles, which are the relative amount of transcripts per gene in a specific time (<xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>). Additionally, being able to recover genomes from metagenomic raw reads is crucial for an optimal computation of gene expression levels and provides a more accurate ecological description of the community&#x2019;s functioning (<xref ref-type="bibr" rid="B59">Tamames and Puente-S&#xe1;nchez, 2018</xref>).</p>
<p>Here, we introduce MIntO (Microbiome Integrated meta-Omics), a pipeline that includes state of the art tools to integrate microbiome metagenomic and metatranscriptomic data in a scalable way for read pre-processing, species composition profiling, MAG generation, gene and function expression profiling, as well as the visualization of the results and comparison of multiple samples. Optionally, MIntO can combine long-read sequences for more contiguous assemblies and short-read sequences for higher accuracy, which helps recover more accurate as well as complete MAGs (<xref ref-type="bibr" rid="B8">Bertrand et al., 2019</xref>; <xref ref-type="bibr" rid="B45">Overholt et al., 2020</xref>; <xref ref-type="bibr" rid="B10">Brown et al., 2021</xref>). Depending on the data availability and research question, the pipeline can be run in three modes: (A) <italic>genome-based assembly-dependent</italic>, (B) <italic>genome-based assembly-free</italic> and (C) <italic>gene-catalog-based assembly-free</italic> (<xref ref-type="fig" rid="F1">Figure 1A</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Schematic overview of metagenomic and metatranscriptomic integration to quantify gene expression levels. <bold>(A)</bold> Three modes are available based on the input data and the experiment design: the <italic>genome-based assembly-dependent</italic> mode (1, in dark purple) recovers MAGs from metagenomic samples, while the <italic>genome-based assembly-free</italic> (2, in dark green) and the <italic>gene-catalog-based assembly-free</italic> (3, in red) modes use publicly available genomes or a gene catalog, respectively, provided by the user. In the three modes, the pipeline workflow includes quality control and preprocessing; assembly-free taxonomy profiling of high-quality metagenomic reads (in orange) by identifying phylogenetic markers (coloured); alignment of the high-quality reads to the selected reference and normalization; integration: gene and functional profiling; and visualization and reporting. The gene prediction and functional annotation step is run using the recovered MAGs (mode 1) or publicly available genomes (mode 2). <bold>(B)</bold> The variation of gene expression depends on the abundance of transcripts from the organisms in the community and/or by changes in the abundance of these members and their related genes (community turnover).</p>
</caption>
<graphic xlink:href="fbinf-02-846922-g001.tif"/>
</fig>
<p>MIntO enables the study of microbial ecology by linking functions to genomes and environmental context, helping to understand the dynamics of the molecular activities captured by the whole community-level changes in composition and gene expression (<xref ref-type="fig" rid="F1">Figure 1B</xref>).</p>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<p>MIntO v1.0.1 has been developed using R software (v4.0.3) (<xref ref-type="bibr" rid="B61">The R Project for Statistical Computing, 2021</xref>), Python 3 (<xref ref-type="bibr" rid="B65">Van Rossum and Drake, 2009</xref>) and Perl (<xref ref-type="bibr" rid="B69">Wall, Christiansen and Orwant, 2000</xref>) programming languages, and has been tested on a 64-bit Linux server with 2 &#xd7; AMD EPYC 7742 64-Core Processors and 2 terabytes of memory.</p>
<sec id="s2-1">
<title>Conda Environment and Singularity Containers</title>
<p>MIntO has been designed to use publicly available software that are available as conda environments (<xref ref-type="bibr" rid="B3">Anaconda Inc, 2020</xref>) or singularity containers (<xref ref-type="bibr" rid="B31">Kurtzer, Sochat and Bauer, 2017</xref>) to minimize the installation of individual software packages by the user. All software dependencies are tied to specific versions in conda or singularity containers to ensure reproducibility and record-keeping of versions of the different libraries. It is encapsulated within a user-friendly framework using Snakemake (<xref ref-type="bibr" rid="B39">M&#xf6;lder, 2021</xref>) to facilitate the scalability of the pipeline by optimizing the number of parallel processes from a single-core workstation to compute clusters. This pipeline enables consistency of the results and straightforward application by users with basic informatics skills to analyze complex omics data.</p>
</sec>
<sec id="s2-2">
<title>Pipeline Inputs</title>
<p>MIntO requires a configuration file as an input indicating the metagenomic (metaG) and/or metatranscriptomic (metaT) sample names and the corresponding raw FASTQ files location together with the path of the pipeline dependencies, currently only FetchMGs (<xref ref-type="bibr" rid="B30">Kultima et al., 2012</xref>). MIntO generates the necessary directories and outputs the required files for further analysis, including the configuration files needed in each step of the pipeline, but they should be filled out by the user. Optionally, the required databases can be downloaded and installed by MIntO.</p>
<p>In addition, if MIntO is run under <italic>genome-based assembly-free</italic> mode, the user should provide input genomes as FASTA files, genome features as GFF files, and amino acid sequences of protein-coding genes as FASTA files, while in the case of <italic>gene-catalog based assembly-free</italic> mode the user should provide a multi FASTA file with the nucleotide sequences of the genes, such as the one published with the Integrated Gene Catalog (IGC) (<xref ref-type="bibr" rid="B33">Li et al., 2014</xref>) (<xref ref-type="fig" rid="F1">Figure 1A</xref>, user-provided input).</p>
</sec>
<sec id="s2-3">
<title>Pre-Processing of Metagenomic and Metatranscriptomic Short Reads</title>
<p>MIntO pre-processes metagenomic and metatranscriptomic short reads independently of each other. The pre-processing step can be subdivided into three different steps: quality and read length, host genome and ribosomal RNA (rRNA) filtering.</p>
<p>1. Quality and read length filtering.</p>
<p>We use Trimmomatic v0.39 (<xref ref-type="bibr" rid="B9">Bolger, Lohse and Usadel, 2014</xref>) to first remove sequencing adapters and low quality bases from raw reads and a second time to remove reads that are too short.<list list-type="simple">
<list-item>
<p>a. In the first step, the option <italic>TRAILING:5 LEADING:5 SLIDINGWINDOW:4:20 ILLUMINACLIP:{adapters.fa}:2:30:10</italic> is used if a sequence adapters file is provided by the user (<italic>trimmomatic_adaptors &#x3d; &#x3c;PathTo&#x3e;/adapters.fa</italic>). Otherwise, a custom script retrieves the adapters by selecting the most abundant index in the first 10,000 headers of the raw FASTQ files (<italic>trimmomatic_adaptors &#x3d; False</italic>). The user can decide to skip this step if adapter sequences have already been removed (<italic>trimmomatic_adaptors &#x3d; Skip</italic>).</p>
</list-item>
<list-item>
<p>b. For the second filtering, the <italic>MINLEN</italic> parameter in Trimmomatic is used to remove reads that are too short. This cutoff is estimated as the maximum length above which a predefined percentage of the reads from the previous step are retained (default parameter is 95% of the reads, <italic>perc_remaining_reads: 95</italic>). If the estimated read length cutoff is below 50bp, trimmomatic will use 50bp as the minimum sequence length (<xref ref-type="sec" rid="s10">Supplementary Figure S3</xref>).</p>
</list-item>
</list>
</p>
<p>2. Host genome filtering.</p>
<p>In the second step to remove putative host-derived sequences, the filtered read-pairs are aligned to a reference genome given by the user. The BWA aligner (<xref ref-type="bibr" rid="B66">Vasimuddin et al., 2019</xref>) version 2.2.1 is used to generate the index (<italic>bwa-mem2 index</italic>) and to map the read-pairs to the host genome (<italic>bwa-mem2 mem -a</italic>). Read-pairs aligned to this reference genome are identified by msamtools v1.1.0 (<xref ref-type="bibr" rid="B78">Arumugam, 2022</xref>) (<italic>filter -S -l 30</italic>) and excluded from the FASTQ files by mseqtools (<ext-link ext-link-type="uri" xlink:href="https://github.com/arumugamlab/mseqtools">https://github.com/arumugamlab/mseqtools</ext-link>) version 0.9.1, even if only one end is mapped (<italic>subset --exclude --paired --list {listfile}</italic>).</p>
<p>3. Ribosomal RNA filtering.</p>
<p>Prior to sequencing, it is recommended to deplete the rRNA in the metatranscriptomic samples. Nevertheless, it is common that metatranscriptomic sequence data still contains rRNA after such a depletion step. MIntO uses SortMeRNA v4.3.4 (<xref ref-type="bibr" rid="B28">Kopylova, No&#xe9; and Touzet, 2012</xref>) to map the metatranscriptomic reads to an rRNA sequence database consisting prokaryotic (16S and 23S) and eukaryotic (18S and 28S) rRNA sequences (<italic>--paired_in --fastx --blast 1 --sam --other --ref</italic>). Reads classified as rRNA by SortMeRNA are excluded from the FASTQ files using mseqtools (<italic>subset --exclude --paired --list {listfile}</italic>).</p>
<p>The remaining high-quality filtered (host-free for metagenomic and host- and rRNA-free for metatranscriptomic) reads are then passed to the sequence analysis and post processing steps.</p>
</sec>
<sec id="s2-4">
<title>Assembly-Free Taxonomic Profiling From High-Quality Filtered Reads</title>
<p>High-quality filtered reads can be profiled by the default program, MetaPhlAn3 v3.0.13 (<xref ref-type="bibr" rid="B7">Beghini et al., 2021</xref>) (<italic>--input_type fastq --bowtie2out -t rel_ab_w_read_stats</italic>). Alternatively, users can choose to run mOTUs2 v2.1.1 (<xref ref-type="bibr" rid="B37">Milanese et al., 2019</xref>) in two different modes to generate a taxonomic profile as relative abundance (taxa_profile: <italic>motus_rel</italic>, <italic>profile -u -q</italic>) or as counts (taxa_profile: <italic>motus_raw</italic>, <italic>profile -c -u -q</italic>). If the latter one is chosen, MIntO estimates the relative abundance of the taxonomic profile. To explore the similarities and dissimilarities of the data, the relative abundance of the species composition is used to generate two visual outputs: 1) the 15 most abundant genera across the samples, and 2) a principal coordinate analysis (PCoA) using Bray-Curtis distance. These visualizations provide users with a general idea of the microbial composition in the different samples. For a more detailed downstream analysis, MIntO outputs the combined table of the taxonomy profiles of all samples in CSV format and as a phyloseq object (<xref ref-type="bibr" rid="B36">McMurdie and Holmes, 2013</xref>), the latter including the abundance of the species, taxonomic classification and metadata tables.</p>
</sec>
<sec id="s2-5">
<title>Retrieving MAGs From Metagenomic High-Quality Host-Free Reads</title>
<p>MIntO&#x2019;s approach to reconstruct MAGs from high-quality host-free reads exploits metagenomic assembly of single samples as well as co-assembly of pre-defined sample groups followed by binning preparation and contig binning.</p>
<p>1. Assembly:<list list-type="simple">
<list-item>
<p>a. Long-read assembly: If available, Nanopore reads are assembled individually using metaFlye assembler (<xref ref-type="bibr" rid="B27">Kolmogorov et al., 2020</xref>) v2.9 (<italic>--nano-raw &#x3c;FASTQ&#x3e; --meta --min-overlap 3000 --iterations 3</italic>)</p>
</list-item>
<list-item>
<p>b. Short-read assembly: MetaSPAdes assembler v3.15.3 (<xref ref-type="bibr" rid="B44">Nurk et al., 2017</xref>) is used to correct paired-end short reads from individual samples (<italic>--only-error-correction,</italic> the default <italic>--phred-offset</italic> is auto) followed by their single-assembly (<italic>--meta --only-assembler</italic>, the default kmer option is <italic>k &#x3d; 21,33,55,77,99,127</italic>).</p>
</list-item>
<list-item>
<p>c. Hybrid assembly: Optionally, we can combine metagenomic Nanopore-based long reads and Illumina paired-end short reads to perform hybrid assembly by MetaSPAdes using the parameters as step (b) with an additional <italic>--nanopore</italic> option.</p>
</list-item>
<list-item>
<p>d. Co-assembly: MEGAHIT (<xref ref-type="bibr" rid="B32">Li et al., 2015</xref>) v1.2.9 is run with two different parameters (--<italic>meta-sensitive</italic> and <italic>--meta-large</italic>) per co-assembly, where by default all samples used in the single-assembly are assembled together. Users can also define their own subsets of samples that should be co-assembled in the configuration file.</p>
</list-item>
</list>
</p>
<p>2. Binning preparation:</p>
<p>Contigs longer than 2,500&#xa0;bp from all the combinations of assemblies above are combined together in preparation for binning. Metagenomic reads from individual short-read metagenomes are first mapped to this set of contigs using BWA aligner (<xref ref-type="bibr" rid="B66">Vasimuddin et al., 2019</xref>) v2.2.1 (<italic>bwa-mem2 mem -a</italic>) in paired-end mode. Sequencing depth of the contigs in each sample is estimated by <italic>jgi_summarize_bam_contig_depths</italic> program included in MetaBAT2 (<xref ref-type="bibr" rid="B24">Kang et al., 2019</xref>).</p>
<p>3. Contig binning:</p>
<p>Contig binning is then performed by executing VAMB (<xref ref-type="bibr" rid="B43">Nissen et al., 2021</xref>), a binner using an unsupervised deep learning approach in the form of variational autoencoders that can be run with or without GPUs. GPU use is highly recommended if available in order to speed up the binning process, especially if working with a large number of samples. By default, MIntO runs VAMB four times, each time with a different set of parameters <italic>-l 16 -n 256,256</italic>; <italic>-l 24 -n 384,384</italic>; <italic>-l 32 -n 512,512</italic>; and -<italic>l 40 -n 768,768</italic>. However the user(s) can choose to perform just one run or a set of runs of their choice.</p>
<p>4. Non-redundant MAGs:</p>
<p>Bins generated by VAMB are split into MAGs derived from individual metagenomic samples. Only the MAGs that pass quality control using CheckM (<xref ref-type="bibr" rid="B46">Parks et al., 2015</xref>) (completeness &#x3e; 90% and contamination &#x3c; 5%) are kept. The MAGs are then subjected to cluster analysis performed with CoverM v0.6.0 (<ext-link ext-link-type="uri" xlink:href="https://github.com/wwood/CoverM">https://github.com/wwood/CoverM&#x23;usage</ext-link>, module cluster) in order to dereplicate them at 99% average nucleotide identity (ANI) (<xref ref-type="bibr" rid="B22">Jain et al., 2018</xref>). For each genome, a score is retrieved with the formula below.<disp-formula id="equ1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>y</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="italic">log</mml:mi>
<mml:mi mathvariant="italic">10</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mo>/</mml:mo>
<mml:mo>&#x23;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">log</mml:mi>
<mml:mn>10</mml:mn>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mn>50</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>L</mml:mi>
<mml:mn>50</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.1</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>y</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>Then for each cluster the genome with the highest score is chosen, generating a unique set of non-redundant MAGs which will be used in the next step.</p>
</sec>
<sec id="s2-6">
<title>Taxonomic Assignment of MAGs</title>
<p>Once the unique set of MAGs is retrieved, taxonomy is assigned using the module <italic>phylophlan_metagenomic</italic> in PhyloPhlAn3 (<xref ref-type="bibr" rid="B5">Asnicar et al., 2020</xref>). MIntO uses SGB.Jul20 or SGB.Dec20 databases depending on user&#x2019;s choice (<italic>--database</italic>) which will be automatically downloaded in the program folder if no other location is specified. Additionally, if the users have previously downloaded one of the PhyloPhlAn3 databases of their interest, they can use that by giving their path.</p>
</sec>
<sec id="s2-7">
<title>Genome Annotation on the Retrieved MAGs</title>
<p>First, Prokka (<xref ref-type="bibr" rid="B54">Seemann, 2014</xref>) (version 1.14) (with options <italic>--addgenes --centre X --compliant</italic>) is used to identify and annotate the genes from the recovered MAGs, retrieving the corresponding nucleotide and amino acid sequences.</p>
<p>Next, predicted genes are annotated with several databases:<list list-type="bullet">
<list-item>
<p>eggNOG database (<xref ref-type="bibr" rid="B19">Huerta-Cepas et al., 2019</xref>) (COG ids) with eggNOG-mapper v2.1.6 (<xref ref-type="bibr" rid="B18">Huerta-Cepas et al., 2017</xref>; <xref ref-type="bibr" rid="B11">Cantalapiedra et al., 2021</xref>) (<italic>--no_annot --no_file_comments --report_no_hits --override -m diamond</italic> and <italic>--annotate_hits_table -m no_search --no_file_comments --override</italic>, emapperdb v5.0.2).</p>
</list-item>
<list-item>
<p>KEGG functions (<xref ref-type="bibr" rid="B23">Kanehisa and Goto, 2000</xref>) (<italic>-k -p prokaryote.hal --create-alignment -f mapper</italic>, Kofam_scan (<xref ref-type="bibr" rid="B4">Aramaki et al., 2020</xref>) version 1.3.0 and ko_list from November 2021).</p>
</list-item>
<list-item>
<p>Carbohydrate-active enzyme database [CAZyme, (<xref ref-type="bibr" rid="B17">Huang et al., 2018</xref>; <xref ref-type="bibr" rid="B76">Zhang et al., 2018</xref>)] with dbCAN annotation tool v2.0.11 (<xref ref-type="bibr" rid="B76">Zhang et al., 2018</xref>) (<italic>run_dbcan.py protein</italic>).</p>
</list-item>
<list-item>
<p>Pfam database (<xref ref-type="bibr" rid="B38">Mistry et al., 2021</xref>) with eggNOG-mapper (<xref ref-type="bibr" rid="B18">Huerta-Cepas et al., 2017</xref>; <xref ref-type="bibr" rid="B11">Cantalapiedra et al., 2021</xref>).</p>
</list-item>
</list>
</p>
<p>These databases are installed locally by the user. The pipeline integrates the different gene annotations: Gene ID, eggNOG, KEGG_ko, KEGG_Pathway, KEGG_Module, dbCAN.mod, dbCAN.enzclass and Pfam.</p>
</sec>
<sec id="s2-8">
<title>Functional Profiling</title>
<p>The high-quality filtered (host-free for metagenomic and host- and rRNA-free for metatranscriptomic) reads are used to generate the functional profiles following four steps: metagenomic and metatranscriptomic read alignments, mappability ratio, read count normalization, and gene and function expression computation.</p>
<sec id="s2-9">
<title>Metagenomic and Metatranscriptomic Reads Alignment</title>
<p>To estimate gene and transcript abundances, the high-quality filtered reads can be aligned to 1) genomes such as the recovered MAGs or publicly available genomes (<italic>genome-based</italic>) or 2) a gene catalog (<italic>gene-based</italic>), depending on the mode that the pipeline is run.<list list-type="simple">
<list-item>
<p>1. <italic>Genome-based</italic> alignment: The retrieved MAGs or the reference genomes are concatenated and indexed using the BWA aligner (<xref ref-type="bibr" rid="B66">Vasimuddin et al., 2019</xref>) v2.2.1 (<italic>bwa-mem2 index</italic>). Mapping reads to the reference (<italic>bwa-mem2 mem -a</italic>) is followed by highest-scoring alignment(s) filtering for each read with msamtools v1.1.0 (<xref ref-type="bibr" rid="B78">Arumugam, 2022</xref>) (<italic>filter -S -b -l 50 -p 95 -z 80 --besthit</italic>)<italic>.</italic> The filtered BAM files are indexed by samtools v1.14 (<xref ref-type="bibr" rid="B12">Danecek et al., 2021</xref>) (<italic>sort --output-fmt &#x3d; BAM</italic>; <italic>index</italic>) and the GFF file with the genome features is used to quantify the raw number of aligned reads to each gene by bedtools <italic>multicov</italic> v2.29.2 (<xref ref-type="bibr" rid="B50">Quinlan and Hall, 2010</xref>).</p>
</list-item>
<list-item>
<p>2. <italic>Gene-based</italic> alignment: As an alternative, the gene catalog given by the user is indexed using <italic>bwa-mem2 index</italic> [BWA aligner v2.2.1 (<xref ref-type="bibr" rid="B66">Vasimuddin et al., 2019</xref>)]. The aligned reads (<italic>bwa-mem2 mem -a</italic>) are filtered for highest-scoring alignment(s) per read with msamtools v1.1.0 (<xref ref-type="bibr" rid="B78">Arumugam, 2022</xref>) (<italic>filter -S -b -l 50 -p 95 -z 80 --besthit</italic>)<italic>.</italic>
</p>
</list-item>
</list>
</p>
<p>Optionally, the user can filter the aligned reads by establishing the minimum number of mapped reads to a gene, using the <italic>MIN_mapped_reads</italic> parameter. While the default value for this parameter is 0, for metagenomes with sequencing depth higher than 10 million paired-end reads, we recommend setting this threshold at 10 mapped reads to a gene (MIN_mapped_reads: 10), which is what we used for IBDMDB dataset.</p>
</sec>
<sec id="s2-10">
<title>Mappability Ratio</title>
<p>In addition, to estimate how representative the gene or genome databases are of the metagenomic and metatranscriptomic samples, the filtered BAM files are used to calculate the mappability ratio by msamtools v1.1.0 (<xref ref-type="bibr" rid="B78">Arumugam, 2022</xref>) (<italic>profile --total {total_reads} --multi prop --unit all --nolen</italic>). Here, we used the IGC (<xref ref-type="bibr" rid="B33">Li et al., 2014</xref>) and recovered MAGs as references.</p>
</sec>
<sec id="s2-11">
<title>Read Count Normalization</title>
<p>Normalization of read counts makes possible the comparison within or between different samples. Based on the users&#x2019; selection, TPM (Transcripts Per Kilobase Million) or MGs (Marker Genes) normalized gene and transcript abundance profiles are generated from the metagenomic and metatranscriptomic read alignments, respectively.<list list-type="simple">
<list-item>
<p>1. <italic>TPM normalization.</italic> Sequencing depth and gene length are used to obtain the relative abundance of genes or transcripts (<xref ref-type="bibr" rid="B68">Wagner, Kin and Lynch, 2012</xref>). The TPM value of the gene <italic>i</italic>, TPM(<italic>i</italic>), is calculated by employing the equation:</p>
</list-item>
</list>
<disp-formula id="equ2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mn>6</mml:mn>
</mml:msup>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>&#xa0;&#xa0;&#x3a3;</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mn>6</mml:mn>
</mml:msup>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>&#x2003;where <italic>n</italic>
<sub>
<italic>i</italic>
</sub> is the number of reads mapped to the gene <italic>i</italic>, <italic>l</italic>
<sub>
<italic>i</italic>
</sub> is the length of that gene and <italic>j</italic> iterates over all genes identified in the sample.<list list-type="simple">
<list-item>
<p>2. <italic>MGs normalization.</italic> In a similar approach to Salazar et al. study, but more customized to MAG-based analysis, the gene or transcript abundances of a MAG are divided by the median abundance of 10 universal single-copy phylogenetic MGs from the corresponding MAG (<xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>). These MGs are identified in each MAG by FetchMGs v1.2 (available at <ext-link ext-link-type="uri" xlink:href="http://motu-tool.org/fetchMG.html">http://motu-tool.org/fetchMG.html</ext-link>) as OGs: COG0012, COG0016, COG0018, COG0172, COG0215, COG0495, COG0525, COG0533, COG0541, and COG0552. In addition, these MGs are constitutively expressed housekeeping genes across many different conditions (<xref ref-type="bibr" rid="B58">Sunagawa et al., 2013</xref>; <xref ref-type="bibr" rid="B37">Milanese et al., 2019</xref>; <xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>). Thus, the MGs-normalized metagenomic and metatranscriptomic profiles can be interpreted as the gene and transcript abundances in a MAG relative to housekeeping MGs abundance and transcript, respectively. The MGs value of the gene i, MGs(i), is calculated by employing the equation:</p>
</list-item>
</list>
<disp-formula id="equ3">
<mml:math id="m3">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>10</mml:mn>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>G</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>m</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mn>6</mml:mn>
</mml:msup>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>G</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mn>6</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>n</italic>
<sub>
<italic>i</italic>
</sub> is the number of reads mapped to the gene <italic>i</italic> in the gene&#x2019;s MAG, <italic>l</italic>
<sub>
<italic>i</italic>
</sub> is the length of that gene and <italic>M</italic>(<italic>MGs</italic>) is the median abundance of the 10&#xa0;MGs from the gene&#x2019;s genome.</p>
<p>When the reads are mapped to a gene database, msamtools v1.1.0 (<xref ref-type="bibr" rid="B78">Arumugam, 2022</xref>) is used to normalize the number of aligned reads per gene to TPM (<italic>profile --total {total_reads} --multi prop --unit tpm</italic>). However, if the reads are mapped to a set of MAGs or publicly available genome(s), the user can choose to obtain TPM or MGs normalized abundances.</p>
</sec>
<sec id="s2-12">
<title>Computing Gene and Function Expression Profiles</title>
<p>The levels of gene expression are computed by the integration of gene and transcript abundance profiles, which is, the relative amount of RNA molecules per DNA copy of that gene (TPM normalization):<disp-formula id="equ4">
<mml:math id="m4">
<mml:mrow>
<mml:mtext>gene&#xa0;expression</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>transcript&#xa0;abundance/gene&#xa0;copy&#xa0;number</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>Or gene expression in that MAG relative to housekeeping MGs expression (MGs normalization):<disp-formula id="equ5">
<mml:math id="m5">
<mml:mrow>
<mml:mtext>MGs</mml:mtext>
<mml:mo>&#x2010;</mml:mo>
<mml:mtext>normalized&#xa0;gene&#xa0;expression</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>gene&#xa0;expression&#xa0;/median&#xa0;MGs&#xa0;gene&#xa0;expression</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>Finally, functional profiles are obtained by grouping the genes into functions.</p>
</sec>
</sec>
<sec id="s2-13">
<title>Visualization</title>
<p>All the visualization outputs are generated in R software (v4.0.3) (<xref ref-type="bibr" rid="B61">The R Project for Statistical Computing, 2021</xref>), using the following packages: BiocManager (v1.30.16) (<xref ref-type="bibr" rid="B40">Morgan, 2021</xref>), data.table (v1.14.2) (<xref ref-type="bibr" rid="B14">Dowle and Srinivasan, 2021</xref>), reshape2 (v1.4.4) (<xref ref-type="bibr" rid="B73">Wickham, 2007</xref>), phyloseq (v1.34.0) (<xref ref-type="bibr" rid="B36">McMurdie and Holmes, 2013</xref>), tidyverse (v1.3.1) (<xref ref-type="bibr" rid="B71">Wickham et al., 2019</xref>), ggplot2 (v3.3.5) (<xref ref-type="bibr" rid="B72">Wickham, 2016</xref>), ggrepel (v0.9.1) (<xref ref-type="bibr" rid="B73">Wickham, 2007</xref>; <xref ref-type="bibr" rid="B56">Slowikowski, 2021</xref>), dplyr (v1.0.7) (<xref ref-type="bibr" rid="B1">Wickham et al., 2021</xref>), tidyr (v1.1.4) (<xref ref-type="bibr" rid="B77">Wickham and Girlich, 2021</xref>), stringr (v1.4.0), rlang (v0.4.11) (<xref ref-type="bibr" rid="B15">Henry and Wickham, 2021</xref>), haven (v2.4.3) (<xref ref-type="bibr" rid="B21">Wickham and Miller, 2021</xref>), vegan (v2.5-7) (<xref ref-type="bibr" rid="B67">Oksanen et al., 2020</xref>), keggrest (v1.30.1) (<xref ref-type="bibr" rid="B60">Tenenbaum, 2017</xref>), and pfam.db (v3.12.0). To have a better representation of the result, it is recommended to provide a metadata table by including the file path in the config file (<italic>METADATA</italic>) with sample ID, conditions and sample alias columns. If no metadata are provided, the sample IDs are used to generate the plots. However, the user can always use MIntO outputs for further downstream analysis.</p>
</sec>
<sec id="s2-14">
<title>Data</title>
<sec id="s2-14-1">
<title>Inflammatory Bowel Disease Multi&#x2019;Omics Database Samples</title>
<p>We used 91 human fecal metagenomes from the Inflammatory Bowel Disease Multi&#x2019;omics Database [IBDMDB, (<xref ref-type="bibr" rid="B34">Lloyd-Price et al., 2019</xref>)]. The IBDMDB study provides matching Illumina metagenomic and metatranscriptomic data. We selected six participants diagnosed as non-IBD [P6018 (nIBD1), M2072 (nIBD2)]; Crohn&#x2019;s disease [H4006 (CD1) and H4020 (CD2)]; and ulcerative colitis [H4019 (UC1) and H4035 (UC2)] that were followed for 1&#xa0;year each (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). Sample H4019_20 was not included due to a parsing error. Sequence data were retrieved from NCBI Short Read Archive under BioProject identifier PRJNA398089.</p>
</sec>
<sec id="s2-14-2">
<title>Paired-End Illumina and Nanopore-Based Metagenomic Data From Head and Neck Cancer Patients</title>
<p>We used human fecal metagenomes from head and neck cancer (HNC) patients (Wongsurawat et al., 2019), where samples were sequenced using Illumina and Nanopore technologies. We selected a subset of five patients: PatientHNC_03, PatientHNC_05, PatientHNC_06, PatientHNC_08 and PatientHNC_10. These were obtained from NCBI Short Read Archive under the accession numbers SRR7947170, SRR7947175, SRR7947177, SRR7947178, SRR7947179, SRR7947181, SRR7947184, SRR7947185, SRR7947186 and SRR7947187.</p>
</sec>
<sec id="s2-14-3">
<title>Human Genome</title>
<p>During MIntO pre-processing, the human genome (build hg38) was used to remove putative host-derived sequences (host genome filtering step).</p>
</sec>
</sec>
<sec id="s2-15">
<title>Implementation of the Pipeline</title>
<p>MIntO implementation and automation are achieved by Snakemake (<xref ref-type="bibr" rid="B39">M&#xf6;lder, 2021</xref>), a user-friendly framework that facilitates the scalability of the pipeline by optimizing the number of parallel processes from a single-core workstation to compute clusters. MIntO leverages singularity containers (<xref ref-type="bibr" rid="B31">Kurtzer, Sochat and Bauer, 2017</xref>) and Conda environments (<xref ref-type="bibr" rid="B3">Anaconda Inc, 2020</xref>) to ensure version control of the different libraries and implements a pipeline connecting several state of the art bioinformatic tools. In this way, MIntO enables consistency of the results and straightforward application by users with basic informatics skills to analyze complex omics data. The only dependencies are FetchMGs and Conda.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>MIntO can be run in three different modes, thanks to its modular design, depending on the user&#x2019;s preference and available data: <italic>genome-based assembly-free</italic>, <italic>gene-catalog-based assembly-free</italic> and <italic>genome-based assembly-dependent</italic>. For all the three modes, users have to input FASTQ files from metagenomic and/or metatranscriptomic paired-end raw short reads and optionally, nanopore-based long reads, as well as a configuration file indicating the metagenomic and/or metatranscriptomic sample names and the corresponding location of raw FASTQ files. In the <italic>genome-based assembly-dependent</italic> mode, the given metagenomes are used to retrieve MAGs, while in the two <italic>assembly-free</italic> modes, <italic>genome-based</italic> or <italic>gene-catalog-based</italic>, the user also has to provide a set of reference genomes or a gene-catalog database, respectively, to generate the gene and functional profiles. These two options could be used when the user is working with a defined community or when there are not enough metagenomic samples to generate representative MAGs. These three modalities are illustrated in <xref ref-type="fig" rid="F1">Figure 1A</xref>.</p>
<p>MIntO can be divided into seven major steps, which will be discussed in the next paragraphs using our analysis of example data (<xref ref-type="fig" rid="F1">Figure 1A</xref>):<list list-type="simple">
<list-item>
<p>1. Quality control and pre-processing</p>
</list-item>
<list-item>
<p>2. Assembly-free taxonomy profiling</p>
</list-item>
<list-item>
<p>3. Recovery of MAGs and taxonomic annotation (only run in <italic>genome-based assembly-dependent</italic> mode)</p>
</list-item>
<list-item>
<p>4. Gene prediction and functional annotation (only run in <italic>genome-based</italic> modes)</p>
</list-item>
<list-item>
<p>5. Alignment and normalization</p>
</list-item>
<list-item>
<p>a. g<italic>enome-based</italic> mode: recovered MAGs or publicly available genomes</p>
</list-item>
<list-item>
<p>b. <italic>gene-based</italic> mode: gene catalog</p>
</list-item>
<list-item>
<p>6. Integration: Gene and functional profiling</p>
</list-item>
<list-item>
<p>7. Visualization and reporting</p>
</list-item>
</list>
</p>
<p>The third step is skipped if an assembly-free mode is selected, and the fourth step is skipped when <italic>gene catalog-based assembly-free</italic> mode is chosen (<xref ref-type="fig" rid="F1">Figure 1A</xref>). An overview of the directories generated can be seen in <xref ref-type="sec" rid="s10">Supplementary Figure S1</xref>
<italic>.</italic>
</p>
<p>To illustrate the use of MIntO, a set of 91 human fecal metagenomes from the Inflammatory Bowel Disease Multi&#x2019;omics Database (IBDMDB) was selected (<xref ref-type="bibr" rid="B34">Lloyd-Price et al., 2019</xref>). These samples correspond to six participants diagnosed as non-IBD (nIBD1 and nIBD2), Crohn&#x2019;s disease, (CD1 and CD2) and ulcerative colitis, (UC1 and UC2), which were followed for 1&#xa0;year each (<xref ref-type="sec" rid="s10">Supplementary Figure S2</xref>, <xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). The IBDMDB study provides matching Illumina metagenomic and metatranscriptomic data. The subset of samples used here correspond to 933.4 and 612 million read-pairs (2 &#xd7; 101&#xa0;bp) from metagenomic and metatranscriptomic sequencing, respectively (mean 10.85 million read-pairs, ranging from 0.26 to 21.04 million for metagenomic; mean 6.18 million read-pairs, ranging from 0.01 to 15.72 million for metatranscriptomic).</p>
<p>Here, we present the results from the <italic>genome-based assembly-dependent</italic> and <italic>gene catalog-based assembly-free</italic> modes, where we used recovered MAGs and the Integrated Gene Catalog (IGC) (<xref ref-type="bibr" rid="B33">Li et al., 2014</xref>), respectively, as reference to profile genes and functions.</p>
<sec id="s3-1">
<title>Quality Control and Pre-Processing</title>
<p>The IBDMDB dataset was already filtered by quality and sequence adapters, therefore the first step in the pre-processing of the 91 samples was skipped (<italic>trimmomatic_adaptors &#x3d; Skip</italic>, see Methods). We then used a minimum read length cutoff of 53&#xa0;bp for metagenomic and 54&#xa0;bp for metatranscriptomic to keep 95% of the longest sequences using Trimmomatic (<xref ref-type="bibr" rid="B9">Bolger, Lohse and Usadel, 2014</xref>) (<xref ref-type="sec" rid="s10">Supplementary Figure S3</xref>).</p>
<p>Subsequently, putative host-derived sequences were removed using the human genome (build hg38). In silico rRNA sequences screening was exclusively applied to metatranscriptomic reads using SortMeRNA (<xref ref-type="bibr" rid="B28">Kopylova, No&#xe9; and Touzet, 2012</xref>). This resulted in a total number of 599.4 million high-quality read-pairs for metagenomic and 910.9 million high-quality read-pairs for metatranscriptomic data (<xref ref-type="table" rid="T2">Table 2</xref>, <xref ref-type="sec" rid="s10">Supplementary Figure S4</xref>).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Median (minimum and maximum) of raw and high-quality million read-pairs in the 91 human fecal microbiome samples from the IBDMDB.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">metagenomic</th>
<th align="center">metatranscriptomic</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Raw read-pairs (millions)</td>
<td align="center">10.85 (10.15&#x2013;21.04)</td>
<td align="center">6.18 (6.65&#x2013;15.72)</td>
</tr>
<tr>
<td align="left">High quality read-pairs (millions)</td>
<td align="center">10.56 (9.9&#x2013;20.58)</td>
<td align="center">6.04 (6.52&#x2013;15.45)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>Assembly-Free Taxonomy Profiling</title>
<p>Once the reads were pre-processed, high-quality reads were profiled at species level using MetaPhlAn3 (<xref ref-type="bibr" rid="B7">Beghini et al., 2021</xref>) (<xref ref-type="fig" rid="F1">Figure 1A</xref>, assembly-free taxonomy profiling step). In <xref ref-type="fig" rid="F2">Figure 2A</xref>, we can see the temporal shifts and dynamics exhibited by microbes over the course of 1&#xa0;year and the difference of microbial composition between the six participants focusing on the 15 most abundant genera across the samples. In general, the most predominant genera are <italic>Bacteroides</italic>, <italic>Faecalibacterium</italic> and <italic>Roseburia</italic>. The constitution of a separate cluster by samples from participant nIBD2 in <xref ref-type="fig" rid="F2">Figure 2B</xref> cannot be explained by the 15 most abundant genera across the samples (<xref ref-type="fig" rid="F2">Figure 2A</xref>), but it could be due to the difference in composition of lower-abundance bacteria.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Taxonomic profiles. <bold>(A)</bold> Relative abundance for the 91 samples for the 15 most abundant genera across the samples using MetaPhlAn3 (<xref ref-type="bibr" rid="B7">Beghini et al., 2021</xref>). <bold>(B)</bold> Projection of the first two principal coordinates based on Bray&#x2013;Curtis dissimilarity from the microbiome composition using MetaPhlAn3 (<xref ref-type="bibr" rid="B7">Beghini et al., 2021</xref>). <bold>(C)</bold> Taxonomy tree representing the 131 SGBs taxonomies after running PhyloPhlAn3 (<xref ref-type="bibr" rid="B5">Asnicar et al., 2020</xref>) on the retrieved MAGs. The first six rings mark MAGs that were retrieved in the 6 patients with the different conditions used in this work (nIBD, CD and UC), while the last ring marks the MAGs obtained from co-assembly. <bold>(D)</bold> Distribution of the SGBs in the 6 patients: 51 SGBs taxonomies were retrieved from just one sample, 13 from two samples, 3 from three samples, 2 from five samples and 1 in all the samples. The last bar represents the 61 taxonomies that were found only by having performed co-assembly.</p>
</caption>
<graphic xlink:href="fbinf-02-846922-g002.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>Recovery of MAGs and Taxonomic Annotation</title>
<p>In parallel, the pre-processed reads underwent the assembly step in the <italic>genome-based assembly-dependent</italic> mode (<xref ref-type="fig" rid="F1">Figure 1A</xref>, recovery of MAGs and taxonomic annotation step). As this dataset consists of short-read metagenomes only, we used two assembly approaches to recover high-quality scaffolds: 1) assembly of each metagenome individually (single-assembly) using MetaSPAdes assembler (<xref ref-type="bibr" rid="B44">Nurk et al., 2017</xref>) and 2) assembly of all metagenomes together (co-assembly) using MEGAHIT (<xref ref-type="bibr" rid="B32">Li et al., 2015</xref>) assembler. Genome bins were generated from assembled scaffolds that were at least 2,500&#xa0;bp long by mapping the 91 samples individually to the scaffolds, calculating the sequence depth of each scaffold in the 91 samples, and finally running VAMB (<xref ref-type="bibr" rid="B43">Nissen et al., 2021</xref>) four times with different parameters and GPU mode (see Methods).</p>
<p>After binning, 5,048 MAGs were retrieved from the 91 metagenomic samples. Using CheckM (<xref ref-type="bibr" rid="B46">Parks et al., 2015</xref>), we identified high-quality (HQ) MAGs (completeness &#x3e; 95% and contamination &#x3c; 5%) and kept 957 MAGs. We then obtained unique high-quality MAGs when clustering the HQ MAGs at 99% ANI distance (<xref ref-type="bibr" rid="B22">Jain et al., 2018</xref>) with CoverM (<ext-link ext-link-type="uri" xlink:href="https://github.com/wwood/CoverM">https://github.com/wwood/CoverM&#x23;usage</ext-link>) and choosing the best genome in a given cluster using a genome quality score (see Methods). This de-replication process resulted in 163 MAGs which constituted a set of non-redundant genomes (available at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.6360083">10.5281/zenodo.6360083</ext-link>). These MAGs are useful to collectively explain the ecological description and biodiversity in the samples, and to capture sample-specific variation at functional and abundance level without relying on publicly available reference genomes. Additionally, working with a restricted number of genomes is helpful to speed up the next steps of the pipeline.</p>
<p>The taxonomic annotation of the 163 MAGs was performed by <italic>phylophlan_metagenomic</italic> module in PhyloPhlAn3 (<xref ref-type="bibr" rid="B5">Asnicar et al., 2020</xref>), which also provides taxonomic lineage information about the 10 nearest genomes in the PhyloPhlAn3 genome database. Each MAG was assigned to a species-level genome bin (SGB) if its closest genome in the database was within 5% average nucleotide identity. This resulted in the 163 MAGs falling into 131 SGBs (<xref ref-type="fig" rid="F2">Figure 2C</xref>). In general, MAGs with a distance higher than 5% to the closest genome in the database can be considered as putative novel species (<xref ref-type="bibr" rid="B35">Manara et al., 2019</xref>; <xref ref-type="bibr" rid="B47">Pasolli et al., 2019</xref>). However, we did not recover any MAGs from putative novel species in this dataset.</p>
<p>By default, MIntO performs co-assembly, which although time consuming, is an extremely important step. In fact, we obtained the highest number of unique taxa from the co-assembled samples compared to any single-sample assembly (<xref ref-type="table" rid="T3">Table 3</xref>). Remarkably, 61 of the 131 taxonomies (&#x223c;46%) could be retrieved only by performing co-assembly (<xref ref-type="fig" rid="F2">Figure 2D</xref>). With single-sample assembly we still retrieved 31 (&#x223c;23%) unique taxonomies not covered by the co-assembled samples, of which 13 (&#x223c;10% of the total) are only found in one sample (<xref ref-type="fig" rid="F2">Figure 2C</xref>). This is helpful to better distinguish sample-specific composition, as for example <italic>Akkermansia muciniphila</italic> SGB9228, which is the second <italic>Akkermansiacae</italic> species by presence in the human population (<xref ref-type="bibr" rid="B25">Karcher et al., 2021</xref>) can only be found in patient CD1. These results are achievable only by performing both single and co-assembly.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Number of SGB taxonomies retrieved per sample.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Sample/Method</th>
<th align="center">Number of Taxa</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">nIBD1</td>
<td align="char" char=".">18</td>
</tr>
<tr>
<td align="left">nIBD2</td>
<td align="char" char=".">21</td>
</tr>
<tr>
<td align="left">CD1</td>
<td align="char" char=".">24</td>
</tr>
<tr>
<td align="left">CD2</td>
<td align="char" char=".">15</td>
</tr>
<tr>
<td align="left">UC1</td>
<td align="char" char=".">21</td>
</tr>
<tr>
<td align="left">UC2</td>
<td align="char" char=".">24</td>
</tr>
<tr>
<td align="left">Co-assembly</td>
<td align="char" char=".">100</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, we performed our own benchmark to show that combining long and short reads improves the assembly contiguity. MIntO assembled paired-end metagenomes from the gut microbiota of five patients with head and neck cancer (Wongsurawat et al., 2019), which were generated by 1) Illumina-only, or 2) Illumina and Nanopore sequencing platforms. The number of generated scaffolds (127,315 and 172,888 for Illumina and Illumina &#x2b; Nanopore, respectively), and their mean length (9.44&#xa0;kb and 9.72&#xa0;kb for Illumina and Illumina &#x2b; Nanopore, respectively), were greater when long-reads were included in the assembly. Furthermore, Illumina &#x2b; Nanopore assembly generated 13 scaffolds longer than 600&#xa0;kb with a maximum of 1,119&#xa0;kb, whereas the assembly of Illumina-only data generated 2 scaffolds longer than 600&#xa0;kb with a maximum of 736&#xa0;kb. Finally, the scaffold length distribution shows that scaffolds from Illumina &#x2b; Nanopore assemblies are more contiguous than Illumina-only assemblies (<xref ref-type="sec" rid="s10">Supplementary Figure S6</xref>).</p>
</sec>
<sec id="s3-4">
<title>Gene Prediction and Functional Annotation</title>
<p>The unique set of MAGs recovered in the previous step underwent gene prediction and functional annotation (<xref ref-type="fig" rid="F1">Figure 1A</xref>, gene prediction and functional annotation). Prokka (<xref ref-type="bibr" rid="B54">Seemann, 2014</xref>) was used to identify and annotate the genes, retrieving the corresponding nucleotide and amino acid sequences. A total of 412,394 genes were predicted in the 163 recovered MAGs. These were annotated with seven different functional databases: eggNOG (<xref ref-type="bibr" rid="B74">Yin et al., 2012</xref>; <xref ref-type="bibr" rid="B19">Huerta-Cepas et al., 2019</xref>), KEGG Pathways, Modules and KOs (<xref ref-type="bibr" rid="B23">Kanehisa and Goto, 2000</xref>), dbCAN modules and enzyme classes (<xref ref-type="bibr" rid="B74">Yin et al., 2012</xref>), and Pfam (<xref ref-type="bibr" rid="B38">Mistry et al., 2021</xref>) (<xref ref-type="fig" rid="F3">Figure 3</xref>). The same process could also be applied to user-provided genome sequences under <italic>genome-based assembly-free</italic> mode.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Comparison of number of genes and features per function database between non-redundant high quality 163 MAGs and IGC.</p>
</caption>
<graphic xlink:href="fbinf-02-846922-g003.tif"/>
</fig>
<p>The gene and function annotation step was skipped in the <italic>gene catalog-based assembly-free</italic> mode as we used existing eggNOG, KEGG Pathways, KEGG Modules, KEGG KO, dbCAN modules, dbCAN enzymes class, Pfam function annotation for IGC (available at <ext-link ext-link-type="uri" xlink:href="https://db.cngb.org/microbiome/genecatalog/genecatalog_human/">https://db.cngb.org/microbiome/genecatalog/genecatalog_human/</ext-link>). The number of expressed genes and functions for both modes are summarized in <xref ref-type="fig" rid="F3">Figure 3</xref>. Even though we detected &#x3e; 5 &#xd7; genes by mapping the metagenomes to IGC compared to genes encoded in the 163 MAGs, genes from the MAGs covered the vast majority of the functions detected via IGC. In some cases such as Pfam and CAZy databases, MAGs recovered more functions suggesting that contiguous assemblies and more complete genes could improve the quality of functional annotations.</p>
</sec>
<sec id="s3-5">
<title>Alignment and Normalization</title>
<p>The metagenomic and metatranscriptomic high-quality reads were mapped to a reference database followed by TPM normalization to obtain the relative abundance of genes from metagenomic read alignments (i.e., gene abundance profile) and transcripts from metatranscriptomic read alignments (i.e., gene transcript profile) (<xref ref-type="fig" rid="F1">Figure 1A</xref>, alignment, normalization and integration). We used as a reference database the 163 recovered MAGs for the <italic>genome-based</italic> mapping and the IGC (<xref ref-type="bibr" rid="B33">Li et al., 2014</xref>) for the <italic>gene-based</italic> alignment. Overall, the mappability rate at 95% of sequence identity for MAGs (median 72.26%) was lower than for IGC (median 92.47%) with the highest difference for participant CD2 (<xref ref-type="sec" rid="s10">Supplementary Figure S5</xref>), which could be due also to the lower number of taxonomies retrieved for the samples (<xref ref-type="table" rid="T3">Table 3</xref>). However, this difference was not as remarkable when using metatranscriptomic reads (77.61 and 73.9% median, respectively).</p>
</sec>
<sec id="s3-6">
<title>Integration: Gene and Function Expression Profiling</title>
<p>The variation of microbial community transcript levels may be affected by the changes in gene expression and/or by the community turnover. To disentangle the individual contributions of these mechanisms across the different samples, we integrated gene abundance and transcript abundance profiles (<xref ref-type="bibr" rid="B52">Salazar et al., 2019</xref>) (see Methods). The obtained levels of gene expression represent the relative amount of expressed transcripts per gene (<xref ref-type="fig" rid="F1">Figure 1A</xref>, integration: gene and functional profiling). From the 412,394 predicted genes in the 163 recovered MAGs, 219,133 genes were expressed in at least one sample, while we detected the expression of 1,260,394 genes from the 9.9 million genes in IGC.</p>
<p>Furthermore, the corresponding gene profiles were used to generate the function abundance, transcript and expression profiles by grouping the annotated genes into functions. The highest number of features detected in the samples corresponded to the eggNOG database on both modes, followed by Pfam or KEGG KO (<xref ref-type="fig" rid="F3">Figure 3</xref>). We identified 5,734 and 6,131 KEGG KO expressed features when we used the recovered MAGs and IGC as a reference, respectively. Among the 7,217 KEGG KO functions identified between the two profiles, 64.4% (4,651 features) were found in both. The 15% of features (1,086) uniquely identified in the MAGs could correspond to genomes not included in the database and the 20.5% of the functions (1,481) detected in IGC could belong to low abundant bacteria whose genomes could not be retrieved or were missed due to MAGs filtered out based on our quality criteria.</p>
<p>We used MIntO&#x2019;s visualization features to perform principal coordinate analysis (PCoA) on the different gene and functional profiles to observe the longitudinal compositional changes and to compare the dissimilarities between participants. In <xref ref-type="fig" rid="F4">Figure 4A</xref> we show the gene expression PCoA plot for the <italic>assembly-free gene catalog</italic> mode using IGC (<xref ref-type="bibr" rid="B33">Li et al., 2014</xref>). In general, the samples were clustered by Crohn&#x2019;s disease and Ulcerative colitis diagnosis suggesting a similar bacterial abundance and expressed genes due to the presence of the disease (<xref ref-type="bibr" rid="B29">Kostic, Xavier and Gevers, 2014</xref>; <xref ref-type="bibr" rid="B34">Lloyd-Price et al., 2019</xref>). Samples from participants used as control (nIBD1 and nIBD2) were clustered separately, probably due to the inter-individual variations in the microbiome composition. In fact, the most abundant genus in all participants was <italic>Bacteroides</italic>, with the exception of nIBD1 where <italic>Roseburia</italic> and <italic>Faecalibacterium</italic> were predominant. At transcript level (<xref ref-type="sec" rid="s10">Supplementary Figure S7</xref>), the dissimilarity between the samples explained by the first two principal coordinates (18.7% and 12.2%) was higher than at gene expression level (8.9% and 7.2%). The transcript abundance changes might be mainly attributed either to differences in the expression of genes encoded by the microbes in the community or changes in the abundance of these members and their related genes or a combination of these mechanisms. Hence, the computation of gene expression profiles by the integration of abundances of genes and the respective transcripts is of crucial importance to obtain a more accurate representation of ecologically relevant processes that are occurring.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Projection of the first two principal coordinates based on expression profiles Bray&#x2013;Curtis dissimilarity at <bold>(A)</bold> gene and <bold>(B)</bold> function KEGG KO levels using a subset of 91 samples from IBDMDB. Labels correspond to the <italic>sample alias</italic> and are colored by <italic>condition</italic> (patients diagnosis).</p>
</caption>
<graphic xlink:href="fbinf-02-846922-g004.tif"/>
</fig>
<p>Overall, the dissimilarities between the samples were visible at the gene expression, gene abundance and transcript abundance profiles (<xref ref-type="fig" rid="F4">Figure 4A</xref> and <xref ref-type="sec" rid="s10">Supplementary Figure S7</xref>). However, at function expression level (<xref ref-type="fig" rid="F4">Figure 4B</xref>) the clusters were not as well defined, suggesting that genes from different species could harbor the same functions in different microbial communities. Although the taxonomic composition differed between the six participants and consequently the gene composition and expression, the functional profiles across individuals and time were more conserved (functional redundancy) (<xref ref-type="bibr" rid="B62">Tian et al., 2020</xref>). Differences in functional profiles between nonIBD and IBD diagnosed participants could provide insights into the functions involved in microbiome&#x2013;host interactions at states of health or disease (<xref ref-type="bibr" rid="B16">Heintz-Buschart and Wilmes, 2018</xref>).</p>
</sec>
<sec id="s3-7">
<title>Visualization and Reporting</title>
<p>Further analyses can be done using the output files (<xref ref-type="fig" rid="F1">Figure 1A</xref>, visualization and reporting; <xref ref-type="sec" rid="s10">Supplementary Figure S1</xref>). MIntO generates three different types of table: 1) assembly-free and assembly-based taxonomic profiles; 2) gene profiles, including the gene IDs [generated by Prokka (<xref ref-type="bibr" rid="B54">Seemann, 2014</xref>; <xref ref-type="bibr" rid="B7">Beghini et al., 2021</xref>) when selecting <italic>assembly-dependent</italic> mode or sequence IDs when choosing <italic>assembly-free</italic> mode] and normalized gene abundance, transcript or expression; and 3) functional profiles per database, including the function IDs, function description and function abundance, transcript or expression normalized counts. For an easier downstream analysis of these data, phyloseq objects are generated for the taxonomic, gene and functional profiles.</p>
<p>MIntO also outputs the shown plots as preliminary results to help the user in the downstream analysis (<xref ref-type="fig" rid="F2">Figures 2A,B</xref>, <xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="sec" rid="s10">Supplementary Figures S3, S7</xref>).</p>
<p>The metadata provided in IBDMDB (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>) was given as an input to the pipeline, which colored the samples by <italic>sample_alias</italic> (participant&#x2019;s ID) in the output plots.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>MIntO is a versatile pipeline that integrates metagenomic and metatranscriptomic data, beyond a comparison of the gene and transcript abundances, in order to quantify gene and function expression in a very straightforward way. The modular design of MIntO enables the user to run the pipeline using three available modes based on the input data and the experimental design.</p>
<p>In order to illustrate the pipeline, a subset of 91 human fecal microbiome samples from the IBDMDB (Illumina metagenomic and metatranscriptomic paired-reads) was used to run the full version of the pipeline with default parameters. Here, we show the complementary results from two of the three available modes, <italic>genome-based assembly-dependent</italic> and <italic>gene catalog-based assembly-free</italic>. In the former, MIntO retrieved 163 high-quality non-redundant MAGs that encoded 412,394 genes, among which 219,133 genes were expressed in at least one sample, while 1,260,394 genes from IGC were expressed in the <italic>gene catalog-based assembly-free</italic> mode. Overall, the dissimilarities between the samples were visible at the taxonomic and gene levels, while the functional profiles across individuals and time were more conserved (functional redundancy), indicating that strain-specific genes from different microbiomes represented similar functions. Interestingly, among the 7,217 KEGG KO functions identified between the two profiles, 15% of the features were uniquely identified in the MAGs and 20.5% of the functions were detected in IGC.</p>
<p>The distinctive feature of this pipeline is the integration of the metagenomic and metatranscriptomic data, to obtain the expression profiles and furthermore the functional profiles by annotating the sequences with several databases. This enables us to study in detail the variation in expression of the genes and functions in the different samples across time and experiment conditions, thus the community behavior. Overall, the IBDMDB-samples clustered by the participant ID using the genes and transcript abundances and gene expression. However, using the KEGG KO annotations at function expression level, the clusters are not as well defined, due to the functional redundancy (<xref ref-type="bibr" rid="B62">Tian et al., 2020</xref>).</p>
<p>Another important feature of MIntO is performing <italic>de novo</italic> assembly and contig binning to recover high-quality MAGs from metagenomic reads, which compared to other methods utilizes an accurate unsupervised deep learning approach in the form of variational autoencoders (<xref ref-type="bibr" rid="B43">Nissen et al., 2021</xref>). The <italic>assembly-dependent</italic> mode could be helpful to retrieve novel genomes that are missed by reference-dependent profiling methods (<xref ref-type="bibr" rid="B47">Pasolli et al., 2019</xref>). The recovery of MAGs is indispensable to uncover the diversity of bacteria in an environment and it is crucial for an optimal calculation of the variation of gene expression, including unknown or functional genes from biosynthetic gene clusters (<xref ref-type="bibr" rid="B75">Youngblut et al., 2020</xref>). Additionally, new putative genomes can increase the number of known species in the available databases, especially when the analyses are performed on metagenomes coming from new environmental sources.</p>
<p>In conclusion, in this paper we show how MIntO can be a useful tool to analyze metagenomic and metatranscriptomic data in a standardized way, enabling the study of microbial ecology by linking functions to genomes and environmental context. We foresee that this pipeline will contribute to the understanding of the dynamics of the molecular activities captured by the community turnover and gene expression alterations as the cause that shapes community transcript levels. Elucidating the functions and characterizing the specific strains of a community will be crucial to increase our knowledge of the microbiome&#x2019;s contribution to human health and environment.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. Matching Illumina metagenomic and metatranscriptomic data from IBDMDB can be found here: <ext-link ext-link-type="uri" xlink:href="https://ibdmdb.org/tunnel/public/summary.html">https://ibdmdb.org/tunnel/public/summary.html</ext-link>, IBDMDB, BioProject identifier PRJNA398089. Matching shotgun metagenomic data generated from both Illumina and Nanopore technologies can be found in NCBI Short Read Archive under the accession numbers SRR7947170, SRR7947175, SRR7947177, SRR7947178, SRR7947179, SRR7947181, SRR7947184, SRR7947185, SRR7947186 and SRR7947187. Non-redundant MAGs constructed by MIntO from 91 metagenomes from IBDMDB are available at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.6360083">https://doi.org/10.5281/zenodo.6360083</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>CS and MA conceived and designed the tool. CS, EN, VG, and MA created the software. CS, EN, and MA wrote the manuscript and performed all necessary testing. All authors read, revised, and approved the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>Novo Nordisk Foundation Center for Basic Metabolic Research is an independent Research Center, based at the University of Copenhagen, Denmark, and partially funded by an unconditional donation from the Novo Nordisk Foundation (<ext-link ext-link-type="uri" xlink:href="http://www.cbmr.ku.dk/">www.cbmr.ku.dk</ext-link>) (Grant no. NNF18CC0034900). This work was supported by the Danish Council of Independent Research (Grant no. 6111-00471B). VG and EN were supported by the European Union&#x2019;s Horizon 2020 research and innovation program (GALAXY: Grant no. 668031).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We are thankful to the members of the Arumugam group and Gabriel Fernandes for inspiring discussions and feedback on the manuscript.</p>
</ack>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fbinf.2022.846922/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fbinf.2022.846922/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Table2.XLSX" id="SM1" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image3.JPEG" id="SM2" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table3.XLSX" id="SM3" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image1.JPEG" id="SM4" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image4.JPEG" id="SM5" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image7.JPEG" id="SM6" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image2.JPEG" id="SM7" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image5.JPEG" id="SM8" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table4.XLSX" id="SM9" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table1.XLSX" id="SM10" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image6.JPEG" id="SM11" mimetype="application/JPEG" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Almeida</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mitchell</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Boland</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Forster</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Gloor</surname>
<given-names>G. B.</given-names>
</name>
<name>
<surname>Tarkowska</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>A New Genomic Blueprint of the Human Gut Microbiota</article-title>. <source>Nature</source> <volume>568</volume> (<issue>7753</issue>), <fpage>499</fpage>&#x2013;<lpage>504</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-0965-1</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="web">
<collab>Anaconda Inc</collab> (<year>2020</year>). <article-title>Anaconda Software Distribution, <italic>Anaconda Documentation</italic> [Preprint]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://docs.anaconda.com/">https://docs.anaconda.com/</ext-link>
</comment>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aramaki</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Blanc-Mathieu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Endo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ohkubo</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kanehisa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Goto</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>KofamKOALA: KEGG Ortholog Assignment Based on Profile HMM and Adaptive Score Threshold</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>7</issue>), <fpage>2251</fpage>&#x2013;<lpage>2252</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz859</pub-id> </citation>
</ref>
<ref id="B78">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Arumugam</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>msamtools: Microbiome-Related Extension to Samtools</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/arumugamlab/msamtools">https://github.com/arumugamlab/msamtools</ext-link> (Accessed: March 31, 2022)</comment>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Asnicar</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Beghini</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mengoni</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Manara</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Manghi</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Precise Phylogenetic Analysis of Microbial Isolates and Genomes from Metagenomes Using PhyloPhlAn 3.0</article-title>. <source>Nat. Commun.</source> <volume>11</volume> (<issue>1</issue>), <fpage>2500</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-16366-7</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bashan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gibson</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Friedman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Carey</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Weiss</surname>
<given-names>S. T.</given-names>
</name>
<name>
<surname>Hohmann</surname>
<given-names>E. L.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Universality of Human Microbial Dynamics</article-title>. <source>Nature</source> <volume>534</volume> (<issue>7606</issue>), <fpage>259</fpage>&#x2013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1038/nature18301</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beghini</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>McIver</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Blanco-M&#xed;guez</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dubois</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Asnicar</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Maharjan</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Integrating Taxonomic, Functional, and Strain-Level Profiling of Diverse Microbial Communities with bioBakery 3</article-title>. <source>eLife</source> <volume>10</volume>, <fpage>e65088</fpage>. <pub-id pub-id-type="doi">10.1101/2020.11.19.388223</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bertrand</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Shaw</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kalathiyappan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ng</surname>
<given-names>A. H. Q.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Hybrid Metagenomic Assembly Enables High-Resolution Analysis of Resistance Determinants and mobile Elements in Human Microbiomes</article-title>. <source>Nat. Biotechnol.</source> <volume>37</volume> (<issue>8</issue>), <fpage>937</fpage>&#x2013;<lpage>944</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-019-0191-2</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bolger</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Lohse</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Usadel</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Trimmomatic: a Flexible Trimmer for Illumina Sequence Data</article-title>. <source>Bioinformatics</source> <volume>30</volume> (<issue>15</issue>), <fpage>2114</fpage>&#x2013;<lpage>2120</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu170</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brown</surname>
<given-names>C. L.</given-names>
</name>
<name>
<surname>Keenum</surname>
<given-names>I. M.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Vikesland</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Pruden</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Critical Evaluation of Short, Long, and Hybrid Assembly for Contextual Analysis of Antibiotic Resistance Genes in Complex Environmental Metagenomes</article-title>. <source>Sci. Rep.</source> <volume>11</volume> (<issue>1</issue>), <fpage>3753</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-83081-8</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cantalapiedra</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Hern&#xe1;ndez-Plaza</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Letunic</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bork</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Huerta-Cepas</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>EggNOG-Mapper V2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale</article-title>. <source>Mol. Biol. Evol.</source> <volume>38</volume>, <fpage>5825</fpage>&#x2013;<lpage>5829</lpage>. <comment>[Preprint]</comment>. <pub-id pub-id-type="doi">10.1093/molbev/msab293</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Danecek</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bonfield</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Liddle</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Marshall</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ohan</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Pollard</surname>
<given-names>M. O.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Twelve Years of SAMtools and BCFtools</article-title>. <source>GigaScience</source> <volume>10</volume> (<issue>2</issue>), <fpage>giab008</fpage>. <pub-id pub-id-type="doi">10.1093/gigascience/giab008</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Donia</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Fischbach</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>HUMAN MICROBIOTA. Small Molecules from the Human Microbiota</article-title>. <source>Science</source> <volume>349</volume> (<issue>6246</issue>), <fpage>1254766</fpage>. <pub-id pub-id-type="doi">10.1126/science.1254766</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Dowle</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Srinivasan</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>data.table: Extension of &#x201c;data.frame&#x201d; [R Package data.table Version 1.14.2]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=data.table">https://CRAN.R-project.org/package&#x3d;data.table</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heintz-Buschart</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wilmes</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Human Gut Microbiome: Function Matters</article-title>. <source>Trends Microbiol.</source> <volume>26</volume> (<issue>7</issue>), <fpage>563</fpage>&#x2013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1016/j.tim.2017.11.002</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Henry</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>rlang: Functions for Base Types and Core R and &#x201c;Tidyverse&#x201d; Features [R Package rlang Version 0.4.11]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=rlang">https://CRAN.R-project.org/package&#x3d;rlang</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Entwistle</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yohe</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>dbCAN-Seq: a Database of Carbohydrate-Active Enzyme (CAZyme) Sequence and Annotation</article-title>. <source>Nucleic Acids Res.</source> <volume>46</volume> (<issue>D1</issue>), <fpage>D516</fpage>&#x2013;<lpage>D521</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkx894</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huerta-Cepas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Forslund</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Coelho</surname>
<given-names>L. P.</given-names>
</name>
<name>
<surname>Szklarczyk</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>von Mering</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Fast Genome-wide Functional Annotation through Orthology Assignment by eggNOG-Mapper</article-title>. <source>Mol. Biol. Evol.</source> <volume>34</volume> (<issue>8</issue>), <fpage>2115</fpage>&#x2013;<lpage>2122</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msx148</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huerta-Cepas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Szklarczyk</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Heller</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hern&#xe1;ndez-Plaza</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Forslund</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Cook</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>eggNOG 5.0: a Hierarchical, Functionally and Phylogenetically Annotated Orthology Resource Based on 5090 Organisms and 2502 Viruses</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>D309</fpage>&#x2013;<lpage>D314</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1085</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<collab>Human Microbiome Project Consortium</collab> (<year>2012</year>). <article-title>A Framework for Human Microbiome Research</article-title>. <source>Nature</source> <volume>486</volume> (<issue>7402</issue>), <fpage>215</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1038/nature11209</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jain</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rodriguez-R</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Phillippy</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Konstantinidis</surname>
<given-names>K. T.</given-names>
</name>
<name>
<surname>Aluru</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>High Throughput ANI Analysis of 90K Prokaryotic Genomes Reveals clear Species Boundaries</article-title>. <source>Nat. Commun.</source> <volume>9</volume> (<issue>1</issue>), <fpage>5114</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-018-07641-9</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kanehisa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Goto</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>KEGG: Kyoto Encyclopedia of Genes and Genomes</article-title>. <source>Nucleic Acids Res.</source> <volume>28</volume> (<issue>1</issue>), <fpage>27</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.27</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kang</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kirton</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Egan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>An</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>MetaBAT 2: an Adaptive Binning Algorithm for Robust and Efficient Genome Reconstruction from Metagenome Assemblies</article-title>. <source>PeerJ</source> <volume>7</volume>, <fpage>e7359</fpage>. <pub-id pub-id-type="doi">10.7717/peerj.7359</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Karcher</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Nigro</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Pun&#x10d;och&#xe1;&#x159;</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Blanco-M&#xed;guez</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ciciani</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Manghi</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Genomic Diversity and Ecology of Human-Associated Akkermansia Species in the Gut Microbiome Revealed by Extensive Metagenomic Assembly</article-title>. <source>Genome Biol.</source> <volume>22</volume> (<issue>1</issue>), <fpage>209</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-021-02427-7</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Koh</surname>
<given-names>A. Y.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhan</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>FMAP: Functional Mapping and Analysis Pipeline for Metagenomics and Metatranscriptomics Studies</article-title>. <source>BMC bioinformatics</source> <volume>17</volume> (<issue>1</issue>), <fpage>420</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-016-1278-0</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kolmogorov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bickhart</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Behsaz</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gurevich</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rayko</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>S. B.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>metaFlye: Scalable Long-Read Metagenome Assembly Using Repeat Graphs</article-title>. <source>Nat. Methods</source> <volume>17</volume> (<issue>11</issue>), <fpage>1103</fpage>&#x2013;<lpage>1110</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-020-00971-x</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kopylova</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>No&#xe9;</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Touzet</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>SortMeRNA: Fast and Accurate Filtering of Ribosomal RNAs in Metatranscriptomic Data</article-title>. <source>Bioinformatics</source> <volume>28</volume> (<issue>24</issue>), <fpage>3211</fpage>&#x2013;<lpage>3217</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts611</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kostic</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Xavier</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Gevers</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>The Microbiome in Inflammatory Bowel Disease: Current Status and the Future Ahead</article-title>. <source>Gastroenterology</source> <volume>146</volume> (<issue>6</issue>), <fpage>1489</fpage>&#x2013;<lpage>1499</lpage>. <pub-id pub-id-type="doi">10.1053/j.gastro.2014.02.009</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kultima</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Sunagawa</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mende</surname>
<given-names>D. R.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>MOCAT: a Metagenomics Assembly and Gene Prediction Toolkit</article-title>. <source>PloS one</source> <volume>7</volume> (<issue>10</issue>), <fpage>e47656</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0047656</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kurtzer</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Sochat</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Bauer</surname>
<given-names>M. W.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Singularity: Scientific Containers for Mobility of Compute</article-title>. <source>PloS one</source> <volume>12</volume> (<issue>5</issue>), <fpage>e0177459</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0177459</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C.-M.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sadakane</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lam</surname>
<given-names>T.-W.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>1674</fpage>&#x2013;<lpage>1676</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv033</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Sunagawa</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>An Integrated Catalog of Reference Genes in the Human Gut Microbiome</article-title>. <source>Nat. Biotechnol.</source> <volume>32</volume> (<issue>8</issue>), <fpage>834</fpage>&#x2013;<lpage>841</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.2942</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lloyd-Price</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Arze</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ananthakrishnan</surname>
<given-names>A. N.</given-names>
</name>
<name>
<surname>Schirmer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Avila-Pacheco</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Poon</surname>
<given-names>T. W.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Multi-omics of the Gut Microbial Ecosystem in Inflammatory Bowel Diseases</article-title>. <source>Nature</source> <volume>569</volume> (<issue>7758</issue>), <fpage>655</fpage>&#x2013;<lpage>662</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-1237-9</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manara</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Asnicar</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Beghini</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bazzani</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cumbo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zolfo</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Microbial Genomes from Non-human Primate Gut Metagenomes Expand the Primate-Associated Bacterial Tree of Life with over 1000 Novel Species</article-title>. <source>Genome Biol.</source> <volume>20</volume> (<issue>1</issue>), <fpage>299</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1923-9</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McMurdie</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Holmes</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Phyloseq: an R Package for Reproducible Interactive Analysis and Graphics of Microbiome Census Data</article-title>. <source>PloS one</source> <volume>8</volume> (<issue>4</issue>), <fpage>e61217</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0061217</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Milanese</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mende</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Paoli</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Salazar</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ruscheweyh</surname>
<given-names>H. J.</given-names>
</name>
<name>
<surname>Cuenca</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Microbial Abundance, Activity and Population Genomic Profiling with mOTUs2</article-title>. <source>Nat. Commun.</source> <volume>10</volume> (<issue>1</issue>), <fpage>1014</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-08844-4</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mistry</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chuguransky</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Qureshi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salazar</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Sonnhammer</surname>
<given-names>E. L. L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Pfam: The Protein Families Database in 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume> (<issue>D1</issue>), <fpage>D412</fpage>&#x2013;<lpage>D419</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa913</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>M&#xf6;lder</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Sustainable Data Analysis with Snakemake</article-title>. <source>F1000Research</source> <volume>10</volume>, <fpage>33</fpage>. <pub-id pub-id-type="doi">10.12688/f1000research.29032.2</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Morgan</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Access the Bioconductor Project Package Repository [R Package BiocManager Version 1.30.16]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=BiocManager">https://CRAN.R-project.org/package&#x3d;BiocManager</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Narayanasamy</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jarosz</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Muller</surname>
<given-names>E. E.</given-names>
</name>
<name>
<surname>Heintz-Buschart</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Herold</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kaysen</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>IMP: a Pipeline for Reproducible Reference-independent Integrated Metagenomic and Metatranscriptomic Analyses</article-title>. <source>Genome Biol.</source> <volume>17</volume> (<issue>1</issue>), <fpage>260</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-016-1116-8</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nicholson</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Holmes</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Kinross</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Burcelin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gibson</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Host-gut Microbiota Metabolic Interactions</article-title>. <source>Science</source> <volume>336</volume> (<issue>6086</issue>), <fpage>1262</fpage>&#x2013;<lpage>1267</lpage>. <pub-id pub-id-type="doi">10.1126/science.1223813</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nissen</surname>
<given-names>J. N.</given-names>
</name>
<name>
<surname>Johansen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Alles&#xf8;e</surname>
<given-names>R. L.</given-names>
</name>
<name>
<surname>S&#xf8;nderby</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Armenteros</surname>
<given-names>J. J. A.</given-names>
</name>
<name>
<surname>Gr&#xf8;nbech</surname>
<given-names>C. H.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Improved Metagenome Binning and Assembly Using Deep Variational Autoencoders</article-title>. <source>Nat. Biotechnol.</source> <volume>39</volume> (<issue>5</issue>), <fpage>555</fpage>&#x2013;<lpage>560</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-020-00777-4</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nurk</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Meleshko</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Korobeynikov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pevzner</surname>
<given-names>P. A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>metaSPAdes: a New Versatile Metagenomic Assembler</article-title>. <source>Genome Res.</source> <volume>27</volume> (<issue>5</issue>), <fpage>824</fpage>&#x2013;<lpage>834</lpage>. <pub-id pub-id-type="doi">10.1101/gr.213959.116</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Oksanen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Blanchet</surname>
<given-names>F. G.</given-names>
</name>
<name>
<surname>Friendly</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kindt</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Legendre</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>McGlinn</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>vegan: Community Ecology Package [R Package vegan Version 2.5-7]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=vegan">https://CRAN.R-project.org/package&#x3d;vegan</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Overholt</surname>
<given-names>W. A.</given-names>
</name>
<name>
<surname>H&#xf6;lzer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Geesink</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Diezel</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Marz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>K&#xfc;sel</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Inclusion of Oxford Nanopore Long Reads Improves All Microbial and Viral Metagenome-Assembled Genomes from a Complex Aquifer System</article-title>. <source>Environ. Microbiol.</source> <volume>22</volume> (<issue>9</issue>), <fpage>4000</fpage>&#x2013;<lpage>4013</lpage>. <pub-id pub-id-type="doi">10.1111/1462-2920.15186</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Parks</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Imelfort</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Skennerton</surname>
<given-names>C. T.</given-names>
</name>
<name>
<surname>Hugenholtz</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tyson</surname>
<given-names>G. W.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>CheckM: Assessing the Quality of Microbial Genomes Recovered from Isolates, Single Cells, and Metagenomes</article-title>. <source>Genome Res.</source> <volume>25</volume> (<issue>7</issue>), <fpage>1043</fpage>&#x2013;<lpage>1055</lpage>. <pub-id pub-id-type="doi">10.1101/gr.186072.114</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pasolli</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Asnicar</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Manara</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zolfo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Karcher</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Armanini</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Extensive Unexplored Human Microbiome Diversity Revealed by Over 150,000 Genomes from Metagenomes Spanning Age, Geography, and Lifestyle</article-title>. <source>Cell</source> <volume>176</volume> (<issue>3</issue>), <fpage>649</fpage>&#x2013;<lpage>e20</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.01.001</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qin</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Raes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Arumugam</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Burgdorf</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Manichanh</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>A Human Gut Microbial Gene Catalogue Established by Metagenomic Sequencing</article-title>. <source>Nature</source> <volume>464</volume> (<issue>7285</issue>), <fpage>59</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1038/nature08821</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Quince</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Walker</surname>
<given-names>A. W.</given-names>
</name>
<name>
<surname>Simpson</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Loman</surname>
<given-names>N. J.</given-names>
</name>
<name>
<surname>Segata</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Shotgun Metagenomics, from Sampling to Analysis</article-title>. <source>Nat. Biotechnol.</source> <volume>35</volume>, <fpage>833</fpage>&#x2013;<lpage>844</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3935</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Quinlan</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>Hall</surname>
<given-names>I. M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>BEDTools: a Flexible Suite of Utilities for Comparing Genomic Features</article-title>. <source>Bioinformatics</source> <volume>26</volume> (<issue>6</issue>), <fpage>841</fpage>&#x2013;<lpage>842</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq033</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saheb Kashaf</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Proctor</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Deming</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Saary</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>H&#xf6;lzer</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Integrating Cultivation and Metagenomics for a Multi-Kingdom View of Skin Microbiome Diversity and Functions</article-title>. <source>Nat. Microbiol.</source> <volume>7</volume> (<issue>1</issue>), <fpage>169</fpage>&#x2013;<lpage>179</lpage>. <pub-id pub-id-type="doi">10.1038/s41564-021-01011-w</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salazar</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Paoli</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Alberti</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Huerta-Cepas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ruscheweyh</surname>
<given-names>H. J.</given-names>
</name>
<name>
<surname>Cuenca</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Gene Expression Changes and Community Turnover Differentially Shape the Global Ocean Metatranscriptome</article-title>. <source>Cell</source> <volume>179</volume> (<issue>5</issue>), <fpage>1068</fpage>&#x2013;<lpage>e21</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.10.014</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Satinsky</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Crump</surname>
<given-names>B. C.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>C. B.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zielinski</surname>
<given-names>B. L.</given-names>
</name>
<name>
<surname>Doherty</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Microspatial Gene Expression Patterns in the Amazon River Plume</article-title>. <source>Proc. Natl. Acad. Sci. U S A.</source> <volume>111</volume> (<issue>30</issue>), <fpage>11085</fpage>&#x2013;<lpage>11090</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1402782111</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seemann</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Prokka: Rapid Prokaryotic Genome Annotation</article-title>. <source>Bioinformatics</source> <volume>30</volume> (<issue>14</issue>), <fpage>2068</fpage>&#x2013;<lpage>2069</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu153</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sequeira</surname>
<given-names>J. C.</given-names>
</name>
<name>
<surname>Rocha</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Madalena Alves</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salvador</surname>
<given-names>A. F.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>MOSCA: An Automated Pipeline for Integrated Metagenomics and Metatranscriptomics Data Analysis</article-title>,&#x201d; in <conf-name>Practical Applications of Computational Biology and Bioinformatics, 12th International Conference</conf-name>, <fpage>183</fpage>&#x2013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-98702-6_22</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Slowikowski</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Automatically Position Non-Overlapping Text Labels with &#x201c;ggplot2&#x201d; [R Package ggrepel Version 0.9.1]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=ggrepel">https://CRAN.R-project.org/package&#x3d;ggrepel</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stewart</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Auffret</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Warr</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Walker</surname>
<given-names>A. W.</given-names>
</name>
<name>
<surname>Roehe</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Watson</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Compendium of 4,941 Rumen Metagenome-Assembled Genomes for Rumen Microbiome Biology and Enzyme Discovery</article-title>. <source>Nat. Biotechnol.</source> <volume>37</volume> (<issue>8</issue>), <fpage>953</fpage>&#x2013;<lpage>961</lpage>. <pub-id pub-id-type="doi">10.1038/s41587-019-0202-3</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sunagawa</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mende</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Zeller</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Izquierdo-Carrasco</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Berger</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Kultima</surname>
<given-names>J. R.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Metagenomic Species Profiling Using Universal Phylogenetic Marker Genes</article-title>. <source>Nat. Methods</source> <volume>10</volume> (<issue>12</issue>), <fpage>1196</fpage>&#x2013;<lpage>1199</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2693</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tamames</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Puente-S&#xe1;nchez</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>SqueezeMeta, A Highly Portable, Fully Automatic Metagenomic Analysis Pipeline</article-title>. <source>Front. Microbiol.</source> <volume>9</volume>, <fpage>3349</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2018.03349</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tenenbaum</surname>
<given-names>D.</given-names>
</name>
</person-group>
<collab>Bioconductor Package Maintainer</collab> (<year>2017</year>). <article-title>KEGGREST: Client-Side REST Access to the Kyoto Encyclopedia of Genes and Genomes (KEGG)</article-title>. <comment>R package version 1.30.1</comment>. <pub-id pub-id-type="doi">10.18129/B9.bioc.KEGGREST</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="web">
<collab>The R Project for Statistical Computing</collab> (<year>2021</year>). <article-title>The R Project for Statistical Computing</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.R-project.org/">https://www.R-project.org/</ext-link>
</comment>(<comment>Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X. W.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Friedman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dahlin</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Deciphering Functional Redundancy in the Human Microbiome</article-title>. <source>Nat. Commun.</source> <volume>11</volume> (<issue>1</issue>), <fpage>6217</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-19940-1</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tl&#xe1;skal</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Metagenomes, Metatranscriptomes and Microbiomes of Naturally Decomposing deadwood</article-title>. <source>Scientific data</source> <volume>8</volume> (<issue>1</issue>), <fpage>198</fpage>. <pub-id pub-id-type="doi">10.6084/m9.figshare.14821752</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Van Damme</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>H&#xf6;lzer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Viehweger</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>M&#xfc;ller</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bongcam-Rudloff</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Brandt</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Metagenomics Workflow for Hybrid Assembly, Differential Coverage Binning, Metatranscriptomics and Pathway Analysis (MUFFIN)</article-title>. <source>Plos Comput. Biol.</source> <volume>17</volume> (<issue>2</issue>), <fpage>e1008716</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1008716</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Van Rossum</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Drake</surname>
<given-names>F. L.</given-names>
</name>
</person-group> (<year>2009</year>). <source>Python 3 Reference Manual: (Python Documentation Manual Part 2)</source>. <publisher-loc>Scotts Valley, CA</publisher-loc>: <publisher-name>CreateSpace</publisher-name>. </citation>
</ref>
<ref id="B66">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vasimuddin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Misra</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Aluru</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Efficient Architecture-Aware Acceleration of BWA-MEM for Multicore Systems</article-title>,&#x201d; in <conf-name>2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)</conf-name>. <comment>[Preprint]</comment>. <pub-id pub-id-type="doi">10.1109/ipdps.2019.00041</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wagner</surname>
<given-names>G. P.</given-names>
</name>
<name>
<surname>Kin</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lynch</surname>
<given-names>V. J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Measurement of mRNA Abundance Using RNA-Seq Data: RPKM Measure Is Inconsistent Among Samples</article-title>. <source>Theor. Biosci</source> <volume>131</volume> (<issue>4</issue>), <fpage>281</fpage>&#x2013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1007/s12064-012-0162-3</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wall</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Christiansen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Orwant</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2000</year>). <source>Programming Perl</source>. <publisher-loc>Sebastopol, CA</publisher-loc>: <publisher-name>O&#x2019;Reilly Media</publisher-name>. </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lv</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Integrated Metagenomic and Metatranscriptomic Profiling Reveals Differentially Expressed Resistomes in Human, Chicken, and Pig Gut Microbiomes</article-title>. <source>Environ. Int.</source> <volume>138</volume>, <fpage>105649</fpage>. <pub-id pub-id-type="doi">10.1016/j.envint.2020.105649</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Averick</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bryan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>McGowan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fran&#xe7;ois</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Welcome to the Tidyverse</article-title>. <source>Joss</source> <volume>4</volume> (<issue>43</issue>), <fpage>1686</fpage>. <pub-id pub-id-type="doi">10.21105/joss.01686</pub-id> </citation>
</ref>
<ref id="B72">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2016</year>). <source>ggplot2: Elegant Graphics for Data Analysis</source>. <publisher-loc>Springer-Verlag New York</publisher-loc>: <publisher-name>Springer</publisher-name>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://ggplot2.tidyverse.org">https://ggplot2.tidyverse.org</ext-link>
</comment>. </citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Reshaping Data with thereshapePackage</article-title>. <source>J. Stat. Soft.</source> <volume>21</volume> (<issue>12</issue>), <fpage>1</fpage>. <pub-id pub-id-type="doi">10.18637/jss.v021.i12</pub-id> </citation>
</ref>
<ref id="B1">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fran&#xe7;ois</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Henry</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>M&#xfc;ller</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>dplyr: A Grammar of Data Manipulation [R Package dplyr Version 1.0.7]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=dplyr">https://CRAN.R-project.org/package&#x3d;dplyr</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B77">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Girlich</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>tidyr: Tidy Messy Data [R Package tidyr Version 1.1.4]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=tidyr">https://CRAN.R-project.org/package&#x003D;tidyr</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B21">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wickham</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>haven: Import and Export &#x2018;SPSS&#x2019;, &#x2018;Stata&#x2019; and &#x2018;SAS&#x2019; Files [R Package haven Version 2.4.3]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=haven">https://CRAN.R-project.org/package&#x3d;haven</ext-link> (Accessed: December 6, 2021)</comment>. </citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Mao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>dbCAN: a Web Resource for Automated Carbohydrate-Active Enzyme Annotation</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>W445</fpage>&#x2013;<lpage>W451</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks479</pub-id> </citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Youngblut</surname>
<given-names>N. D.</given-names>
</name>
<name>
<surname>de la Cuesta-Zuluaga</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Reischer</surname>
<given-names>G. H.</given-names>
</name>
<name>
<surname>Dauser</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schuster</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Walzer</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Large-Scale Metagenome Assembly Reveals Novel Animal-Associated Microbial Genomes, Biosynthetic Gene Clusters, and Other Genetic Diversity</article-title>. <source>mSystems</source> <volume>5</volume> (<issue>6</issue>), <fpage>e01045-20</fpage>. <pub-id pub-id-type="doi">10.1128/mSystems.01045-20</pub-id> </citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yohe</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Entwistle</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>dbCAN2: a Meta Server for Automated Carbohydrate-Active Enzyme Annotation</article-title>. <source>Nucleic Acids Res.</source> <volume>46</volume> (<issue>W1</issue>), <fpage>W95</fpage>&#x2013;<lpage>W101</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky418</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>