<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2023.1247119</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Microbial dark matter sequences verification in amplicon sequencing and environmental metagenomics data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Barak</surname> <given-names>Hana</given-names></name><xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1011528/overview"/>
</contrib>
<contrib contrib-type="author"><name><surname>Fuchs</surname> <given-names>Naomi</given-names></name><xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author"><name><surname>Liddor-Naim</surname> <given-names>Michal</given-names></name><xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author"><name><surname>Nir</surname> <given-names>Irit</given-names></name><xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author"><name><surname>Sivan</surname> <given-names>Alex</given-names></name><xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes"><name><surname>Kushmaro</surname> <given-names>Ariel</given-names></name><xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="aff3" ref-type="aff"><sup>3</sup></xref>
<xref rid="aff4" ref-type="aff"><sup>4</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/93373/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Civil and Environmental Engineering, Ben-Gurion University of the Negev</institution>, <addr-line>Beer-Sheva</addr-line>, <country>Israel</country></aff>
<aff id="aff2"><sup>2</sup><institution>Avram and Stella Goldstein-Goren Department of Biotechnology Engineering, Ben-Gurion University of the Negev</institution>, <addr-line>Beer-Sheva</addr-line>, <country>Israel</country></aff>
<aff id="aff3"><sup>3</sup><institution>The Ilse Katz Center for Nanoscale Science and Technology, Ben-Gurion University of the Negev</institution>, <addr-line>Beer-Sheva</addr-line>, <country>Israel</country></aff>
<aff id="aff4"><sup>4</sup><institution>School of Sustainability and Climate Change, Ben-Gurion University of the Negev</institution>, <addr-line>Beer-Sheva</addr-line>, <country>Israel</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0002">
<p>Edited by: George Tsiamis, University of Patras, Greece</p>
</fn>
<fn fn-type="edited-by" id="fn0003">
<p>Reviewed by: Lucas Auer, Institut National de recherche pour l&#x2019;agriculture, l&#x2019;alimentation et l&#x2019;environnement (INRAE), France; Federica Chiappori, National Research Council (CNR), Italy; Daljeet Singh Dhanjal, Lovely Professional University, India</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Ariel Kushmaro, <email>arielkus@bgu.ac.il</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>11</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1247119</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Barak, Fuchs, Liddor-Naim, Nir, Sivan and Kushmaro.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Barak, Fuchs, Liddor-Naim, Nir, Sivan and Kushmaro</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Although microorganisms constitute the most diverse and abundant life form on Earth, in many environments, the vast majority of them remain uncultured. As it is based on information gleaned mainly from cultivated microorganisms, our current body of knowledge regarding microbial life is partial and does not reflect actual microbial diversity. That diversity is hidden in the uncultured microbial majority, termed by microbiologists as &#x201C;microbial dark matter&#x201D; (MDM), a term borrowed from astrophysics. Metagenomic sequencing analysis techniques (both 16S rRNA gene and shotgun sequencing) compare gene sequences to reference databases, each of which represents only a small fraction of the existing microorganisms. Unaligned sequences lead to groups of &#x201C;unknown microorganisms&#x201D; that are usually ignored and rarefied from diversity analysis. To address this knowledge gap, we analyzed the 16S rRNA gene sequences of microbial communities from four different environments&#x2014;a living organism, a desert environment, a natural aquatic environment, and a membrane bioreactor for wastewater treatment. From those datasets, we chose representative sequences of potentially unknown bacteria for additional examination as &#x201C;microbial dark matter sequences&#x201D; (MDMS). Sequence existence was validated by specific amplification and re-sequencing. These sequences were screened against databases and aligned to the Genome Taxonomy Database to build a comprehensive phylogenetic tree for additional sequence classification, revealing potentially new candidate phyla and other lineages. These putative MDMS were also screened against metagenome-assembled genomes from the explored environments for additional validation and for taxonomic and metabolic characterizations. This study shows the immense importance of MDMS in environmental metataxonomic analyses of 16S rRNA gene sequences and provides a simple and readily available methodology for the examination of MDM hidden behind amplicon sequencing results.</p>
</abstract>
<kwd-group>
<kwd>metagenomics</kwd>
<kwd>microbial dark matter</kwd>
<kwd>microbial community</kwd>
<kwd>amplicon sequencing</kwd>
<kwd>bacteria</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="63"/>
<page-count count="12"/>
<word-count count="8018"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Systems Microbiology</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1"><label>1.</label>
<title>Introduction</title>
<p>The most diverse and abundant life form on planet Earth, microorganisms play a fundamental role in the planet&#x2019;s ecosystem health by cycling nutrients, degrading environmental pollutants, facilitating primary production, and providing essential nutrients and chemicals such as oxygen and different vitamins that humans and animals cannot produce themselves (<xref ref-type="bibr" rid="ref32">Morowitz et al., 2011</xref>; <xref ref-type="bibr" rid="ref45">Rinke et al., 2013</xref>; <xref ref-type="bibr" rid="ref50">Solden et al., 2016</xref>). The conventional methods of studying these microorganisms and to elucidate their capabilities have, in the past, relied on already well-developed, classical laboratory techniques, in particular the use of cultivation methods. Nonetheless, in many environments only limited numbers of microorganisms have been cultivated to date (<xref ref-type="bibr" rid="ref50">Solden et al., 2016</xref>; <xref ref-type="bibr" rid="ref60">Zamkovaya et al., 2021</xref>). The famous &#x201C;great plate count anomaly&#x201D; is one of the earliest depictions of the gap between the actual number of bacteria present in a given sample and the much smaller number that can be effectively cultivated (<xref ref-type="bibr" rid="ref51">Staley and Konopka, 1985</xref>). The extent of microorganism diversity was further elucidated by analyzing microbial ribosomal RNA (rRNA) gene sequences directly collected from environmental samples (<xref ref-type="bibr" rid="ref3">Baker and Dick, 2013</xref>). During the last few decades, the 16S rRNA gene has emerged as the most sequenced taxonomic marker (<xref ref-type="bibr" rid="ref53">Tringe and Hugenholtz, 2008</xref>), forming a cornerstone for systematic classification that is also exploited as a genetic marker to infer the phylogenetic relationships among prokaryotes.</p>
<p>The use of metabarcoding based on short variable region sequencing of the 16S rRNA gene has revolutionized microbial ecology, allowing for the rapid and high-throughput identification of complex microbial communities (<xref ref-type="bibr" rid="ref46">Santos et al., 2020</xref>). However, due to the short amplicon lengths used in this analysis, this approach has limitations in the extent to which it can accurately affiliate microbial taxa to species or even genus levels, a resolution that is insufficient for differentiating closely related taxa. In addition, this method is prone to PCR amplification biases, sequencing errors, and variations in the copy number of the 16S rRNA gene across different taxa. To address these limitations, recent strategies have been developed that enable nearly full-length sequencing of the 16S rRNA gene, improving the accuracy of microbial identification and facilitating the discovery of novel taxa. Included among these approaches are long-read sequencing technologies such as PacBio and Oxford Nanopore and hybrid sequencing approaches that combine short-read and long-read sequencing technologies. These methods provide higher resolution and more accurate taxonomic classification, thereby increasing the reliability of microbial identification in various research fields. Nevertheless, Illumina short variable region sequencing is still the standard sequencing technology and the most frequently used method in microbial ecology studies. The importance of 16S rRNA gene sequences to the field notwithstanding, an exclusive reliance on this analytical method may fail to provide complete information about bacterial classification. According to <xref ref-type="bibr" rid="ref57">Yarza et al. (2014)</xref>, a sequence identity of 94.5% or lower for two 16S rRNA genes provides strong evidence that they belong to distinct genera, while lower sequence identities of 86.5% correspond to families, 82% to orders, 78.5% to classes and 75% to phyla. Analyses of the 16S rRNA gene from environmental samples revealed that fewer than half of the known microbial phyla are represented by at least one cultivated representative. Moreover, among all microbial isolates, more than 88% belong to only four bacterial phyla (from among the more than 1,500 estimated phyla): Proteobacteria, Firmicutes, Actinobacteria and Bacteroidetes (<xref ref-type="bibr" rid="ref45">Rinke et al., 2013</xref>; <xref ref-type="bibr" rid="ref50">Solden et al., 2016</xref>). To date, the phyla that contain only uncultured representatives, identified via the phylogenetic analysis of rRNA genes recovered from environmental samples, have been referred to as candidate phyla. Lacking the support of bacterial culture results, rRNA based sequence analysis alone is unable to classify the majority of the microbial population. Microbiologists have therefore compared the problem of this &#x201C;uncultured microbial majority&#x201D; to that of &#x201C;dark matter&#x201D; in astrophysics, adopting similar terms such as &#x201C;microbial dark matter&#x201D; (MDM) to describe the uncultivated microbes (<xref ref-type="bibr" rid="ref15">Hedlund et al., 2014</xref>; <xref ref-type="bibr" rid="ref20">Jiao et al., 2021</xref>). Among the MDM, one prominent group of candidate phyla radiation (CPR) is known by the super-phylum name Patescibacteria (<xref ref-type="bibr" rid="ref14">Harris et al., 2004</xref>; <xref ref-type="bibr" rid="ref33">Nakai, 2020</xref>).</p>
<p>Genomic analyses of CPR representatives showed that metabolic limitations have prevented our ability to cultivate these organisms, which are typically smaller than cultivated bacteria (&#x223C;0.2 microns) (<xref ref-type="bibr" rid="ref54">Vigneron et al., 2020</xref>) and with shorter genomes (&#x223C;1&#x2009;Mbp). Moreover, they often have unusual ribosome compositions that contain self-splicing introns and proteins encoded within their rRNA genes, a feature rarely reported in bacteria (<xref ref-type="bibr" rid="ref6">Brown et al., 2015</xref>). Many are thought to be unable to produce their own nucleotides and are believed to possess minimal amino acid contents and limited cofactor biosynthetic capacity. Indeed, analyses of their genomes showed that they lack CRISPR (<xref ref-type="bibr" rid="ref52">Tian et al., 2020</xref>) and the components necessary to synthesize membrane lipids (<xref ref-type="bibr" rid="ref8">Castelle and Banfield, 2018</xref>). Nevertheless, their genomes have been recovered from diverse environments ranging from the human microbiome to drinking water to marine and deep subsurface sediments and soil (<xref ref-type="bibr" rid="ref30">M&#x00E9;heust et al., 2019</xref>). A recent phylogenetic study found that protein family presence/absence patterns cluster the Patescibacteria super-phyla together and separate from all other bacteria and archaea.</p>
<p>Debate over the extent of the MDM diversity has led to estimates that it could account for as much as 25&#x2013;50% of all bacterial diversity (<xref ref-type="bibr" rid="ref17">Hug et al., 2016</xref>; <xref ref-type="bibr" rid="ref40">Parks et al., 2017</xref>; <xref ref-type="bibr" rid="ref48">Schulz et al., 2017</xref>). The inability to definitively determine its contribution to diversity may be because some of its groups are not detected in 16S rRNA gene taxonomic and diversity surveys due to primer mismatch and/or to the presence of introns within their 16S rRNA gene that may interfere with polymerase chain reaction (PCR) amplification (<xref ref-type="bibr" rid="ref8">Castelle and Banfield, 2018</xref>). There is accumulating evidence that these uncultivated microorganisms account for a larger portion of the Earth&#x2019;s biomass and biodiversity than was previously thought, reflecting the profound bias of the current body of knowledge about microbial life.</p>
<p>In metagenomic sequencing analysis (both 16S rRNA gene and shotgun sequencing), sequences are compared to reference databases that contain only a small part of the existing microorganisms. This results in uncovering of groups of yet unclassified microorganisms. Despite the increasing awareness of their immense importance, these unclassified amplicon sequences, designated by us as &#x201C;microbial dark matter sequences&#x201D; (MDMS), are usually ignored or discarded during typical microbial community profiling studies.</p>
<p>The aim of this study, therefore, was to provide additional support for the immense importance of MDMS in environmental metataxonomic analyses using the 16S rRNA gene. To that end, we analyzed 16S rRNA gene sequences collected from four different, highly diverse environments&#x2014;a living organism, rocks from a desert environment, natural aquatic environments and a membrane bioreactor for wastewater treatment. Our ongoing studies of the varied microbiomes of these environments availed us of the necessary samples from each environment. Of the sequences collected, 163 16S rRNA representative gene sequences, obtained from amplicon sequencing, were chosen for additional examination as potential MDMS. These sequences were screened against various databases and aligned to the GTDB (Genome Taxonomy Database) to build a comprehensive phylogenetic tree for additional sequence classifications. The putative MDMS were screened against metagenome-assembled genomes from the explored environments for additional validation and for taxonomic and metabolic capacity characterization. Using a relatively simple, currently available methodology, this study sheds additional light on MDMS that will improve our conceptualization of the bacterial diversity in any environment.</p>
</sec>
<sec sec-type="materials|methods" id="sec2"><label>2.</label>
<title>Materials and methods</title>
<sec id="sec3"><label>2.1.</label>
<title>Total genomic DNA extraction</title>
<p>For the purposes of this study, we used total gDNA obtained from four vastly different environments:<list list-type="bullet">
<list-item>
<p>A membrane bioreactor (MBR) used to treat chemical industry wastewater; system description and DNA extractions described in <xref ref-type="bibr" rid="ref4">Barak et al. (2020)</xref>.</p>
</list-item>
<list-item>
<p>Larvae of the beetle <italic>Capnodis tenebrionis</italic> (CT); experiment described in <xref ref-type="bibr" rid="ref5">Barak et al. (2019)</xref>.</p>
</list-item>
<list-item>
<p>Surfaces of Negev desert rocks (NDR)&#x2014;12 rock samples from two petroglyph sites in the Negev desert of Israel from the Ramat Matred and Har Michya sites; experiment and DNA extraction procedure described in <xref ref-type="bibr" rid="ref19">Irit et al. (2019)</xref> and <xref ref-type="bibr" rid="ref36">Nir et al. (2019)</xref>.</p>
</list-item>
<list-item>
<p>Confined and unconfined aquifers&#x2014;five biomass samples scraped from different coupons made of glass, steel and stainless steel that had been deployed in water wells in the Arava Valley.</p>
</list-item>
<list-item>
<p>In addition, 20 biomass samples were obtained By sterile filtering 50&#x2009;L of water from The wells In The Arava Valley using The Stericup-GP sterile vacuum filtration system containing a polyethersulfone membrane with a pore size of 0.22&#x2009;&#x03BC;m (Merck, Gillingham, United Kingdom). Extraction of total genomic DNA from The biomass samples Was carried Out using The MoBio PowerWater isolation Kit (MoBio laboratories Inc. Carlsbad, CA, United States) and The DNeasy PowerSoil Kit (Qiagen, United States).</p>
</list-item>
</list></p>
</sec>
<sec id="sec4"><label>2.2.</label>
<title>Next generation amplicon sequencing</title>
<p>
<list list-type="simple">
<list-item>
<p>The total genomic DNA that was extracted from the samples was submitted to the DNA Services facility (DNAS) of the Research Resources Center at the University of Illinois Chicago (UIC) for gene sequencing of the bacterial small subunit (16S) of ribosomal RNA (rRNA) using the Illumina MiSeq platform with a sequencing length of 300&#x2009;bps. Prior to sequencing, two PCR amplification steps were performed. During the first PCR reaction, fragments of the V3&#x2013;V4 (environments 1&#x2013;3) and V1&#x2013;V3 (aquifers) regions of the 16S rRNA gene were amplified using universal primers (341F/806R and 27F/534R, respectively) (<xref ref-type="bibr" rid="ref21">Jumpstart Consortium Human Microbiome Project Data Generation Working Group, 2012</xref>; <xref ref-type="bibr" rid="ref18">Hugerth et al., 2014</xref>; <xref ref-type="bibr" rid="ref10">Elovitz et al., 2019</xref>) to which were attached the 5&#x2032; linker sequences CS1 and CS2 (known as common sequence 1 and 2). The second PCR reaction was done to prepare the library as described by <xref ref-type="bibr" rid="ref13">Green et al. (2015)</xref>.</p>
</list-item>
</list>
</p>
</sec>
<sec id="sec5"><label>2.3.</label>
<title>Metataxonomic data analysis</title>
<p>
<list list-type="simple">
<list-item>
<p>Raw reads were merged using the PEAR software package (v0.9.10) (<xref ref-type="bibr" rid="ref61">Zhang et al., 2014</xref>), with a quality score threshold of 25 for trimming and a base PHRED quality score of 33. Sequence data were screened to remove low-quality sequences and potentially chimeric sequences with the Mothur software package (v1.36.1) (<xref ref-type="bibr" rid="ref47">Schloss et al., 2009</xref>). Sequences that contained more than eight bases homopolymers or any ambiguous bases were removed, and a length cutoff of 250&#x2009;bp was used. The resultant sequences file was screened against the phix 174 genome (ID&#x2014;MN385565) using BLASTN (<xref ref-type="bibr" rid="ref9">Chen et al., 2015</xref>) to remove sequencing/processing artifacts. The quality-controlled sequences were then processed with the Qiime software package (v1.9.1) (<xref ref-type="bibr" rid="ref7">Caporaso et al., 2010</xref>). Briefly, sequence data were clustered into operational taxonomic units (OTU) at 97% similarity. Representative sequences from each OTU were extracted and classified using the &#x201C;assign_taxonomy.py&#x201D; script with the UCLUST assignment method, utilizing the SILVA database (<xref ref-type="bibr" rid="ref44">Quast et al., 2012</xref>).</p>
</list-item>
<list-item>
<p>Representative sequences were also aligned using the &#x201C;align_seqs.py&#x201D; script with percent identity thresholds of 75 and 90% to the Silva alignment reference file (<xref ref-type="bibr" rid="ref44">Quast et al., 2012</xref>). The aligned sequences were filtered using the silva_lanemask_mothur file and then used to produce a phylogenetic tree. Four biological observation matrices (BIOM) (<xref ref-type="bibr" rid="ref29">McDonald et al., 2012</xref>) were generated at taxonomic levels from phylum to genus using the &#x201C;make_OTU_table.py&#x201D; script. Sequences that failed to align with the Silva DB for the above-mentioned thresholds were not included in the BIOM tables. An additional BIOM table was also generated in which no alignment-based sequence filter was applied. The &#x201C;filter_otus_from_otu_table.py&#x201D; script ensured that only OTUs with minimum total observation counts of 50 reads were retained. All data analysis was done using the Silva database (v.138) as a reference. BIOM tables were converted from read counts to relative abundances and the relative abundances of the unassigned OTUs from each dataset were plotted to present the differences between 75 and 90% alignment thresholds.</p>
</list-item>
<list-item>
<p>Beta diversity (pairwise sample dissimilarity) was calculated using Bray-Curtis, and a 2D nMDS plot was generated using R.</p>
</list-item>
<list-item>
<p>The OTU table (based on all representative sequences, without eliminating alignment failures) was converted from read numbers to relative abundance values, and OTUs that were not assigned to any known lineage (not even at the phylum level) and that had relative abundance summaries higher than 0.5% were chosen for further observation as putative MDMS (<xref ref-type="supplementary-material" rid="SM7">Supplementary Table S6</xref> presents a summarized overview of taxa at the phylum level, derived from the biome table).</p>
</list-item>
</list>
</p>
</sec>
<sec id="sec6"><label>2.4.</label>
<title>Taxonomic analysis of putative MDMS</title>
<p>For a more comprehensive taxonomic classification, the 163 putative MDMS were compared to four different databases using BLASTN (<xref ref-type="bibr" rid="ref2">Altschul et al., 1990</xref>): the Silva database (v.138) (<xref ref-type="bibr" rid="ref44">Quast et al., 2012</xref>), EzBioCloud&#x2019;s 16S database (updated in May 2018) (<xref ref-type="bibr" rid="ref59">Yoon et al., 2017</xref>), the GTDB (r89) (<xref ref-type="bibr" rid="ref39">Parks et al., 2018</xref>) and the nucleotide collection database (nt) of the NCBI last accessed in February 2020 (<xref ref-type="bibr" rid="ref1001">NCBI Resource Coordinators, 2013</xref>). Manual observation of the similarity percentage and query cover of the obtained hits for each putative MDMS provide a more accurate taxonomic classification based on similarity percentage as described in <xref ref-type="bibr" rid="ref57">Yarza et al. (2014)</xref>.</p>
</sec>
<sec id="sec7"><label>2.5.</label>
<title>Phylogenetic analysis of putative MDMS</title>
<p>To generate a phylogenetic tree that integrates our putative MDMS with the known bacteria, we used the SSU rRNA sequences with lengths of 600&#x2013;2,000 bases from the GTDB repository (bac120 ssu r89). First, the GTDB SSU rRNA sequences were aligned using the SSU-ALIGN v.0.1 software (<xref ref-type="bibr" rid="ref34">Nawrocki, 2009</xref>). The aligned sequences were then masked based on posterior probability (PP) annotation at the default value of 0.95 for aligned residues and as a value of 0.70 for the gap threshold based on the frequency of gap characters in each column. Numerous candidates of the CPR super-phylum known to encode insertions were clustered in several locations of these MDMS 16S rRNA genes. The SSU-ALIGN algorithm that was used in the secondary structure-and function-based multiple sequence alignment (MSA) analysis only included parts of the gene that lacked the insertions.</p>
<p>The putative MDMS were added to the GTDB MSA using the MAFFT v7.464 software (with the Addfragments option) (<xref ref-type="bibr" rid="ref23">Katoh and Standley, 2013</xref>). The full phylogenetic tree was generated based on the merged alignment using FastTree_v2.1.10 (<xref ref-type="bibr" rid="ref42">Price et al., 2010</xref>). Visualization was carried out using the Interactive Tree of Life (iTOL) online interface (<xref ref-type="bibr" rid="ref26">Letunic and Bork, 2016</xref>).</p>
</sec>
<sec id="sec8"><label>2.6.</label>
<title>MDMS existence validation</title>
<p>Specific primers were designed for about 30 MDMS using Primer-BLAST (<xref ref-type="bibr" rid="ref58">Ye et al., 2012</xref>). Primers suggested by Primer-BLAST were examined through the Amplifx software for GC content, self-dimer, Tm and annealing to the target sequence. Primers were synthesized by SIGMA-ALDRICH Co., LLC (Rehovot, Israel). The primers were attached to the 5&#x2032; linker sequences CS1 and CS2 and the samples originated each MDMS of interest were sent for sequencing using Illumina MiSeq platform by the DNA Services facility (DNAS) of the Research Resources Center at the University of Illinois Chicago (UIC). The obtained sequencing data was analyzed as described in the &#x201C;Metataxonomic data analysis&#x201D; section previously. If the amplification was not specific, it was ignored. If it did provide specific OTU, the representative OTU sequence was compared to the original MDMS sequence using blast. Only sequences with high levels of similarity (&#x003E;95%) and 100% query cover are shown.</p>
<p>Furthermore, targeted chimera check was conducted for all MDMS, utilizing the DECIPHER web tool (v2.27.2) (<xref ref-type="bibr" rid="ref12">Firth et al., 2009</xref>).</p>
</sec>
<sec id="sec9"><label>2.7.</label>
<title>Metagenomic analysis, putative MDMS screening, and genome characterization</title>
<p>Genomic DNA from 17 representative samples from the two environments with abundances of MDMS (NDR and aquifers) were sequenced by the Illumina NextSeq500 platform in the DNA Services (DNAS) Facility of the Research Resources Center at the University of Illinois at Chicago (UIC).</p>
<p>Metagenomic data were processed by the metaWRAP pipeline v1.2.1. Raw reads were subjected to quality control (QC) using TrimGalore v0.5.0 (<xref ref-type="bibr" rid="ref25">Krueger, 2012</xref>) and low-quality reads were removed. The QC-passed sequences were assembled using MetaSPAdes v3.13.0 (<xref ref-type="bibr" rid="ref37">Nurk et al., 2017</xref>) (or MEGAHIT v1.1.3 (<xref ref-type="bibr" rid="ref27">Li et al., 2015</xref>) in Ramat-Matred samples due to memory limitation). The assemblies and the QC-passed sequences were used for metagenomic binning using three different algorithms: MaxBin v2.2.6 (<xref ref-type="bibr" rid="ref22">Kang et al., 2015</xref>), metaBAT v2.12.1 (<xref ref-type="bibr" rid="ref22">Kang et al., 2015</xref>), and CONCOCT v1.0.0 (<xref ref-type="bibr" rid="ref1">Alneberg et al., 2013</xref>). The resulting three bin sets were consolidated to obtain a single, strong bin set with a minimum completion of 50% and a maximum contamination of 10%. The consolidated bin set was reassembled using both &#x201C;strict&#x201D; and &#x201C;permissive&#x201D; algorithms, and once the reassembled bin had been improved, it replaced the original bin.</p>
<p>The chosen 163 putative MDMS were screened against both the assembly results and the final bins using BLASTN. The results were examined manually based on the percent similarity (&#x003E;96%) and cover and on the overlap locations.</p>
<p>The consolidated matched MDMS bins were functionally annotated using Prokka v1.13 (<xref ref-type="bibr" rid="ref49">Seemann, 2014</xref>) with metaWRAP&#x2019;s Annotate_bins module. Additional metabolic and biogeochemical functional trait profiling was carried out using the METABOLIC profiler software (<xref ref-type="bibr" rid="ref62">Zhou et al., 2019</xref>) with METABOLIC-C.pl. version 4.0.</p>
<p>See <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S2</xref> for an outline of the methodology pipeline.</p>
</sec>
</sec>
<sec sec-type="results-and-discussion" id="sec10"><label>3.</label>
<title>Results and discussion</title>
<p>Microbial community analysis based on 16S rRNA gene amplicon sequencing is a widespread and important technique in microbiological research (<xref ref-type="bibr" rid="ref43">Prodan et al., 2020</xref>) that allows researchers to characterize the environment and to determine which microorganisms, both cultured and uncultured, are present in an environmental sample. General analyses of the 16S rRNA gene sequences should compare them to relevant databases. Based mainly on laboratory-cultured bacteria, however, these databases (and indeed, most of our knowledge of microorganisms) are relatively limited in scope, thus rendering the resulting notion of the tree of life unable to present a comprehensive picture of the microbial world. Shedding light on the &#x201C;dark matter&#x201D; inhabiting the tree of life may therefore improve our understanding of explored environments and contribute to reshaping the microbial world&#x2019;s taxonomy.</p>
<p>Today&#x2019;s whole genome shotgun sequencing studies, especially those focused on single-cell sequencing, constitute the leading methods used to explore uncultured microorganisms and expand our knowledge of the microbial world (<xref ref-type="bibr" rid="ref20">Jiao et al., 2021</xref>; <xref ref-type="bibr" rid="ref55">Wiegand et al., 2021</xref>). Indeed, this technique has illuminated the understudied &#x201C;microbial dark matter&#x201D; (MDM), thereby helping to fill the gaps in the growing tree of life and eventually explain those microorganisms&#x2019; roles in the environment. To date, however, phylogenetic studies rely mostly on 16S rRNA sequences and metagenomic shotgun sequencing.</p>
<p>The objective of this study is to fortify our ability to discover the hidden potential of the &#x201C;microbial dark matter&#x201D; by using 16S rRNA amplicon sequencing. To achieve this, we performed bioinformatic analyses of 16S rRNA gene sequences obtained from four very different environments representing diverse conditions: (1) A contaminated industrial environment, i.e., a membrane bioreactor used to treat chemical industry wastewater; (2) <italic>Capnodis tenebrionis</italic> as a living habitat; and two natural desert environments, (3) desert rocks with petroglyphs, and (4) water wells (confined and unconfined aquifers) in the Arava Valley. These environments, varied habitats that have yet to be rigorously explored, demonstrate their potential as sources for the discovery of new, unculturable bacteria.</p>
<p>As expected, non-metric multidimensional scaling (nMDS) analysis (<xref rid="fig1" ref-type="fig">Figure 1</xref>) showed high variance between the datasets (confined and unconfined aquifers treated as two separate groups). Anosim and Permanova tests supported this result with a <italic>p</italic>-value of 0.0001 and test statistics of 0.998 and 11.394, respectively.</p>
<fig position="float" id="fig1"><label>Figure 1</label>
<caption>
<p>Non-metric multidimensional scaling (nMDS) based on Bray-Curtis, with normal data ellipses (stress level: 0.09).</p>
</caption>
<graphic xlink:href="fmicb-14-1247119-g001.tif"/>
</fig>
<p>Using a set of bioinformatic filters, we generated a total of 5,174,233 high-quality reads obtained from 61 samples. These reads originated from an initial dataset comprising approximately 14 million raw reads. Among the 5,558 representative OTUs with a minimum of 50 repeated observations, 529 OTUs (~9.5%) were not assigned to any known lineage. We found a major difference in the number of unassigned OTUs when data were rarified based on 75 and 90% identity thresholds for alignment (<xref rid="fig2" ref-type="fig">Figure 2</xref>), a finding which may indicate that the &#x201C;dark&#x201D; part of the microbial environment is located in the gap between the 75 and 90% similarity thresholds. These cutoffs (75 and 90%) were chosen based on the recommended minimum percent similarities to include a sequence in an alignment and to consider a database match a hit, respectively.<xref rid="fn0001" ref-type="fn"><sup>1</sup></xref> Interestingly, natural aquifer water and desert rocks contained higher number of unassigned OTUs in both relative abundance and absolute numbers compared to the engineered environment of the wastewater treatment system. Indeed, according to previous works, unclassified sequences are commonly found in less studied natural environments such as natural water habitats (<xref ref-type="bibr" rid="ref24">Keshri et al., 2015</xref>; <xref ref-type="bibr" rid="ref38">Panda et al., 2017</xref>) and semiarid endoliths (<xref ref-type="bibr" rid="ref17">Hug et al., 2016</xref>). Since aquifer samples were sequenced for the variable regions V1-V3 and all other samples were sequenced for V3&#x2013;V4, it could also explain part of the differences in the portion of unclassified sequences between the different environments.</p>
<fig position="float" id="fig2"><label>Figure 2</label>
<caption>
<p>Dots represent the relative abundance of each unassigned OTU for the studied environments: Aquifers (confined and unconfined), CT, Capnodis Tenebrionis; MBR, industrial wastewater (membrane bioreactor); NDR, Negev desert rocks. The identity thresholds are 75% (left) and 90% (right).</p>
</caption>
<graphic xlink:href="fmicb-14-1247119-g002.tif"/>
</fig>
<p>The relative abundances of the unassigned OTUs ranged from minor to as high as 40% of the reads obtained from a confined aquifer sample. Indeed, our results together with those of recent works (<xref ref-type="bibr" rid="ref60">Zamkovaya et al., 2021</xref>) demonstrate that &#x201C;microbial dark matter&#x201D; are key ecological players within their respective communities. While <xref ref-type="bibr" rid="ref28">Lynch et al. (2012)</xref> emphasize the importance of novel phylogenetic diversity in what has been dubbed the &#x201C;rare biosphere,&#x201D; wherein they examine low relative abundance sequences, the present study focuses on the highly abundant but uncharacterized sequences. Rare biosphere sequences are liable to be missed by metagenomic sequencing due to the lack of a PCR amplification step (<xref ref-type="bibr" rid="ref41">Pascoal et al., 2021</xref>).</p>
<p>Based on their relative abundances, 163 of the unassigned sequences were chosen to represent putative MDMS, and these were screened against four different updated databases: Silva, EZ, NCBI, and GTDB. The best match for each MDMS after manual observation is presented in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>. To enable assumptions about their taxonomic attributions, the putative MDMS were also aligned to the GTDB to build a phylogenetic tree (<xref rid="fig3" ref-type="fig">Figure 3</xref>) that was pruned into four smaller trees (<xref rid="fig4" ref-type="fig">Figure 4</xref>) to facilitate a more comprehensive perspective of MDMS distribution across the tree of life. A substantial number of the MDMS (40 out of 163) were found to be part of the Patescibacteria super-phylum (<xref rid="fig4" ref-type="fig">Figure 4A</xref>). Indeed, it is reasonable that a relatively large portion of the MDMS belongs to the Patescibacteria super-phylum, since they are largely uncultured and therefore understudied. Interestingly, it is still not known whether the distinct phylogenetic position of Patescibacteria in the tree of life is due to rapid evolution by genome reduction or to its early evolutionary split from the non-Patescibacteria (<xref ref-type="bibr" rid="ref30">M&#x00E9;heust et al., 2019</xref>; <xref ref-type="bibr" rid="ref55">Wiegand et al., 2021</xref>).</p>
<fig position="float" id="fig3"><label>Figure 3</label>
<caption>
<p>Phylogenetic tree. One hundred and sixty three representatives of the unassigned group are marked with black dots and integrated within the bacteria sequences of the GTDB (bac120_ssu_r89.fna). Branches, strips, and labels are uniquely colored according to phyla.</p>
</caption>
<graphic xlink:href="fmicb-14-1247119-g003.tif"/>
</fig>
<fig position="float" id="fig4"><label>Figure 4</label>
<caption>
<p>Pruned phylogenetic trees of the <bold>(A)</bold> Patescibacteria super-phylum [Candidate phyla radiation (CPR)], combines 40 representative unassigned OTUs (bold); <bold>(B)</bold> Elusimicrobiota; <bold>(C)</bold> Nitrospirota; and <bold>(D)</bold> Planctomycetota. The tree is pruned from the phylogenetic tree in <xref rid="fig3" ref-type="fig">Figure 3</xref>. Branches, strips, and labels are uniquely colored according to phyla.</p>
</caption>
<graphic xlink:href="fmicb-14-1247119-g004.tif"/>
</fig>
<p>Eight MDMS were related to the Elusimicrobiota and four were related to the Planctomycetota phylum. A group of 53 MDMS, all obtained from aquifer samples, was situated near the Nitrospirota phylum. A tree of putative MDMS from aquifers (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S1A</xref>) suggests that the members of this group do not necessarily belong together. Comparisons of their BLAST results with the GTDB also yielded similarities of 75&#x2013;85% to different phyla such as Bacteroidota, Methylomirabilota, Desulfuromonadota, Actinobacteriota, Planctomycetota, etc. Nitrospirota have been shown to consistently coexist with Patescibacteria, after which they are the most common phylum in the groundwater population (<xref ref-type="bibr" rid="ref16">Herrmann et al., 2019</xref>; <xref ref-type="bibr" rid="ref56">Yan et al., 2021</xref>). Nevertheless, it seems that in our case, not all of the 53 MDMS are part of the Nitrospirota phylum, which may be due to their misclassification.</p>
<p>In the general phylogenetic tree (<xref rid="fig3" ref-type="fig">Figure 3</xref>), MDMS were also integrated within different phyla, including the Gammaproteobacteria, Firmicutes, Bacteroidota, Cyanobacteria, etc. We also validated the existence of the putative 16S rRNA MDMS by specific PCR amplification and Mi-Seq Illumina re-sequencing using specific self-designed primers for a few representative MDMS (<xref rid="tab1" ref-type="table">Table 1</xref>). Comparisons of the re-sequenced fragments to the original putative MDMS yielded similarity percentages of 95.91&#x2013;100%, indicating appropriate primer design and the existence of these sequences in our sequencing data. In the present work, each MDMS is a representative sequence of a group of similar sequences (97% similarity) that constitute an OTU. Previous works found that distinct taxa may be found within a single OTU (<xref ref-type="bibr" rid="ref35">Needham et al., 2017</xref>). Therefore, when validating the putative MDMS against resequencing results, we treat similarity percentages higher than 94.5% as relevant because they may indicate sequences of the same genus (<xref ref-type="bibr" rid="ref57">Yarza et al., 2014</xref>).</p>
<table-wrap position="float" id="tab1"><label>Table 1</label>
<caption>
<p>Six representative sequences used for validation using re-sequencing by MiSeq Illumina with self-designed specific primers.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Seq ID</th>
<th align="center" valign="top">Original seq length</th>
<th align="center" valign="top">Amplified seq length</th>
<th align="center" valign="top">% similarity</th>
<th align="left" valign="top">F-primer</th>
<th align="left" valign="top">R-primer</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">R001</td>
<td align="center" valign="top">449</td>
<td align="center" valign="top">225</td>
<td align="center" valign="top">100</td>
<td align="left" valign="top">CGTAGGCGGTTTCTTAAGTTTTGA</td>
<td align="left" valign="top">ACTCGGGTTTCTAATCCTCTTCG</td>
</tr>
<tr>
<td align="left" valign="top">R003</td>
<td align="center" valign="top">449</td>
<td align="center" valign="top">271</td>
<td align="center" valign="top">98.15</td>
<td align="left" valign="top">AAAGCCTGATCCAGCCACAT</td>
<td align="left" valign="top">ACTCTCCTCTCCCTTCCTCT</td>
</tr>
<tr>
<td align="left" valign="top">A016</td>
<td align="center" valign="top">495</td>
<td align="center" valign="top">439</td>
<td align="center" valign="top">95.91</td>
<td align="left" valign="top">TCAGGGTGAACGCTGGTAAC</td>
<td align="left" valign="top">TCCACCGGTACAGTCAACCT</td>
</tr>
<tr>
<td align="left" valign="top">A054</td>
<td align="center" valign="top">469</td>
<td align="center" valign="top">392</td>
<td align="center" valign="top">100</td>
<td align="left" valign="top">GCAAGTCAAACCCCGCTTAT</td>
<td align="left" valign="top">CCGGTGCTATTTGCAGGAGT</td>
</tr>
<tr>
<td align="left" valign="top">A073</td>
<td align="center" valign="top">521</td>
<td align="center" valign="top">318</td>
<td align="center" valign="top">97.17</td>
<td align="left" valign="top">ACCGGATAGGATGGCTCTCT</td>
<td align="left" valign="top">CGTCAGGTACCGTCATACCAG</td>
</tr>
<tr>
<td align="left" valign="top">A080</td>
<td align="center" valign="top">498</td>
<td align="center" valign="top">458</td>
<td align="center" valign="top">100</td>
<td align="left" valign="top">GGCTCAGAATGAACGCTTGAAA</td>
<td align="left" valign="top">GCCAGGGCTTCTTCTTTAGGT</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Only sequences with high levels of similarity (&#x003E;95%) and 100% query cover are shown.</p>
</table-wrap-foot>
</table-wrap>
<p>The MDMS were compared against the draft genomes that were generated from the metagenomic analyses of samples obtained from the natural aquifers and desert rocks. The metagenomics study of aquifers included nine samples with a sequencing depth of 120 million sequences, leading to the generation of a total of 106 consolidated bins (with a minimum completion of 50% and a maximum contamination of 10%). In parallel, the analysis of desert rocks involved eight samples with a sequencing depth of 181 million sequences, resulting in the identification of 45 bins. Nine of the draft genomes presented similarities to the MDMS higher than 96% (<xref rid="tab2" ref-type="table">Table 2</xref>). The estimated level of completeness for those genomes ranged from 54.38 to 96.47%. 15 of the MDMS were present in the assembly results of the same samples (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S3</xref>). Finding only 9 matches corresponds to the discovery that ribosomal protein genes may be absent in over 20&#x2013;40% of nearly complete metagenome-assembled genomes (<xref ref-type="bibr" rid="ref31">Mise and Iwasaki, 2022</xref>).</p>
<table-wrap position="float" id="tab2"><label>Table 2</label>
<caption>
<p>Bins (draft genomes) with high similarity (genus level) to the MDMS (blastn results) and bin information (completeness and contamination level).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Environment</th>
<th align="center" valign="top">MDMS seq ID</th>
<th align="left" valign="top">seqid</th>
<th align="center" valign="top">% similarity</th>
<th align="center" valign="top">Overlap length</th>
<th align="center" valign="top">Seq length</th>
<th align="center" valign="top">Node length</th>
<th align="center" valign="top">Completeness</th>
<th align="center" valign="top">Contamination</th>
<th align="center" valign="top">Size</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle" rowspan="2">Confined aquifers</td>
<td align="center" valign="middle">A034</td>
<td align="left" valign="middle">bin.2.permissive_NODE_252</td>
<td align="center" valign="middle">100</td>
<td align="center" valign="middle">370</td>
<td align="center" valign="middle">514</td>
<td align="center" valign="middle">2,362</td>
<td align="center" valign="middle">83.51</td>
<td align="center" valign="middle">0.959</td>
<td align="center" valign="middle">1,528,181</td>
</tr>
<tr>
<td align="center" valign="middle">A010</td>
<td align="left" valign="middle">bin.8.permissive_NODE_72</td>
<td align="center" valign="middle">99.707</td>
<td align="center" valign="middle">341</td>
<td align="center" valign="middle">542</td>
<td align="center" valign="middle">14,242</td>
<td align="center" valign="middle">94.64</td>
<td align="center" valign="middle">4.444</td>
<td align="center" valign="middle">5,203,795</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="3">Unconfined aquifers</td>
<td align="center" valign="middle">A020</td>
<td align="left" valign="middle">bin.10.permissive_NODE_12</td>
<td align="center" valign="middle">99.603</td>
<td align="center" valign="middle">504</td>
<td align="center" valign="middle">504</td>
<td align="center" valign="middle">23,198</td>
<td align="center" valign="middle">55.87</td>
<td align="center" valign="middle">0.094</td>
<td align="center" valign="middle">686,180</td>
</tr>
<tr>
<td align="center" valign="middle">A014</td>
<td align="left" valign="middle">bin.40.orig_NODE_559</td>
<td align="center" valign="middle">99.788</td>
<td align="center" valign="middle">471</td>
<td align="center" valign="middle">471</td>
<td align="center" valign="middle">29,535</td>
<td align="center" valign="middle">61.46</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">954,011</td>
</tr>
<tr>
<td align="center" valign="middle">A078</td>
<td align="left" valign="middle">bin.32.strict_NODE_9</td>
<td align="center" valign="middle">98.11</td>
<td align="center" valign="middle">529</td>
<td align="center" valign="middle">530</td>
<td align="center" valign="middle">155,909</td>
<td align="center" valign="middle">96.47</td>
<td align="center" valign="middle">0.352</td>
<td align="center" valign="middle">2,729,365</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="3">Biofilm from aquifers</td>
<td align="center" valign="middle">A083</td>
<td align="left" valign="middle">bin.7.orig_NODE_10041</td>
<td align="center" valign="middle">96.603</td>
<td align="center" valign="middle">471</td>
<td align="center" valign="middle">471</td>
<td align="center" valign="middle">2,996</td>
<td align="center" valign="middle">54.38</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">529,863</td>
</tr>
<tr>
<td align="center" valign="middle">A054</td>
<td align="left" valign="middle">bin.26.permissive_NODE_64</td>
<td align="center" valign="middle">99.787</td>
<td align="center" valign="middle">469</td>
<td align="center" valign="middle">469</td>
<td align="center" valign="middle">4,606</td>
<td align="center" valign="middle">59.33</td>
<td align="center" valign="middle">0.854</td>
<td align="center" valign="middle">834,322</td>
</tr>
<tr>
<td align="center" valign="middle">A146</td>
<td align="left" valign="middle">bin.34.permissive_NODE_1</td>
<td align="center" valign="middle">99.656</td>
<td align="center" valign="middle">291</td>
<td align="center" valign="middle">558</td>
<td align="center" valign="middle">52,977</td>
<td align="center" valign="middle">56.38</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">621,998</td>
</tr>
<tr>
<td align="left" valign="middle">Har Michya</td>
<td align="center" valign="middle">R008</td>
<td align="left" valign="middle">bin.15.strict_NODE_164</td>
<td align="center" valign="middle">98.795</td>
<td align="center" valign="middle">166</td>
<td align="center" valign="middle">445</td>
<td align="center" valign="middle">656</td>
<td align="center" valign="middle">95.95</td>
<td align="center" valign="middle">3.636</td>
<td align="center" valign="middle">6,110,660</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To ensure the integrity of the MDMS data, we performed an additional chimera check, specifically targeted to the 163 MDMS. Among the sequences analyzed, 20 sequences exhibited potential chimeric features (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S2</xref>). Although the low false-positive chimera detection was reported by DECIPHER (<xref ref-type="bibr" rid="ref12">Firth et al., 2009</xref>), some of the 20 MDMS sequences which were suspected as chimeric were found to be similar to sequences in the metagenomic data in the validation process. Due to the limited overlap of the reads and low coverage percentages observed in some of the validated sequences, drawing definitive conclusions about the suspected chimeric sequences poses challenges. Thus, it is without doubt that several of the putative MDMS might be chimeric, which suggests that their taxonomic and phylogenetic analysis are incorrect.</p>
<p>The nine draft genome matches to putative MDMS were characterized in terms of metabolic capacities based on their genes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S4</xref>). <xref rid="tab3" ref-type="table">Table 3</xref> provides the assumed taxonomic attribution for the 9 MDMS. A034 is probably a new class within the Nitrospirota phylum, A010 is related to the Desulfuromonadota and it may be a new class within this phylum or a new, separate phylum. A078 and R008 belong to the Gammaproteobacteria and Chloroflexota, respectively. Five of the MDMS genomes were identified as part of the Patescibacteria super-phylum, such that A020 and A083 are apparently Paceibacteria, A014 is Microgenomatia, and A054 and A0146 are putative new candidate phyla. MDMS that were identified as part of the Patescibacteria super-phylum have fewer features than the other MDMS (<xref rid="tab3" ref-type="table">Table 3</xref> and <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S4</xref>). Such a discrepancy could be caused by the typically small genome size, relatively small percentage of completeness, and lack of basic metabolic capacities that characterized the members of this group (<xref ref-type="bibr" rid="ref52">Tian et al., 2020</xref>), but it could also be due to the lack of information about the functional genes of these uncultured microorganisms. <xref rid="fig5" ref-type="fig">Figure 5</xref> presents some of the metabolic capacities of A010 (related to the Desulfuromonadota) and demonstrates the large amount of information that can be tapped about a prevalent MDMS (A010 constituted 40% of the reads in one sample) but that may be ignored due to their low similarity to existing databases. Bin A010 was assembled with a completion level of 94.6%. In addition to the comprehensive information about bacterial transport systems, we found genes whose expression controls morphology properties such as gram negativity, rod shape and basal body flagella. Moreover, it also contained genes for twitching mobility, sporulation, gluconeogenesis and glycolysis, chitin degradation, formate oxidation, selenate and arsenate reduction, and parts of the nitrogen and sulfur cycles. This metagenome-assembled genome also contained genes such as OmcS (outer-membrane hexaheme c-type cytochrome) and PilA (pilin monomers) that are typical in members of the Desulfuromonadota group and indicate their potential to transfer electrons extracellularly either to iron mineral particles or to microbial syntrophs, including methanogens (<xref ref-type="bibr" rid="ref11">Elul et al., 2021</xref>). Given its origin from water aquifers, this bacterium could play a crucial part in carbon cycling and nutrient transformations within aquatic ecosystems.</p>
<table-wrap position="float" id="tab3"><label>Table 3</label>
<caption>
<p>MDMS attribution to phyla based on 16S rRNA BLAST comparison to databases, location on the GTDB phylogenetic tree and information from the matchings with the draft genomes.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">SeqID</th>
<th align="left" valign="top">Phylum</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">A034</td>
<td align="left" valign="middle">Nitrospirota (new class)</td>
</tr>
<tr>
<td align="left" valign="middle">A010</td>
<td align="left" valign="middle">Desulfuromonadota (new class/phylum)</td>
</tr>
<tr>
<td align="left" valign="middle">A020</td>
<td align="left" valign="middle">Paceibacteria</td>
</tr>
<tr>
<td align="left" valign="middle">A014</td>
<td align="left" valign="middle">Microgenomatia</td>
</tr>
<tr>
<td align="left" valign="middle">A078</td>
<td align="left" valign="middle">Gammaproteobacteria</td>
</tr>
<tr>
<td align="left" valign="middle">A083</td>
<td align="left" valign="middle">Paceibacteria</td>
</tr>
<tr>
<td align="left" valign="middle">A054</td>
<td align="left" valign="middle">Putative new candidate phyla</td>
</tr>
<tr>
<td align="left" valign="middle">A146</td>
<td align="left" valign="middle">Putative new candidate phyla</td>
</tr>
<tr>
<td align="left" valign="middle">R008</td>
<td align="left" valign="middle">Chloroflexota</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig5"><label>Figure 5</label>
<caption>
<p>Visualized map of the metabolic capacities of the putative MDMS A010 based on METABOLIC results and Prokka. Information for other genomes is in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S4</xref>.</p>
</caption>
<graphic xlink:href="fmicb-14-1247119-g005.tif"/>
</fig>
</sec>
<sec sec-type="conclusions" id="sec11"><label>4.</label>
<title>Conclusion</title>
<p>Microbial dark matter (MDM) comprises an immense diversity of yet-uncultivated bacteria. While cultivation independent techniques have been exploited in recent years to expand our knowledge about MDM, the bulk of microbial ecology studies continue to use 16S rRNA gene amplicon sequencing to characterize the microbial communities in a wide range of environments. When using this technique, researchers encounter groups of sequences that cannot be classified under known lineages in the existing databases, sequences that are now identified as belonging to the group of microbial dark matter sequences (MDMS). While these sequences are discarded from most analytical pipelines, they may still play important roles in environmental functioning. Furthermore, while in some well-studied environments, the ecological contribution of the MDMS may be negligible, their presence in the community in certain under-studied environments may be essential. Illuminating their functional contribution in these cases may facilitate a more robust and better understanding of the unique microbial community structures of these environments.</p>
<p>Here, in addition to demonstrating that microbial dark matter indeed present in amplicon sequencing, we present a pipeline to examine the MDM hidden in amplicon sequencing analysis. This study demonstrates that these abundant unidentified OTUs might be an essential part of their ecosystems. Therefore, we encourage researchers to retain these sequences and examine them as they might correspond to complete genomes containing metabolic functions critical to their ecosystems. Though they must be treated carefully, the results of MDMS investigations can be used to expand microbial databases and to situate these microorganisms in the tree of life, which together will promote a better comprehension of their evolution and contribute to the evolving taxonomy of the microbial world.</p>
</sec>
<sec sec-type="data-availability" id="sec12">
<title>Data availability statement</title>
<p>The datasets associated with this study have been deposited in the National Center for Biotechnology Information (NCBI) database. A comprehensive overview of these datasets, including their corresponding accession numbers and types, is provided in <xref ref-type="supplementary-material" rid="SM6">Supplementary Table S5</xref>.</p>
</sec>
<sec sec-type="author-contributions" id="sec13">
<title>Author contributions</title>
<p>HB and NF implemented all bioinformatic analyses and wrote the main manuscript text. HB prepared all figures. IN and ML-N performed samples of rocks and aquifers samples collection and DNA extraction. AK and AS supervised the project. All authors reviewed the manuscript.</p>
</sec>
</body>
<back>
<ack>
<p>The authors gratefully acknowledge the support of the Ministry of Science and Technology (MOST), Israel Fund, Mekorot (Israel National Water Company) and the Ministry of Agriculture for partial funding. The authors also thank the Israel Nature and Parks Authorities (INPA) for granting them permission to sample, and Liran Bugoslavsky and Pradeep Kumar for sharing their data. Furthermore, the authors thank the Avram and Stella Goldstein-Goren fund for partial support.</p>
</ack>
<sec id="sec14">
<title>In memoriam</title>
<p>This paper is dedicated to the memory of Prof. Alex Sivan, who participated in this study.</p>
</sec>
<sec sec-type="COI-statement" id="sec15">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="sec16">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fmicb.2023.1247119/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fmicb.2023.1247119/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table_1.xlsx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table_5.xlsx" id="SM6" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table_6.xlsx" id="SM7" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="fn0001">
<p><sup>1</sup><ext-link xlink:href="http://qiime.org/" ext-link-type="uri">http://qiime.org/</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="ref1">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Alneberg</surname> <given-names>J.</given-names></name> <name><surname>Bjarnason</surname> <given-names>B. S.</given-names></name> <name><surname>De Bruijn</surname> <given-names>I.</given-names></name> <name><surname>Schirmer</surname> <given-names>M.</given-names></name> <name><surname>Quick</surname> <given-names>J.</given-names></name> <name><surname>Ijaz</surname> <given-names>U. Z.</given-names></name> <etal/></person-group>. (<year>2013</year>) <article-title>CONCOCT: clustering contigs on coverage and composition</article-title>. <comment>arXiv preprint arXiv:1312.4038</comment> (<year>2013</year>).</citation></ref>
<ref id="ref2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume>, <fpage>403</fpage>&#x2013;<lpage>410</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id>, PMID: <pub-id pub-id-type="pmid">2231712</pub-id></citation></ref>
<ref id="ref3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>B. J.</given-names></name> <name><surname>Dick</surname> <given-names>G. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Omic approaches in microbial ecology: charting the unknown</article-title>. <source>Microbe</source> <volume>8</volume>, <fpage>353</fpage>&#x2013;<lpage>359</lpage>. doi: <pub-id pub-id-type="doi">10.1128/microbe.8.353.1</pub-id></citation></ref>
<ref id="ref4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barak</surname> <given-names>H.</given-names></name> <name><surname>Brenner</surname> <given-names>A.</given-names></name> <name><surname>Sivan</surname> <given-names>A.</given-names></name> <name><surname>Kushmaro</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Temporal distribution of microbial community in an industrial wastewater treatment system following crash and during recovery periods</article-title>. <source>Chemosphere</source> <volume>258</volume>:<fpage>127271</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.chemosphere.2020.127271</pub-id>, PMID: <pub-id pub-id-type="pmid">32535444</pub-id></citation></ref>
<ref id="ref5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barak</surname> <given-names>H.</given-names></name> <name><surname>Kumar</surname> <given-names>P.</given-names></name> <name><surname>Zaritsky</surname> <given-names>A.</given-names></name> <name><surname>Mendel</surname> <given-names>Z.</given-names></name> <name><surname>Ment</surname> <given-names>D.</given-names></name> <name><surname>Kushmaro</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Diversity of bacterial biota in Capnodis tenebrionis (Coleoptera: Buprestidae) larvae</article-title>. <source>Pathogens</source> <volume>8</volume>:<fpage>4</fpage>. doi: <pub-id pub-id-type="doi">10.3390/pathogens8010004</pub-id>, PMID: <pub-id pub-id-type="pmid">30621355</pub-id></citation></ref>
<ref id="ref6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>C. T.</given-names></name> <name><surname>Hug</surname> <given-names>L. A.</given-names></name> <name><surname>Thomas</surname> <given-names>B. C.</given-names></name> <name><surname>Sharon</surname> <given-names>I.</given-names></name> <name><surname>Castelle</surname> <given-names>C. J.</given-names></name> <name><surname>Singh</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Unusual biology across a group comprising more than 15% of domain Bacteria</article-title>. <source>Nature</source> <volume>523</volume>, <fpage>208</fpage>&#x2013;<lpage>211</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature14486</pub-id>, PMID: <pub-id pub-id-type="pmid">26083755</pub-id></citation></ref>
<ref id="ref7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caporaso</surname> <given-names>J. G.</given-names></name> <name><surname>Kuczynski</surname> <given-names>J.</given-names></name> <name><surname>Stombaugh</surname> <given-names>J.</given-names></name> <name><surname>Bittinger</surname> <given-names>K.</given-names></name> <name><surname>Bushman</surname> <given-names>F. D.</given-names></name> <name><surname>Costello</surname> <given-names>E. K.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>QIIME allows analysis of high-throughput community sequencing data</article-title>. <source>Nat. Methods</source> <volume>7</volume>, <fpage>335</fpage>&#x2013;<lpage>336</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nmeth.f.303</pub-id>, PMID: <pub-id pub-id-type="pmid">20383131</pub-id></citation></ref>
<ref id="ref8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Castelle</surname> <given-names>C. J.</given-names></name> <name><surname>Banfield</surname> <given-names>J. F.</given-names></name></person-group> (<year>2018</year>). <article-title>Major new microbial groups expand diversity and alter our understanding of the tree of life</article-title>. <source>Cells</source> <volume>172</volume>, <fpage>1181</fpage>&#x2013;<lpage>1197</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2018.02.016</pub-id>, PMID: <pub-id pub-id-type="pmid">29522741</pub-id></citation></ref>
<ref id="ref9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Ye</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>High speed BLASTN: an accelerated MegaBLAST search tool</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>7762</fpage>&#x2013;<lpage>7768</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkv784</pub-id>, PMID: <pub-id pub-id-type="pmid">26250111</pub-id></citation></ref>
<ref id="ref10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elovitz</surname> <given-names>M. A.</given-names></name> <name><surname>Gajer</surname> <given-names>P.</given-names></name> <name><surname>Riis</surname> <given-names>V.</given-names></name> <name><surname>Brown</surname> <given-names>A. G.</given-names></name> <name><surname>Humphrys</surname> <given-names>M. S.</given-names></name> <name><surname>Holm</surname> <given-names>J. B.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Cervicovaginal microbiota and local immune response modulate the risk of spontaneous preterm delivery</article-title>. <source>Nat. Commun.</source> <volume>10</volume>:<fpage>1305</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41467-019-09285-9</pub-id></citation></ref>
<ref id="ref11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elul</surname> <given-names>M.</given-names></name> <name><surname>Rubin-Blum</surname> <given-names>M.</given-names></name> <name><surname>Ronen</surname> <given-names>Z.</given-names></name> <name><surname>Bar-Or</surname> <given-names>I.</given-names></name> <name><surname>Eckert</surname> <given-names>W.</given-names></name> <name><surname>Sivan</surname> <given-names>O.</given-names></name></person-group> (<year>2021</year>). <article-title>Metagenomic insights into the metabolism of microbial communities that mediate iron and methane cycling in Lake Kinneret iron-rich methanic sediments</article-title>. <source>Biogeosciences</source> <volume>18</volume>, <fpage>2091</fpage>&#x2013;<lpage>2106</lpage>. doi: <pub-id pub-id-type="doi">10.5194/bg-18-2091-2021</pub-id></citation></ref>
<ref id="ref12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Firth</surname> <given-names>H. V.</given-names></name> <name><surname>Richards</surname> <given-names>S. M.</given-names></name> <name><surname>Bevan</surname> <given-names>A. P.</given-names></name> <name><surname>Clayton</surname> <given-names>S.</given-names></name> <name><surname>Corpas</surname> <given-names>M.</given-names></name> <name><surname>Rajan</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>DECIPHER: database of chromosomal imbalance and phenotype in humans using ensembl resources</article-title>. <source>Am. J. Hum. Genet.</source> <volume>84</volume>, <fpage>524</fpage>&#x2013;<lpage>533</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ajhg.2009.03.010</pub-id>, PMID: <pub-id pub-id-type="pmid">19344873</pub-id></citation></ref>
<ref id="ref13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Green</surname> <given-names>S. J.</given-names></name> <name><surname>Venkatramanan</surname> <given-names>R.</given-names></name> <name><surname>Naqib</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Deconstructing the polymerase chain reaction: understanding and correcting bias associated with primer degeneracies and primer-template mismatches</article-title>. <source>PLoS One</source> <volume>10</volume>:<fpage>e0128122</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0128122</pub-id>, PMID: <pub-id pub-id-type="pmid">25996930</pub-id></citation></ref>
<ref id="ref14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harris</surname> <given-names>J. K.</given-names></name> <name><surname>Kelley</surname> <given-names>S. T.</given-names></name> <name><surname>Pace</surname> <given-names>N. R.</given-names></name></person-group> (<year>2004</year>). <article-title>New perspective on uncultured bacterial phylogenetic division OP11</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>70</volume>, <fpage>845</fpage>&#x2013;<lpage>849</lpage>. doi: <pub-id pub-id-type="doi">10.1128/AEM.70.2.845-849.2004</pub-id>, PMID: <pub-id pub-id-type="pmid">14766563</pub-id></citation></ref>
<ref id="ref15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hedlund</surname> <given-names>B. P.</given-names></name> <name><surname>Dodsworth</surname> <given-names>J. A.</given-names></name> <name><surname>Murugapiran</surname> <given-names>S. K.</given-names></name> <name><surname>Rinke</surname> <given-names>C.</given-names></name> <name><surname>Woyke</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Impact of single-cell genomics and metagenomics on the emerging view of extremophile &#x201C;microbial dark matter&#x201D;</article-title>. <source>Extremophiles</source> <volume>18</volume>, <fpage>865</fpage>&#x2013;<lpage>875</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00792-014-0664-7</pub-id>, PMID: <pub-id pub-id-type="pmid">25113821</pub-id></citation></ref>
<ref id="ref16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herrmann</surname> <given-names>M.</given-names></name> <name><surname>Wegner</surname> <given-names>C.</given-names></name> <name><surname>Taubert</surname> <given-names>M.</given-names></name> <name><surname>Geesink</surname> <given-names>P.</given-names></name> <name><surname>Lehmann</surname> <given-names>K.</given-names></name> <name><surname>Yan</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Predominance of cand. Patescibacteria in groundwater is caused by their preferential mobilization from soils and flourishing under oligotrophic conditions</article-title>. <source>Front. Microbiol.</source> <volume>10</volume>:<fpage>1407</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fmicb.2019.01407</pub-id></citation></ref>
<ref id="ref17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hug</surname> <given-names>L. A.</given-names></name> <name><surname>Baker</surname> <given-names>B. J.</given-names></name> <name><surname>Anantharaman</surname> <given-names>K.</given-names></name> <name><surname>Brown</surname> <given-names>C. T.</given-names></name> <name><surname>Probst</surname> <given-names>A. J.</given-names></name> <name><surname>Castelle</surname> <given-names>C. J.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>A new view of the tree of life</article-title>. <source>Nat. Microbiol.</source> <volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nmicrobiol.2016.48</pub-id></citation></ref>
<ref id="ref18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hugerth</surname> <given-names>L. W.</given-names></name> <name><surname>Wefer</surname> <given-names>H. A.</given-names></name> <name><surname>Lundin</surname> <given-names>S.</given-names></name> <name><surname>Jakobsson</surname> <given-names>H. E.</given-names></name> <name><surname>Lindberg</surname> <given-names>M.</given-names></name> <name><surname>Rodin</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>DegePrime, a program for degenerate primer design for broad-taxonomic-range PCR in microbial ecology studies</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>80</volume>, <fpage>5116</fpage>&#x2013;<lpage>5123</lpage>. doi: <pub-id pub-id-type="doi">10.1128/AEM.01403-14</pub-id>, PMID: <pub-id pub-id-type="pmid">24928874</pub-id></citation></ref>
<ref id="ref19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Irit</surname> <given-names>N.</given-names></name> <name><surname>Hana</surname> <given-names>B.</given-names></name> <name><surname>Yifat</surname> <given-names>B.</given-names></name> <name><surname>Esti</surname> <given-names>K.</given-names></name> <name><surname>Ariel</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Insights into bacterial communities associated with petroglyph sites from the Negev Desert, Israel</article-title>. <source>J. Arid. Environ.</source> <volume>166</volume>, <fpage>79</fpage>&#x2013;<lpage>82</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jaridenv.2019.04.010</pub-id></citation></ref>
<ref id="ref20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiao</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Hua</surname> <given-names>Z.</given-names></name> <name><surname>Fang</surname> <given-names>B.</given-names></name> <name><surname>Zhou</surname> <given-names>E.</given-names></name> <name><surname>Salam</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Microbial dark matter coming to light: challenges and opportunities</article-title>. <source>Natl. Sci. Rev.</source> <volume>8</volume>:<fpage>nwaa280</fpage>. doi: <pub-id pub-id-type="doi">10.1093/nsr/nwaa280</pub-id>, PMID: <pub-id pub-id-type="pmid">34691599</pub-id></citation></ref>
<ref id="ref21">
<citation citation-type="journal"><person-group person-group-type="author"><collab id="coll1">Jumpstart Consortium Human Microbiome Project Data Generation Working Group</collab></person-group> (<year>2012</year>). <article-title>Evaluation of 16S rDNA-based community profiling for human microbiome research</article-title>. <source>PLoS One</source> <volume>7</volume>:<fpage>e39315</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0051204</pub-id>, PMID: <pub-id pub-id-type="pmid">23349657</pub-id></citation></ref>
<ref id="ref22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>D. D.</given-names></name> <name><surname>Froula</surname> <given-names>J.</given-names></name> <name><surname>Egan</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>MetaBAT, an efficient tool for accurately reconstructing single genomes from complex microbial communities</article-title>. <source>PeerJ</source> <volume>3</volume>:<fpage>e1165</fpage>. doi: <pub-id pub-id-type="doi">10.7717/peerj.1165</pub-id>, PMID: <pub-id pub-id-type="pmid">26336640</pub-id></citation></ref>
<ref id="ref23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoh</surname> <given-names>K.</given-names></name> <name><surname>Standley</surname> <given-names>D. M.</given-names></name></person-group> (<year>2013</year>). <article-title>MAFFT multiple sequence alignment software version 7: improvements in performance and usability</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>772</fpage>&#x2013;<lpage>780</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/mst010</pub-id>, PMID: <pub-id pub-id-type="pmid">23329690</pub-id></citation></ref>
<ref id="ref24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keshri</surname> <given-names>J.</given-names></name> <name><surname>Mankazana</surname> <given-names>B. B.</given-names></name> <name><surname>Momba</surname> <given-names>M. N.</given-names></name></person-group> (<year>2015</year>). <article-title>Profile of bacterial communities in south African mine-water samples using Illumina next-generation sequencing platform</article-title>. <source>Appl. Microbiol. Biotechnol.</source> <volume>99</volume>, <fpage>3233</fpage>&#x2013;<lpage>3242</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00253-014-6213-6</pub-id>, PMID: <pub-id pub-id-type="pmid">25416590</pub-id></citation></ref>
<ref id="ref25">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Krueger</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>) <source>Trim galore: a wrapper tool around Cutadapt and FastQC to consistently apply quality and adapter trimming to FastQ files, with some extra functionality for MspI-digested RRBS-type (reduced representation Bisufite-Seq) libraries</source>. <comment>Available at: </comment><ext-link xlink:href="http://www.bioinformatics.babraham.ac.uk/projects/trim_galore/" ext-link-type="uri">http://www.bioinformatics.babraham.ac.uk/projects/trim_galore/</ext-link> (Accessed April 28, 2016).</citation></ref>
<ref id="ref26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Letunic</surname> <given-names>I.</given-names></name> <name><surname>Bork</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Interactive tree of life (iTOL) v3: an online tool for the display and annotation of phylogenetic and other trees</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>W242</fpage>&#x2013;<lpage>W245</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkw290</pub-id>, PMID: <pub-id pub-id-type="pmid">27095192</pub-id></citation></ref>
<ref id="ref27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Luo</surname> <given-names>R.</given-names></name> <name><surname>Sadakane</surname> <given-names>K.</given-names></name> <name><surname>Lam</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>1674</fpage>&#x2013;<lpage>1676</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btv033</pub-id>, PMID: <pub-id pub-id-type="pmid">25609793</pub-id></citation></ref>
<ref id="ref28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lynch</surname> <given-names>M. D.</given-names></name> <name><surname>Bartram</surname> <given-names>A. K.</given-names></name> <name><surname>Neufeld</surname> <given-names>J. D.</given-names></name></person-group> (<year>2012</year>). <article-title>Targeted recovery of novel phylogenetic diversity from next-generation sequence data</article-title>. <source>ISME J.</source> <volume>6</volume>, <fpage>2067</fpage>&#x2013;<lpage>2077</lpage>. doi: <pub-id pub-id-type="doi">10.1038/ismej.2012.50</pub-id>, PMID: <pub-id pub-id-type="pmid">22791239</pub-id></citation></ref>
<ref id="ref29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDonald</surname> <given-names>D.</given-names></name> <name><surname>Clemente</surname> <given-names>J. C.</given-names></name> <name><surname>Kuczynski</surname> <given-names>J.</given-names></name> <name><surname>Rideout</surname> <given-names>J. R.</given-names></name> <name><surname>Stombaugh</surname> <given-names>J.</given-names></name> <name><surname>Wendel</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The biological observation matrix (BIOM) format or: how I learned to stop worrying and love the ome-ome</article-title>. <source>Gigascience</source> <volume>1</volume>:<fpage>7</fpage>. doi: <pub-id pub-id-type="doi">10.1186/2047-217X-1-7</pub-id></citation></ref>
<ref id="ref30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x00E9;heust</surname> <given-names>R.</given-names></name> <name><surname>Burstein</surname> <given-names>D.</given-names></name> <name><surname>Castelle</surname> <given-names>C. J.</given-names></name> <name><surname>Banfield</surname> <given-names>J. F.</given-names></name></person-group> (<year>2019</year>). <article-title>The distinction of CPR bacteria from other bacteria based on protein family content</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41467-019-12171-z</pub-id></citation></ref>
<ref id="ref31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mise</surname> <given-names>K.</given-names></name> <name><surname>Iwasaki</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Unexpected absence of ribosomal protein genes from metagenome-assembled genomes</article-title>. <source>ISME Commun.</source> <volume>2</volume>:<fpage>118</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s43705-022-00204-6</pub-id></citation></ref>
<ref id="ref32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morowitz</surname> <given-names>M. J.</given-names></name> <name><surname>Carlisle</surname> <given-names>E. M.</given-names></name> <name><surname>Alverdy</surname> <given-names>J. C.</given-names></name></person-group> (<year>2011</year>). <article-title>Contributions of intestinal bacteria to nutrition and metabolism in the critically ill</article-title>. <source>Surg. Clin.</source> <volume>91</volume>, <fpage>771</fpage>&#x2013;<lpage>785</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.suc.2011.05.001</pub-id>, PMID: <pub-id pub-id-type="pmid">21787967</pub-id></citation></ref>
<ref id="ref33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakai</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Size matters: ultra-small and filterable microorganisms in the environment</article-title>. <source>Microbes Environ.</source> <volume>35</volume>:<fpage>ME20025</fpage>. doi: <pub-id pub-id-type="doi">10.1264/jsme2.ME20025</pub-id></citation></ref>
<ref id="ref34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nawrocki</surname> <given-names>E. P.</given-names></name></person-group> (<year>2009</year>). <source>Structural RNA homology search and alignment using covariance models [dissertation/master&#x2019;s thesis]</source> <publisher-name>Washington University in Saint Louis</publisher-name>.</citation></ref>
<ref id="ref1001">
<citation citation-type="journal"><person-group person-group-type="author"><collab id="coll1001">NCBI Resource Coordinators</collab></person-group> (<year>2013</year>). <article-title>Database resources of the National Center for Biotechnology Information</article-title>. <source>Nucleic Acids Res.</source> <volume>41</volume>, <fpage>D8</fpage>&#x2013;<lpage>D20</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gks1189</pub-id></citation></ref>
<ref id="ref35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Needham</surname> <given-names>D. M.</given-names></name> <name><surname>Sachdeva</surname> <given-names>R.</given-names></name> <name><surname>Fuhrman</surname> <given-names>J. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Ecological dynamics and co-occurrence among marine phytoplankton, bacteria and myoviruses shows microdiversity matters</article-title>. <source>ISME J.</source> <volume>11</volume>, <fpage>1614</fpage>&#x2013;<lpage>1629</lpage>. doi: <pub-id pub-id-type="doi">10.1038/ismej.2017.29</pub-id>, PMID: <pub-id pub-id-type="pmid">28398348</pub-id></citation></ref>
<ref id="ref36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nir</surname> <given-names>I.</given-names></name> <name><surname>Barak</surname> <given-names>H.</given-names></name> <name><surname>Kramarsky-Winter</surname> <given-names>E.</given-names></name> <name><surname>Kushmaro</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Seasonal diversity of the bacterial communities associated with petroglyphs sites from the Negev Desert, Israel</article-title>. <source>Ann. Microbiol.</source> <volume>69</volume>, <fpage>1079</fpage>&#x2013;<lpage>1086</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s13213-019-01509-z</pub-id></citation></ref>
<ref id="ref37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nurk</surname> <given-names>S.</given-names></name> <name><surname>Meleshko</surname> <given-names>D.</given-names></name> <name><surname>Korobeynikov</surname> <given-names>A.</given-names></name> <name><surname>Pevzner</surname> <given-names>P. A.</given-names></name></person-group> (<year>2017</year>). <article-title>metaSPAdes: a new versatile metagenomic assembler</article-title>. <source>Genome Res.</source> <volume>27</volume>, <fpage>824</fpage>&#x2013;<lpage>834</lpage>. doi: <pub-id pub-id-type="doi">10.1101/gr.213959.116</pub-id>, PMID: <pub-id pub-id-type="pmid">28298430</pub-id></citation></ref>
<ref id="ref38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Panda</surname> <given-names>A. K.</given-names></name> <name><surname>Bisht</surname> <given-names>S. S.</given-names></name> <name><surname>Kaushal</surname> <given-names>B. R.</given-names></name> <name><surname>De Mandal</surname> <given-names>S.</given-names></name> <name><surname>Kumar</surname> <given-names>N. S.</given-names></name> <name><surname>Basistha</surname> <given-names>B. C.</given-names></name></person-group> (<year>2017</year>). <article-title>Bacterial diversity analysis of Yumthang hot spring, North Sikkim, India by Illumina sequencing</article-title>. <source>Big Data Anal.</source> <volume>2</volume>, <fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s41044-017-0022-8</pub-id></citation></ref>
<ref id="ref39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parks</surname> <given-names>D. H.</given-names></name> <name><surname>Chuvochina</surname> <given-names>M.</given-names></name> <name><surname>Waite</surname> <given-names>D. W.</given-names></name> <name><surname>Rinke</surname> <given-names>C.</given-names></name> <name><surname>Skarshewski</surname> <given-names>A.</given-names></name> <name><surname>Chaumeil</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life</article-title>. <source>Nat. Biotechnol.</source> <volume>36</volume>, <fpage>996</fpage>&#x2013;<lpage>1004</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nbt.4229</pub-id>, PMID: <pub-id pub-id-type="pmid">30148503</pub-id></citation></ref>
<ref id="ref40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parks</surname> <given-names>D. H.</given-names></name> <name><surname>Rinke</surname> <given-names>C.</given-names></name> <name><surname>Chuvochina</surname> <given-names>M.</given-names></name> <name><surname>Chaumeil</surname> <given-names>P.</given-names></name> <name><surname>Woodcroft</surname> <given-names>B. J.</given-names></name> <name><surname>Evans</surname> <given-names>P. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life</article-title>. <source>Nat. Microbiol.</source> <volume>2</volume>, <fpage>1533</fpage>&#x2013;<lpage>1542</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41564-017-0012-7</pub-id>, PMID: <pub-id pub-id-type="pmid">28894102</pub-id></citation></ref>
<ref id="ref41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pascoal</surname> <given-names>F.</given-names></name> <name><surname>Costa</surname> <given-names>R.</given-names></name> <name><surname>Magalh&#x00E3;es</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>The microbial rare biosphere: current concepts, methods and ecological principles</article-title>. <source>FEMS Microbiol. Ecol.</source> <volume>97</volume>:<fpage>fiaa227</fpage>. doi: <pub-id pub-id-type="doi">10.1093/femsec/fiaa227</pub-id></citation></ref>
<ref id="ref42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Price</surname> <given-names>M. N.</given-names></name> <name><surname>Dehal</surname> <given-names>P. S.</given-names></name> <name><surname>Arkin</surname> <given-names>A. P.</given-names></name></person-group> (<year>2010</year>). <article-title>FastTree 2&#x2013;approximately maximum-likelihood trees for large alignments</article-title>. <source>PLoS One</source> <volume>5</volume>:<fpage>e9490</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0009490</pub-id>, PMID: <pub-id pub-id-type="pmid">20224823</pub-id></citation></ref>
<ref id="ref43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prodan</surname> <given-names>A.</given-names></name> <name><surname>Tremaroli</surname> <given-names>V.</given-names></name> <name><surname>Brolin</surname> <given-names>H.</given-names></name> <name><surname>Zwinderman</surname> <given-names>A. H.</given-names></name> <name><surname>Nieuwdorp</surname> <given-names>M.</given-names></name> <name><surname>Levin</surname> <given-names>E.</given-names></name></person-group> (<year>2020</year>). <article-title>Comparing bioinformatic pipelines for microbial 16S rRNA amplicon sequencing</article-title>. <source>PLoS One</source> <volume>15</volume>:<fpage>e0227434</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0227434</pub-id>, PMID: <pub-id pub-id-type="pmid">31945086</pub-id></citation></ref>
<ref id="ref44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quast</surname> <given-names>C.</given-names></name> <name><surname>Pruesse</surname> <given-names>E.</given-names></name> <name><surname>Yilmaz</surname> <given-names>P.</given-names></name> <name><surname>Gerken</surname> <given-names>J.</given-names></name> <name><surname>Schweer</surname> <given-names>T.</given-names></name> <name><surname>Yarza</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The SILVA ribosomal RNA gene database project: improved data processing and web-based tools</article-title>. <source>Nucleic Acids Res.</source> <volume>41</volume>, <fpage>D590</fpage>&#x2013;<lpage>D596</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gks1219</pub-id>, PMID: <pub-id pub-id-type="pmid">23193283</pub-id></citation></ref>
<ref id="ref45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rinke</surname> <given-names>C.</given-names></name> <name><surname>Schwientek</surname> <given-names>P.</given-names></name> <name><surname>Sczyrba</surname> <given-names>A.</given-names></name> <name><surname>Ivanova</surname> <given-names>N. N.</given-names></name> <name><surname>Anderson</surname> <given-names>I. J.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Insights into the phylogeny and coding potential of microbial dark matter</article-title>. <source>Nature</source> <volume>499</volume>, <fpage>431</fpage>&#x2013;<lpage>437</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature12352</pub-id>, PMID: <pub-id pub-id-type="pmid">23851394</pub-id></citation></ref>
<ref id="ref46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santos</surname> <given-names>A.</given-names></name> <name><surname>van Aerle</surname> <given-names>R.</given-names></name> <name><surname>Barrientos</surname> <given-names>L.</given-names></name> <name><surname>Martinez-Urtaza</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Computational methods for 16S metabarcoding studies using nanopore sequencing data</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>18</volume>, <fpage>296</fpage>&#x2013;<lpage>305</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.csbj.2020.01.005</pub-id>, PMID: <pub-id pub-id-type="pmid">32071706</pub-id></citation></ref>
<ref id="ref47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schloss</surname> <given-names>P. D.</given-names></name> <name><surname>Westcott</surname> <given-names>S. L.</given-names></name> <name><surname>Ryabin</surname> <given-names>T.</given-names></name> <name><surname>Hall</surname> <given-names>J. R.</given-names></name> <name><surname>Hartmann</surname> <given-names>M.</given-names></name> <name><surname>Hollister</surname> <given-names>E. B.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Introducing mothur: open-source, platform-independent, community-supported software for describing and comparing microbial communities</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>75</volume>, <fpage>7537</fpage>&#x2013;<lpage>7541</lpage>. doi: <pub-id pub-id-type="doi">10.1128/AEM.01541-09</pub-id>, PMID: <pub-id pub-id-type="pmid">19801464</pub-id></citation></ref>
<ref id="ref48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schulz</surname> <given-names>F.</given-names></name> <name><surname>Eloe-Fadrosh</surname> <given-names>E. A.</given-names></name> <name><surname>Bowers</surname> <given-names>R. M.</given-names></name> <name><surname>Jarett</surname> <given-names>J.</given-names></name> <name><surname>Nielsen</surname> <given-names>T.</given-names></name> <name><surname>Ivanova</surname> <given-names>N. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Towards a balanced view of the bacterial tree of life</article-title>. <source>Microbiome</source> <volume>5</volume>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s40168-017-0360-9</pub-id></citation></ref>
<ref id="ref49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seemann</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Prokka: rapid prokaryotic genome annotation</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>2068</fpage>&#x2013;<lpage>2069</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btu153</pub-id>, PMID: <pub-id pub-id-type="pmid">24642063</pub-id></citation></ref>
<ref id="ref50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Solden</surname> <given-names>L.</given-names></name> <name><surname>Lloyd</surname> <given-names>K.</given-names></name> <name><surname>Wrighton</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>The bright side of microbial dark matter: lessons learned from the uncultivated majority</article-title>. <source>Curr. Opin. Microbiol.</source> <volume>31</volume>, <fpage>217</fpage>&#x2013;<lpage>226</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.mib.2016.04.020</pub-id>, PMID: <pub-id pub-id-type="pmid">27196505</pub-id></citation></ref>
<ref id="ref51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staley</surname> <given-names>J. T.</given-names></name> <name><surname>Konopka</surname> <given-names>A.</given-names></name></person-group> (<year>1985</year>). <article-title>Measurement of in situ activities of nonphotosynthetic microorganisms in aquatic and terrestrial habitats</article-title>. <source>Annu. Rev. Microbiol.</source> <volume>39</volume>, <fpage>321</fpage>&#x2013;<lpage>346</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.mi.39.100185.001541</pub-id>, PMID: <pub-id pub-id-type="pmid">3904603</pub-id></citation></ref>
<ref id="ref52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>R.</given-names></name> <name><surname>Ning</surname> <given-names>D.</given-names></name> <name><surname>He</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name> <name><surname>Spencer</surname> <given-names>S. J.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Small and mighty: adaptation of superphylum Patescibacteria to groundwater environment drives their genome simplicity</article-title>. <source>Microbiome</source> <volume>8</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s40168-020-00825-w</pub-id></citation></ref>
<ref id="ref53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tringe</surname> <given-names>S. G.</given-names></name> <name><surname>Hugenholtz</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>A renaissance for the pioneering 16S rRNA gene</article-title>. <source>Curr. Opin. Microbiol.</source> <volume>11</volume>, <fpage>442</fpage>&#x2013;<lpage>446</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.mib.2008.09.011</pub-id>, PMID: <pub-id pub-id-type="pmid">18817891</pub-id></citation></ref>
<ref id="ref54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vigneron</surname> <given-names>A.</given-names></name> <name><surname>Cruaud</surname> <given-names>P.</given-names></name> <name><surname>Langlois</surname> <given-names>V.</given-names></name> <name><surname>Lovejoy</surname> <given-names>C.</given-names></name> <name><surname>Culley</surname> <given-names>A. I.</given-names></name> <name><surname>Vincent</surname> <given-names>W. F.</given-names></name></person-group> (<year>2020</year>). <article-title>Ultra-small and abundant: candidate phyla radiation bacteria are potential catalysts of carbon transformation in a thermokarst lake ecosystem</article-title>. <source>Limnol. Oceanogr. Lett.</source> <volume>5</volume>, <fpage>212</fpage>&#x2013;<lpage>220</lpage>. doi: <pub-id pub-id-type="doi">10.1002/lol2.10132</pub-id></citation></ref>
<ref id="ref55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wiegand</surname> <given-names>S.</given-names></name> <name><surname>Dam</surname> <given-names>H. T.</given-names></name> <name><surname>Riba</surname> <given-names>J.</given-names></name> <name><surname>Vollmers</surname> <given-names>J.</given-names></name> <name><surname>Kaster</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Printing microbial dark matter: using single cell dispensing and genomics to investigate the patescibacteria/candidate phyla radiation</article-title>. <source>Front. Microbiol.</source> <volume>12</volume>:<fpage>1512</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fmicb.2021.635506</pub-id></citation></ref>
<ref id="ref56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>L.</given-names></name> <name><surname>Hermans</surname> <given-names>S. M.</given-names></name> <name><surname>Totsche</surname> <given-names>K. U.</given-names></name> <name><surname>Lehmann</surname> <given-names>R.</given-names></name> <name><surname>Herrmann</surname> <given-names>M.</given-names></name> <name><surname>K&#x00FC;sel</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>Groundwater bacterial communities evolve over time in response to recharge</article-title>. <source>Water Res.</source> <volume>201</volume>:<fpage>117290</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.watres.2021.117290</pub-id>, PMID: <pub-id pub-id-type="pmid">34186289</pub-id></citation></ref>
<ref id="ref57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yarza</surname> <given-names>P.</given-names></name> <name><surname>Yilmaz</surname> <given-names>P.</given-names></name> <name><surname>Pruesse</surname> <given-names>E.</given-names></name> <name><surname>Gl&#x00F6;ckner</surname> <given-names>F. O.</given-names></name> <name><surname>Ludwig</surname> <given-names>W.</given-names></name> <name><surname>Schleifer</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Uniting the classification of cultured and uncultured bacteria and archaea using 16S rRNA gene sequences</article-title>. <source>Nat. Rev. Microbiol.</source> <volume>12</volume>, <fpage>635</fpage>&#x2013;<lpage>645</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nrmicro3330</pub-id>, PMID: <pub-id pub-id-type="pmid">25118885</pub-id></citation></ref>
<ref id="ref58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ye</surname> <given-names>J.</given-names></name> <name><surname>Coulouris</surname> <given-names>G.</given-names></name> <name><surname>Zaretskaya</surname> <given-names>I.</given-names></name> <name><surname>Cutcutache</surname> <given-names>I.</given-names></name> <name><surname>Rozen</surname> <given-names>S.</given-names></name> <name><surname>Madden</surname> <given-names>T. L.</given-names></name></person-group> (<year>2012</year>). <article-title>Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction</article-title>. <source>BMC Bioinformatics</source> <volume>13</volume>:<fpage>134</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2105-13-134</pub-id></citation></ref>
<ref id="ref59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>S.</given-names></name> <name><surname>Ha</surname> <given-names>S.</given-names></name> <name><surname>Kwon</surname> <given-names>S.</given-names></name> <name><surname>Lim</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>Y.</given-names></name> <name><surname>Seo</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Introducing EzBioCloud: a taxonomically united database of 16S rRNA gene sequences and whole-genome assemblies</article-title>. <source>Int. J. Syst. Evol. Microbiol.</source> <volume>67</volume>:<fpage>1613</fpage>. doi: <pub-id pub-id-type="doi">10.1099/ijsem.0.001755</pub-id>, PMID: <pub-id pub-id-type="pmid">28005526</pub-id></citation></ref>
<ref id="ref60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zamkovaya</surname> <given-names>T.</given-names></name> <name><surname>Foster</surname> <given-names>J. S.</given-names></name> <name><surname>de Cr&#x00E9;cy-Lagard</surname> <given-names>V.</given-names></name> <name><surname>Conesa</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>A network approach to elucidate and prioritize microbial dark matter in microbial communities</article-title>. <source>ISME J.</source> <volume>15</volume>, <fpage>228</fpage>&#x2013;<lpage>244</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41396-020-00777-x</pub-id>, PMID: <pub-id pub-id-type="pmid">32963345</pub-id></citation></ref>
<ref id="ref61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Kobert</surname> <given-names>K.</given-names></name> <name><surname>Flouri</surname> <given-names>T.</given-names></name> <name><surname>Stamatakis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>PEAR: a fast and accurate Illumina paired-end reAd mergeR</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>614</fpage>&#x2013;<lpage>620</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btt593</pub-id>, PMID: <pub-id pub-id-type="pmid">24142950</pub-id></citation></ref>
<ref id="ref62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Tran</surname> <given-names>P.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Kieft</surname> <given-names>K.</given-names></name> <name><surname>Anantharaman</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>METABOLIC: a scalable high-throughput metabolic and biogeochemical functional trait profiler based on microbial genomes</article-title>. <source>bioRxiv</source></citation></ref>
</ref-list>
</back>
</article>