<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2022.1079279</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Comparative genomics reveals cellobiose hydrolysis mechanism of <italic>Ruminiclostridium thermocellum</italic> M3, a cellulosic saccharification bacterium</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Tao</surname>
<given-names>Sheng</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1583677/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Qingbin</surname>
<given-names>Meng</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Zhiling</surname>
<given-names>Li</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c002" ref-type="corresp"><sup>&#x002A;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Caiyu</surname>
<given-names>Sun</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Lixin</surname>
<given-names>Li</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Lilai</surname>
<given-names>Liu</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>College of Environmental and Chemical Engineering, Heilongjiang University of Science and Technology</institution>, <addr-line>Harbin</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>State Key Lab of Urban Water Resource and Environment, Harbin Institute of Technology</institution>, <addr-line>Harbin</addr-line>, <country>China</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Mamoru Yamada, Yamaguchi University, Japan</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Yejun Han, Institute of Process Engineering (CAS), China; Rosa Estela Quiroz Casta&#x00F1;eda,Instituto Nacional de Investigaciones Forestales, Agr&#x00ED;colas y Pecuarias (INIFAP), Mexico</p></fn>
<corresp id="c001">&#x002A;Correspondence: Sheng Tao, <email>tsheng@usth.edu.cn</email></corresp>
<corresp id="c002">Li Zhiling, <email>lzlhit@163.com</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Evolutionary and Genomic Microbiology, a section of the journal Frontiers in Microbiology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>1079279</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Tao, Qingbin, Zhiling, Caiyu, Lixin and Lilai.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Tao, Qingbin, Zhiling, Caiyu, Lixin and Lilai</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The cellulosome of <italic>Ruminiclostridium thermocellum</italic> was one of the most efficient cellulase systems in nature. However, the product of cellulose degradation by <italic>R. thermocellum</italic> is cellobiose, which leads to the feedback inhibition of cellulosome, and it limits the <italic>R. thermocellum</italic> application in the field of cellulosic biomass consolidated bioprocessing (CBP) industry. In a previous study, <italic>R. thermocellum</italic> M3, which can hydrolyze cellulosic feedstocks into monosaccharides, was isolated from horse manure. In this study, the complete genome of <italic>R. thermocellum</italic> M3 was sequenced and assembled. The genome of <italic>R. thermocellum</italic> M3 was compared with the other <italic>R. thermocellum</italic> to reveal the mechanism of cellulosic saccharification by <italic>R. thermocellum</italic> M3. In addition, we predicted the key genes for the elimination of feedback inhibition of cellobiose in <italic>R. thermocellum</italic>. The results indicated that the whole genome sequence of <italic>R. thermocellum</italic> M3 consisted of 3.6&#x2009;Mb of chromosomes with a 38.9% of GC%. To be specific, eight gene islands and 271 carbohydrate-active enzyme-encoded proteins were detected. Moreover, the results of gene function annotation showed that 2,071, 2,120, and 1,246 genes were annotated into the Clusters of Orthologous Groups (COG), Gene Ontology (GO), and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases, respectively, and most of the genes were involved in carbohydrate metabolism and enzymatic catalysis. Different from other <italic>R. thermocellum</italic>, strain M3 has three proteins related to &#x03B2;-glucosidase, and the cellobiose hydrolysis was enhanced by the synergy of gene <italic>BglA</italic> and <italic>BglX</italic>. Meanwhile, the GH42 family, CBM36 family, and AA8 family might participate in cellobiose degradation.</p>
</abstract>
<kwd-group>
<kwd>thermocellum</kwd>
<kwd>genome</kwd>
<kwd>cellobiose</kwd>
<kwd><bold>&#x03B2;</bold>-glucosidase</kwd>
<kwd>CAZyme</kwd>
</kwd-group>
<contract-num rid="cn1">51908200</contract-num>
<contract-num rid="cn2">UNPYSCT-2020028</contract-num>
<contract-num rid="cn3">2021-KYYWF-1465</contract-num>
<contract-sponsor id="cn1">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<contract-sponsor id="cn2">Heilongjiang Provincial College Youth Innovation Talent Project</contract-sponsor>
<contract-sponsor id="cn3">Basic Scientific Research Operating Expenses</contract-sponsor>
<counts>
<fig-count count="6"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="58"/>
<page-count count="11"/>
<word-count count="7492"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>1. Introduction</title>
<p>Lignocellulosic biomass was considered an ideal sustainable resource for potential feedstock for various value-added chemicals and biofuels (<xref ref-type="bibr" rid="ref18">Khare et al., 2015</xref>; <xref ref-type="bibr" rid="ref48">Usmani et al., 2021</xref>). As a result, the rational use of lignocellulosic biomass not only relieves the pressure of fossil energy shortage but also mitigates natural environment damage caused by improper treatment (<xref ref-type="bibr" rid="ref46">Staples et al., 2017</xref>). Converting lignocellulosic feedstocks into high-value chemical products includes three steps: pretreatment, saccharification, and fermentation (<xref ref-type="bibr" rid="ref52">Yadav et al., 2020</xref>). Saccharification refers to the hydrolysis of holocellulose into monosaccharides/oligosaccharides, which facilitates the downstream process. The saccharification of lignocellulose feedstocks was one of the bottlenecks of lignocellulosic biomass utilization (<xref ref-type="bibr" rid="ref10">Guo et al., 2018</xref>; <xref ref-type="bibr" rid="ref49">Usmani et al., 2020</xref>). Cellulosic biomass can be saccharified by acids or cellulases. Cellulose hydrolysis by cellulase is considered more environmentally friendly than cellulose hydrolysis by acid when conducted under mild conditions. The degradation of cellulose is achieved by the synergistic action of endoglucosidase, extranosidase, and &#x03B2;-glucosidase (<xref ref-type="bibr" rid="ref42">Sheng et al., 2016</xref>). Both bacteria and fungi can synthesize cellulase in nature; <italic>Trichoderma</italic> and <italic>Aspergillus</italic> sp. are known for their potential to produce cellulases, while these fungi lack a complete cellulase system, which leads to a decrease in the catalytic efficiency of cellulase (<xref ref-type="bibr" rid="ref45">Srivastava et al., 2018</xref>). Anaerobic bacteria degrade cellulose by synthesizing cellulosomes, which gather different cellulases in a narrow space and anchor them on the cell surface, the cellulosome of <italic>R. thermocellum</italic> is the most efficient cellulase system found at present (<xref ref-type="bibr" rid="ref31">Mazzoli and Olson, 2020</xref>).</p>
<p>However, the hydrolysis of lignocellulosic biomass by <italic>R. thermocellum</italic> has not been industrialized for the low activity of &#x03B2;-glucosidase, which is insufficient for the lignocellulosic feedstocks saccharification (<xref ref-type="bibr" rid="ref42">Sheng et al., 2016</xref>). Moreover, the feedback inhibition of exoglucanosidases induced by the accumulation of cellobiose severely reduced the catalytic efficiency of the cellulosome (<xref ref-type="bibr" rid="ref23">Lamed et al., 1985</xref>; <xref ref-type="bibr" rid="ref47">Tian et al., 2016</xref>; <xref ref-type="bibr" rid="ref12">Haldar and Purkait, 2020</xref>). The addition of exogenous &#x03B2;-glucosidase is a way to improve the hydrolysis efficiency of cellulosome, but additional &#x03B2;-glucosidase directly increases the cost and complexity of the saccharification process. Accordingly, enhancing the activity of &#x03B2;-glucosidase activity of wild <italic>R. thermocellum</italic> by building recombinant strains that secrete large amounts of &#x03B2;-glucosidase was a promising solution. Meki et al. fused the <italic>E. coli</italic> plasmid containing cloned bglA into wild <italic>R. thermocellum</italic> ATCC 27405 to construct a recombinant strain <italic>R. thermocellum</italic> ATCC 27405 (+<italic>McbglA</italic>), the result indicated that the &#x03B2;-glucosidase activity expressed by recombinant strain was 2.3 times higher than that of the wild strain at the late logarithmic growth stage (<xref ref-type="bibr" rid="ref30">Maki et al., 2013</xref>). Waeonukul et al. found that the addition of <italic>BglB</italic> from <italic>R. thermocellum</italic> S14 to cellulosome was observed to increase the saccharification rate of cellulose compared to that of Novazyme-188 and cellulosome alone (<xref ref-type="bibr" rid="ref50">Waeonukul et al., 2012</xref>). Nevertheless, the activity and stability of &#x03B2;-glucosidase secreted by recombinant strain decreased during the hydrolysis process (<xref ref-type="bibr" rid="ref57">Zhang et al., 2017</xref>). Isiam et al. found that when cellobiose was used as a carbon source, seven cellulosome structural proteins, 31 cellulosome-related glycosidases, and 19 non-cellulosome glycoside hydrolases were expressed in <italic>R. thermocellum</italic>, which suggests that the degradation of cellobiose by <italic>R. thermocellum</italic> was not only related to the &#x03B2;-glucosidase gene but also associated with genes other than &#x03B2;-glucosidase (<xref ref-type="bibr" rid="ref17">Islam et al., 2006</xref>). Therefore, it is particularly important to find strains with a stable ability to degrade cellobiose and reveal the genes involved in stable cellobiose degradation by <italic>R. thermocellum</italic>.</p>
<p>In previous studies, we isolated an <italic>R. thermocellum</italic> M3 that can efficiently degrade lignocellulosic biomass from horse manure. Different from other <italic>R. thermocellum</italic>, 97% of the cellulosic saccharification products of <italic>R. thermocellum</italic> M3 were monosaccharides. More importantly, <italic>R. thermocellum</italic> M3 inherited the ability of cellobiose degradation stably that conducts <italic>R. thermocellum</italic> M3 being an excellent sample for stable expression of exogenous &#x03B2;-glucosidase in the genus of <italic>R. thermocellum</italic>. In this study, we report the whole genome sequence of <italic>R. thermocellum</italic> M3 and compare the high-quality complete genome sequence of <italic>R. thermocellum</italic> M3 with both intra- and inter-generically to those of its close or distant phylogenetic relatives. Moreover, genes related to cellobiose degradation were comprehensively analyzed.</p>
</sec>
<sec id="sec2" sec-type="materials|methods">
<title>2. Materials and methods</title>
<sec id="sec3">
<title>2.1. Bacterial strain and cultivation</title>
<p><italic>Ruminiclostridium thermocellum</italic> M3 strain was isolated and enriched from horse manure by <xref ref-type="bibr" rid="ref42">Sheng et al. (2016)</xref> and deposited in the Microbiology Laboratory of the School of Environment and Chemical Engineering, Heilongjiang University of Science and Technology. The seed was stored in a constant temperature incubator at &#x2212;20&#x00B0;C and cultivated in an anaerobic bottle (filled with nitrogen) containing modified ATCC 1191 (MA) medium. The main components of culture medium were K<sub>2</sub>HPO<sub>4,</sub> 1.5&#x2009;g/L; MgSO<sub>4</sub>&#x00B7;7H<sub>2</sub>O, 0.2&#x2009;g/L; (NH<sub>4</sub>)<sub>2</sub>SO<sub>4,</sub> 1.0&#x2009;g KCl, 0.2&#x2009;g/L; L-cysteine, 0.5&#x2009;g/L; KH<sub>2</sub>PO<sub>4,</sub> 3.0&#x2009;g/L; CaCl<sub>2</sub>&#x2022;2H<sub>2</sub>O, 0.025&#x2009;g/L; NaCl, 1.0&#x2009;g/L; Yeast, 1.5&#x2009;g/L; and Avicel, 5.0&#x2009;g/L. The temperature of the culture was maintained at 60&#x00B0;C at 120&#x2009;rpm.</p>
</sec>
<sec id="sec4">
<title>2.2. DNA extraction and whole genome sequencing</title>
<p>The <italic>R. thermocellum</italic> samples for whole genome sequencing analysis were cultured in MA medium for 24&#x2009;h at 60&#x00B0;C, centrifuged for 5&#x2009;min at 4&#x00B0;C, 12,000&#x2009;&#x00D7;&#x2009;<italic>g</italic>, then the cell pellet was washed twice with normal saline (NS), and crushed with the FastPrep-24 instrument in lysing matrix B tubes (MP Biomedical) for 40&#x2009;s to release the genomic DNA from cells. The extraction of genomic DNA was performed with the Tiangen bacterial DNA mini kit (Tiangen Biotech Co. Ltd., Beijing, China) according to the manufacturer&#x2019;s protocol. The harvested DNA was detected using agarose gel electrophoresis and then quantified by Qubit 4.0 (ThermoFisher, Q33226). The genomic sequencing was conducted by the Pacbio sequencing platform with <italic>de novo</italic> assembly (SMRT portal; <xref ref-type="bibr" rid="ref3">Berlin et al., 2015</xref>).</p>
</sec>
<sec id="sec5">
<title>2.3. Genome annotation and component prediction</title>
<p>The genome annotation and gene function were predicted in the Gene Ontology (GO) database, the Kyoto Encyclopedia of Genes and Genomes (KEGG) database, the Clusters of Orthologous Groups (COG) database, and the Non-Redundant Protein databases (NR). The whole genome BLAST search (E-value below 1<sup>e&#x2212;5</sup>, minimal alignment length percentage above 40%) was performed with the above four databases. The prediction of carbohydrate-active enzymes was conducted with the Carbohydrate-Active enZYmes Database (<xref ref-type="bibr" rid="ref29">Lombard et al., 2013</xref>).</p>
<p>Genome component prediction of <italic>R. thermocellum</italic> M3 strains included the coding gene, repetitive sequences, signal peptide, genomic islands, prophage, lipoprotein, prophage, and clustered regularly interspaced short palindromic repeat sequences (CRISPR). The GeneMarkS program was conducted to retrieve the coding genes (<ext-link xlink:href="http://topaz.gatech.edu/" ext-link-type="uri">http://topaz.gatech.edu/</ext-link>; <xref ref-type="bibr" rid="ref54">Yang et al., 2020</xref>). The tRNA was predicted by the Aragorn program (<xref ref-type="bibr" rid="ref24">Laslett and Canback, 2004</xref>), the rRNA was predicted by the RNAmmer program (<xref ref-type="bibr" rid="ref22">Lagesen et al., 2007</xref>), and the miscRNA was predicted by Infernal (v1.1.2; <xref ref-type="bibr" rid="ref35">Mostajo Berrospi et al., 2019</xref>). Repeat Modeler was used to predict the repeat sequence <italic>de novo</italic> of the assembly results, and the RepeatMasker program was used for identifying repetitive elements in nucleotide sequences.</p>
<p>Genomic islands were predicted by the Island Path-DIOMB program (<xref ref-type="bibr" rid="ref16">Hsiao et al., 2003</xref>). The CRISPR was predicted by the CRISPR recognition tool (CRT; <xref ref-type="bibr" rid="ref9">Grissa et al., 2007</xref>). Prophages were predicted using PhiSpy (<xref ref-type="bibr" rid="ref11">H&#x00E4;hnke et al., 2009</xref>). RepeatMasker was used to identify the location and frequency of repeats on the genome (<xref ref-type="bibr" rid="ref40">Saha et al., 2008</xref>). NCBI Blast+ was used to compare the protein sequences with CDD, KOG, COG, NR, NT, PFAM, Swissprot, TrEMBL, and other databases to obtain the functional annotation information.</p>
</sec>
<sec id="sec6">
<title>2.4. Whole genome-based comparative genomic analysis</title>
<p>The core genes, specific genes, the gene family phylogenetic tree, single nucleotide polymorphism (SNP), and genome visualization were analyzed to reveal the result of comparative genomics (<xref ref-type="bibr" rid="ref58">Zhong et al., 2018</xref>). The genome sequences of <italic>R. thermocellum</italic> DSM 2360, <italic>R. thermocellum</italic> ATCC 27405, <italic>R. thermocellum</italic> DSM 1313, and <italic>R. thermocellum</italic> AD2 were obtained from the NCBI database. Genomic alignments among <italic>R. thermocellum</italic> M3 and other genomes of <italic>R. thermocellum</italic> were performed using the MUMmer and LASTZ tools (<xref ref-type="bibr" rid="ref21">Kurtz et al., 2004</xref>). Core genes and specific genes were analyzed using the CDHIT rapid clustering of similar proteins software with a threshold of 50% pairwise identity and a 0.7 length difference cutoff for amino acids (<xref ref-type="bibr" rid="ref27">Li et al., 2001</xref>, <xref ref-type="bibr" rid="ref28">2002</xref>; <xref ref-type="bibr" rid="ref26">Li and Godzik, 2006</xref>). The relationships between five <italic>R. thermocellum</italic> strains were analyzed, and the results were represented in a Venn diagram. NCBI Blast+ was used to compare the predicted 16S rRNA sequence with the NCBI 16S database to obtain its homologous strain information and a phylogenetic tree was constructed using mega software (<xref ref-type="bibr" rid="ref13">Hall, 2013</xref>).</p>
</sec>
</sec>
<sec id="sec7" sec-type="results">
<title>3. Results</title>
<sec id="sec8">
<title>3.1. Feature of the whole genome of <italic>Ruminiclostridium thermocellum</italic> M3 genome</title>
<p>The whole genome of <italic>R. thermocellum</italic> M3 was sequenced and analyzed with regard to the predictions of coding genes. The identified total size of the genome <italic>R. thermocellum</italic> M3 was 3,602,270&#x2009;bp with 39% GC content using SPAdes, and the number of coding genes was 3,195 with an average gene length of 973.56&#x2009;bp (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S1</xref>). The major characteristics of the genome of the <italic>R. thermocellum</italic> M3 and four <italic>R. thermocellum</italic> are summarized to acquire more perceptions concerning the genetic information (<xref rid="tab1" ref-type="table">Table 1</xref>). The final assemblies indicated that the genomes of five <italic>R. thermocellum</italic> were similar in size and G&#x2009;+&#x2009;C contents, genome annotation yielded 3,062, 3,196, 2,949, 2,959, and 3,077 genes for strains DSM 2360, ATCC 27405, DSM 1313, AD2, and M3, respectively.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>General features of <italic>Ruminiclostridium thermocellum</italic> M3 genome and comparison with other closely related species.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Features</th>
<th align="center" valign="top">DSM 2360</th>
<th align="center" valign="top">ATCC 27405</th>
<th align="center" valign="top">DSM 1313</th>
<th align="center" valign="top">AD2</th>
<th align="center" valign="top">M3</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Genome Size (Mb)</td>
<td align="center" valign="top">3.57758</td>
<td align="center" valign="top">3.8433</td>
<td align="center" valign="top">3.56162</td>
<td align="center" valign="top">3.55485</td>
<td align="center" valign="top">3.60227</td>
</tr>
<tr>
<td align="left" valign="top">G&#x2009;+&#x2009;C content (%)</td>
<td align="center" valign="top">39.2</td>
<td align="center" valign="top">39</td>
<td align="center" valign="top">39.1</td>
<td align="center" valign="top">39.2</td>
<td align="center" valign="top">38.9</td>
</tr>
<tr>
<td align="left" valign="top">No. of Scaffolds</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">1</td>
</tr>
<tr>
<td align="left" valign="top">No. of CDS</td>
<td align="center" valign="top">3,026</td>
<td align="center" valign="top">3,196</td>
<td align="center" valign="top">2,949</td>
<td align="center" valign="top">2,959</td>
<td align="center" valign="top">3,077</td>
</tr>
<tr>
<td align="left" valign="top">No. of rRNA operons</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top">12</td>
</tr>
<tr>
<td align="left" valign="top">No. of tRNA operons</td>
<td align="center" valign="top">56</td>
<td align="center" valign="top">56</td>
<td align="center" valign="top">56</td>
<td align="center" valign="top">56</td>
<td align="center" valign="top">56</td>
</tr>
<tr>
<td align="left" valign="top">ANI (%)</td>
<td align="center" valign="top">99.65</td>
<td align="center" valign="top">99.65</td>
<td align="center" valign="top">99.58</td>
<td align="center" valign="top">99.65</td>
<td align="center" valign="top">100</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The result of genome annotation indicated that a series of genes encoding virulence factors (VFDB), antibiotic resistance (CARD), and pathogen-host interactions (PHI-base) were present in <italic>R. thermocellum</italic> (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>). Meanwhile, <italic>R. thermocellum</italic> strains had several genes related to carbohydrate degradation, biosynthesis, and modification enzymes. The whole-genome visualization map clearly identified the composition, location, and function of the <italic>R. thermocellum</italic> M3 genome, which is demonstrated in <xref rid="fig1" ref-type="fig">Figure 1</xref>, with 100% coverage of sequencing. Detailed information includes analysis of CDS, non-coding RNA, COG, functional classification of genes, and the size of genome and GC%. Additionally, among the 3,077 predicted genes of <italic>R. thermocellum</italic> M3, only 2,071 CDSs were assigned any COG, which accounted for 67.3% of the predicted genes, and an additional 171 CDSs were assigned to group S with an unknown function. There are 137 genes involved in amino acid production and transformation (group E) in the aligned COG database genes, accounting for 6.62% of the COG annotations; the number of carbohydrate transport and metabolism-related genes (group G) was 130, accounting for 6.28% of the COG-annotated genes; the quantity of energy production and conversion-related genes (group C) was 107, taking up 5.17% of the COG-annotated genes; and only 16 genes fall into group Q (secondary metabolites biosynthesis, transport, and catabolism; see <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S2</xref>).</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Whole-genome visualization map of <italic>Ruminiclostridium thermocellum</italic> M3.</p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g001.tif"/>
</fig>
</sec>
<sec id="sec9">
<title>3.2. Genome functional studies of <italic>Ruminiclostridium thermocellum</italic> M3</title>
<p>Gene Ontology and KEGG databases were used to obtain the elucidation of the &#x201C;character&#x201D; of the coding gene in the bacteria from a macroscopic perspective. In the genome of <italic>R. thermocellum</italic> M3, 2021 genes were annotated in the GO database (<xref rid="tab2" ref-type="table">Table 2</xref>). In the biological process, the metabolic and cellular processes account for the highest proportion; in molecular function, it is mainly related to catalytic activity and binding; in the cellular component, it is closely compared to cells and cell parts, which indicated that strain M3 has more proteins involved in metabolism, cell composition, and enzymatic catalysis. These results are in accordance with the biochemical characteristics of M3, which possesses an outstanding catalytic capacity for cellulose substrate. In the KEGG database, 1,246 genes were annotated, which can be divided into five branches according to the metabolic pathways involved in the genes: cellular processes, environmental information processing, genetic information processing, metabolism, and organismal systems. As shown in <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S3</xref>, aligning these annotated genes into metabolic pathways yielded a total of 147 metabolic pathways. It can be clearly seen that the predominant pathway was carbohydrate metabolism and overview with 329 and 229 unigenes, followed by amino acid metabolism and energy metabolism with 197 and 153 unigenes.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Gene function analysis of <italic>Ruminiclostridium thermocellum</italic> M3 based on Gene Ontology (GO) annotation.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Ontology</th>
<th align="left" valign="top">GO-ID</th>
<th align="left" valign="top">Term</th>
<th align="center" valign="top">Gene-Num</th>
<th align="center" valign="top">Ratio</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="5">Molecular function</td>
<td align="left" valign="top">GO:0003824</td>
<td align="left" valign="top">Catalytic activity</td>
<td align="left" valign="top">1,432</td>
<td align="left" valign="top">46.54</td>
</tr>
<tr>
<td align="left" valign="top">GO:0005488</td>
<td align="left" valign="top">Binding</td>
<td align="left" valign="top">1,293</td>
<td align="left" valign="top">42.02</td>
</tr>
<tr>
<td align="left" valign="top">GO:0005215</td>
<td align="left" valign="top">Transporter activity</td>
<td align="left" valign="top">147</td>
<td align="left" valign="top">4.78</td>
</tr>
<tr>
<td align="left" valign="top">GO:0060089</td>
<td align="left" valign="top">Molecular transducer activity</td>
<td align="left" valign="top">106</td>
<td align="left" valign="top">3.44</td>
</tr>
<tr>
<td align="left" valign="top">GO:0001071</td>
<td align="left" valign="top">Nucleic acid binding transcription factor activity</td>
<td align="left" valign="top">73</td>
<td align="left" valign="top">2.37</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="6">Cellular component</td>
<td align="left" valign="top">GO:0005623</td>
<td align="left" valign="top">Cell</td>
<td align="left" valign="top">1,035</td>
<td align="left" valign="top">33.64</td>
</tr>
<tr>
<td align="left" valign="top">GO:0044464</td>
<td align="left" valign="top">Cell part</td>
<td align="left" valign="top">1,035</td>
<td align="left" valign="top">33.64</td>
</tr>
<tr>
<td align="left" valign="top">GO:0016020</td>
<td align="left" valign="top">Membrane</td>
<td align="left" valign="top">482</td>
<td align="left" valign="top">15.66</td>
</tr>
<tr>
<td align="left" valign="top">GO:0044425</td>
<td align="left" valign="top">Membrane part</td>
<td align="left" valign="top">363</td>
<td align="left" valign="top">11.8</td>
</tr>
<tr>
<td align="left" valign="top">GO:0043226</td>
<td align="left" valign="top">Organelle</td>
<td align="left" valign="top">141</td>
<td align="left" valign="top">4.58</td>
</tr>
<tr>
<td align="left" valign="top">GO:0032991</td>
<td align="left" valign="top">Macromolecular complex</td>
<td align="left" valign="top">118</td>
<td align="left" valign="top">3.83</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="8">Biological process</td>
<td align="left" valign="top">GO:0008152</td>
<td align="left" valign="top">Metabolic process</td>
<td align="left" valign="top">1,468</td>
<td align="left" valign="top">47.71</td>
</tr>
<tr>
<td align="left" valign="top">GO:0009987</td>
<td align="left" valign="top">Cellular process</td>
<td align="left" valign="top">1,419</td>
<td align="left" valign="top">46.12</td>
</tr>
<tr>
<td align="left" valign="top">GO:0044699</td>
<td align="left" valign="top">Single-organism process</td>
<td align="left" valign="top">597</td>
<td align="left" valign="top">19.4</td>
</tr>
<tr>
<td align="left" valign="top">GO:0050896</td>
<td align="left" valign="top">Response to stimulus</td>
<td align="left" valign="top">313</td>
<td align="left" valign="top">10.17</td>
</tr>
<tr>
<td align="left" valign="top">GO:0065007</td>
<td align="left" valign="top">Biological regulation</td>
<td align="left" valign="top">245</td>
<td align="left" valign="top">7.96</td>
</tr>
<tr>
<td align="left" valign="top">GO:0050789</td>
<td align="left" valign="top">Regulation of biological process</td>
<td align="left" valign="top">235</td>
<td align="left" valign="top">7.64</td>
</tr>
<tr>
<td align="left" valign="top">GO:0051179</td>
<td align="left" valign="top">Localization</td>
<td align="left" valign="top">200</td>
<td align="left" valign="top">6.5</td>
</tr>
<tr>
<td align="left" valign="top">GO:0051234</td>
<td align="left" valign="top">Establishment of localization</td>
<td align="left" valign="top">176</td>
<td align="left" valign="top">5.72</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Meanwhile, the predicted numbers of SignaIP-TM and SignaIP-noTM were 24 and 106, respectively, among the total Signa proteins. A total of four CRISPR arrays were found by CRT (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S2</xref>). The number of repeat counts in four arrays was 51, 90, 134, and 144. Moreover, several cas or cas-like genes were found in their neighborhoods. This suggests that R<italic>. thermocellum</italic> M3 has a defense against phage contamination, as CRISPR is very important in prokaryotes and is involved in resisting foreign phages and plasmids and recognizing and silencing invading functional elements. The results of gene-island (GI) prediction obtained eight GI with an average G&#x2009;+&#x2009;C content of about 36.4%, which is slightly lower than the G&#x2009;+&#x2009;C content of the M3 genome. It is worth noting that the G&#x2009;+&#x2009;C content of GI5 (33.7%) is significantly different than that of other GI, which indicates GI5 may be an exogenous sequence by horizontal transfer. The main components of the exogenous sequence are the transposase protein and the hypothetical protein (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S3</xref>).</p>
</sec>
<sec id="sec10">
<title>3.3. Comparative genomics of <italic>Ruminiclostridium thermocellum</italic></title>
<p>To reveal the genetic and evolutionary relationships between <italic>R. thermocellum</italic> M3 and other typical <italic>R. thermocellum</italic>, the analysis in view of the core-pan gene of the whole genome sequence was conducted. The Gene family boxplot (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S4</xref>) revealed that the pan-genome trend of <italic>R. thermocellum</italic> is open with the increasing number of <italic>R. thermocellum</italic> strains sequenced. The open pan-genome indicated that there is a significant capacity for discovering novel genes with the evolution and development of strains. Among the five <italic>R. thermocellum</italic> strains, the number of genes in the pan-genome is 3,042 and includes a core gene set (2,544 genes; <xref rid="fig2" ref-type="fig">Figure 2</xref>). Dispensable genes and unique genes existed across all genomes of five <italic>R. thermocellum</italic> strains. Among them, strain ATCC 27405 had the largest number of specific genes (267 genes), followed by strain M3 (192 genes). The high resemblance of five <italic>R. thermocellum</italic> strains was reflected by the large proportion of core genes.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Venn diagram of core and specific genes among five <italic>R. thermocellum</italic> strains. Each circle represents an <italic>Ruminiclostridium thermocellum</italic> strain. The number of orthologous coding sequences (core genome) shared by all strains is shown in the center circle, and the number of specific genes is shown in non-overlapping portions of each oval.</p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g002.tif"/>
</fig>
<p>To identify the functional classes of the <italic>R. thermocellum</italic> pan-genome, the COG database was used to classify the functional genes. In the pan-genome of <italic>R. thermocellum</italic>, the most core, dispensable, and specific gene clusters fell in the metabolism category (<xref rid="fig3" ref-type="fig">Figure 3</xref>). Meanwhile, the results indicated that both gene clusters for general function prediction and unknown functions were also more abundant. Compared with the dispensable and specific gene clusters, the majority of genes in the core gene clusters were involved in translation, ribosomal structure, biogenesis (J), cell wall/membrane/envelope biogenesis (M), carbohydrate transport and metabolism (G), amino acid transport, and metabolism (E). By comparison, the majority of genes in the no-core gene clusters was concerned with housekeeping functions, for example, replication, recombination and repair (L), and defense mechanisms (V). Single nucleotide polymorphism (SNP) represents the variation situation of bacteria in the evolutionary process. We constructed the phylogenetic tree based on SNP, which revealed the similarity in bacterial strain variation in adaptation to the natural environment. Compared with <italic>R. thermocellum</italic> ATCC 27405, M3 showed stronger evolutionary relationships with the other three <italic>R. thermocellum</italic> strains (<xref rid="fig4" ref-type="fig">Figure 4</xref>). Orthologous protein linear analysis based on MCScanX software was performed to further understand the differences in protein homology and amino acid arrangement between M3 and the other four <italic>R. thermocellum</italic> strains (<xref rid="fig5" ref-type="fig">Figure 5</xref>). Between M3 and other three <italic>R. thermocellum</italic> strains (DSM 1313, DSM 2360 and AD2), it was found that the proteins are not only homologous but also have a good linear relationship in sequence, while <italic>R. thermocellum</italic> ATCC 27405 was related to M3 but showed a large number of inversions. At the same time, we performed a genome-linear analysis of M3 and <italic>R. thermocellum</italic> ATCC 27405 (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S5</xref>). The results were similar to the results of the orthologous protein analysis, and it was more clearly evident that nearly half of the genes were inversions.</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>Distribution of core, dispensable, and specific genes on the Clusters of Orthologous Groups (COG) category of <italic>Ruminiclostridium thermocellum</italic> M3.</p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g003.tif"/>
</fig>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption>
<p>Phylogenetic relationship of <italic>Ruminiclostridium thermocellum</italic> M3 and other four <italic>R. thermocellum</italic> strains. Numbers along branches indicate bootstrap values with 1,000 times.</p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g004.tif"/>
</fig>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption>
<p>Plot of protein linear analysis between <italic>Ruminiclostridium thermocellum</italic> M3 and other four <italic>R. thermocellum.</italic></p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g005.tif"/>
</fig>
</sec>
<sec id="sec11">
<title>3.4. Protein-encoding genes related to the CAZyme system of <italic>Ruminiclostridium thermocellum M3</italic></title>
<p>The results of the CAZyme system analysis indicated that the multi-modular enzyme system of <italic>R. thermocellum</italic> M3 consisted of 75 dockerin and eight cohesin, which construct scaffolding structural proteins offering a large number of binding sites for cellulase. The quantity of enzyme protein and composition across dissimilar CAZy families in <italic>R. thermocellum</italic> M3 were analyzed and compared to those in the other four <italic>R. thermocellum</italic> to evaluate the inclination for lignocellulose saccharification. In the genome of <italic>R. thermocellum</italic>, most genes fell into glycoside hydrolases (GHs), carbohydrate binding molecules (CBMs), and glycosyltransferases (GTs), whereas a few genes were annotated in carbohydrate esterases (CEs), polysaccharide lyases (PLs), and auxiliary activities (AAs). To be specific, 271 cellulase proteins were detected in <italic>R. thermocellum</italic> M3 using dbCAN2 (DIAMOND algorithm; <xref rid="tab3" ref-type="table">Table 3</xref>). Most proteins were detected to be GHs (108 candidates), with GH124 (<italic>n</italic>&#x2009;=&#x2009;71), GH9 (<italic>n</italic>&#x2009;=&#x2009;17), and GH5 (<italic>n</italic>&#x2009;=&#x2009;13) being the most abundant families. In addition, there were 75 carbohydrate-binding module (CBM) proteins, 55 glycosyl transferases (GTs) proteins, 22 carbohydrate esterases (CEs) proteins, seven polysaccharide lyases (PLs) proteins, and four auxiliary activities (AAs) proteins. It is worth noting that there were nine CBM genes and 24 GH genes that had not been identified in <italic>R. thermocellum</italic> before (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S4</xref>).</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Comparison of the number of enzyme protein in <italic>Ruminiclostridium thermocellum</italic> M3 and other <italic>R. thermocellum</italic>.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">GH</th>
<th align="center" valign="top">CBM</th>
<th align="center" valign="top">GT</th>
<th align="center" valign="top">CE</th>
<th align="center" valign="top">PL</th>
<th align="center" valign="top">AA</th>
<th align="center" valign="top">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top"><italic>R. thermocellum</italic> DSM 2360</td>
<td align="center" valign="top">108</td>
<td align="center" valign="top">72</td>
<td align="center" valign="top">55</td>
<td align="center" valign="top">22</td>
<td align="center" valign="top">8</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">269</td>
</tr>
<tr>
<td align="left" valign="top"><italic>R. thermocellum</italic> ATCC 27405</td>
<td align="center" valign="top">102</td>
<td align="center" valign="top">70</td>
<td align="center" valign="top">56</td>
<td align="center" valign="top">22</td>
<td align="center" valign="top">7</td>
<td align="center" valign="top">3</td>
<td align="center" valign="top">260</td>
</tr>
<tr>
<td align="left" valign="top"><italic>R. thermocellum</italic> DSM 1313</td>
<td align="center" valign="top">110</td>
<td align="center" valign="top">75</td>
<td align="center" valign="top">54</td>
<td align="center" valign="top">22</td>
<td align="center" valign="top">8</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">273</td>
</tr>
<tr>
<td align="left" valign="top"><italic>R. thermocellum</italic> AD2</td>
<td align="center" valign="top">108</td>
<td align="center" valign="top">72</td>
<td align="center" valign="top">55</td>
<td align="center" valign="top">22</td>
<td align="center" valign="top">8</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">269</td>
</tr>
<tr>
<td align="left" valign="top"><italic>R. thermocellum</italic> M3</td>
<td align="center" valign="top">108</td>
<td align="center" valign="top">75</td>
<td align="center" valign="top">55</td>
<td align="center" valign="top">22</td>
<td align="center" valign="top">7</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">271</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>GH, glycoside hydrolase; GT, glycoside transferase; AA, auxiliary activity; CE, carbohydrate esterase; PL, polysaccharide lyase; and CBM, carbohydrate-binding module. All CAZy domains were identified using dbCAN2 (DIAMOND algorithm).</p>
</table-wrap-foot>
</table-wrap>
<p>Four encoded proteins associated with cellobiose degradation were found among the cellulase-encoded proteins of strain M3 by Conserved Domain Database (CDD) analysis (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S6</xref>), including <italic>BglA</italic> (Protein: PROKKA_02182), <italic>BglB</italic> (Protein: PROKKA_00862), and <italic>BglX</italic> (Protein: PROKKA_01050 and PROKKA_02059). <italic>BglA</italic> encoded glycoside hydrolase family 1 (GH1) protein (&#x03B2;-glucosidase). <italic>BglB</italic> encoded glycoside hydrolase family 1 (GH1) protein (6-phospho-&#x03B2;-glucosidase). <italic>BglX</italic> (PROKKA_01050) encoded glycoside hydrolase family 3 (GH3) protein (&#x03B2;-glucosidase). <italic>BglX</italic> (PROKKA_02059) encoded glycoside hydrolase family 3 (GH3) protein (&#x03B2;-glucosidase). The result of protein sequence alignment indicated that the ratio of similarity of pairwise comparison of proteins encoded by three genes (PROKKA_02182, PROKKA_01050, and PROKKA_02059), which correlated with &#x03B2;-glucosidase, were lower than 28.39%. To further demonstrate the diversity among the three <italic>Bgls</italic>, we constructed phylogenetic trees of <italic>Bgls</italic> from different strains by MEGA software. Not surprisingly, the phylogenetic tree based on &#x03B2;-glucosidase showed that each of the three <italic>Bgls</italic> was located in one of three different clusters (<xref rid="fig6" ref-type="fig">Figure 6</xref>). Furthermore, a gene encoding &#x03B2;-galactosidase (PROKKA_03012) was present in the M3 genome, which was not found in the other wild <italic>R. thermocellum</italic> before.</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption>
<p>Phylogenetic tree of &#x03B2;-glucosidase from <italic>Ruminiclostridium thermocellum</italic> M3. Numbers along branches indicate bootstrap values with 1,000 times.</p>
</caption>
<graphic xlink:href="fmicb-13-1079279-g006.tif"/>
</fig>
<p>In addition, 11 coding proteins about ABC sugar transport protein and two coding proteins (PROKKA_01086 and PROKKA_02112) about cellobiose phosphorylase coming from glycoside hydrolase family 94 (GH94) were detected in the genome of <italic>R. thermocellum</italic> M3 (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S5</xref>). Two group coding proteins (MglA and UpgA) relating polysaccharide transportation (cellobiose, fructose, and arabinose) ware found in 11 ABC sugar transport proteins by CDD. MglA (PROKKA_00080 and PROKKA_01980) is the ATPase component of ABC sugar transporter, which mainly promotes sugar transport; the primary function of UpgA (PROKKA_01330 and PROKKA_02933) is the transmembrane of sugar substrates, which is a permease component of the ABC transporter. Meanwhile, two unique ABC sugar transport proteins (PROKKA_02931 and PROKKA_02932), which had not been reported in other <italic>R. thermocellum</italic>, were found in strain M3.</p>
</sec>
</sec>
<sec id="sec12" sec-type="discussions">
<title>4. Discussion</title>
<sec id="sec13">
<title>4.1. Genome specificity of cellobiose hydrolysis by <italic>Ruminiclostridium thermocellum</italic> M3</title>
<p><italic>Ruminiclostridium thermocellum</italic>, as a typical genus of thermophagic cellulosic-degrading bacteria, was a high-utility candidate in lignocellulosic biomass refinement by means of a powerful cellulosome system (<xref ref-type="bibr" rid="ref1">Akinosho et al., 2014</xref>; <xref ref-type="bibr" rid="ref42">Sheng et al., 2016</xref>). However, the activity of &#x03B2;-glucosidase (<italic>BglA</italic>) was low in wild-type <italic>R. thermocellum</italic> strains (<xref ref-type="bibr" rid="ref43">Shinoda et al., 2019</xref>), and due to the adsorption of cellulosome to cellulose, only a small fraction of <italic>BglA</italic> was available on cellulosome. In nature, the efficient decomposition of cellulosic biomass strongly depends on the fast adaptation of <italic>R. thermocellum</italic> to enable the regulation of its cellulosomal enzymes for a specific substrate composition. Compared to other <italic>R. thermocellum</italic>, such as genus <italic>R. thermocellum</italic> ATCC 27045 and AD 2, genus <italic>R. thermocellum</italic> M3 not only has high cellulosic saccharification ability but also possesses the properties of cellubiose-resistance, which contributes to M3 being a potential industrial option for cellulosic biofuels refine.</p>
<p>In general, the wild type <italic>R. thermocellum</italic> containing &#x03B2;-glucosidase A (<italic>BglA</italic>) and &#x03B2;-glucosidase B (<italic>BglB</italic>), gene <italic>BglX</italic> (&#x03B2;-glucosidase) was first found in <italic>R. thermocellum</italic> in this study. In general, &#x03B2;-glucosidase A was sensitive to temperature and thus inactivated under the optimum temperature of <italic>R. thermocellum</italic> (<xref ref-type="bibr" rid="ref56">Yoav et al., 2019</xref>). Different from &#x03B2;-glucosidase A, <italic>BglB</italic> encodes a novel thermostable &#x03B2;-glucosidase B, which is more thermally stable than &#x03B2;-glucosidase A, whereas the biosynthesis of &#x03B2;-glucosidase B is repressed by cellobiose. Moreover, the &#x03B2;-glucosidase B of <italic>R. thermocellum</italic> was inhibited by low glucose concentration (<xref ref-type="bibr" rid="ref50">Waeonukul et al., 2012</xref>). Different from <italic>BglA</italic> and <italic>BglB</italic>, the gene <italic>BglX</italic> (Protein:PROKKA_01050) may encode &#x03B2;-xylosidase or &#x03B2;-glucosidase, since <italic>R. thermocellum</italic> cannot utilize xylose as a carbon source as the previous report (<xref ref-type="bibr" rid="ref42">Sheng et al., 2016</xref>) suggests that the <italic>BglX</italic> (Protein: PROKKA_01050) gene mainly plays a role in the hydrolysis of cellobiose in strain M3 and that the <italic>BglX</italic> (PROKKA_01050) gene encodes an enzyme with &#x03B2;-glucosidase activity predominantly. It is reported that the &#x03B2;-glucosidase encoded by <italic>BglX</italic> was a periplasmic cellulase that hydrolysis cellobiose into glucose (<xref ref-type="bibr" rid="ref44">Souto et al., 2021</xref>). In addition, a signal peptide that anchors &#x03B2;-glucosidase <italic>BglX</italic> to the periplasmatic space was found in the conserved domain of <italic>BglX</italic>, which might help to reduce the cellobiose concentration in a restricted area of the cell surface (<xref ref-type="bibr" rid="ref44">Souto et al., 2021</xref>). Kim also found that the &#x03B2;-glucosidase of <italic>Aspergillus aculeatus</italic> can bind to the yeast surface without any modification (<xref ref-type="bibr" rid="ref19">Kim et al., 2013</xref>), which suggests that the &#x03B2;-glucosidase of strain M3 could be directly connected to the cell membrane to hydrolyze cellobiose without the involvement of the CBM module.</p>
<p>It is believed that the genome of <italic>R. thermocellum</italic> does not encode any &#x03B2;-galactosidase (GH42; <xref ref-type="bibr" rid="ref38">Ravachol et al., 2016</xref>). Surprisingly, GH42 &#x03B2;-galactosidase (PROKKA_03012) was found in the genome of <italic>R. thermocellum</italic> M3. Some studies found that in addition to &#x03B2;-glucosidase, &#x03B2;-1,4 glucosidic can also be cleaved by &#x03B2;-galactosidase (<xref ref-type="bibr" rid="ref34">Morita et al., 2008</xref>; <xref ref-type="bibr" rid="ref55">Yang et al., 2018</xref>). Therefore, GH42 might be another key factor in cellobiose hydrolysis by <italic>R. thermocellum</italic> M3. Similar to GH42, GH116 was not found in <italic>R. thermocellum</italic> in the previous study; the GH116 family was reported as a thermophilic &#x03B2;-d-glucosidase, which was found in animals, plants, archaea, and bacteria (<xref ref-type="bibr" rid="ref41">Sansenya et al., 2015</xref>; <xref ref-type="bibr" rid="ref39">Rohman et al., 2019</xref>). The GH116 protein from thermophilic bacterial was reported for high hydrolytic activity toward &#x03B2;-1,3- and &#x03B2;-1,4-linked gluco-oligosaccharides and 4-nitrophenyl &#x03B2;-D-glucopyranoside (4NPGlc) artificial substrate (<xref ref-type="bibr" rid="ref41">Sansenya et al., 2015</xref>); therefore, GH116 protein might be another key factor involved in the hydrolysis of cellobiose by <italic>R. thermocellum</italic> M3.</p>
<p>The combination of ABC transport protein and cellobiose phosphorylase is a common strategy concerning cellobiose transportation in <italic>R. thermocellum</italic> (<xref ref-type="bibr" rid="ref37">Parisutham et al., 2017</xref>). Five putative cellodextrin-specific ABC transporters, labeled as CbpA-D and Lbp, had been identified in <italic>R. thermocellum</italic> DSM 1313 (<xref ref-type="bibr" rid="ref36">Nataf et al., 2009</xref>); Yan et al. found that only CbpB plays a key role in cellobiose transport using the functional verification of CbpA-D by genetic inactivation (<xref ref-type="bibr" rid="ref53">Yan et al., 2022</xref>). We identified four ABC sugar transporters that specifically transport polysaccharides by gene annotation and conserved domain database analysis, but the high affinity for binding to cellobiose needs further proof in future works.</p>
</sec>
<sec id="sec14">
<title>4.2. Genome specificity of carbohydrate binding module of <italic>Ruminiclostridium thermocellum</italic></title>
<p>It is believed that the cellulase system of <italic>R. thermocellum</italic>-cellulosomes contains cellulose-binding modules (CBMs), which leads to catalytic activity varying greatly in different regions. As a result, there are higher local cellobiose concentrations at particular sites. However, the &#x03B2;-glucosidase of wild type <italic>R. thermocellum</italic> was not bound to the CBM which led to the ineffective hydrolysis of local cellobiose (<xref ref-type="bibr" rid="ref56">Yoav et al., 2019</xref>). Therefore, the construction of an artificial chimeric cohesion-containing scaffold in which binding the &#x03B2;-glucosidase to the cellulosome and mimicking the enzymatic synergism of native cellulosome systems was feasible to enhance the hydrolysis efficiency of cellobiose (<xref ref-type="bibr" rid="ref7">Fierobe et al., 2002</xref>, <xref ref-type="bibr" rid="ref8">2005</xref>; <xref ref-type="bibr" rid="ref33">Morais et al., 2010</xref>). Different from other <italic>R. thermocellum</italic> strains, nine CBM family genes were unique to strain M3, including CBM23, CBM36, CBM37, CBM40, CBM47, CBM53, CBM61, CBM70, and CBM75. It is worth noting that the CBM37s were initially discovered in <italic>R. albus</italic> on the basis of adhesion-defective strains that lacked specific surface proteins. The CBM37s were collectively shown to exhibit a broad specificity pattern, which indicated a mechanism for binding the parent enzymes to cellulosic substrates (<xref ref-type="bibr" rid="ref14">Himmel et al., 2010</xref>). It is reported that CBM37 was responsible for anchoring substrate with enzymes to the cell surface. For example, CBM37 was reported as the mode that binds cellobiohydrolase and endoglucanase to the bacterial cell surface by the C terminus of glycoside hydrolases (<xref ref-type="bibr" rid="ref5">Ezer et al., 2008</xref>). In addition, binding the substrate with enzymes (<xref ref-type="bibr" rid="ref14">Himmel et al., 2010</xref>) might enhance the linkage between the cellulosome and enzymes related to the cellobiose degradation of <italic>R. thermocellum</italic> M3.</p>
<p>Meanwhile, CBM13 and CBM15 were also found in <italic>R. thermocellum</italic> M3. CBM13s acquire a larger variety of carbohydrate binding specificities including endo-&#x03B2;-glucanase (EC 3.2.1.6; <xref ref-type="bibr" rid="ref6">Ferrer et al., 1996</xref>), &#x03B1;-galactosidase (EC 3.2.1.22; <xref ref-type="bibr" rid="ref15">Holan et al., 1993</xref>), and some other glycoside hydrolases (<xref ref-type="bibr" rid="ref51">White et al., 1995</xref>). Different from CBM13, CBM35 was reported to have conserved ligand specificity, which is often appended to plant cell wall-degrading enzymes (<xref ref-type="bibr" rid="ref32">Montanier et al., 2009</xref>) and xylan-degrading enzymes (<xref ref-type="bibr" rid="ref4">Chen et al., 1995</xref>), but it is often found in &#x03B2;-galactosidase (EC 3.2.1.23; <xref ref-type="bibr" rid="ref2">Ali-Ahmad et al., 2017</xref>). Therefore, it is suggested that the &#x03B2;-galactosidase of strain M3 is mainly related to CBM13 and CBM35 modules during the degradation of cellulosic feedstocks.</p>
</sec>
<sec id="sec15">
<title>4.3. Genome specificity of auxiliary activities of <italic>Ruminiclostridium thermocellum</italic></title>
<p>AA6 and AA8 families were first identified in <italic>R. thermocellum</italic> in this study. In general, members of the AA8 family (cellobiose dehydrogenase, CDH) can be isolated or appended to a CBM. Proteins contain iron reductase domains and may generate reactive oxygen species that could contribute to the non-enzymatic degradation of cellulose chains by the generation of highly reactive hydroxyl radicals (OH&#x2022;) <italic>via</italic> Fenton&#x2019;s reaction (<xref ref-type="bibr" rid="ref25">Levasseur et al., 2013</xref>). It is reported that the poor cellobiose availability of the substrate was a limiting factor to CDH activity; on the surface of <italic>R. thermocellum</italic>, cellobiose was accumulated in a restricted area, which was a sufficient substrate for the CDH. Good contact with CDH might play an important role in reducing the concentration of cellobiose in the environment (<xref ref-type="bibr" rid="ref20">Kittl et al., 2012</xref>) and potentially releasing the strain M3 from the feedback inhibition of cellobiose.</p>
</sec>
</sec>
<sec id="sec16" sec-type="conclusions">
<title>5. Conclusion</title>
<p>The genome of <italic>R. thermocellum</italic> M3 harbored a high level of genomic uniqueness compared to other wild <italic>R. thermocellum</italic>. The majority of genes of <italic>R. thermocellum</italic> M3 fell into GHs and CBMs. Moreover, some unique genes were found in <italic>R. thermocellum</italic> M3 which were not found in <italic>R. thermocellum</italic>, which belong to the Auxiliary Activity Family (AA), cellulose-binding modules Family (CBM), Carbohydrate Esterase Family (CE), Glycosyl Transferases Family (GT), and Glycoside Hydrolase Family (GH). The hydrolysis of cellobiose by <italic>R. thermocellum</italic> M3 was conducted by the synergy of BglA and BglX, which not only hydrolyze cellobiose but also increase the affinity of Bgl for cellulosic feedstocks, increasing catalytic activity. Meanwhile, the GH42, GH116, CBM37, and AA8 families might participate in the cellobiose degradation that released the <italic>R. thermocellum</italic> M3 from the feedback inhibition of cellobiose.</p>
</sec>
<sec id="sec17" sec-type="data-availability">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
</sec>
<sec id="sec18">
<title>Author contributions</title>
<p>ST, MQ, LZ, SC, LiL, and LiuL contributed jointly to all aspects of the work reported in the manuscript. ST and LZ designed the experiment. MQ performed the experiments. ST, MQ, LiL, SC, and LiuL contributed to the data analysis. ST and MQ drafted the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec19" sec-type="funding-information">
<title>Funding</title>
<p>This study was supported by the National Natural Science Foundation of China (No. 51908200), the Heilongjiang Provincial College Youth Innovation Talent Project (No. UNPYSCT-2020028), and the Basic Scientific Research Operating Expenses (No. 2021-KYYWF-1465).</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="sec21" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fmicb.2022.1079279/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fmicb.2022.1079279/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Akinosho</surname> <given-names>H.</given-names></name> <name><surname>Yee</surname> <given-names>K.</given-names></name> <name><surname>Close</surname> <given-names>D.</given-names></name> <name><surname>Ragauskas</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>The emergence of clostridium thermocellum as a high utility candidate for consolidated bioprocessing applications</article-title>. <source>Front. Chem.</source> <volume>2</volume>:<fpage>66</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fchem.2014.00066</pub-id>, PMID: <pub-id pub-id-type="pmid">25207268</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ali-Ahmad</surname> <given-names>A.</given-names></name> <name><surname>Garron</surname> <given-names>M.-L.</given-names></name> <name><surname>Zamboni</surname> <given-names>V.</given-names></name> <name><surname>Lenfant</surname> <given-names>N.</given-names></name> <name><surname>Nurizzo</surname> <given-names>D.</given-names></name> <name><surname>Henrissat</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Structural insights into a family 39 glycoside hydrolase from the gut symbiont Bacteroides cellulosilyticus WH2</article-title>. <source>J. Struct. Biol.</source> <volume>197</volume>, <fpage>227</fpage>&#x2013;<lpage>235</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jsb.2016.11.004</pub-id>, PMID: <pub-id pub-id-type="pmid">27890857</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berlin</surname> <given-names>K.</given-names></name> <name><surname>Koren</surname> <given-names>S.</given-names></name> <name><surname>Chin</surname> <given-names>C.-S.</given-names></name> <name><surname>Drake</surname> <given-names>J. P.</given-names></name> <name><surname>Landolin</surname> <given-names>J. M.</given-names></name> <name><surname>Phillippy</surname> <given-names>A. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Assembling large genomes with single-molecule sequencing and locality-sensitive hashing</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>623</fpage>&#x2013;<lpage>630</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nbt.3238</pub-id>, PMID: <pub-id pub-id-type="pmid">26006009</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Leipprandt</surname> <given-names>J. R.</given-names></name> <name><surname>Traviss</surname> <given-names>C. E.</given-names></name> <name><surname>Sopher</surname> <given-names>B. L.</given-names></name> <name><surname>Jones</surname> <given-names>M. Z.</given-names></name> <name><surname>Cavanagh</surname> <given-names>K. T.</given-names></name> <etal/></person-group>. (<year>1995</year>). <article-title>Molecular cloning and characterization of bovine &#x03B2;-mannosidase</article-title>. <source>J. Biol. Chem.</source> <volume>270</volume>, <fpage>3841</fpage>&#x2013;<lpage>3848</lpage>. doi: <pub-id pub-id-type="doi">10.1074/jbc.270.8.3841</pub-id>, PMID: <pub-id pub-id-type="pmid">7876128</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ezer</surname> <given-names>A.</given-names></name> <name><surname>Matalon</surname> <given-names>E.</given-names></name> <name><surname>Jindou</surname> <given-names>S.</given-names></name> <name><surname>Borovok</surname> <given-names>I.</given-names></name> <name><surname>Atamna</surname> <given-names>N.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Cell surface enzyme attachment is mediated by family 37 carbohydrate-binding modules, unique to Ruminococcus albus</article-title>. <source>J. Bacteriol.</source> <volume>190</volume>, <fpage>8220</fpage>&#x2013;<lpage>8222</lpage>. doi: <pub-id pub-id-type="doi">10.1128/JB.00609-08</pub-id>, PMID: <pub-id pub-id-type="pmid">18931104</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrer</surname> <given-names>P.</given-names></name> <name><surname>Halkier</surname> <given-names>T.</given-names></name> <name><surname>Hedegaard</surname> <given-names>L.</given-names></name> <name><surname>Savva</surname> <given-names>D.</given-names></name> <name><surname>Diers</surname> <given-names>I.</given-names></name> <name><surname>Asenjo</surname> <given-names>J. A.</given-names></name></person-group> (<year>1996</year>). <article-title>Nucleotide sequence of a beta-1,3-glucanase isoenzyme IIA gene of Oerskovia xanthineolytica LL G109 (Cellulomonas cellulans) and initial characterization of the recombinant enzyme expressed in Bacillus subtilis</article-title>. <source>J. Bacteriol.</source> <volume>178</volume>, <fpage>4751</fpage>&#x2013;<lpage>4757</lpage>. doi: <pub-id pub-id-type="doi">10.1128/jb.178.15.4751-4757.1996</pub-id>, PMID: <pub-id pub-id-type="pmid">8755914</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fierobe</surname> <given-names>H. P.</given-names></name> <name><surname>Bayer</surname> <given-names>E. A.</given-names></name> <name><surname>Tardif</surname> <given-names>C.</given-names></name> <name><surname>Czjzek</surname> <given-names>M.</given-names></name> <name><surname>Mechaly</surname> <given-names>A.</given-names></name> <name><surname>Belaich</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2002</year>). <article-title>Degradation of cellulose substrates by cellulosome chimeras - substrate targeting versus proximity of enzyme components</article-title>. <source>J. Biol. Chem.</source> <volume>277</volume>, <fpage>49621</fpage>&#x2013;<lpage>49630</lpage>. doi: <pub-id pub-id-type="doi">10.1074/jbc.M207672200</pub-id>, PMID: <pub-id pub-id-type="pmid">12397074</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fierobe</surname> <given-names>H. P.</given-names></name> <name><surname>Mingardon</surname> <given-names>F.</given-names></name> <name><surname>Mechaly</surname> <given-names>A.</given-names></name> <name><surname>Belaich</surname> <given-names>A.</given-names></name> <name><surname>Rincon</surname> <given-names>M. T.</given-names></name> <name><surname>Pages</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Action of designer cellulosomes on homogeneous versus complex substrates - controlled incorporation of three distinct enzymes into a defined trifunctional scaffoldin</article-title>. <source>J. Biol. Chem.</source> <volume>280</volume>, <fpage>16325</fpage>&#x2013;<lpage>16334</lpage>. doi: <pub-id pub-id-type="doi">10.1074/jbc.M414449200</pub-id>, PMID: <pub-id pub-id-type="pmid">15705576</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grissa</surname> <given-names>I.</given-names></name> <name><surname>Vergnaud</surname> <given-names>G.</given-names></name> <name><surname>Pourcel</surname> <given-names>C.</given-names></name></person-group> (<year>2007</year>). <article-title>CRISPRFinder: a web tool to identify clustered regularly interspaced short palindromic repeats</article-title>. <source>Nucleic Acids Res.</source> <volume>35</volume>, <fpage>W52</fpage>&#x2013;<lpage>W57</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkm360</pub-id>, PMID: <pub-id pub-id-type="pmid">17537822</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Chang</surname> <given-names>Y.</given-names></name> <name><surname>Lee</surname> <given-names>D.-J.</given-names></name></person-group> (<year>2018</year>). <article-title>Enzymatic saccharification of lignocellulosic biorefinery: research focuses</article-title>. <source>Bioresour. Technol.</source> <volume>252</volume>, <fpage>198</fpage>&#x2013;<lpage>215</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.biortech.2017.12.062</pub-id>, PMID: <pub-id pub-id-type="pmid">29329774</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>H&#x00E4;hnke</surname> <given-names>V.</given-names></name> <name><surname>Hofmann</surname> <given-names>B.</given-names></name> <name><surname>Proschak</surname> <given-names>E.</given-names></name> <name><surname>Steinhilber</surname> <given-names>D.</given-names></name> <name><surname>Schneider</surname> <given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>PhAST: pharmacophore alignment search tool</article-title>. <source>Chem. Cent. J.</source> <volume>3</volume>:<fpage>P67</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1752-153X-3-S1-P67</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haldar</surname> <given-names>D.</given-names></name> <name><surname>Purkait</surname> <given-names>M. K.</given-names></name></person-group> (<year>2020</year>). <article-title>Lignocellulosic conversion into value-added products: a review</article-title>. <source>Process Biochem.</source> <volume>89</volume>, <fpage>110</fpage>&#x2013;<lpage>133</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.procbio.2019.10.001</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hall</surname> <given-names>B. G.</given-names></name></person-group> (<year>2013</year>). <article-title>Building phylogenetic trees from molecular data with MEGA</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>1229</fpage>&#x2013;<lpage>1235</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/mst012</pub-id>, PMID: <pub-id pub-id-type="pmid">23486614</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Himmel</surname> <given-names>M. E.</given-names></name> <name><surname>Xu</surname> <given-names>Q.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Ding</surname> <given-names>S.-Y.</given-names></name> <name><surname>Lamed</surname> <given-names>R.</given-names></name> <name><surname>Bayer</surname> <given-names>E. A.</given-names></name></person-group> (<year>2010</year>). <article-title>Microbial enzyme systems for biomass conversion: emerging paradigms</article-title>. <source>Biofuels</source> <volume>1</volume>, <fpage>323</fpage>&#x2013;<lpage>341</lpage>. doi: <pub-id pub-id-type="doi">10.4155/bfs.09.25</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holan</surname> <given-names>Z. R.</given-names></name> <name><surname>Volesky</surname> <given-names>B.</given-names></name> <name><surname>Prasetyo</surname> <given-names>I.</given-names></name></person-group> (<year>1993</year>). <article-title>Biosorption of cadmium by biomass of marine algae</article-title>. <source>Biotechnol. Bioeng.</source> <volume>41</volume>, <fpage>819</fpage>&#x2013;<lpage>825</lpage>. doi: <pub-id pub-id-type="doi">10.1002/bit.260410808</pub-id>, PMID: <pub-id pub-id-type="pmid">18609626</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hsiao</surname> <given-names>W.</given-names></name> <name><surname>Wan</surname> <given-names>I.</given-names></name> <name><surname>Jones</surname> <given-names>S. J.</given-names></name> <name><surname>Brinkman</surname> <given-names>F. S. L.</given-names></name></person-group> (<year>2003</year>). <article-title>IslandPath: aiding detection of genomic islands in prokaryotes</article-title>. <source>Bioinformatics</source> <volume>19</volume>, <fpage>418</fpage>&#x2013;<lpage>420</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btg004</pub-id>, PMID: <pub-id pub-id-type="pmid">12584130</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Islam</surname> <given-names>R.</given-names></name> <name><surname>Cicek</surname> <given-names>N.</given-names></name> <name><surname>Sparling</surname> <given-names>R.</given-names></name> <name><surname>Levin</surname> <given-names>D.</given-names></name></person-group> (<year>2006</year>). <article-title>Effect of substrate loading on hydrogen production during anaerobic fermentation by clostridium thermocellum 27405</article-title>. <source>Appl. Microbiol. Biotechnol.</source> <volume>72</volume>, <fpage>576</fpage>&#x2013;<lpage>583</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00253-006-0316-7</pub-id>, PMID: <pub-id pub-id-type="pmid">16685495</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khare</surname> <given-names>S. K.</given-names></name> <name><surname>Pandey</surname> <given-names>A.</given-names></name> <name><surname>Larroche</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Current perspectives in enzymatic saccharification of lignocellulosic biomass</article-title>. <source>Biochem. Eng. J.</source> <volume>102</volume>, <fpage>38</fpage>&#x2013;<lpage>44</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.bej.2015.02.033</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Baek</surname> <given-names>S.-H.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Hahn</surname> <given-names>J.-S.</given-names></name></person-group> (<year>2013</year>). <article-title>Cellulosic ethanol production using a yeast consortium displaying a minicellulosome and &#x03B2;-glucosidase</article-title>. <source>Microb. Cell Factories</source> <volume>12</volume>:<fpage>14</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1475-2859-12-14</pub-id>, PMID: <pub-id pub-id-type="pmid">23383678</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kittl</surname> <given-names>R.</given-names></name> <name><surname>Kracher</surname> <given-names>D.</given-names></name> <name><surname>Burgstaller</surname> <given-names>D.</given-names></name> <name><surname>Haltrich</surname> <given-names>D.</given-names></name> <name><surname>Ludwig</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>Production of four Neurospora crassa lytic polysaccharide monooxygenases in Pichia pastoris monitored by a fluorimetric assay</article-title>. <source>Biotechnol. Biofuels</source> <volume>5</volume>:<fpage>79</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1754-6834-5-79</pub-id>, PMID: <pub-id pub-id-type="pmid">23102010</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kurtz</surname> <given-names>S.</given-names></name> <name><surname>Phillippy</surname> <given-names>A.</given-names></name> <name><surname>Delcher</surname> <given-names>A. L.</given-names></name> <name><surname>Smoot</surname> <given-names>M.</given-names></name> <name><surname>Shumway</surname> <given-names>M.</given-names></name> <name><surname>Antonescu</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>Versatile and open software for comparing large genomes</article-title>. <source>Genome Biol.</source> <volume>5</volume>:<fpage>R12</fpage>. doi: <pub-id pub-id-type="doi">10.1186/gb-2004-5-2-r12</pub-id>, PMID: <pub-id pub-id-type="pmid">14759262</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lagesen</surname> <given-names>K.</given-names></name> <name><surname>Hallin</surname> <given-names>P.</given-names></name> <name><surname>R&#x00F8;dland</surname> <given-names>E. A.</given-names></name> <name><surname>St&#x00E6;rfeldt</surname> <given-names>H.-H.</given-names></name> <name><surname>Rognes</surname> <given-names>T.</given-names></name> <name><surname>Ussery</surname> <given-names>D. W.</given-names></name></person-group> (<year>2007</year>). <article-title>RNAmmer: consistent and rapid annotation of ribosomal RNA genes</article-title>. <source>Nucleic Acids Res.</source> <volume>35</volume>, <fpage>3100</fpage>&#x2013;<lpage>3108</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkm160</pub-id>, PMID: <pub-id pub-id-type="pmid">17452365</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lamed</surname> <given-names>R.</given-names></name> <name><surname>Kenig</surname> <given-names>R.</given-names></name> <name><surname>Setter</surname> <given-names>E.</given-names></name> <name><surname>Bayer</surname> <given-names>E. A.</given-names></name></person-group> (<year>1985</year>). <article-title>Major characteristics of the cellulolytic system of clostridium thermocellum coincide with those of the purified cellulosome</article-title>. <source>Enzym. Microb. Technol.</source> <volume>7</volume>, <fpage>37</fpage>&#x2013;<lpage>41</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0141-0229(85)90008-0</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laslett</surname> <given-names>D.</given-names></name> <name><surname>Canback</surname> <given-names>B.</given-names></name></person-group> (<year>2004</year>). <article-title>ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>11</fpage>&#x2013;<lpage>16</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkh152</pub-id>, PMID: <pub-id pub-id-type="pmid">14704338</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levasseur</surname> <given-names>A.</given-names></name> <name><surname>Drula</surname> <given-names>E.</given-names></name> <name><surname>Lombard</surname> <given-names>V.</given-names></name> <name><surname>Coutinho</surname> <given-names>P. M.</given-names></name> <name><surname>Henrissat</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>Expansion of the enzymatic repertoire of the CAZy database to integrate auxiliary redox enzymes</article-title>. <source>Biotechnol. Biofuels</source> <volume>6</volume>:<fpage>41</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1754-6834-6-41</pub-id>, PMID: <pub-id pub-id-type="pmid">23514094</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Godzik</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences</article-title>. <source>Bioinformatics</source> <volume>22</volume>, <fpage>1658</fpage>&#x2013;<lpage>1659</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btl158</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Jaroszewski</surname> <given-names>L.</given-names></name> <name><surname>Godzik</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>Clustering of highly homologous sequences to reduce the size of large protein databases</article-title>. <source>Bioinformatics</source> <volume>17</volume>, <fpage>282</fpage>&#x2013;<lpage>283</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/17.3.282</pub-id>, PMID: <pub-id pub-id-type="pmid">11294794</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Jaroszewski</surname> <given-names>L.</given-names></name> <name><surname>Godzik</surname> <given-names>A.</given-names></name></person-group> (<year>2002</year>). <article-title>Tolerating some redundancy significantly speeds up clustering of large protein databases</article-title>. <source>Bioinformatics</source> <volume>18</volume>, <fpage>77</fpage>&#x2013;<lpage>82</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/18.1.77</pub-id>, PMID: <pub-id pub-id-type="pmid">11836214</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lombard</surname> <given-names>V.</given-names></name> <name><surname>Golaconda Ramulu</surname> <given-names>H.</given-names></name> <name><surname>Drula</surname> <given-names>E.</given-names></name> <name><surname>Coutinho</surname> <given-names>P. M.</given-names></name> <name><surname>Henrissat</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>The carbohydrate-active enzymes database (CAZy) in 2013</article-title>. <source>Nucleic Acids Res.</source> <volume>42</volume>, <fpage>D490</fpage>&#x2013;<lpage>D495</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkt1178</pub-id>, PMID: <pub-id pub-id-type="pmid">24270786</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maki</surname> <given-names>M. L.</given-names></name> <name><surname>Armstrong</surname> <given-names>L.</given-names></name> <name><surname>Leung</surname> <given-names>K. T.</given-names></name> <name><surname>Qin</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Increased expression of beta-glucosidase a in clostridium thermocellum 27405 significantly increases cellulase activity</article-title>. <source>Bioengineered</source> <volume>4</volume>, <fpage>15</fpage>&#x2013;<lpage>20</lpage>. doi: <pub-id pub-id-type="doi">10.4161/bioe.21951</pub-id>, PMID: <pub-id pub-id-type="pmid">22922214</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mazzoli</surname> <given-names>R.</given-names></name> <name><surname>Olson</surname> <given-names>D. G.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Chapter three&#x2014;clostridium thermocellum: a microbial platform for high-value chemical production from lignocellulose</article-title>&#x201D; in <source>Advances in Applied Microbiology</source>. eds. <person-group person-group-type="editor"><name><surname>Gadd</surname> <given-names>G. M.</given-names></name> <name><surname>Sariaslani</surname> <given-names>S.</given-names></name></person-group> (<publisher-loc>San Diego, Calif</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>111</fpage>&#x2013;<lpage>161</lpage>.</citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Montanier</surname> <given-names>C.</given-names></name> <name><surname>van Bueren</surname> <given-names>A. L.</given-names></name> <name><surname>Dumon</surname> <given-names>C.</given-names></name> <name><surname>Flint</surname> <given-names>J. E.</given-names></name> <name><surname>Correia</surname> <given-names>M. A.</given-names></name> <name><surname>Prates</surname> <given-names>J. A.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Evidence that family 35 carbohydrate binding modules display conserved specificity but divergent function</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>106</volume>, <fpage>3065</fpage>&#x2013;<lpage>3070</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.0808972106</pub-id>, PMID: <pub-id pub-id-type="pmid">19218457</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morais</surname> <given-names>S.</given-names></name> <name><surname>Barak</surname> <given-names>Y.</given-names></name> <name><surname>Caspi</surname> <given-names>J.</given-names></name> <name><surname>Hadar</surname> <given-names>Y.</given-names></name> <name><surname>Lamed</surname> <given-names>R.</given-names></name> <name><surname>Shoham</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Cellulase-xylanase synergy in designer Cellulosomes for enhanced degradation of a complex cellulosic substrate</article-title>. <source>MBio</source> <volume>1</volume>:<fpage>e00285-10</fpage>. doi: <pub-id pub-id-type="doi">10.1128/mBio.00285-10</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morita</surname> <given-names>T.</given-names></name> <name><surname>Ozawa</surname> <given-names>M.</given-names></name> <name><surname>Ito</surname> <given-names>H.</given-names></name> <name><surname>Kimio</surname> <given-names>S.</given-names></name> <name><surname>Kiriyama</surname> <given-names>S.</given-names></name></person-group> (<year>2008</year>). <article-title>Cellobiose is extensively digested in the small intestine by beta-galactosidase in rats</article-title>. <source>Nutrition</source> <volume>24</volume>, <fpage>1199</fpage>&#x2013;<lpage>1204</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.nut.2008.06.029</pub-id>, PMID: <pub-id pub-id-type="pmid">18752931</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Mostajo Berrospi</surname> <given-names>N.</given-names></name> <name><surname>Lataretu</surname> <given-names>M.</given-names></name> <name><surname>Krautwurst</surname> <given-names>S.</given-names></name> <name><surname>Mock</surname> <given-names>F.</given-names></name> <name><surname>Desir&#x00F2;</surname> <given-names>D.</given-names></name> <name><surname>Lamkiewicz</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2019</year>). A comprehensive annotation and differential expression analysis of short and long non-coding RNAs in 16 bat genomes. bioRxiv [Preprint]. doi: <pub-id pub-id-type="doi">10.1101/738526</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nataf</surname> <given-names>Y.</given-names></name> <name><surname>Yaron</surname> <given-names>S.</given-names></name> <name><surname>Stahl</surname> <given-names>F.</given-names></name> <name><surname>Lamed</surname> <given-names>R.</given-names></name> <name><surname>Bayer Edward</surname> <given-names>A.</given-names></name> <name><surname>Scheper</surname> <given-names>T.-H.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Cellodextrin and Laminaribiose ABC transporters in clostridium thermocellum</article-title>. <source>J. Bacteriol.</source> <volume>191</volume>, <fpage>203</fpage>&#x2013;<lpage>209</lpage>. doi: <pub-id pub-id-type="doi">10.1128/JB.01190-08</pub-id>, PMID: <pub-id pub-id-type="pmid">18952792</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parisutham</surname> <given-names>V.</given-names></name> <name><surname>Chandran</surname> <given-names>S. P.</given-names></name> <name><surname>Mukhopadhyay</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>S. K.</given-names></name> <name><surname>Keasling</surname> <given-names>J. D.</given-names></name></person-group> (<year>2017</year>). <article-title>Intracellular cellobiose metabolism and its applications in lignocellulose-based biorefineries</article-title>. <source>Bioresour. Technol.</source> <volume>239</volume>, <fpage>496</fpage>&#x2013;<lpage>506</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.biortech.2017.05.001</pub-id>, PMID: <pub-id pub-id-type="pmid">28535986</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ravachol</surname> <given-names>J.</given-names></name> <name><surname>de Philip</surname> <given-names>P.</given-names></name> <name><surname>Borne</surname> <given-names>R.</given-names></name> <name><surname>Mansuelle</surname> <given-names>P.</given-names></name> <name><surname>Mate</surname> <given-names>M. J.</given-names></name> <name><surname>Perret</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Mechanisms involved in xyloglucan catabolism by the cellulosome-producing bacterium Ruminiclostridium cellulolyticum</article-title>. <source>Sci. Rep.</source> <volume>6</volume>:<fpage>22770</fpage>. doi: <pub-id pub-id-type="doi">10.1038/srep22770</pub-id>, PMID: <pub-id pub-id-type="pmid">26946939</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rohman</surname> <given-names>A.</given-names></name> <name><surname>Dijkstra</surname> <given-names>B. W.</given-names></name> <name><surname>Puspaningsih</surname> <given-names>N. N. T.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x03B2;-Xylosidases: structural diversity, catalytic mechanism, and inhibition by monosaccharides</article-title>. <source>Int. J. Mol. Sci.</source> <volume>20</volume>:<fpage>5524</fpage>. doi: <pub-id pub-id-type="doi">10.3390/ijms20225524</pub-id>, PMID: <pub-id pub-id-type="pmid">31698702</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saha</surname> <given-names>S.</given-names></name> <name><surname>Bridges</surname> <given-names>S.</given-names></name> <name><surname>Magbanua</surname> <given-names>Z. V.</given-names></name> <name><surname>Peterson</surname> <given-names>D. G.</given-names></name></person-group> (<year>2008</year>). <article-title>Empirical comparison of ab initio repeat finding programs</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>2284</fpage>&#x2013;<lpage>2294</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkn064</pub-id>, PMID: <pub-id pub-id-type="pmid">18287116</pub-id></citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sansenya</surname> <given-names>S.</given-names></name> <name><surname>Mutoh</surname> <given-names>R.</given-names></name> <name><surname>Charoenwattanasatien</surname> <given-names>R.</given-names></name> <name><surname>Kurisu</surname> <given-names>G.</given-names></name> <name><surname>Ketudat Cairns</surname> <given-names>J. R.</given-names></name></person-group> (<year>2015</year>). <article-title>Expression and crystallization of a bacterial glycoside hydrolase family 116 [beta]-glucosidase from Thermoanaerobacterium xylanolyticum</article-title>. <source>Acta Crystallogr. Sect. F</source> <volume>71</volume>, <fpage>41</fpage>&#x2013;<lpage>44</lpage>. doi: <pub-id pub-id-type="doi">10.1107/S2053230X14025461</pub-id>, PMID: <pub-id pub-id-type="pmid">25615966</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sheng</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Gao</surname> <given-names>L. F.</given-names></name> <name><surname>Liu</surname> <given-names>W. Z.</given-names></name> <name><surname>Cui</surname> <given-names>M. H.</given-names></name> <name><surname>Guo</surname> <given-names>Z. C.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Lignocellulosic saccharification by a newly isolated bacterium, Ruminiclostridium thermocellum M3 and cellular cellulase activities for high ratio of glucose to cellobiose</article-title>. <source>Biotechnol. Biofuels</source> <volume>9</volume>:<fpage>172</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13068-016-0585-z</pub-id>, PMID: <pub-id pub-id-type="pmid">27525041</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shinoda</surname> <given-names>S.</given-names></name> <name><surname>Kurosaki</surname> <given-names>M.</given-names></name> <name><surname>Kokuzawa</surname> <given-names>T.</given-names></name> <name><surname>Hirano</surname> <given-names>K.</given-names></name> <name><surname>Takano</surname> <given-names>H.</given-names></name> <name><surname>Ueda</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Comparative biochemical analysis of Cellulosomes isolated from clostridium clariflavum DSM 19732 and clostridium thermocellum ATCC 27405 grown on plant biomass</article-title>. <source>Appl. Biochem. Biotechnol.</source> <volume>187</volume>, <fpage>994</fpage>&#x2013;<lpage>1010</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s12010-018-2864-6</pub-id>, PMID: <pub-id pub-id-type="pmid">30136170</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Souto</surname> <given-names>B. M.</given-names></name> <name><surname>de Araujo</surname> <given-names>A. C. B.</given-names></name> <name><surname>Hamann</surname> <given-names>P. R. V.</given-names></name> <name><surname>Bastos</surname> <given-names>A. R.</given-names></name> <name><surname>Cunha</surname> <given-names>I. S.</given-names></name> <name><surname>Peixoto</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Functional screening of a Caatinga goat (Capra hircus) rumen metagenomic library reveals a novel GH3 beta-xylosidase</article-title>. <source>PLoS One</source> <volume>16</volume>:<fpage>e0245118</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0245118</pub-id>, PMID: <pub-id pub-id-type="pmid">33449963</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srivastava</surname> <given-names>N.</given-names></name> <name><surname>Srivastava</surname> <given-names>M.</given-names></name> <name><surname>Mishra</surname> <given-names>P. K.</given-names></name> <name><surname>Gupta</surname> <given-names>V. K.</given-names></name> <name><surname>Molina</surname> <given-names>G.</given-names></name> <name><surname>Rodriguez-Couto</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Applications of fungal cellulases in biofuel production: advances and limitations</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>82</volume>, <fpage>2379</fpage>&#x2013;<lpage>2386</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rser.2017.08.074</pub-id></citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staples</surname> <given-names>M. D.</given-names></name> <name><surname>Malina</surname> <given-names>R.</given-names></name> <name><surname>Barrett</surname> <given-names>S. R. H.</given-names></name></person-group> (<year>2017</year>). <article-title>The limits of bioenergy for mitigating global life-cycle greenhouse gas emissions from fossil fuels</article-title>. <source>Nat. Energy</source> <volume>2</volume>:<fpage>16202</fpage>. doi: <pub-id pub-id-type="doi">10.1038/nenergy.2016.202</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>L.</given-names></name> <name><surname>Papanek</surname> <given-names>B.</given-names></name> <name><surname>Olson</surname> <given-names>D. G.</given-names></name> <name><surname>Rydzak</surname> <given-names>T.</given-names></name> <name><surname>Holwerda</surname> <given-names>E. K.</given-names></name> <name><surname>Zheng</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Simultaneous achievement of high ethanol yield and titer in clostridium thermocellum</article-title>. <source>Biotechnol. Biofuels</source> <volume>9</volume>:<fpage>116</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13068-016-0528-8</pub-id>, PMID: <pub-id pub-id-type="pmid">27257435</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Usmani</surname> <given-names>Z.</given-names></name> <name><surname>Sharma</surname> <given-names>M.</given-names></name> <name><surname>Awasthi</surname> <given-names>A. K.</given-names></name> <name><surname>Lukk</surname> <given-names>T.</given-names></name> <name><surname>Tuohy</surname> <given-names>M. G.</given-names></name> <name><surname>Gong</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Lignocellulosic biorefineries: the current state of challenges and strategies for efficient commercialization</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>148</volume>:<fpage>111258</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rser.2021.111258</pub-id></citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Usmani</surname> <given-names>Z.</given-names></name> <name><surname>Sharma</surname> <given-names>M.</given-names></name> <name><surname>Karpichev</surname> <given-names>Y.</given-names></name> <name><surname>Pandey</surname> <given-names>A.</given-names></name> <name><surname>Chander Kuhad</surname> <given-names>R.</given-names></name> <name><surname>Bhat</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Advancement in valorization technologies to improve utilization of bio-based waste in bioeconomy context</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>131</volume>:<fpage>109965</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rser.2020.109965</pub-id></citation></ref>
<ref id="ref50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Waeonukul</surname> <given-names>R.</given-names></name> <name><surname>Kosugi</surname> <given-names>A.</given-names></name> <name><surname>Tachaapaikoon</surname> <given-names>C.</given-names></name> <name><surname>Pason</surname> <given-names>P.</given-names></name> <name><surname>Ratanakhanokchai</surname> <given-names>K.</given-names></name> <name><surname>Prawitwong</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Efficient saccharification of ammonia soaked rice straw by combination of clostridium thermocellum cellulosome and Thermoanaerobacter brockii &#x03B2;-glucosidase</article-title>. <source>Bioresour. Technol.</source> <volume>107</volume>, <fpage>352</fpage>&#x2013;<lpage>357</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.biortech.2011.12.126</pub-id>, PMID: <pub-id pub-id-type="pmid">22257861</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>White</surname> <given-names>T.</given-names></name> <name><surname>Bennett</surname> <given-names>E. P.</given-names></name> <name><surname>Takio</surname> <given-names>K.</given-names></name> <name><surname>Sorensen</surname> <given-names>T.</given-names></name> <name><surname>Bonding</surname> <given-names>N.</given-names></name> <name><surname>Clausen</surname> <given-names>H.</given-names></name></person-group> (<year>1995</year>). <article-title>Purification and cDNA cloning of a human UDP-N-acetyl-alpha- D-galactosamine: polypeptide N-acetylgalactosaminyltransferase</article-title>. <source>J. Biol. Chem.</source> <volume>270</volume>, <fpage>24156</fpage>&#x2013;<lpage>24165</lpage>. doi: <pub-id pub-id-type="doi">10.1074/jbc.270.41.24156</pub-id>, PMID: <pub-id pub-id-type="pmid">7592619</pub-id></citation></ref>
<ref id="ref52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yadav</surname> <given-names>M.</given-names></name> <name><surname>Paritosh</surname> <given-names>K.</given-names></name> <name><surname>Vivekanand</surname> <given-names>V.</given-names></name></person-group> (<year>2020</year>). <article-title>Lignocellulose to bio-hydrogen: an overview on recent developments</article-title>. <source>Int. J. Hydrog. Energy</source> <volume>45</volume>, <fpage>18195</fpage>&#x2013;<lpage>18210</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ijhydene.2019.10.027</pub-id></citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>F.</given-names></name> <name><surname>Dong</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>Y.-J.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Deciphering Cellodextrin and glucose uptake in clostridium thermocellum</article-title>. <source>MBio</source> <volume>13</volume>:<fpage>e0147622</fpage>. doi: <pub-id pub-id-type="doi">10.1128/mbio.01476-22</pub-id>, PMID: <pub-id pub-id-type="pmid">36069444</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Nie</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Characterization of a hypervirulent multidrug-resistant ST23 Klebsiella pneumoniae carrying a Bla CTX-M-24 IncFII plasmid and a pK2044-like plasmid</article-title>. <source>J. Glob. Antimicrob. Resist.</source> <volume>22</volume>, <fpage>674</fpage>&#x2013;<lpage>679</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jgar.2020.05.004</pub-id>, PMID: <pub-id pub-id-type="pmid">32439569</pub-id></citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Jiang</surname> <given-names>C.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name> <name><surname>Xue</surname> <given-names>C.</given-names></name> <name><surname>Mao</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>A novel Agaro-oligosaccharide-lytic &#x03B2;-galactosidase from Agarivorans gilvus WH 0801</article-title>. <source>Appl. Microbiol. Biotechnol.</source> <volume>102</volume>, <fpage>5165</fpage>&#x2013;<lpage>5172</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00253-018-8999-0</pub-id>, PMID: <pub-id pub-id-type="pmid">29682702</pub-id></citation></ref>
<ref id="ref56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoav</surname> <given-names>S.</given-names></name> <name><surname>Stern</surname> <given-names>J.</given-names></name> <name><surname>Salama-Alber</surname> <given-names>O.</given-names></name> <name><surname>Frolow</surname> <given-names>F.</given-names></name> <name><surname>Anbar</surname> <given-names>M.</given-names></name> <name><surname>Karpol</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Directed evolution of clostridium thermocellum &#x03B2;-glucosidase a towards enhanced Thermostability</article-title>. <source>Int. J. Mol. Sci.</source> <volume>20</volume>:<fpage>4701</fpage>. doi: <pub-id pub-id-type="doi">10.3390/ijms20194701</pub-id>, PMID: <pub-id pub-id-type="pmid">31547488</pub-id></citation></ref>
<ref id="ref57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Hong</surname> <given-names>W.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name> <name><surname>Feng</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Efficient whole-cell-catalyzing cellulose saccharification using engineered clostridium thermocellum</article-title>. <source>Biotechnol. Biofuels</source> <volume>10</volume>:<fpage>124</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13068-017-0796-y</pub-id>, PMID: <pub-id pub-id-type="pmid">28507596</pub-id></citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>C.</given-names></name> <name><surname>Han</surname> <given-names>M.</given-names></name> <name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Ning</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Pan-genome analyses of 24 Shewanella strains re-emphasize the diversification of their functions yet evolutionary dynamics of metal-reducing pathway</article-title>. <source>Biotechnol. Biofuels</source> <volume>11</volume>:<fpage>193</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13068-018-1201-1</pub-id>, PMID: <pub-id pub-id-type="pmid">30026808</pub-id></citation></ref>
</ref-list>
</back>
</article>