<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2024.1387428</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Mining local and global spatiotemporal features for tactile object recognition</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Qian</surname> <given-names>Xiaoliang</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/636477/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Deng</surname> <given-names>Wei</given-names></name>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>Wei</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Yucui</given-names></name>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Jiang</surname> <given-names>Liying</given-names></name>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>College of Electrical and Information Engineering, Zhengzhou University of Light Industry</institution>, <addr-line>Zhengzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ganesh R. Naik, Flinders University, Australia</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jin-Gang Yu, South China University of Technology, China</p>
<p>Li Zhang, Fudan University, China</p>
<p>Lv ZhiYong, Xi&#x00027;an University of Technology, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Wei Wang <email>wangwei-zzuli&#x00040;zzuli.edu.cn</email></corresp>
<corresp id="c002">Liying Jiang <email>jiangliying&#x00040;zzuli.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>05</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>18</volume>
<elocation-id>1387428</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>02</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>04</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Qian, Deng, Wang, Liu and Jiang.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Qian, Deng, Wang, Liu and Jiang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The tactile object recognition (TOR) is highly important for environmental perception of robots. The previous works usually utilize single scale convolution which cannot simultaneously extract local and global spatiotemporal features of tactile data, which leads to low accuracy in TOR task. To address above problem, this article proposes a local and global residual (LGR-18) network which is mainly consisted of multiple local and global convolution (LGC) blocks. An LGC block contains two pairs of local convolution (LC) and global convolution (GC) modules. The LC module mainly utilizes a temporal shift operation and a 2D convolution layer to extract local spatiotemporal features. The GC module extracts global spatiotemporal features by fusing multiple 1D and 2D convolutions which can expand the receptive field in temporal and spatial dimensions. Consequently, our LGR-18 network can extract local-global spatiotemporal features without using 3D convolutions which usually require a large number of parameters. The effectiveness of LC module, GC module and LGC block is verified by ablation studies. Quantitative comparisons with state-of-the-art methods reveal the excellent capability of our method.</p></abstract>
<kwd-group>
<kwd>tactile object recognition</kwd>
<kwd>LGR-18 network</kwd>
<kwd>local convolution module</kwd>
<kwd>global convolution module</kwd>
<kwd>local and global spatiotemporal features</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="5"/>
<equation-count count="3"/>
<ref-count count="36"/>
<page-count count="11"/>
<word-count count="5231"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Robots perceive objects around them mainly through touch and vision. Although vision can intuitively capture the appearance of an object, it cannot capture the basic object properties, such as mass, hardness and texture. In addition, many limitations exist for vision perception (Lv et al., <xref ref-type="bibr" rid="B16">2023a</xref>). Consequently, the tactile object recognition (TOR) task is proposed to predict the category of the object being grasped by robot, which can provide support for subsequent grasping operations without being constrained by the aforementioned conditions.</p>
<p>TOR has a broad range of applications in life, including descriptive analysis in the food industry (Philippe et al., <xref ref-type="bibr" rid="B19">2004</xref>), electronic skin (Liu et al., <xref ref-type="bibr" rid="B14">2020</xref>) and embedded prostheses in the biomedical field (Wu et al., <xref ref-type="bibr" rid="B33">2018</xref>), and postdisaster rescue (Gao et al., <xref ref-type="bibr" rid="B7">2021</xref>), etc. The rapid development of deep learning (Lv et al., <xref ref-type="bibr" rid="B18">2023c</xref>) has led to tremendous progress in various fields (Qian et al., <xref ref-type="bibr" rid="B22">2020</xref>, <xref ref-type="bibr" rid="B20">2023a</xref>,<xref ref-type="bibr" rid="B21">b</xref>,<xref ref-type="bibr" rid="B24">d</xref>,<xref ref-type="bibr" rid="B25">e</xref>,<xref ref-type="bibr" rid="B26">f</xref>; Huo et al., <xref ref-type="bibr" rid="B9">2023</xref>; Li et al., <xref ref-type="bibr" rid="B12">2023</xref>; Xie et al., <xref ref-type="bibr" rid="B34">2024</xref>), and deep learning based TOR are the mainstream methods (Liu et al., <xref ref-type="bibr" rid="B13">2016</xref>; Ibrahim et al., <xref ref-type="bibr" rid="B10">2022</xref>; Yi et al., <xref ref-type="bibr" rid="B35">2022</xref>). Currently, TORs using deep learning can be divided into two categories: one uses a 2D CNN to extract features from each frame of tactile data and then fuses the features of each frame to recognize the grasped object based on the fused features and the other uses a 3D CNN to extract spatial and temporal features from tactile frames for recognition.</p>
<p>Traditional TOR methods mostly adopt the methods in the first category, which use sensors (primarily pressure sensor arrays) to acquire tactile information, then the tactile data are sent to a 2D CNN for feature extraction and category prediction. Gandarias et al. (<xref ref-type="bibr" rid="B6">2017</xref>) used a 2D CNN to extract high-resolution tactile features and trained a support vector machine (SVM) using these features. The trained SVM was used to predict the object category. Bottcher et al. (<xref ref-type="bibr" rid="B1">2021</xref>) collected tactile data via two different tactile sensors and subsequently input the data into a 2D CNN to extract features and infer results. Other related works include Sundaram et al. (<xref ref-type="bibr" rid="B31">2019</xref>), Chung et al. (<xref ref-type="bibr" rid="B5">2020</xref>), and Carvalho et al. (<xref ref-type="bibr" rid="B4">2022</xref>), etc.</p>
<p>Recently, the another category of methods has achieved remarkable results and has become mainstream. Qian et al. (<xref ref-type="bibr" rid="B23">2023c</xref>). used a gradient adaptive sampling (GAS) strategy to process the acquired tactile data and subsequently fed the data into a 3D CNN network to extract multiple scale temporal features. The features were fused at the fully connected and outputted prediction category. Inspired by the optical flow method (Cao et al., <xref ref-type="bibr" rid="B2">2018</xref>) used not only original tactile data but also tactile flow and intensity differences as input data. These data underwent convolution, weighting, and other operations on different branches and were ultimately fused at the fully connected layer to infer the results. Other related works include Kirby et al. (<xref ref-type="bibr" rid="B11">2022</xref>) and Lu et al. (<xref ref-type="bibr" rid="B15">2023</xref>), etc.</p>
<p>The first category of methods extracts only the features of each frame and does not use temporal information between frames, therefore, their overall performance is limited. In the second category, spatiotemporal features are extracted via 3D CNNs, and the overall performance is better than that of the first category of methods. However, existing methods utilize single scale convolution operations to extract features in spatial and temporal dimensions, and they cannot simultaneously extract local and global features very well.</p>
<p>This paper proposes a local-global spatiotemporal feature extraction scheme to solve the above problems. First, a local convolution (LC) module is proposed, which utilizes the interaction of adjacent temporal information and 2D convolution to extract local spatiotemporal features. Next, the global convolution (GC) module is proposed to extract global spatiotemporal features and the module combines multiple 1D and 2D [(1&#x0002B;2)D] convolutions which can extend the receptive field (Lv et al., <xref ref-type="bibr" rid="B17">2023b</xref>) in spatial and temporal dimensions. Finally, this article achieves accurate TOR by comprehensively utilizing local and global spatiotemporal features.</p>
<p>The main contributions are as follows:</p>
<p>1. This paper proposes a local convolution (LC) module which extracts spatiotemporal features by using the interaction of adjacent temporal information and 2D convolution operation.</p>
<p>2. This paper proposes a global convolution (GC) module which extracts global spatiotemporal features by fusing multiple 1D and 2D convolutions which can expand the receptive field in temporal and spatial dimensions.</p>
<p>3. Our method achieves the highest object recognition accuracy on two public datasets by comprehensively using local-global spatiotemporal features.</p></sec>
<sec id="s2">
<title>2 Related works</title>
<p>Gradient adaptive sampling (GAS) (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>) and MR3D-18 network (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>) have strong relevance to this paper, therefore, they are introduced here.</p>
<sec>
<title>2.1 Gradient adaptive sampling</title>
<p>Unlike uniform sampling and sparse sampling, GAS uses the pressure gradient to guide the adaptive sampling. The specific approach is to normalize the accumulated gradient over the <italic>T</italic> period and then divide it into multiple intervals, randomly selecting one point from each interval. This process obtains multiple data frames, which are subsequently fed into the network.</p></sec>
<sec>
<title>2.2 MR3D-18 network</title>
<p>The MR3D-18 network is proposed to address the problems that the size of tactile frames is small and overfitting is easily occur, which removes a pooling operation and adds a dropout layer to the ResNet3D-18 (Hara et al., <xref ref-type="bibr" rid="B8">2018</xref>) network for handling above problems.</p></sec></sec>
<sec id="s3">
<title>3 Proposed method</title>
<sec>
<title>3.1 Overview</title>
<p>As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, the P frames of tactile data are adaptively selected from all frames by using GAS. Then, they are fed into local and global residual (LGR-18) network to extract local and global spatiotemporal features. Finally, the features are imported into a fully connected layer and a softmax classifier to predict categories. Our LGR-18 network is primarily composed of multiple local and global convolution (LGC) blocks, which is proposed in this paper, and the LGC block consists of two pairs of LC and GC modules. The LC and GC modules, LGC block and LGR-18 network will be precisely explained in the following sections.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Framework of our method.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0001.tif"/>
</fig></sec>
<sec>
<title>3.2 Local convolution module</title>
<p>The LC module focuses on extracting the local spatiotemporal features of tactile data. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the size of the input features <italic>X</italic> is [<italic>N</italic>, <italic>T</italic>, <italic>C</italic>, <italic>H</italic>, <italic>W</italic>], where <italic>N</italic> denotes the batch size, <italic>T</italic> and <italic>C</italic> separately denote the quantity of input frames and feature channels, <italic>H</italic> and <italic>W</italic> denote the height and width of the features respectively. First, the LC module utilizes the temporal shift (TS) operation to extract local temporal features.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Illustration of LC module.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0002.tif"/>
</fig>
<p>The TS operation is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. A tensor with <italic>C</italic> channels and <italic>T</italic> frames is also shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The features of the different time stamps are shown in different colors in each row. Along the temporal dimension, the TS operation shifts one channel in forward direction and one channel in backward direction. We utilize the abandoning and zero-padding operation to address the problems of excessive and missing features. It is worth noting that replacing the zero padding features by the abandoning features is infeasible because it will destroy the temporal sequence. After the TS operation, a 2D convolution layer is utilized to extract the local spatial features. Next, we utilize a global average pooling (GAP) and a sigmoid function to extract the local spatiotemporal weights of each channel. Finally, we utilize a simple method to extract local spatiotemporal features by performing channel-wise product between the input feature <italic>X</italic> and the local spatiotemporal weight <italic>S</italic>. A residual connection is employed to prevent the loss of crucial information in the original features. Finally, the LC module, which relies on TS operations and traditional 2D convolutions as its core, extracts local features.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Illustration of temporal shift operation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0003.tif"/>
</fig></sec>
<sec>
<title>3.3 Global convolution module</title>
<p>Inspired by the conventional depthwise separable convolution, multiple (1&#x0002B;2)D convolutions are utilized to extract global spatiotemporal features from tactile data, where the 1D and 2D convolutions are used to extract temporal and spatial features, respectively. However, the innovation of GC module does not lie in the depthwise separable convolution. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the core idea of GC module is that the receptive field of convolution kernels is continuously enlarged to extract global spatiotemporal features via iterative residual connections and convolutions. The detailed procedure can be seen in <xref ref-type="disp-formula" rid="E1">Equations 1</xref>, <xref ref-type="disp-formula" rid="E2">2</xref> and <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M5"><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>Y</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>Y</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>Y</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>H</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Illustration of GC module.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0004.tif"/>
</fig><p>In <xref ref-type="disp-formula" rid="E1">Equation 1</xref>, <italic>Y</italic><sub>1</sub> is the output of the first branch and is identical to the input of the first branch. <italic>Y</italic><sub>2</sub>, <italic>Y</italic><sub>3</sub> and <italic>Y</italic><sub>4</sub> represent the outputs of the 2nd, 3rd and 4th branches, respectively, in the GC module. The <italic>conv</italic><sub>(1&#x0002B;2)<italic>D</italic></sub> is the same as (1&#x0002B;2)D convolutions. The parameters for <italic>conv</italic><sub>(1&#x0002B;2)<italic>D</italic></sub> are 3 and 3&#x02013;3. As shown in <xref ref-type="disp-formula" rid="E2">Equation 2</xref>, the final output of GC module, denoted as Z, is obtained by fusing the outputs of four branches:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>Z</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x02003;</mml:mtext><mml:mi>Z</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>H</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>,</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>3.4 Architecture of the LGR-18 network</title>
<p>As shown in <xref ref-type="table" rid="T1">Table 1</xref>, our LGR-18 network is modified from the MR3D-18 network. First, in the LGR-18 network, the (1&#x0002B;2)D convolutions replaced the 7 &#x000D7; 7 &#x000D7; 7 convolution layer in the MR3D-18 network. Second, in the LGR-18 network, multiple 3 &#x000D7; 3 &#x000D7; 3 3D convolution layers are replaced with LGC blocks, which is proposed in this paper, and two 3D convolution layers are approximately equivalent to an LGC block. Third, the LGR-18 network adds (1&#x0002B;2)D convolutions with size of 1 in the Res<sub>3</sub>, Res<sub>4</sub>, and Res<sub>5</sub> layers to change the number of channels. Finally, the LGC block consists of two pairs of LC and GC modules connected by residual connections. <xref ref-type="table" rid="T1">Table 1</xref> shows that the LGR-18 network has two advantages over the MR3D-18 network.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Comparison between MR3D-18 and LGR-18 network, where the size of input tactile data is 32 &#x000D7; 32 &#x000D7; 32.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Layers</bold></th>
<th valign="top" align="left" colspan="2"><bold>MR3D-18</bold></th>
<th valign="top" align="left" colspan="2"><bold>LGR-18</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="left"><bold>Filters</bold></td>
<td valign="top" align="left"><bold>Output size</bold></td>
<td valign="top" align="left"><bold>Filters</bold></td>
<td valign="top" align="left"><bold>Output size</bold></td>
</tr> <tr>
<td valign="top" align="left">Conv<sub>1</sub></td>
<td valign="top" align="left">7 &#x000D7; 7 &#x000D7; 7, 64 stride 1, 2<sup>2</sup></td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
<td valign="top" align="left">7 7 &#x000D7; 7 stride 1, 2<sup>2</sup></td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
</tr> <tr>
<td valign="top" align="left">Res<sub>2</sub></td>
<td valign="top" align="left"><inline-formula><mml:math id="M1"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> stride 1,1<sup>2</sup></td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
<td valign="top" align="left">LGC, 64 LGC, 64 stride 1,1<sup>2</sup></td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
</tr> <tr>
<td valign="top" align="left">Dropout</td>
<td valign="top" align="left">0.3</td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
<td valign="top" align="left">0.3</td>
<td valign="top" align="left">32 &#x000D7; 16<sup>2</sup>&#x000D7;64</td>
</tr> <tr>
<td valign="top" align="left">Res<sub>3</sub></td>
<td valign="top" align="left"><inline-formula><mml:math id="M2"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>stride 2,2<sup>2</sup></td>
<td valign="top" align="left">16 &#x000D7; 8<sup>2</sup>&#x000D7;128</td>
<td valign="top" align="left">LGC, 64 1, 128 1 &#x000D7; 1, 128 LGC, 128 stride 2,2<sup>2</sup></td>
<td valign="top" align="left">16 &#x000D7; 8<sup>2</sup>&#x000D7;128</td>
</tr> <tr>
<td valign="top" align="left">Res<sub>4</sub></td>
<td valign="top" align="left"><inline-formula><mml:math id="M3"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>256</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>256</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>stride 2,2<sup>2</sup></td>
<td valign="top" align="left">8 &#x000D7; 4<sup>2</sup>&#x000D7;256</td>
<td valign="top" align="left">LGC, 128 1, 256 1 &#x000D7; 1, 256 LGC, 256 stride 2,2<sup>2</sup></td>
<td valign="top" align="left">8 &#x000D7; 4<sup>2</sup>&#x000D7;256</td>
</tr> <tr>
<td valign="top" align="left">Res<sub>5</sub></td>
<td valign="top" align="left"><inline-formula><mml:math id="M4"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>512</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>512</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>stride 2,2<sup>2</sup></td>
<td valign="top" align="left">4 &#x000D7; 2<sup>2</sup>&#x000D7;512</td>
<td valign="top" align="left">LGC, 256 1, 512 1 &#x000D7; 1, 512 LGC, 512 stride 2,2<sup>2</sup></td>
<td valign="top" align="left">4 &#x000D7; 2<sup>2</sup>&#x000D7;512</td>
</tr> <tr>
<td valign="top" align="left" colspan="5">Global average pooling</td>
</tr></tbody>
</table>
</table-wrap>
<p>1. The LGR-18 network abandons all 3D convolution layers, reducing computational complexity.</p>
<p>2. The LGR-18 network can extract local and global spatiotemporal features through the LGC blocks.</p></sec>
<sec>
<title>3.5 Training scheme</title>
<p>To enhance the performance of the LGR-18 network, we utilize the large-scale Kinetics-400 dataset (Carreira and Zisserman, <xref ref-type="bibr" rid="B3">2017</xref>), which consists of 400 human action categories and at least 400 video clips in each category, to pre-train our method. Subsequently, we utilized the UCF101 (Soomro et al., <xref ref-type="bibr" rid="B30">2012</xref>) and target datasets to pretrain the LGR-18 network.</p>
<p>To address the problem of large dataset sizes, the size of input data is adjusted to 32 &#x000D7; 32 when the LGR-18 network pre-trained on the Kinetics-400 and UCF101 datasets. The traditional cross-entropy loss, denoted as <italic>L</italic>, is employed to optimize the LGR-18 network, which is formulated <xref ref-type="disp-formula" rid="E3">Equation 3</xref>:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mtext class="textrm" mathvariant="normal">&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>K</italic> denotes the number of categories, <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the prediction score of category <italic>k</italic>, and <italic>v</italic><sub><italic>k</italic></sub> denotes the label of category <italic>k</italic>.</p></sec></sec>
<sec id="s4">
<title>4 Experiment</title>
<sec>
<title>4.1 Experiment setup</title>
<sec>
<title>4.1.1 Datasets</title>
<p>The LGR-18 network is verified on the MIT-STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>) and iCub datasets (Soh et al., <xref ref-type="bibr" rid="B29">2012</xref>). The MIT-STAG dataset includes 26 categories of common objects and empty hands, i.e., allen key set, ball, battery, board eraser, bracket, stress toy, cat, chain, clip, coin, gel, kiwano, lotion, mug, multimeter, pen, safety glasses, scissors, screwdriver, spoon, spray can, stapley, tape, tea box, full cola can, and empty cola can. A total of 88,269 valid frames are collected, each with a size of 32 &#x000D7; 32. To achieve accurate prediction results under fair conditions, 1,353 frames (totaling 36,531 frames) are selected from each category as the training set, and 597 frames (totaling 16,119 frames) are selected as the testing set. The MIT-STAG dataset includes too many categories with similar appearance characteristics, making accurate object recognition highly challenging on this dataset.</p>
<p>The iCub dataset is acquired by two anthropomorphic dexterous hands of the iCub humanoid robot platform. Each anthropomorphic hand is equipped with five fingers, each with 20 movable joints. Additionally, each finger is equipped with pressure sensors to acquire tactile data. The iCub dataset includes 2,200 frames with 10 categories, i.e., monkey toy, med vitamin water, med coke, lotion, vitamin water, full cola, empty vitamin water, empty coke, book and blue bear (toy), and the size of each frame is 5 &#x000D7; 12. For each category in the iCub dataset, 132 frames (totaling 1,320 frames) are selected as the training set, and 88 frames are selected (totaling 880 frames) as the testing set.</p></sec>
<sec>
<title>4.1.2 Implementation details</title>
<p>This paper uses the top 1 score, kappa coefficient (KC) and confusion matrix for evaluation. The stochastic gradient descent is used to optimize our module. The momentum and decay rate are 0.9 and 0.0001, respectively. The initial learning rate is 0.002 and the quantity of epochs is 50. The learning rate decreases to 10% of the previous stage after every 10 epochs. The batch sizes are 32 and 8 for the MIT-STAG and iCub datasets, respectively.</p>
<p>The experiments were all performed on the PyTorch framework and run on a workstation with two NVIDIA GeForce RTX 2080 Ti (2 &#x000D7; 11 GB).</p></sec></sec>
<sec>
<title>4.2 Ablation study</title>
<sec>
<title>4.2.1 Ablation study of LGC block</title>
<p>As <xref ref-type="table" rid="T2">Table 2</xref> shows, the LGC block is compared with the other four convolution blocks to verify its effectiveness. Architecture A is used as a baseline and is composed of two 3D convolution layers connected by residual connection. Architecture B replaces one of the 3D convolutions layers in architecture A with an LC module, and architecture C uses the GC module to replace one of the 3D convolutions layers in architecture A. In architecture D, a pair of LC and GC modules are utilized to replace one of two 3D convolution layers in architecture A. Architecture E is our method, and it uses two pairs of LC and GC modules to replace all the 3D convolution layers in architecture A.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Ablation study of LGC block on the MIT-STAG dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Architecture</bold></th>
<th valign="top" align="left"><bold>Illustration of architecture</bold></th>
<th valign="top" align="left"><bold>Top 1 score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">A(baseline)</td>
<td valign="top" align="left"><inline-graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-i0001.tif"/></td>
<td valign="top" align="left">84.98</td>
</tr> <tr>
<td valign="top" align="left">B</td>
<td valign="top" align="left"><inline-graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-i0002.tif"/></td>
<td valign="top" align="left">85.32</td>
</tr> <tr>
<td valign="top" align="left">C</td>
<td valign="top" align="left"><inline-graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-i0003.tif"/></td>
<td valign="top" align="left">88.07</td>
</tr> <tr>
<td valign="top" align="left">D</td>
<td valign="top" align="left"><inline-graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-i0004.tif"/></td>
<td valign="top" align="left">89.05</td>
</tr> <tr>
<td valign="top" align="left">E</td>
<td valign="top" align="left"><inline-graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-i0005.tif"/></td>
<td valign="top" align="left">91.07</td>
</tr></tbody>
</table>
</table-wrap>
<p>Ablation studies reveal the effectiveness of the LC module, GC module, and combination of the LC and GC modules to compare architectures B, C, and D with A. By comparing architecture E with A, ablation studies strongly demonstrate that the best performance can be achieved by using a combination of LC and GC modules.</p></sec>
<sec>
<title>4.2.2 Ablation study of spatial pooling operation in LC module</title>
<p>As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the spatial pooling operation is involved in the LC module, therefore, the max pooling and GAP are compared with each other to determine who is more appropriate for LC module. As shown in <xref ref-type="table" rid="T3">Table 3</xref>, the top 1 score and KC of GAP are higher than the ones of max pooling, consequently, the GAP is adopted by LC module.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Ablation study of LGC block on the MIT-STAG dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Spatial pooling operation</bold></th>
<th valign="top" align="left"><bold>Top 1 score</bold></th>
<th valign="top" align="left"><bold>KC</bold></th>
</tr></thead>
<tbody>
<tr>
<td valign="top" align="left">Max pooling</td>
<td valign="top" align="left">90.8</td>
<td valign="top" align="left">89.7</td>
</tr> <tr>
<td valign="top" align="left">Global average pooling</td>
<td valign="top" align="left">91.1</td>
<td valign="top" align="left">90.7</td>
</tr></tbody>
</table>
</table-wrap></sec></sec>
<sec>
<title>4.3 Parameter analysis</title>
<sec>
<title>4.3.1 Parameter analysis for the number of shifting channels</title>
<p>As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the number of shifting channels is an important hyperparameter for the LC module, therefore, it is quantitatively analyzed in this section. It is worth noting that the number of shifting channels must be even because the shifting operation is bidirectional. As shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, the top 1 score achieves the highest value when the number of shifting channels is set to 2, which means that shifting many channels is not suitable for the LC module because it can induce the information reduction.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The top 1 score (%) of different number of shifting channels on the MIT-STAG dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0005.tif"/>
</fig></sec>
<sec>
<title>4.3.2 Parameter analysis for dropout rate</title>
<p>As shown in <xref ref-type="table" rid="T1">Table 1</xref>, the dropout layer is used to prevent the overfitting problem, therefore, the dropout rate is quantitatively analyzed in this section. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, the top 1 score achieves the highest value when the dropout rate is set to 0.3.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The top 1 score (%) of different dropout rate on the MIT-STAG dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0006.tif"/>
</fig>
</sec></sec><sec>
<title>4.4 Comparisons with state-of-the-art methods</title>
<p>To demonstrate the overall effectiveness of our model, we conducted a comprehensive quantitative comparison with 5 methods on the MIT-STAG dataset (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>), i.e., STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>), Smart-hand (Wang et al., <xref ref-type="bibr" rid="B32">2021</xref>), ResNet10-v1 (Zhang et al., <xref ref-type="bibr" rid="B36">2021</xref>), Tactile-ViewGCN (Sharma et al., <xref ref-type="bibr" rid="B27">2022</xref>), and GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>), and 5 methods on the iCub dataset (Soh et al., <xref ref-type="bibr" rid="B29">2012</xref>), i.e., DS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>), GS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>), STORK-GP (Soh et al., <xref ref-type="bibr" rid="B29">2012</xref>), STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>), and GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>).</p>
<p>As shown in <xref ref-type="table" rid="T4">Table 4</xref>, our method achieved the highest top 1 score and KC. This demonstrates that our model has the best prediction accuracy and the lowest level of confusion on the MIT-STAG dataset. A comparison of the confusion matrices in <xref ref-type="fig" rid="F7">Figure 7</xref> further supports the conclusions above. As <xref ref-type="table" rid="T5">Table 5</xref> and <xref ref-type="fig" rid="F8">Figure 8</xref> show, both our method and Qian et al. achieved a recognition accuracy of 100%. Our method and that of Qian et al. yield the highest detection accuracy and the lowest level of confusion on the iCub dataset.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Comparisons with 5 methods in terms of the top 1 score (%) and KC (%) on the MIT-STAG dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Top 1 score</bold></th>
<th valign="top" align="left"><bold>KC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>)</td>
<td valign="top" align="left">72.4</td>
<td valign="top" align="left">71.4</td>
</tr> <tr>
<td valign="top" align="left">Smart-hand (Wang et al., <xref ref-type="bibr" rid="B32">2021</xref>)</td>
<td valign="top" align="left">72.0</td>
<td valign="top" align="left">71.0</td>
</tr> <tr>
<td valign="top" align="left">ResNet10-v1 (Zhang et al., <xref ref-type="bibr" rid="B36">2021</xref>)</td>
<td valign="top" align="left">80.1</td>
<td valign="top" align="left">79.3</td>
</tr> <tr>
<td valign="top" align="left">Tactile-ViewGCN (Sharma et al., <xref ref-type="bibr" rid="B27">2022</xref>)</td>
<td valign="top" align="left">81.8</td>
<td valign="top" align="left">81.1</td>
</tr> <tr>
<td valign="top" align="left">GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>)</td>
<td valign="top" align="left">88.8</td>
<td valign="top" align="left">88.5</td>
</tr> <tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="left">91.1</td>
<td valign="top" align="left">90.7</td>
</tr></tbody>
</table>
</table-wrap><fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Comparisons with 5 methods in terms of confusion matrix on the MIT-STAG dataset. <bold>(A)</bold> STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>), <bold>(B)</bold> Smart-hand (Wang et al., <xref ref-type="bibr" rid="B32">2021</xref>), <bold>(C)</bold> ResNet10-v1 (Zhang et al., <xref ref-type="bibr" rid="B36">2021</xref>), <bold>(D)</bold> Tactile-ViewGCN (Sharma et al., <xref ref-type="bibr" rid="B27">2022</xref>), <bold>(E)</bold> GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>), <bold>(F)</bold> Ours.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0007.tif"/>
</fig><table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Comparisons with 5 methods in terms of the top 1 score (%) and KC (%) on the iCub dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Top 1 score</bold></th>
<th valign="top" align="left"><bold>KC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>)</td>
<td valign="top" align="left">98.5</td>
<td valign="top" align="left">98.4</td>
</tr> <tr>
<td valign="top" align="left">GS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>)</td>
<td valign="top" align="left">98.9</td>
<td valign="top" align="left">98.8</td>
</tr> <tr>
<td valign="top" align="left">STORK-GP (Soh et al., <xref ref-type="bibr" rid="B29">2012</xref>)</td>
<td valign="top" align="left">99.3</td>
<td valign="top" align="left">99.2</td>
</tr> <tr>
<td valign="top" align="left">STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>)</td>
<td valign="top" align="left">99.5</td>
<td valign="top" align="left">99.4</td>
</tr> <tr>
<td valign="top" align="left">GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>)</td>
<td valign="top" align="left">100</td>
<td valign="top" align="left">1</td>
</tr> <tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="left">100</td>
<td valign="top" align="left">1</td>
</tr></tbody>
</table>
</table-wrap><fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Comparisons with 5 methods in terms of confusion matrix on the iCub dataset. <bold>(A)</bold> DS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>), <bold>(B)</bold> GS (Soh and Demiris, <xref ref-type="bibr" rid="B28">2014</xref>), <bold>(C)</bold> STORK-GP (Soh et al., <xref ref-type="bibr" rid="B29">2012</xref>), <bold>(D)</bold> STAG (Sundaram et al., <xref ref-type="bibr" rid="B31">2019</xref>), <bold>(E)</bold> GAS-MR3D (Qian et al., <xref ref-type="bibr" rid="B23">2023c</xref>), <bold>(F)</bold> Ours.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-18-1387428-g0008.tif"/>
</fig><p>In summary, our method is superior to 5 advanced methods.</p></sec></sec>
<sec sec-type="conclusions" id="s5">
<title>5 Conclusion</title>
<p>A novel LGR-18 network is proposed to address the problem that the current TOR models cannot simultaneously extract local and global features very well. The LGR-18 network consists primarily of multiple traditional 1D and 2D convolution kernels and LGC blocks, which is proposed in this paper. The LGC block is formed by combining LC and GC modules through residual connections. The LC module mainly utilizes a temporal shift operation and a 2D convolution to extract local spatiotemporal features. The GC module extracts global spatiotemporal features by fusing multiple 1D and 2D convolutions which can expand the receptive field in temporal and spatial dimensions. In this paper, we utilize the LGR-18 network to extract local and global spatiotemporal features while mitigating the issue of large parameter in existing 3D CNN models. Ablation studies verify the validity of the LC module, GC module, and LGC block. A comprehensive quantitative comparison between our method and 5 advanced methods on the MIT-STAG and iCub datasets reveal the excellent capability of our method.</p>
<p>The future work of our team includes two parts. The first part is combining our method with video based object detection method, and the another part is deploying our method on more robots.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. The datasets for this study can be found in the <ext-link ext-link-type="uri" xlink:href="http://humangrap.io">http://humangrap.io</ext-link> and <ext-link ext-link-type="uri" xlink:href="http://https:/github.com/clear-nus/icub_grasp_dataset">https:/github.com/clear-nus/icub_grasp_dataset</ext-link>.</p></sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>XQ: Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. WD: Writing &#x02013; original draft. WW: Writing &#x02013; review &#x00026; editing. YL: Writing &#x02013; review &#x00026; editing. LJ: Writing &#x02013; review &#x00026; editing.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This research was funded by the National Natural Science Foundation of China under Grant (Nos: 62076223 and 62073299), Key Research Project of Henan Province Universities (No: 24ZX005), Key Science and Technology Program of Henan Province (No: 232102211018), Joint Fund of Henan Province Science and Technology R&#x00026;D Program (No: 225200810071), and the Science and Technology Major project of Henan Province (No: 231100220800).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bottcher</surname> <given-names>W.</given-names></name> <name><surname>Machado</surname> <given-names>P.</given-names></name> <name><surname>Lama</surname> <given-names>N.</given-names></name> <name><surname>McGinnity</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Object recognition for robotics from tactile time series data utilising different neural network architectures,&#x0201D;</article-title> in <source>2021 International Joint Conference on Neural Networks (IJCNN)</source>, 1&#x02013;8. <pub-id pub-id-type="doi">10.1109/IJCNN52387.2021.9533388</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>L.</given-names></name> <name><surname>Sun</surname> <given-names>F.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>W.</given-names></name> <name><surname>Kotagiri</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>End-to-end convnet for tactile recognition using residual orthogonal tiling and pyramid convolution ensemble</article-title>. <source>Cogn. Comput</source>. <volume>10</volume>, <fpage>718</fpage>&#x02013;<lpage>736</lpage>. <pub-id pub-id-type="doi">10.1007/s12559-018-9568-7</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carreira</surname> <given-names>J.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Quo vadis, action recognition? A new model and the kinetics dataset,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 6299&#x02013;6308. <pub-id pub-id-type="doi">10.1109/CVPR.2017.502</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carvalho</surname> <given-names>H. N.d</given-names></name> <name><surname>Castro</surname> <given-names>L. P.</given-names></name> <name><surname>Rego</surname> <given-names>T. G. D.</given-names></name> <name><surname>Filho</surname> <given-names>T. M. S.</given-names></name> <name><surname>Barbosa</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;Evaluating data representations for object recognition during pick-and-place manipulation tasks,&#x0201D;</article-title> in <source>2022 IEEE International Systems Conference (SysCon)</source>, 1&#x02013;6. <pub-id pub-id-type="doi">10.1109/SysCon53536.2022.9773911</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>J.</given-names></name> <name><surname>Lim</surname> <given-names>H.</given-names></name> <name><surname>Lim</surname> <given-names>M.</given-names></name> <name><surname>Cha</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Object classification based on piezoelectric actuator-sensor pair on robot hand using neural network</article-title>. <source>Smart Mater. Struct</source>. <volume>29</volume>:<fpage>105020</fpage>. <pub-id pub-id-type="doi">10.1088/1361-665X/aba540</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gandarias</surname> <given-names>J. M.</given-names></name> <name><surname>G&#x000F3;mez-de Gabriel</surname> <given-names>J. M.</given-names></name> <name><surname>Garc&#x000ED;a-Cerezo</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Human and object recognition with a high-resolution tactile sensor,&#x0201D;</article-title> in <source>2017 IEEE Sensors</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>3</lpage>. <pub-id pub-id-type="doi">10.1109/ICSENS.2017.8234203</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;A dynamic priority packet scheduling scheme for post-disaster uav-assisted mobile ad hoc network,&#x0201D;</article-title> in <source>2021 IEEE Wireless Communications and Networking Conference (WCNC)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/WCNC49053.2021.9417537</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hara</surname> <given-names>K.</given-names></name> <name><surname>Kataoka</surname> <given-names>H.</given-names></name> <name><surname>Satoh</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Can spatiotemporal 3D cnns retrace the history of 2D cnns and imagenet?&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 6546&#x02013;6555. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00685</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huo</surname> <given-names>Y.</given-names></name> <name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>Multiple instance complementary detection and difficulty evaluation for weakly supervised object detection in remote sensing images</article-title>. <source>IEEE Geosci. Rem. Sens. Lett</source>. <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2023.3283403</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ibrahim</surname> <given-names>A.</given-names></name> <name><surname>Ali</surname> <given-names>H. H.</given-names></name> <name><surname>Hassan</surname> <given-names>M. H.</given-names></name> <name><surname>Valle</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Convolutional neural networks based tactile object recognition for tactile sensing system,&#x0201D;</article-title> in <source>Applications in Electronics Pervading Industry, Environment and Society: APPLEPIES 2021</source> (<publisher-loc>Springer</publisher-loc>), <fpage>280</fpage>&#x02013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-95498-7_39</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kirby</surname> <given-names>E.</given-names></name> <name><surname>Zenha</surname> <given-names>R.</given-names></name> <name><surname>Jamone</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>Comparing single touch to dynamic exploratory procedures for robotic tactile object recognition</article-title>. <source>IEEE Robot. Autom. Lett</source>. <volume>7</volume>, <fpage>4252</fpage>&#x02013;<lpage>4258</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2022.3151261</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Hong</surname> <given-names>D.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Robust few-shot aerial image object detection via unbiased proposals filtration</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3300071</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Qin</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>F.</given-names></name> <name><surname>Guo</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Extreme kernel sparse learning for tactile object recognition</article-title>. <source>IEEE Trans. Cybern</source>. <volume>47</volume>, <fpage>4509</fpage>&#x02013;<lpage>4520</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2016.2614809</pub-id><pub-id pub-id-type="pmid">27775546</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Bao</surname> <given-names>R.</given-names></name> <name><surname>Tao</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>M.</given-names></name> <name><surname>Pan</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Recent progress in tactile sensors and their applications in intelligent systems</article-title>. <source>Sci. Bull</source>. <volume>65</volume>, <fpage>70</fpage>&#x02013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1016/j.scib.2019.10.021</pub-id><pub-id pub-id-type="pmid">36659072</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>D.</given-names></name> <name><surname>Yin</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>3-D tactile-based object recognition for robot hands using force-sensitive and bend sensor arrays</article-title>. <source>IEEE Trans. Cogn. Dev. Syst</source>. <volume>15</volume>, <fpage>1645</fpage>&#x02013;<lpage>1655</lpage>. <pub-id pub-id-type="doi">10.1109/TCDS.2022.3215021</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lv</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Sun</surname> <given-names>W.</given-names></name> <name><surname>Lei</surname> <given-names>T.</given-names></name> <name><surname>Benediktsson</surname> <given-names>J. A.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2023a</year>). <article-title>Novel enhanced unet for change detection using multimodal remote sensing image</article-title>. <source>IEEE Geosci. Rem. Sens. Lett</source>. <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2023.3325439</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lv</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>W.</given-names></name> <name><surname>Lei</surname> <given-names>T.</given-names></name> <name><surname>Benediktsson</surname> <given-names>J. A.</given-names></name> <name><surname>Jia</surname> <given-names>X.</given-names></name></person-group> (<year>2023b</year>). <article-title>Hierarchical attention feature fusion-based network for land cover change detection with homogeneous and heterogeneous remote sensing images</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3334521</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lv</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Sun</surname> <given-names>W.</given-names></name> <name><surname>Benediktsson</surname> <given-names>J. A.</given-names></name> <name><surname>Lei</surname> <given-names>T.</given-names></name> <name><surname>Falco</surname> <given-names>N.</given-names></name></person-group> (<year>2023c</year>). <article-title>Spatial-contextual information utilization framework for land cover change detection with hyperspectral remote sensed images</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3336791</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Philippe</surname> <given-names>F.</given-names></name> <name><surname>Schacher</surname> <given-names>L.</given-names></name> <name><surname>Adolphe</surname> <given-names>D. C.</given-names></name> <name><surname>Dacremont</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <article-title>Tactile feeling: sensory analysis applied to textile goods</article-title>. <source>Textile Res. J</source>. <volume>74</volume>, <fpage>1066</fpage>&#x02013;<lpage>1072</lpage>. <pub-id pub-id-type="doi">10.1177/004051750407401207</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Huo</surname> <given-names>Y.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Gao</surname> <given-names>C.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name></person-group> (<year>2023a</year>). <article-title>Mining high-quality pseudoinstance soft labels for weakly supervised object detection in remote sensing images</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3266838</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name></person-group> (<year>2023b</year>). <article-title>Semantic segmentation guided pseudo label mining and instance re-detection for weakly supervised object detection in remote sensing images</article-title>. <source>Int. J. Appl. Earth Observ. Geoinf</source>. <volume>119</volume>:<fpage>103301</fpage>. <pub-id pub-id-type="doi">10.1016/j.jag.2023.103301</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Lin</surname> <given-names>S.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Object detection in remote sensing images based on improved bounding box regression and multi-level features fusion</article-title>. <source>Rem. Sens</source>. <volume>12</volume>:<fpage>143</fpage>. <pub-id pub-id-type="doi">10.3390/rs12010143</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Meng</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Jiang</surname> <given-names>L.</given-names></name></person-group> (<year>2023c</year>). <article-title>Gradient adaptive sampling and multiple temporal scale 3d cnns for tactile object recognition</article-title>. <source>Front. Neurorob</source>. <volume>17</volume>:<fpage>1159168</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2023.1159168</pub-id><pub-id pub-id-type="pmid">37180284</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zeng</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2023d</year>). <article-title>Multiscale image splitting based feature enhancement and instance difficulty aware training for weakly supervised object detection in remote sensing images</article-title>. <source>IEEE J. Select. Topics Appl. Earth Observ. Rem. Sens</source>. <volume>16</volume>, <fpage>7497</fpage>&#x02013;<lpage>7506</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2023.3304411</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Wu</surname> <given-names>B.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name></person-group> (<year>2023e</year>). <article-title>Building a bridge of bounding box regression between oriented and horizontal object detection in remote sensing images</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3256373</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Zeng</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name></person-group> (<year>2023f</year>). <article-title>Co-saliency detection guided by group weakly supervised learning</article-title>. <source>IEEE Trans. Multim</source>. <volume>25</volume>, <fpage>1810</fpage>&#x02013;<lpage>1818</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2022.3167805</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharma</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Tactile-viewgcn: Learning shape descriptor from tactile data using graph convolutional network</article-title>. <source>arXiv preprint arXiv:2203.06183</source>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soh</surname> <given-names>H.</given-names></name> <name><surname>Demiris</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Incrementally learning objects by touch: online discriminative and generative models for tactile-based recognition</article-title>. <source>IEEE Trans. Hapt</source>. <volume>7</volume>, <fpage>512</fpage>&#x02013;<lpage>525</lpage>. <pub-id pub-id-type="doi">10.1109/TOH.2014.2326159</pub-id><pub-id pub-id-type="pmid">25532151</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soh</surname> <given-names>H.</given-names></name> <name><surname>Su</surname> <given-names>Y.</given-names></name> <name><surname>Demiris</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). Online spatio-temporal gaussian process experts with application to tactile classification,&#x0201D; in <italic>2012 IEEE/RSJ International Conference on Intelligent Robots and Systems</italic> (IEEE), <fpage>4489</fpage>&#x02013;<lpage>4496</lpage>. <pub-id pub-id-type="doi">10.1109/IROS.2012.6385992</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soomro</surname> <given-names>K.</given-names></name> <name><surname>Zamir</surname> <given-names>A.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Ucf101: a dataset of 101 human actions classes from videos in the wild</article-title>. <source>arXiv, abs/1212.0402</source>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sundaram</surname> <given-names>S.</given-names></name> <name><surname>Kellnhofer</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>J.-Y.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Matusik</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>Learning the signatures of the human grasp using a scalable tactile glove</article-title>. <source>Nature</source> <volume>569</volume>, <fpage>698</fpage>&#x02013;<lpage>702</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-1234-z</pub-id><pub-id pub-id-type="pmid">31142856</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Geiger</surname> <given-names>F.</given-names></name> <name><surname>Niculescu</surname> <given-names>V.</given-names></name> <name><surname>Magno</surname> <given-names>M.</given-names></name> <name><surname>Benini</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Smarthand: towards embedded smart hands for prosthetic and robotic applications,&#x0201D;</article-title> in <source>2021 IEEE Sensors Applications Symposium (SAS)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/SAS51076.2021.9530050</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Man</surname> <given-names>Q.</given-names></name> <name><surname>Hu</surname> <given-names>C.</given-names></name> <name><surname>Asghar</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>A skin-inspired tactile sensor for smart prosthetics</article-title>. <source>Sci. Robot</source>. <volume>3</volume>:<fpage>eaat0429</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.aat0429</pub-id><pub-id pub-id-type="pmid">33141753</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Feng</surname> <given-names>X.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Attention erasing and instance sampling for weakly supervised object detection</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>62</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3339956</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yi</surname> <given-names>Z.</given-names></name> <name><surname>Xu</surname> <given-names>T.</given-names></name> <name><surname>Shang</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name></person-group> (<year>2022</year>). <article-title>Genetic algorithm-based ensemble hybrid sparse elm for grasp stability recognition with multimodal tactile signals</article-title>. <source>IEEE Trans. Industr. Electr</source>. <volume>70</volume>, <fpage>2790</fpage>&#x02013;<lpage>2799</lpage>. <pub-id pub-id-type="doi">10.1109/TIE.2022.3170631</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Bai</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Shen</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Target classification method of tactile perception data with deep learning</article-title>. <source>Entropy</source> <volume>23</volume>:<fpage>1537</fpage>. <pub-id pub-id-type="doi">10.3390/e23111537</pub-id><pub-id pub-id-type="pmid">34828235</pub-id></citation></ref>
</ref-list>
</back>
</article>