<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2023.1204445</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An improved fused feature residual network for 3D point cloud data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Gezawa</surname> <given-names>Abubakar Sulaiman</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2278515/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Liu</surname> <given-names>Chibiao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Jia</surname> <given-names>Heming</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1818539/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Nanehkaran</surname> <given-names>Y. A.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Almutairi</surname> <given-names>Mubarak S.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1709078/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chiroma</surname> <given-names>Haruna</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>College of Information Engineering, Fujian Key Lab of Agriculture IOT Application, Sanming University, Sanming</institution>, <addr-line>Fujian</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Software Engineering, School of Information Engineering, Yancheng Teachers University, Yancheng</institution>, <addr-line>Jiangsu</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>College of Computer Science and Engineering, University of Hafr Al-Batin</institution>, <addr-line>Hafar Al Batin</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff4"><sup>4</sup><institution>College of Computer Science and Engineering Technology, Applied College, University of Hafr Al-Batin</institution>, <addr-line>Hafar Al Batin</addr-line>, <country>Saudi Arabia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Si Wu, Peking University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Laith Abualigah, Amman Arab University, Jordan; Saghir Alfasly, Mayo Clinic, United States</p></fn>

<corresp id="c001">&#x0002A;Correspondence: Chibiao Liu <email>lcbsmc&#x00040;163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>30</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1204445</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>08</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Gezawa, Liu, Jia, Nanehkaran, Almutairi and Chiroma.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Gezawa, Liu, Jia, Nanehkaran, Almutairi and Chiroma</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Point clouds have evolved into one of the most important data formats for 3D representation. It is becoming more popular as a result of the increasing affordability of acquisition equipment and growing usage in a variety of fields. Volumetric grid-based approaches are among the most successful models for processing point clouds because they fully preserve data granularity while additionally making use of point dependency. However, using lower order local estimate functions to close 3D objects, such as the piece-wise constant function, necessitated the use of a high-resolution grid in order to capture detailed features that demanded vast computational resources. This study proposes an improved fused feature network as well as a comprehensive framework for solving shape classification and segmentation tasks using a two-branch technique and feature learning. We begin by designing a feature encoding network with two distinct building blocks: layer skips within, batch normalization (BN), and rectified linear units (ReLU) in between. The purpose of using layer skips is to have fewer layers to propagate across, which will speed up the learning process and lower the effect of gradients vanishing. Furthermore, we develop a robust grid feature extraction module that consists of multiple convolution blocks accompanied by max-pooling to represent a hierarchical representation and extract features from an input grid. We overcome the grid size constraints by sampling a constant number of points in each grid using a simple K-points nearest neighbor (KNN) search, which aids in learning approximation functions in higher order. The proposed method outperforms or is comparable to state-of-the-art approaches in point cloud segmentation and classification tasks. In addition, a study of ablation is presented to show the effectiveness of the proposed method.</p>
</abstract>
<kwd-group>
<kwd>point clouds</kwd>
<kwd>part segmentation</kwd>
<kwd>classification</kwd>
<kwd>shape features</kwd>
<kwd>3D objects recognition</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="8"/>
<equation-count count="12"/>
<ref-count count="75"/>
<page-count count="16"/>
<word-count count="11266"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Three-dimensional (3D) data are a great asset in the computer vision field since it contains detailed information on the whole geometry of detected objects and scenes. With the availability of massive 3D datasets and processing power, it is now possible to apply deep learning to learn specific tasks on 3D data such as segmentation with classification (Varga et al., <xref ref-type="bibr" rid="B58">2020</xref>; Erg&#x000FC;n and Sahillioglu, <xref ref-type="bibr" rid="B15">2023</xref>; Qi et al., <xref ref-type="bibr" rid="B47">2023</xref>), recognition, and correspondence (Long et al., <xref ref-type="bibr" rid="B41">2021</xref>). There are several categories of 3D data representations including point cloud, voxel, mesh, multi views, octree, and many others. A comprehensive overview of point clouds and other 3D data representations may be found in the study by Bello et al. (<xref ref-type="bibr" rid="B5">2020</xref>) and Gezawa et al. (<xref ref-type="bibr" rid="B18">2020</xref>). Point cloud data processing employs a variety of approaches. Following dispatching a point cloud to a voxel grid that is quantized spatially in the grid space, volumetric models use a volumetric convolution to compute (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>; Choy et al., <xref ref-type="bibr" rid="B10">2016</xref>). Volumetric approaches correlate points with grid positions by using grids as data structuring technique and convolutional kernels in 3D to get data from nearby voxels. Although grid data structures are efficient, to maintain the granularity of the data position, a high voxel resolution is essential. The amount of processing and memory used grows in a cubical relationship with the voxel resolution since large point clouds are expensive to process. Furthermore, most point clouds contain &#x0007E;90% empty voxels (Zhou and Tuzel, <xref ref-type="bibr" rid="B75">2018</xref>), processing no data could use a lot of computing power. Point-based models are another type of point cloud data processing paradigm. Unlike volumetric models, point-based models offer effective computation but have poor data organization. For instance, PointNet (Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>) aggregates the data in the network&#x00027;s final stage using the point cloud without quantization, as a result the precise locations of the data are preserved. However, the cost of computation rises in lockstep with the point number. Subsequent studies (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>; Wang et al., <xref ref-type="bibr" rid="B60">2018</xref>; Yifan et al., <xref ref-type="bibr" rid="B73">2018</xref>; Qiangeng et al., <xref ref-type="bibr" rid="B48">2019</xref>; Wang Y. et al., <xref ref-type="bibr" rid="B64">2019</xref>) aggregate information using a downsampling approach at each layer. Graph convolutional networks (GCN) have been used in the network layer to generate a local graph for each point cluster (Simonovsky and Komodakis, <xref ref-type="bibr" rid="B52">2017</xref>; Kuangen et al., <xref ref-type="bibr" rid="B31">2019</xref>; Wang L. et al., <xref ref-type="bibr" rid="B63">2019</xref>; Li et al., <xref ref-type="bibr" rid="B35">2023</xref>) that can be regarded as a variant of the PointNet&#x0002B;&#x0002B; design (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>). This architecture, however, is costly in terms of data structuring [e.g., Random Point Sampling (RPS)]. As reported by Zhijian et al. (<xref ref-type="bibr" rid="B74">2019</xref>), data structuring costs account for up to 88% of the entire computational cost in three common point-based models (Li Y. et al., <xref ref-type="bibr" rid="B37">2018</xref>; Yifan et al., <xref ref-type="bibr" rid="B73">2018</xref>; Wang Y. et al., <xref ref-type="bibr" rid="B64">2019</xref>). Furthermore, SO-Net (Li J. et al., <xref ref-type="bibr" rid="B36">2018</xref>) employs the self-organizing map (SOM; Kohonen, <xref ref-type="bibr" rid="B30">1998</xref>) to create a set of points used to model a point cloud&#x00027;s spatial pattern. Even though SO-Net considers a point cloud&#x00027;s regional correlation, SOM is trained independently. As a result, SOM&#x00027;s spatial modeling and a specific point cloud task are no longer coupled. DGCB-Net (Tian et al., <xref ref-type="bibr" rid="B57">2020</xref>) uses cutting-edge convolutional layers built by weight-shared multiple-layer perceptrons (MLPs), to automatically extract local features from the point cloud graph structure. A feature aggregation is formed by concatenating the features received from all edge convolutional layers. Rather than stacking multiple layers deep, the DGCB-Net adopts a strategy to flatly extend point cloud feature aggregation.</p>
<p>In this study, we utilize deep learning to develop an approach that manage enormous 3D object datasets without compromising shape resolution. The majority of handcrafted 3D features are limited to low 3D resolutions. For example, Chiotellis et al. (<xref ref-type="bibr" rid="B9">2016</xref>) and Zhou and Tuzel (<xref ref-type="bibr" rid="B75">2018</xref>) require each 3D model in the datasets to be down-sampled to 20,000 faces with Meshlab before they can be fed into the system. Additionally, a method is provided that can handle structural variations in 3D objects without the need for data pre-processing. Many machine learning algorithms, such as the support vector machine (SVM), are effective when the datasets are small and well-curated, which implies that the data have been carefully pre-processed and requires human intervention. To address these challenges, this study offers an improved fused feature network, an end-to-end framework that solves shape classification and segmentation tasks using a two-branch technique with feature representation learning. To efficiently simplify the network, we start by developing a feature encoding network with two independent building blocks and layer skips with batch normalization and ReLU in between. Because there are few layers through which to propagate, using the layer skips speeds up learning and lessens the effect of gradients vanishing. <xref ref-type="fig" rid="F1">Figure 1</xref> presents the entire network structure of the approach. In addition, we create a detail grid feature extraction module, which comprises various convolution blocks accompanied by a max-pooling to represent a hierarchical representation of several feature representations and extracts features from the input grid. Max-pooling is used in each of the pooling layers, resulting in each spatial dimension having a smaller grid and helps to manage overfitting by gradually lowering the representation&#x00027;s spatial dimension, the parameters in the network, and the amount of processing. This module includes a regular-structured enclosing volumetric grid that helps capture details and features hierarchically. To extract features of high-resolution inputs, this module is utilized in conjunction with the feature encoding network. To pull through the limitation of the grid size, the local region in every grid sampled a constant number of points using a simple KNN search which aids in learning approximation functions in higher order to better characterize the details of the features.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The complete architecture of the proposed method. The network is divided into three branches. The feature encoding network extract features from the input grid in <bold>(A)</bold>. The DGFE module exploits the detailed shape characteristics in <bold>(B)</bold>. The feature fusion unit which has two consecutive convolutional layers, fuses the features from the two branches to produce a feature with improved contextual representation by exploiting both local and global shape structures in <bold>(C)</bold>. See also Section 3.5.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0001.tif"/>
</fig>
<p>Our major contributions are as follows:</p>
<list list-type="bullet">
<list-item><p>We design an effective module named detail grid feature extraction (DGFE) module. This module aids 3D convolutions to hierarchically capture global information and reduces the grid size in each spatial dimension as well as managing overfitting by gradually lowering the spatial dimension of the representation making it viable for high-resolution 3D objects.</p></list-item>
<list-item><p>We design a feature encoding network that uses two different building blocks with layer skips containing batch normalization and ReLU in between, resulting in fewer layers in the early training phase which helps speed learning and reduces the effect of gradients vanishing since there are few layers through which to propagate.</p></list-item>
<list-item><p>We built a network using the modules that have been proposed, which achieves a notable balance of accuracy and speed.</p></list-item>
</list></sec>
<sec id="s2">
<title>2. Related work</title>
<sec>
<title>2.1. 3D learning using voxel-based methods</title>
<p>To build on the advance of CNN models on images (He et al., <xref ref-type="bibr" rid="B20">2016a</xref>; Huang et al., <xref ref-type="bibr" rid="B24">2017</xref>), Voxnet and its revisions (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>; Wang and Posner, <xref ref-type="bibr" rid="B61">2015</xref>; Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>; Brock et al., <xref ref-type="bibr" rid="B6">2016</xref>) start by converting a point cloud to a grid occupancy and then used convolution in a volumetric form. To overcome the problem of rising memory usage due to cubical expansion, OctNet creates structures like a tree for non-empty voxels to avoid computing in space. While the volumetric approach is effective at structuring data, it suffers from poor computational effectiveness and data granularity loss. Transformers have lately been incorporated into the model designs of many 3D vision approaches in response to the success of transformer-based designs in the two-dimensional (2D) domain. The transformer has improved previous 3D learning techniques because of its ability to read remote input and provide task-specific inductive biases. The point-voxel transformer for single-stage 3D detection (PVT-SSD) proposed by Yang et al. (<xref ref-type="bibr" rid="B69">2023</xref>) uses input-dependent query initialization and voxel-based sparse convolutions for strong feature encoding. The PVT-SSD overcame the drawbacks of both point clouds and voxels by combining their advantages. To reduce farthest point sampling (FPS) runtime, they used sparse convolutions to transform points into a limited number of voxels rather than directly sampling them. They also sampled non-empty voxels. The voxel features were adaptively blended with the point features to make up for the difficulty of quantization.</p></sec>
<sec>
<title>2.2. 3D learning using point cloud-based methods</title>
<p>Charles et al. (<xref ref-type="bibr" rid="B7">2017</xref>); Qi et al. (<xref ref-type="bibr" rid="B46">2017</xref>) pioneered the use of point-based models which used pooling to aggregate the point features to achieve the permutation invariant. To better capture local characteristics, methods such as kernel correlation (Atzmon et al., <xref ref-type="bibr" rid="B2">2018</xref>; Wu et al., <xref ref-type="bibr" rid="B67">2019</xref>) and extended convolutions (Thomas et al., <xref ref-type="bibr" rid="B56">2019</xref>) are proposed. To resolve the ambiguity, the local point order is predicted by PointCNN (Li Y. et al., <xref ref-type="bibr" rid="B37">2018</xref>) while RSNet (Huang et al., <xref ref-type="bibr" rid="B25">2018</xref>) sequentially consumes points from various directions. In methods based on points, the cost of computation grows linearly with the points input. The cost of structuring data, nevertheless, turned out to be a performance bottleneck for large inputs. Recently, a dynamic sparse voxel transformer (DSVT) was presented by Wang et al. (<xref ref-type="bibr" rid="B62">2023</xref>) in an effort to widen the uses of transformers so that they may serve as a solid foundation for outdoor 3D perception just as they do for 2D vision. A number of local regions are split up into smaller ones in each window using DSVT based on sparsity, and each window&#x00027;s attributes are then computed fully in parallel. Another recent point cloud classification framework named point content-based transformer (PointConT) was introduced by Liu et al. (<xref ref-type="bibr" rid="B40">2023</xref>), and it employs local self-attention in the space of features rather than the 3D space. One of the main advantages of PointConT is that it takes advantage of the locality of points in the feature space by clustering sampled points with similar features into the same class and computing self-attention within each class, allowing for an efficient trade-off between collecting long-range dependencies and computational complexity.</p></sec>
<sec>
<title>2.3. Strategies for point data structuring</title>
<p>The majority of point-based methods (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>; Li Y. et al., <xref ref-type="bibr" rid="B37">2018</xref>; Bello et al., <xref ref-type="bibr" rid="B4">2021</xref>; Gezawa et al., <xref ref-type="bibr" rid="B17">2021</xref>) employ FPS (Eldar et al., <xref ref-type="bibr" rid="B13">1997</xref>) to sample uniformly distributed group centers. However, it does not account for the subsequent processing of the sampled points which may result in suboptimal performance. Random point sampling (RPS) has the advantage of having a minimal downtime. It is indeed, nevertheless, sensitive to variation in density. The KNN search we used for sampling the local region in each grid cell combines sampling and neighbor querying in a single step, making it faster than RPS.</p>
<p>SO-Net (Li J. et al., <xref ref-type="bibr" rid="B36">2018</xref>), on the other hand, creates a self-organizing map. To split the spaces, KDNet (Klokov and Lempitsky, <xref ref-type="bibr" rid="B29">2017</xref>) employs kd-tree. Gumble subset sampling is used instead of FPS by Yang et al. (<xref ref-type="bibr" rid="B70">2019</xref>). To create super points, Landrieu and Simonovsky (<xref ref-type="bibr" rid="B32">2018</xref>) employs a clustering algorithm. The majority of these approaches are either too slow or necessitate structure preprocessing. VoxelNet (Le and Duan, <xref ref-type="bibr" rid="B33">2018</xref>; Zhou and Tuzel, <xref ref-type="bibr" rid="B75">2018</xref>), for example, blends point-based and volumetric approaches by performing voxel convolution and employing the study by Charles et al. (<xref ref-type="bibr" rid="B7">2017</xref>) inside each voxel. Similar concepts are used by the fast model (Zhijian et al., <xref ref-type="bibr" rid="B74">2019</xref>), whereas Lu et al. (<xref ref-type="bibr" rid="B42">2022</xref>) made use of ball query with graph convolution layers. However, the number of points is not steadily decreased over all layers. Our DGFE module, however, utilized max-pooling in each of the pooling layers, resulting in each spatial dimension having a smaller grid allowing it to be used for high-resolution 3D objects. Apart from those features, the local region in every grid sampled a constant number of points using a simple KNN search which aids in learning approximation functions in higher order to better characterize the detailed features.</p></sec></sec>
<sec id="s3">
<title>3. The proposed method</title>
<p>In this section, the KNN search for local region sampling is first introduced. Following that, we propose the feature encoding network that serves as the basis of the enhanced fused feature network. The split-transform-merge paradigm, which is based on the residual learning framework, is one of the primary building block we employ to design our feature encoding network (<xref ref-type="fig" rid="F1">Figure 1A</xref>). One of the primary benefits of employing the residual network is its simplicity in training networks with many layers without raising the training error percentage. It also aids in solving the vanishing gradient problem by applying identity mapping. To compensate for structural changes in 3D objects, our feature encoding network employs two different building blocks [feature encoding block (FEB) unit A and feature encoding block (FEB) unit B], with layer skips in between. We begin with 3x3x3 convolutions twice, followed by 1x1x1 convolutions with a stride in each convolution to accommodate both small and large datasets without possible overfitting and to lower the spatial dimension of the representation. Then, we introduce the detail grid feature extraction module and finally the feature fusion unit. The complete framework is presented <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<sec>
<title>3.1. KNN search for local region sampling</title>
<p>Point clouds are typically represented as raw coordinates of points in 3D space. Here, we will go over how our model extracts features from 3D objects when given a point cloud of number of points (N) as input. When provided with an input of N &#x000D7; 3 set of point clouds, the object is then subdivided into equal-sized 3D voxels, such as 64 &#x000D7; 64 &#x000D7; 64, 16 &#x000D7; 16 &#x000D7; 16 or 8 &#x000D7; 8 &#x000D7; 8. Using KNN, K points will be sampled from each grid cell. To avoid extra computation, those with empty points will be padded with zeros. In contrast to standard KNN, in which the search area consists of all points, it just needs to search among non-empty voxels in our situation, making the query much faster. Unlike VoxNet (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>) which represents the 3D structure using an occupancy grid, we build a grid from point clouds and designate the grid&#x00027;s key feature to the points that are inside each grid. Some grids, on the other hand, may contain a different point number. This implies that we need a grid that will share kernels in 3D convolution. Moreover, for addressing this constraint, we utilized a sampling strategy that ensures each grid has an equal point number. In particular, if there are beyond K points in the grid, we use the KNN sampling strategy to choose K points from the total points. K points are sampled with substitution when the points inside a grid are below K. Consequently, each grid will have the same number of points, allowing us to encode the grid feature so that each grid feature has the same feature size vector which enables us to extract hierarchical features of the object using 3D convolutional kernels.</p></sec>
<sec>
<title>3.2. Feature encoding network</title>
<p>We concentrate on developing a robust network for shape classification and segmentation that achieves a notable balance of accuracy and speed. The feature encoding network is one of the key blocks that we create by making use of the split-transform-merge paradigm, inspired by the residual learning framework design in the study by Szegedy et al. (<xref ref-type="bibr" rid="B55">2015</xref>), He et al. (<xref ref-type="bibr" rid="B20">2016a</xref>,<xref ref-type="bibr" rid="B21">b</xref>), and Elhassan et al. (<xref ref-type="bibr" rid="B14">2021</xref>) and leveraging its powerful representational ability. These networks are scalable structures that bundle building units with the same linked shape which are referred to as residual units or blocks. The original blocks in the study by He et al. (<xref ref-type="bibr" rid="B21">2016b</xref>) compute as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In this case, <italic>I</italic><sub><italic>i</italic></sub> represents the i-th block&#x00027;s input feature. <italic>Weights</italic><sub><italic>i</italic></sub> &#x0003D; {<italic>Weights</italic><sub><italic>i</italic></sub>, <italic>k</italic>&#x02223;1 &#x02264; <italic>k</italic> &#x02264; <italic>K</italic>} contains biases and weights connected to block i-th. K stands for total layers in a block. <italic>f</italic> signifies the block function, such as a pile of convolutional layers of two 3x3 in Equation 1. The operation following element-wise addition is represented by the function <italic>f</italic>, which is ReLU in Equation 1. The <italic>h</italic> function is designated as an identity mapping: <italic>h</italic>(<italic>I</italic><sub><italic>i</italic></sub>) &#x0003D; <italic>I</italic><sub><italic>i</italic></sub>. Similarly, if function <italic>f</italic> is identity mapping, <italic>I</italic><sub><italic>i</italic>&#x0002B;1</sub>&#x02261;<italic>O</italic><sub><italic>i</italic></sub>. Putting Equation 2 into Equation 1 yields:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>To efficiently accelerate training and reduce the number of parameters, the feature encoding network uses two separate construction blocks, such as Feature encoding block (FEB unit A) and feature encoding block (FEB unit B), with layer skips containing batch normalization (BN) and ReLU in between. The BN and ReLU are regarded as the weight layers&#x00027; pre-activation, according to He et al. (<xref ref-type="bibr" rid="B21">2016b</xref>). We make some minor changes here by using the ReLu with BN and Conv before the addition of operation. We start with 3x3x3 convolutions twice, followed by 1x1x1 convolutions, and then we apply the BN and ReLu before the addition. We use a stride in each convolution to help manage overfitting by gradually reducing the spatial dimension of the representation. The feature encoding network&#x00027;s design is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Building blocks of the feature encoding network with two different layer skips. <bold>(A)</bold> Feature encoding block (FEB unit A) <bold>(B)</bold> Feature encoding block (FEB unit B).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0002.tif"/>
</fig></sec>
<sec>
<title>3.3. Detail grid feature extraction module</title>
<p>To represent numerous hierarchical feature representations, the detail grid feature extraction module employs several convolution blocks and max-pooling and extracts features from the input grid, as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. Max-pooling is used in each of the pooling layers, resulting in each spatial dimension having a smaller grid and helps to manage overfitting by gradually lowering the representation&#x00027;s spatial dimension, the parameters in the network, and the amount of processing. BN (Ioffe and Szegedy, <xref ref-type="bibr" rid="B26">2015</xref>) can be done to any set of network activations using:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>H</mml:mi><mml:mi>u</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>H</italic> and <italic>p</italic> are model parameters that have been learned, and <italic>g</italic>(.) denotes a non-linearity being ReLU or sigmoid. By normalizing <italic>z</italic> &#x0003D; <italic>Hu</italic>&#x0002B;<italic>p</italic>, the BN transform can be introduced right before the non-linearity. Since <italic>z</italic> is normalized, <italic>y</italic> &#x0003D; <italic>g</italic>(<italic>Hu</italic>&#x0002B;<italic>p</italic>) can be replaced with</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>H</mml:mi><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the BN (Ioffe and Szegedy, <xref ref-type="bibr" rid="B26">2015</xref>) is used separately for each dimension of <italic>z</italic> &#x0003D; <italic>Hu</italic>, with a distinct set of learned parameters for each dimension. We utilized a 3 &#x000D7; 3 &#x000D7; 3 kernel with stride 1 convolution and a ReLU (Nair and Hinton, <xref ref-type="bibr" rid="B45">2010</xref>) in each convolution layer. The initial block employs 32-filter convolutions, which are then doubled in subsequent blocks. This module offers a regular-structured embedding volumetric grid that supports 3D convolutions in hierarchically capturing global information. To extract features of high-resolution inputs, this module is utilized in conjunction with the feature encoding network. To keep local fine details in early encoder layers, at the same spatial resolution, we connect the encoder network&#x00027;s encoded features to equivalent features extracted from the detail grid feature extraction module.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Detail grid feature extraction module (DGFE Module). This module extracts features from the input grid using many convolution blocks. Max-pooling is used in each of the pooling layers, resulting in each spatial dimension having a smaller grid and helps to manage overfitting by gradually lowering the representation&#x00027;s spatial dimension, the parameters in the network, and the amount of processing.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0003.tif"/>
</fig></sec>
<sec>
<title>3.4. Feature fusion unit</title>
<p>The feature fusion unit is made up of two consecutive convolutional layers. We used 3 &#x000D7; 3 &#x000D7; 3 convolutions twice, with BN and ReLU in between, and a stride in each convolution to help manage overfitting. The proposed DGFE module and the encoding network outputs are fused using a cross-product in the feature fusion unit, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, to produce a feature with improved contextual representation.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Illustration of the detailed design of the feature fusion unit, which consists of two consecutive 3x3x3 convolutions with BN and ReLU in between, as well as a stride in each convolution to help manage overfitting.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0004.tif"/>
</fig></sec>
<sec>
<title>3.5. Network overview</title>
<p>We built a 3D convolutional network with fixed points inside each grid cell, which aids in the learning of local approximation functions in high-order that better capture local shape features. <xref ref-type="fig" rid="F1">Figure 1</xref> presents a diagram of the proposed architecture. The network is made up of two major modules. A feature encoding network that serves as the foundation for extracting features from the input grid, as shown in <xref ref-type="fig" rid="F1">Figure 1A</xref> in Section 3.2, and detail grid feature extraction (DGFE) module which comprises various convolution blocks accompanied with an operation of max-pooling to help in representing several relational features and pull out features from the input (Section 3.3). We hierarchically combine these two modules to form the proposed improved fused feature network. The proposed DGFE module and the encoding network outputs are fused in the feature fusion unit containing two consecutive convolutional layers (<xref ref-type="fig" rid="F1">Figure 1C</xref>) to produce a feature with improved contextual representation by utilizing both local and global shape structures.</p>
<p>The point cloud is first normalized within the unit box. In each grid, the coordinates of the points are piled as features. accordingly, given the appropriate x, y, and z coordinates, a K-point grid has features 3K. In theory, by dividing the sum of points (P) by its grid cells, K can be approximated. To acquire classification scores, the resulting fused feature can be categorized using two fully connected layers. Finally, one additional fully connected layer is added, along with a softmax, which aids in regressing the likelihood in every group. The whole layer&#x00027;s nodes correspond to the set of categories of objects inside the dataset. To generate the segmentation, the segmentation network decodes the retrieved features. To create the output, this network upsamples and combines the features. For every cell inside the grid, this network produces K&#x0002B;1 labels, as for K points in that cell equivalent to K labels and one more label level cell. Obtaining ground truth labels of object components, we chose its greater label among the labels of points within every cell. Unoccupied Cells are tagged &#x00022;no label.&#x00022; Before actually acquiring the object part, we perform a deconvolution operation by concatenating the feature obtained from the feature fusion unit, with the feature retrieved out of each block of the feature encoding network.</p></sec></sec>
<sec id="s4">
<title>4. Experiments</title>
<p>In this section, a number of datasets including ModelNet10 and ModelNet40 (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>) for object classification and part segmentation on ShapeNetPart (Yi et al., <xref ref-type="bibr" rid="B72">2016</xref>) were used to assess the performance of the proposed network. We discuss the dataset&#x00027;s specifics and the evaluation metrics in Section 4.1. The implementation protocol discussion presented in Section 4.2. In Sections 4.3, 4.4, and 4.5, we discuss some experimental results from applying the proposed network to classify shapes on ModelNet, measure precision-recall on ModelNet10, and segment parts on ShapeNetPart. In Section 4.6, we demonstrate the advantages of the proposed method by conducting a good set of ablation experimental tests to evaluate various setup adjustments.</p>
<sec>
<title>4.1. Datasets and evaluation metrics</title>
<p><bold>ModelNet dataset:</bold> This is indeed a notable dataset. It comprises two datasets with CAD models in 10 and 40 categories, respectively. ModelNet10 is made up of 4,899 object instances including 2,468 training samples and 909 testing samples. ModelNet40 is made up of 12,311 object instances, 9,843 of which are in the training set and 3,991 samples in the testing set. For object classification on the ModelNet dataset, we employed accuracy as the assessment metric.</p>
<p><bold>ShapeNetPart dataset:</bold> There are 16,881 shapes in this dataset, divided into 16 categories and annotated with a combined amount of 50 components. A considerable share of shape categories is partitioned into 2&#x02013;5 segments. We, then, used mean intersection over union (mIoU) for evaluation. For every part shape within the object category, we calculate the union of prediction and ground truth. The mIoU was computed using Equation 6 as follows:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>G</mml:mi><mml:mo>-</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>G</italic>, <italic>P</italic>, and <italic>X</italic> denote the number of ground truth points, predicted positive points, and true positive points, respectively. The mIoU is obtained by taking the average of each class&#x00027;s IoU.</p></sec>
<sec>
<title>4.2. Implementation protocol</title>
<p>In Python, the proposed method was implemented using the Tensorflow deep learning library. Each experiment is conducted on an Nvidia Geforce Titan GTX GPU, CUDA 10.1, and CuDNN 7.1 with RAM of 12 GB. For the classification task, we test with various parameters setup including different grid sizes and K values. Each point&#x00027;s location is jittered with a standard deviation of 0.02. The batch size is 32, and batch normalization is used for all layers. For both the segmentation and classification tasks, we used the cross-entropy loss to improve the discrimination of the class features. We utilized an initial learning rate of 10<sup>&#x02212;4</sup> and employ Adam optimizer (Kingma and Ba, <xref ref-type="bibr" rid="B28">2015</xref>).</p>
<p><bold>Loss function:</bold> Over the years, a wide range of loss functions have been proposed to perform 3D shape analysis tasks. For example, the cross-entropy loss was already been utilized successfully in many shape analysis tasks. Although the network can be trained using cross-entropy loss alone, we employ a combination of Shape loss (Wei et al., <xref ref-type="bibr" rid="B65">2020</xref>) and modified cross-entropy loss (Huang et al., <xref ref-type="bibr" rid="B23">2019</xref>) to make the class features more discriminatory. The Shape Loss is given as follows:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>M</italic> is the shapes&#x00027;s class label, <italic>L</italic><sub><italic>s</italic></sub> is a cross-entropy loss based on shape feature <italic>S</italic>, and <italic>C</italic> is a classifier.</p>
<p>Moreover, the cross-entropy loss is given as follows:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>-</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>Q</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For each sample, <italic>Q</italic>&#x02208;[0, 1] is the likelihood of the network output and <italic>z</italic> represents the class ground truth. To minimize the weight of easily categorized samples, the cross-entropy function can be reshaped by inserting a hyperparameter that aids in weight balancing.</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>-</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msup><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>Q</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msup><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Once a sample is successfully identified, <inline-formula><mml:math id="M10"><mml:mi>Q</mml:mi><mml:mover class="overset"><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow><mml:mrow></mml:mrow></mml:mover><mml:mn>1</mml:mn></mml:math></inline-formula>, the factor <inline-formula><mml:math id="M11"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mover class="overset"><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow><mml:mrow></mml:mrow></mml:mover><mml:mn>0</mml:mn></mml:math></inline-formula>; Alternatively, when <italic>Q</italic> is small, the factor (1&#x02212;<italic>Q</italic>) approaches 1. Our total loss is the combination of this two losses as follows:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>-</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></sec>
<sec>
<title>4.3. Classification on ModelNet</title>
<p>We use the PointNet (Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>) convention to prepare the data. Input points are set to 1,024 by default. Furthermore, we improve performance by incorporating more points and surface normal. To analyze various models to varying degrees of speed and accuracy, the network is trained with varying settings to balance speed and performance (Section 4.6). The variants are in different grid sizes and K values.</p>
<sec>
<title>4.3.1. Classification on ModelNet10</title>
<p><bold>Comparison:</bold> The proposed improved fused feature residual network approach was compared with a number of state-of-the-art methods, as shown in <xref ref-type="table" rid="T1">Table 1</xref>. The proposed method outperforms the majority of previous voxel-based techniques in terms of &#x00022;overall accuracy&#x00022; including VoxNet (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>), 3DShapeNets (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>), 3DGAN (Wu et al., <xref ref-type="bibr" rid="B66">2016</xref>), VSL (Liu et al., <xref ref-type="bibr" rid="B39">2018</xref>), and BV-CNN&#x00027;s (Ma et al., <xref ref-type="bibr" rid="B43">2017</xref>). Although VRN (Brock et al., <xref ref-type="bibr" rid="B6">2016</xref>), which combines many networks, outperforms our method in ModelNet classification, their network structure is quite complex, with each network being trained separately and taking many days to complete, making them unsuitable for large datasets. When compared with point cloud-based methods, the proposed method outperforms many of them, including Dominguez et al. (<xref ref-type="bibr" rid="B12">2018</xref>), OctNet (Riegler et al., <xref ref-type="bibr" rid="B49">2017</xref>), ECC (Simonovsky and Komodakis, <xref ref-type="bibr" rid="B52">2017</xref>), DGCB-Net (Tian et al., <xref ref-type="bibr" rid="B57">2020</xref>), and VACWGAN-GP (Erg&#x000FC;n and Sahillioglu, <xref ref-type="bibr" rid="B15">2023</xref>). The DGFE module helps 3D convolutions hierarchically acquire global information, allowing the network to capture the contextual neighborhood of points. Despite using viewpoints in a predefined sequence, as opposed to any random views by DeepPano (Shi et al., <xref ref-type="bibr" rid="B51">2015</xref>), Gan classifier (Varga et al., <xref ref-type="bibr" rid="B58">2020</xref>), GPSP-DWRN (Long et al., <xref ref-type="bibr" rid="B41">2021</xref>), OrthographicNet (Kasaei, <xref ref-type="bibr" rid="B27">2019</xref>), PANORAMA-NN (Sfikas et al., <xref ref-type="bibr" rid="B50">2017</xref>), and SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>) both of which are multi-view techniques, the method outperforms these approaches, making it suitable for high resolution input. The proposed method also outperforms PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>), a mesh-based 3D representation network that combined the features in a much smaller dimension using PolyShape&#x00027;s multi-resolution structure.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Object classification accuracy (%) on ModelNet10.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Input</bold></th>
<th valign="top" align="center"><bold>Acc (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">VoxNet (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">92.0</td>
</tr> <tr>
<td valign="top" align="left">3DShapeNet (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">83.5</td>
</tr> <tr>
<td valign="top" align="left">3DGAN (Wu et al., <xref ref-type="bibr" rid="B66">2016</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">91.0</td>
</tr> <tr>
<td valign="top" align="left">VSL (Liu et al., <xref ref-type="bibr" rid="B39">2018</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">91.0</td>
</tr> <tr>
<td valign="top" align="left">BV-CNNs (Ma et al., <xref ref-type="bibr" rid="B43">2017</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">92.3</td>
</tr> <tr>
<td valign="top" align="left">VRN (Brock et al., <xref ref-type="bibr" rid="B6">2016</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">97.1</td>
</tr> <tr>
<td valign="top" align="left">PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>)</td>
<td valign="top" align="center">Mesh</td>
<td valign="top" align="center">94.9</td>
</tr> <tr>
<td valign="top" align="left">DeepPano (Shi et al., <xref ref-type="bibr" rid="B51">2015</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">85.4</td>
</tr> <tr>
<td valign="top" align="left">OrthographicNet (Kasaei, <xref ref-type="bibr" rid="B27">2019</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">88.5</td>
</tr> <tr>
<td valign="top" align="left">PANORAMA-NN (Sfikas et al., <xref ref-type="bibr" rid="B50">2017</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">91.1</td>
</tr> <tr>
<td valign="top" align="left">SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">94.8</td>
</tr> <tr>
<td valign="top" align="left">Geometry-image (Sinha et al., <xref ref-type="bibr" rid="B53">2016</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">88.4</td>
</tr> <tr>
<td valign="top" align="left">Gan Classifier (Varga et al., <xref ref-type="bibr" rid="B58">2020</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">89.2</td>
</tr> <tr>
<td valign="top" align="left">GPSP-DWRN (Long et al., <xref ref-type="bibr" rid="B41">2021</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">92.4</td>
</tr> <tr>
<td valign="top" align="left">G3DNet (Dominguez et al., <xref ref-type="bibr" rid="B12">2018</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">93.1</td>
</tr> <tr>
<td valign="top" align="left">OctNet (Riegler et al., <xref ref-type="bibr" rid="B49">2017</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">90.4</td>
</tr> <tr>
<td valign="top" align="left">ECC (Simonovsky and Komodakis, <xref ref-type="bibr" rid="B52">2017</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">90.0</td>
</tr> <tr>
<td valign="top" align="left">DGCB-Net (Tian et al., <xref ref-type="bibr" rid="B57">2020</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">94.6</td>
</tr> <tr>
<td valign="top" align="left">VACWGAN-GP (Erg&#x000FC;n and Sahillioglu, <xref ref-type="bibr" rid="B15">2023</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">91.7</td>
</tr> <tr>
<td valign="top" align="left">(Ours)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center"><bold>95.6</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.3.2. Classification on ModelNet40</title>
<p><bold>Comparison:</bold> We further tested the effectiveness and applicability of the proposed approach using the ModelNet40 dataset. <xref ref-type="table" rid="T2">Table 2</xref> compares the classification accuracy of the proposed method to that of alternative scalable 3D representations techniques on the ModelNet40 datasets. As observed, the proposed method performs better than VoxNet (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>), 3DGAN (Wu et al., <xref ref-type="bibr" rid="B66">2016</xref>), 3DShapeNets (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>), NormalNet, VACWGAN-GP (Wang et al., <xref ref-type="bibr" rid="B59">2019a</xref>; Erg&#x000FC;n and Sahillioglu, <xref ref-type="bibr" rid="B15">2023</xref>), DPRNet (Arshad et al., <xref ref-type="bibr" rid="B1">2019</xref>), Pointwise (Hua et al., <xref ref-type="bibr" rid="B22">2018</xref>), BV-CNN&#x00027;s (Ma et al., <xref ref-type="bibr" rid="B43">2017</xref>), NPCEM (Song et al., <xref ref-type="bibr" rid="B54">2020</xref>), ECC (Simonovsky and Komodakis, <xref ref-type="bibr" rid="B52">2017</xref>), PointNet (Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>), Geometry image (Sinha et al., <xref ref-type="bibr" rid="B53">2016</xref>), VSL (Liu et al., <xref ref-type="bibr" rid="B39">2018</xref>), GIFT (Bai et al., <xref ref-type="bibr" rid="B3">2016</xref>), FPNN (Li et al., <xref ref-type="bibr" rid="B38">2016</xref>), DGCB-Net (Tian et al., <xref ref-type="bibr" rid="B57">2020</xref>), and DeepNN (Gao et al., <xref ref-type="bibr" rid="B16">2022</xref>) that utilized mesh 3D data. The recent RECON (Qi et al., <xref ref-type="bibr" rid="B47">2023</xref>) and PointConT (Liu et al., <xref ref-type="bibr" rid="B40">2023</xref>) slightly outperformed our technique, which could be attributed to their usage of transformers and pre-train models. The improved fused feature residual network offers a significant advantage over the bulk of voxel and point cloud-based approaches, as shown in <xref ref-type="table" rid="T2">Table 2</xref>. The proposed method performs below VRN (Brock et al., <xref ref-type="bibr" rid="B6">2016</xref>), which makes usage of 24 rotating replicas for training and voting when compared with non-voxel-based approaches. Additionally, the proposed method outperformed PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>), a mesh-based 3D representation network that integrated the features in a much fewer dimension using PolyShape&#x00027;s multi-resolution structure. It is also worth noting that the improved fused feature residual network proposed already has a high level of accuracy, with a score of above 90%. This may be attributed to the fact that our feature encoding network together with the DGFE module, directly extracts features from the input grid and represents an organized structure of numerous feature representations.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Object classification accuracy (%) on ModelNet40.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Input</bold></th>
<th valign="top" align="center"><bold>Acc (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">VoxNet (Maturana and Scherer, <xref ref-type="bibr" rid="B44">2015</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">83.0</td>
</tr> <tr>
<td valign="top" align="left">3DShapeNet (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">77.0</td>
</tr> <tr>
<td valign="top" align="left">3DGAN (Wu et al., <xref ref-type="bibr" rid="B66">2016</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">83.3</td>
</tr> <tr>
<td valign="top" align="left">VSL (Liu et al., <xref ref-type="bibr" rid="B39">2018</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">84.5</td>
</tr> <tr>
<td valign="top" align="left">BV-CNNs (Ma et al., <xref ref-type="bibr" rid="B43">2017</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">85.4</td>
</tr> <tr>
<td valign="top" align="left">VRN (Brock et al., <xref ref-type="bibr" rid="B6">2016</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">95.5</td>
</tr> <tr>
<td valign="top" align="left">NormalNet (Wang et al., <xref ref-type="bibr" rid="B59">2019a</xref>)</td>
<td valign="top" align="center">Volume</td>
<td valign="top" align="center">88.6</td>
</tr> <tr>
<td valign="top" align="left">DeepNN (Gao et al., <xref ref-type="bibr" rid="B16">2022</xref>)</td>
<td valign="top" align="center">Mesh</td>
<td valign="top" align="center">91.0</td>
</tr> <tr>
<td valign="top" align="left">PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>)</td>
<td valign="top" align="center">Mesh</td>
<td valign="top" align="center">82.8</td>
</tr> <tr>
<td valign="top" align="left">GIFT (Bai et al., <xref ref-type="bibr" rid="B3">2016</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">83.1</td>
</tr> <tr>
<td valign="top" align="left">DeepPano (Shi et al., <xref ref-type="bibr" rid="B51">2015</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">77.6</td>
</tr> <tr>
<td valign="top" align="left">OrthographicNet (Kasaei, <xref ref-type="bibr" rid="B27">2019</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">88.5</td>
</tr> <tr>
<td valign="top" align="left">SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">93.0</td>
</tr> <tr>
<td valign="top" align="left">Geometry-image (Sinha et al., <xref ref-type="bibr" rid="B53">2016</xref>)</td>
<td valign="top" align="center">Image</td>
<td valign="top" align="center">83.9</td>
</tr> <tr>
<td valign="top" align="left">PointNet (Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">89.2</td>
</tr> <tr>
<td valign="top" align="left">PointConT (Liu et al., <xref ref-type="bibr" rid="B40">2023</xref>)</td>
<td valign="top" align="center">Points</td>
<td valign="top" align="center">93.5</td>
</tr> <tr>
<td valign="top" align="left">RECON (Qi et al., <xref ref-type="bibr" rid="B47">2023</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">93.9</td>
</tr> <tr>
<td valign="top" align="left">Pointwise (Hua et al., <xref ref-type="bibr" rid="B22">2018</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">86.1</td>
</tr> <tr>
<td valign="top" align="left">NPCEM (Song et al., <xref ref-type="bibr" rid="B54">2020</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">89.4</td>
</tr> <tr>
<td valign="top" align="left">ECC (Simonovsky and Komodakis, <xref ref-type="bibr" rid="B52">2017</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">83.2</td>
</tr> <tr>
<td valign="top" align="left">DGCB-Net (Tian et al., <xref ref-type="bibr" rid="B57">2020</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">92.9</td>
</tr> <tr>
<td valign="top" align="left">3DCTN (Lu et al., <xref ref-type="bibr" rid="B42">2022</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">91.2</td>
</tr> <tr>
<td valign="top" align="left">VACWGAN-GP (Erg&#x000FC;n and Sahillioglu, <xref ref-type="bibr" rid="B15">2023</xref>)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center">81.3</td>
</tr> <tr>
<td valign="top" align="left">(Ours)</td>
<td valign="top" align="center">Point</td>
<td valign="top" align="center"><bold>93.1</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.3.3. ModelNet40 per-class classification accuracy comparison</title>
<p><xref ref-type="table" rid="T3">Table 3</xref> and <xref ref-type="fig" rid="F5">Figure 5</xref> compared the per-class accuracies of the proposed method to PointNet (Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>), Pointwise (Hua et al., <xref ref-type="bibr" rid="B22">2018</xref>), and DPRNet (Arshad et al., <xref ref-type="bibr" rid="B1">2019</xref>) on ModelNet40 dataset. As shown in <xref ref-type="table" rid="T3">Table 3</xref> and <xref ref-type="fig" rid="F5">Figure 5</xref>, using residual learning and extracting detail features improves per class classification accuracy. The proposed method outperforms PointNet, Pointwise, and DPRNet in key classes such as bathhub, car, bottle dresser, flowerpot, cup, and radio. In terms of average class performance, the method outperformed PointNet (1.2%), Pointwise (6%), and DPRNet (5.5%). <xref ref-type="table" rid="T3">Table 3</xref> illustrates it.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>ModelNet40 per-class classification comparison between PointNet, Pointwise, DPRNet, and (ours).</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>Ours</bold></th>
<th valign="top" align="center"><bold>PointNet</bold></th>
<th valign="top" align="center"><bold>Pointwise</bold></th>
<th valign="top" align="center"><bold>DPRNet</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Avg. class</bold></th>
<th valign="top" align="center"><bold>87.4</bold></th>
<th valign="top" align="center"><bold>86.2</bold></th>
<th valign="top" align="center"><bold>81.4</bold></th>
<th valign="top" align="center"><bold>81.9</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Airplane</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
</tr> <tr>
<td valign="top" align="left">Bathtub</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">82.0</td>
<td valign="top" align="center">76.0</td>
</tr> <tr>
<td valign="top" align="left">Bed</td>
<td valign="top" align="center">94.0</td>
<td valign="top" align="center">94.0</td>
<td valign="top" align="center">93.0</td>
<td valign="top" align="center">95.0</td>
</tr> <tr>
<td valign="top" align="left">Bench</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">75.0</td>
<td valign="top" align="center">68.4</td>
<td valign="top" align="center">80.0</td>
</tr> <tr>
<td valign="top" align="left">Bookshelf</td>
<td valign="top" align="center">88.0</td>
<td valign="top" align="center">93.0</td>
<td valign="top" align="center">91.8</td>
<td valign="top" align="center">85.0</td>
</tr> <tr>
<td valign="top" align="left">Bottle</td>
<td valign="top" align="center">98.0</td>
<td valign="top" align="center">94.0</td>
<td valign="top" align="center">93.9</td>
<td valign="top" align="center">95.0</td>
</tr> <tr>
<td valign="top" align="left">Bowl</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">95.0</td>
</tr> <tr>
<td valign="top" align="left">Car</td>
<td valign="top" align="center">99.0</td>
<td valign="top" align="center">97.9</td>
<td valign="top" align="center">95.6</td>
<td valign="top" align="center">91.0</td>
</tr> <tr>
<td valign="top" align="left">Chair</td>
<td valign="top" align="center">97.0</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">97.0</td>
</tr> <tr>
<td valign="top" align="left">Cone</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">90.0</td>
</tr> <tr>
<td valign="top" align="left">Cup</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">70.0</td>
<td valign="top" align="center">60.0</td>
<td valign="top" align="center">70.0</td>
</tr> <tr>
<td valign="top" align="left">Curtain</td>
<td valign="top" align="center">85.0</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">80.0</td>
</tr> <tr>
<td valign="top" align="left">Desk</td>
<td valign="top" align="center">77.0</td>
<td valign="top" align="center">79.0</td>
<td valign="top" align="center">76.7</td>
<td valign="top" align="center">86.0</td>
</tr> <tr>
<td valign="top" align="left">Door</td>
<td valign="top" align="center">92.0</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">75.0</td>
<td valign="top" align="center">85.0</td>
</tr> <tr>
<td valign="top" align="left">Dresser</td>
<td valign="top" align="center">74.0</td>
<td valign="top" align="center">65.1</td>
<td valign="top" align="center">67.4</td>
<td valign="top" align="center">60.5</td>
</tr> <tr>
<td valign="top" align="left">Flowerpot</td>
<td valign="top" align="center">44.6</td>
<td valign="top" align="center">30.0</td>
<td valign="top" align="center">10.0</td>
<td valign="top" align="center">25.0</td>
</tr> <tr>
<td valign="top" align="left">Glassbox</td>
<td valign="top" align="center">91.0</td>
<td valign="top" align="center">94.0</td>
<td valign="top" align="center">80.8</td>
<td valign="top" align="center">86.0</td>
</tr> <tr>
<td valign="top" align="left">Guiter</td>
<td valign="top" align="center">99.0</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">98.0</td>
<td valign="top" align="center">100</td>
</tr> <tr>
<td valign="top" align="left">Keyboard</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100</td>
</tr> <tr>
<td valign="top" align="left">Lamp</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">83.3</td>
<td valign="top" align="center">80.0</td>
</tr> <tr>
<td valign="top" align="left">Laptop</td>
<td valign="top" align="center">86.0</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">100</td>
</tr> <tr>
<td valign="top" align="left">Mental</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">93.9</td>
<td valign="top" align="center">93.0</td>
</tr> <tr>
<td valign="top" align="left">Monitor</td>
<td valign="top" align="center">71.0</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">92.9</td>
<td valign="top" align="center">96.0</td>
</tr> <tr>
<td valign="top" align="left">Nightstand</td>
<td valign="top" align="center">65.0</td>
<td valign="top" align="center">82.6</td>
<td valign="top" align="center">70.2</td>
<td valign="top" align="center">70.9</td>
</tr> <tr>
<td valign="top" align="left">Person</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">85.0</td>
<td valign="top" align="center">89.5</td>
<td valign="top" align="center">90.0</td>
</tr> <tr>
<td valign="top" align="left">Piano</td>
<td valign="top" align="center">91.0</td>
<td valign="top" align="center">88.8</td>
<td valign="top" align="center">84.5</td>
<td valign="top" align="center">83.0</td>
</tr> <tr>
<td valign="top" align="left">Plant</td>
<td valign="top" align="center">91.0</td>
<td valign="top" align="center">73.0</td>
<td valign="top" align="center">78.8</td>
<td valign="top" align="center">83.0</td>
</tr> <tr>
<td valign="top" align="left">Radio</td>
<td valign="top" align="center">88.0</td>
<td valign="top" align="center">70.0</td>
<td valign="top" align="center">65.0</td>
<td valign="top" align="center">55.0</td>
</tr> <tr>
<td valign="top" align="left">Range hood</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">91.0</td>
<td valign="top" align="center">88.9</td>
<td valign="top" align="center">89.9</td>
</tr> <tr>
<td valign="top" align="left">Sink</td>
<td valign="top" align="center">85.0</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">65.0</td>
<td valign="top" align="center">70.0</td>
</tr> <tr>
<td valign="top" align="left">Sofa</td>
<td valign="top" align="center">93.0</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">96.0</td>
<td valign="top" align="center">93.0</td>
</tr> <tr>
<td valign="top" align="left">Stairs</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">85.0</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">75.0</td>
</tr> <tr>
<td valign="top" align="left">Stool</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">83.3</td>
<td valign="top" align="center">70.0</td>
</tr> <tr>
<td valign="top" align="left">Table</td>
<td valign="top" align="center">98.0</td>
<td valign="top" align="center">88.0</td>
<td valign="top" align="center">90.9</td>
<td valign="top" align="center">77.0</td>
</tr> <tr>
<td valign="top" align="left">Tent</td>
<td valign="top" align="center">85.0</td>
<td valign="top" align="center">95.0</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">90.0</td>
</tr> <tr>
<td valign="top" align="left">Toilet</td>
<td valign="top" align="center">98.0</td>
<td valign="top" align="center">99.0</td>
<td valign="top" align="center">94.9</td>
<td valign="top" align="center">95.0</td>
</tr> <tr>
<td valign="top" align="left">TV stand</td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">84.5</td>
<td valign="top" align="center">89.0</td>
</tr> <tr>
<td valign="top" align="left">Vase</td>
<td valign="top" align="center">83.0</td>
<td valign="top" align="center">78.8</td>
<td valign="top" align="center">81.3</td>
<td valign="top" align="center">80.0</td>
</tr> <tr>
<td valign="top" align="left">Wardrobe</td>
<td valign="top" align="center">65.0</td>
<td valign="top" align="center">60.0</td>
<td valign="top" align="center">30.0</td>
<td valign="top" align="center">20.0</td>
</tr> <tr>
<td valign="top" align="left">Xbox</td>
<td valign="top" align="center">90.0</td>
<td valign="top" align="center">70.0</td>
<td valign="top" align="center">75.0</td>
<td valign="top" align="center">80.0</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>ModelNet40 per-class classification accuracy comparisons between PointNet, Pointwise, DPRNet, and (proposed).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0005.tif"/>
</fig></sec></sec>
<sec>
<title>4.4. Precision-recall on ModelNet10</title>
<p>Precision is a metric that assesses the accuracy of predictions, i.e., the percentage of correct predictions. It determines how many of the model&#x00027;s predictions were actually right. The precision was computed using Equation 11 as follows:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>T</italic><sub><italic>P</italic></sub> is true positive while <italic>F</italic><sub><italic>P</italic></sub> is false positive (predicted as positive but was incorrect). In the case of recall, it determines how well all of the positives are found which is given as follows:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>F</italic><sub><italic>N</italic></sub> is false negatives (unable to predict the presence of an object). The mAP is calculated as the average precision of all classes in the dataset while the F1-score is the harmonic mean of the precision and recall. We used these metrics to assess the efficacy and robustness of the proposed method. We used a grid size of 32 &#x000D7; 32 &#x000D7; 32 and kept the value of K at 8. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, the model can learn all 10 object class categories with high precision and recall on the ModelNet10 dataset, with 100% precision on bathtub and chair and 100% recall on bed and toilet. We can also observe that the four classes with the lowest precision and recall (desk, table, nightstand, and dresser) are highly similar which makes them difficult to distinguish even by a human expert. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, we observed that the proposed approach successfully generated results with (1) more than 90% precision on the bed, monitor, sofa, table, and toilet and more than 80% on the remaining classes, (2) 90% or higher recall of bathtub, chair, monitor, sofa, and table with more than 80% on the desk, dresser, and nightstand, and (3) 90% or higher F1-score of the bathtub, bed, chair, monitor, sofa, toilet, and table with more than 80% on the desk, dresser, and nightstand. This demonstrates that our model can learn discriminative features from 3D shapes directly across several classes.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Precision, recall, and F1-score on ModelNet10.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0006.tif"/>
</fig>
<p>To calculate the mAP, we perform several experiments, one of which involved using 16 &#x000D7; 16 &#x000D7; 16 voxel size combined with sampling 8 points per grid. The model was trained using ModelNet10 from scratch, which achieved a 90.2% mAP score. We, then, reduced the learning rate by half (0.5<sup>&#x02212;5</sup>) and retrained the model. The effect of fine-tuning improves the mAP to 90.7%. Another experiment was using a 32 &#x000D7; 32 &#x000D7; 32 grid size with the same points per grid. We train the model using the same procedure in the first experiment. We achieved 92.5% with 0.1<sup>&#x02212;4</sup> learning rate, and after reducing the learning rate to half and retraining the model, the result improves to 93.3%. With mAP scores of 93.3%, our model surpasses 3DShapeNets (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>), PANORAMA-ENN (Sfikas et al., <xref ref-type="bibr" rid="B50">2017</xref>), DeepPano (Shi et al., <xref ref-type="bibr" rid="B51">2015</xref>), PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>), Multimodal (Chen et al., <xref ref-type="bibr" rid="B8">2021</xref>), SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>), Geometry image (Sinha et al., <xref ref-type="bibr" rid="B53">2016</xref>), and GIFT (Bai et al., <xref ref-type="bibr" rid="B3">2016</xref>) on the ModelNet10 dataset, as shown in <xref ref-type="table" rid="T4">Table 4</xref>. Even while SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>) has the advantage of pre-existing 2D networks that have been pre-trained on big datasets such as ImageNet1K, we achieved a higher mean average precious mAP with 1.9% margin on ModelNet10. To further illustrate the effectiveness of the improved fused feature network, <xref ref-type="fig" rid="F7">Figure 7</xref> shows the confusion matrix. The confusion matrix was normalized to 100%. We can see that most objects from all classes are recognized correctly.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Mean average precision mAP (%) on ModelNet10.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>mAP (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">3DShapeNet (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>)</td>
<td valign="top" align="center">68.3</td>
</tr> <tr>
<td valign="top" align="left">DeepPano (Shi et al., <xref ref-type="bibr" rid="B51">2015</xref>)</td>
<td valign="top" align="center">84.1</td>
</tr> <tr>
<td valign="top" align="left">PANORAMA-ENN (Sfikas et al., <xref ref-type="bibr" rid="B50">2017</xref>)</td>
<td valign="top" align="center">93.2</td>
</tr> <tr>
<td valign="top" align="left">SeqViews2SeqLabels (Han et al., <xref ref-type="bibr" rid="B19">2019</xref>)</td>
<td valign="top" align="center">91.4</td>
</tr> <tr>
<td valign="top" align="left">Geometry-image (Sinha et al., <xref ref-type="bibr" rid="B53">2016</xref>)</td>
<td valign="top" align="center">88.4</td>
</tr> <tr>
<td valign="top" align="left">GIFT (Bai et al., <xref ref-type="bibr" rid="B3">2016</xref>)</td>
<td valign="top" align="center">91.1</td>
</tr> <tr>
<td valign="top" align="left">PolyNet (Yavartanoo et al., <xref ref-type="bibr" rid="B71">2021</xref>)</td>
<td valign="top" align="center">84.6</td>
</tr> <tr>
<td valign="top" align="left">(Ours) (16 &#x000D7; 16 &#x000D7; 16&#x02212;<italic>grid</italic>)</td>
<td valign="top" align="center"><bold>90.7</bold></td>
</tr> <tr>
<td valign="top" align="left">(Ours) (32 &#x000D7; 32 &#x000D7; 32&#x02212;<italic>grid</italic>)</td>
<td valign="top" align="center"><bold>93.3</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Confusion matrix on ModelNet10.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0007.tif"/>
</fig></sec>
<sec>
<title>4.5. Part segmentation on ShapeNetPart</title>
<p>Part segmentation seems to be more difficult than classification tasks and is regarded as every-point classification. Given a triangular mesh or point cloud representation of a 3D object, the purpose of part segmentation is to give each point or triangle face a part category which makes it more challenging than object classification because of the fine-grained and dense predictions. We used the metric procedure from PointNet&#x0002B;&#x0002B; (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>). For every part shape within the object category, we calculate the union of prediction and ground truth. <xref ref-type="fig" rid="F8">Figure 8</xref> shows some ShapeNetPart dataset segmentation results from our method. As observed, in most cases, the proposed method results are visually appealing.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>On the ShapeNet-part dataset, we compared the visual results of our object part segmentation with groundthruth.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0008.tif"/>
</fig>
<p><bold>Comparison:</bold> The segmentation performance of the proposed method is compared with that of various deep learning methods, as shown in <xref ref-type="table" rid="T5">Table 5</xref>. Although OCNN and RS-Net (Huang et al., <xref ref-type="bibr" rid="B25">2018</xref>) exceed ours in terms of mIoU of all shapes, the improved fused feature residual network outperforms OCNN in specific categories, such as bag, cap, rocket, lamp, and motorbike, and achieves comparable results in the remaining categories. While OCNN has the best IoU, it also uses a conditional dense random field to rectify their network output which serve as a post-processing step, whereas our approach has no similar strategy.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Segmentation results of different methods on ShapeNet-part dataset (Yi et al., <xref ref-type="bibr" rid="B72">2016</xref>).</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="left"><bold>(Ours)</bold></th>
<th valign="top" align="left"><bold>P.Net</bold></th>
<th valign="top" align="center"><bold>ShapeNet</bold></th>
<th valign="top" align="center"><bold>KD-Net</bold></th>
<th valign="top" align="center"><bold>MRTNet</bold></th>
<th valign="top" align="center"><bold>3DCNN</bold></th>
<th valign="top" align="center"><bold>RS-Net</bold></th>
<th valign="top" align="center"><bold>O-CNN</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">mIoU</td>
<td valign="top" align="left"><bold>84.2</bold></td>
<td valign="top" align="left">83.7</td>
<td valign="top" align="center">81.4</td>
<td valign="top" align="center">77.2</td>
<td valign="top" align="center">83.0</td>
<td valign="top" align="center">79.4</td>
<td valign="top" align="center">84.9</td>
<td valign="top" align="center">85.9</td>
</tr> <tr>
<td valign="top" align="left">Airplane</td>
<td valign="top" align="left">83.8</td>
<td valign="top" align="left">83.4</td>
<td valign="top" align="center">81</td>
<td valign="top" align="center">79.9</td>
<td valign="top" align="center">81.0</td>
<td valign="top" align="center">75.1</td>
<td valign="top" align="center">82.7</td>
<td valign="top" align="center">85.5</td>
</tr> <tr>
<td valign="top" align="left">Bag</td>
<td valign="top" align="left">88.9</td>
<td valign="top" align="left">78.7</td>
<td valign="top" align="center">78.4</td>
<td valign="top" align="center">71.2</td>
<td valign="top" align="center">76.7</td>
<td valign="top" align="center">72.8</td>
<td valign="top" align="center">86.4</td>
<td valign="top" align="center">87.1</td>
</tr> <tr>
<td valign="top" align="left">Cap</td>
<td valign="top" align="left">91.9</td>
<td valign="top" align="left">82.5</td>
<td valign="top" align="center">77.7</td>
<td valign="top" align="center">80.9</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">73.3</td>
<td valign="top" align="center">84.1</td>
<td valign="top" align="center">84.7</td>
</tr> <tr>
<td valign="top" align="left">Car</td>
<td valign="top" align="left">72</td>
<td valign="top" align="left">74.9</td>
<td valign="top" align="center">75.7</td>
<td valign="top" align="center">68.8</td>
<td valign="top" align="center">73.8</td>
<td valign="top" align="center">70.0</td>
<td valign="top" align="center">78.2</td>
<td valign="top" align="center">77.0</td>
</tr> <tr>
<td valign="top" align="left">Chair</td>
<td valign="top" align="left">88</td>
<td valign="top" align="left">89.6</td>
<td valign="top" align="center">87.6</td>
<td valign="top" align="center">88.0</td>
<td valign="top" align="center">89.1</td>
<td valign="top" align="center">87.2</td>
<td valign="top" align="center">90.4</td>
<td valign="top" align="center">91.1</td>
</tr> <tr>
<td valign="top" align="left">Earphone</td>
<td valign="top" align="left">47.0</td>
<td valign="top" align="left">73.0</td>
<td valign="top" align="center">61.9</td>
<td valign="top" align="center">72.4</td>
<td valign="top" align="center">67.6</td>
<td valign="top" align="center">63.5</td>
<td valign="top" align="center">69.3</td>
<td valign="top" align="center">85.1</td>
</tr> <tr>
<td valign="top" align="left">Guitar</td>
<td valign="top" align="left">86.8</td>
<td valign="top" align="left">91.5</td>
<td valign="top" align="center">92</td>
<td valign="top" align="center">88.9</td>
<td valign="top" align="center">90.6</td>
<td valign="top" align="center">88.4</td>
<td valign="top" align="center">91.4</td>
<td valign="top" align="center">91.9</td>
</tr> <tr>
<td valign="top" align="left">Knife</td>
<td valign="top" align="left">86.7</td>
<td valign="top" align="left">85.9</td>
<td valign="top" align="center">85.4</td>
<td valign="top" align="center">86.4</td>
<td valign="top" align="center">85.4</td>
<td valign="top" align="center">79.6</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">87.4</td>
</tr> <tr>
<td valign="top" align="left">Lamp</td>
<td valign="top" align="left">89.8</td>
<td valign="top" align="left">80.8</td>
<td valign="top" align="center">82.5</td>
<td valign="top" align="center">79.8</td>
<td valign="top" align="center">80.6</td>
<td valign="top" align="center">74.4</td>
<td valign="top" align="center">83.5</td>
<td valign="top" align="center">83.3</td>
</tr> <tr>
<td valign="top" align="left">Laptop</td>
<td valign="top" align="left">60.8</td>
<td valign="top" align="left">95.3</td>
<td valign="top" align="center">95.7</td>
<td valign="top" align="center">94.9</td>
<td valign="top" align="center">95.1</td>
<td valign="top" align="center">93.9</td>
<td valign="top" align="center">95.4</td>
<td valign="top" align="center">95.4</td>
</tr> <tr>
<td valign="top" align="left">Motorbike</td>
<td valign="top" align="left">93.7</td>
<td valign="top" align="left">65.2</td>
<td valign="top" align="center">70.6</td>
<td valign="top" align="center">55.8</td>
<td valign="top" align="center">64.4</td>
<td valign="top" align="center">58.7</td>
<td valign="top" align="center">66.0</td>
<td valign="top" align="center">56.9</td>
</tr> <tr>
<td valign="top" align="left">Mug</td>
<td valign="top" align="left">94.4</td>
<td valign="top" align="left">93.0</td>
<td valign="top" align="center">91.9</td>
<td valign="top" align="center">86.5</td>
<td valign="top" align="center">91.8</td>
<td valign="top" align="center">91.8</td>
<td valign="top" align="center">92.6</td>
<td valign="top" align="center">96.2</td>
</tr> <tr>
<td valign="top" align="left">Pistol</td>
<td valign="top" align="left">80</td>
<td valign="top" align="left">81.2</td>
<td valign="top" align="center">85.9</td>
<td valign="top" align="center">79.3</td>
<td valign="top" align="center">79.7</td>
<td valign="top" align="center">76.4</td>
<td valign="top" align="center">81.8</td>
<td valign="top" align="center">81.6</td>
</tr> <tr>
<td valign="top" align="left">Rocket</td>
<td valign="top" align="left">86.1</td>
<td valign="top" align="left">57.9</td>
<td valign="top" align="center">53.1</td>
<td valign="top" align="center">50.4</td>
<td valign="top" align="center">57.0</td>
<td valign="top" align="center">51.2</td>
<td valign="top" align="center">56.1</td>
<td valign="top" align="center">53.5</td>
</tr> <tr>
<td valign="top" align="left">Skateboard</td>
<td valign="top" align="left">70.1</td>
<td valign="top" align="left">72.8</td>
<td valign="top" align="center">69.8</td>
<td valign="top" align="center">71.1</td>
<td valign="top" align="center">69.1</td>
<td valign="top" align="center">65.3</td>
<td valign="top" align="center">75.8</td>
<td valign="top" align="center">74.1</td>
</tr> <tr>
<td valign="top" align="left">Table</td>
<td valign="top" align="left">74.1</td>
<td valign="top" align="left">80.6</td>
<td valign="top" align="center">75.3</td>
<td valign="top" align="center">80.2</td>
<td valign="top" align="center">80.6</td>
<td valign="top" align="center">77.1</td>
<td valign="top" align="center">82.2</td>
<td valign="top" align="center">84.4</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.6. Ablation experiments</title>
<p>Here, we conduct some ablation experimental tests to assess various setup modifications and highlight the benefits of the improved fused feature network. The experiments were carried out using the ModelNet10 (Wu et al., <xref ref-type="bibr" rid="B68">2015</xref>) dataset.</p>
<sec>
<title>4.6.1. Effects of extracted features in the DGFE module</title>
<p>We present an ablation test on ModelNet10 classification to demonstrate the impact of the DGFE module&#x00027;s extracted features. Specifically, we experimented with many variables, including different grid sizes and K values. In the first settings, using a grid size of 16 &#x000D7; 16 &#x000D7; 16 and increasing the value of K from 2 to 8, the classification accuracy increased from 88.1% with K = 2 to 90.5% with K = 8. In the second attempt, we used a grid size of 32 &#x000D7; 32 &#x000D7; 32 and kept the values of K between 2 and 8, and the classification accuracy increased from 90.1% with K = 2 to 91.8% with K = 8. We end up using the later attempt to set the DVFE module in our approach which yields the best model result of 95.6%. <xref ref-type="fig" rid="F9">Figure 9</xref> displays the results. It shows how the proposed DGFE module encourages correlation among different point cloud regions and is useful for modeling the entire point cloud spatial distribution.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>To highlight the influence of both the DGFE module and the feature encoding network, a ModelNet10 classification ablation test is presented. We experimented with some variables including different grid sizes and K values. <bold>(A)</bold> Shows how the feature encoding network performs with 16 &#x000D7; 16 &#x000D7; 16 and 32 &#x000D7; 32 &#x000D7; 32 grid sizes and different values of K; <bold>(B)</bold> demonstrates the performance of the DGFE module&#x00027;s effects of extracted features on 16 &#x000D7; 16 &#x000D7; 16 and 32 &#x000D7; 32 &#x000D7; 32 grid sizes with 2, 3, 6, and 8 K values.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1204445-g0009.tif"/>
</fig></sec>
<sec>
<title>4.6.2. Effects of feature encoding network</title>
<p>This section analyzes the significance of the encoding branch in the proposed approach. After removing the encoding branch, the network is trained using only the DVFE module and KNN search, to sample the local region in each grid cell. We, then, repeated the tests using the same configuration as the previous ablation experiment, with a grid size of 16x16x16 and K = 2. The classification accuracy was 90.1% with K = 2 and 91.1% with K = 8. The classification accuracy improved from 91.3% with K = 2 to 92.4% with K = 8 when utilizing a grid size of 32 &#x000D7; 32 &#x000D7; 32. The results are shown in <xref ref-type="fig" rid="F9">Figure 9</xref>. The model design aids in the efficient encoding of features from the input grid and DVFE module. The output features are combined to complement one another. <xref ref-type="fig" rid="F9">Figure 9</xref> demonstrates the accuracy achieved by inserting the feature encoding network into the whole network, which results in boosting the classification accuracy. The next experiments investigate the sensitivities of the feature encoding units which consist of two units (Feature Encoding Block FEB Unit A and Feature Encoding Block FEB Unit B) with layer skips containing BN and ReLU in between. In each unit, we start with 3 &#x000D7; 3 &#x000D7; 3 convolutions twice, followed by 1 &#x000D7; 1 &#x000D7; 1 convolutions. The main difference between the units is in the application of BN, a regularly used technique to speed up and stabilize the learning process of deep neural networks, and Relu, which has the advantage of allowing complicated correlations in the data to be learned. To test how resilient our approaches are to changes of this type, we swapped the units in different orders. With a 32 &#x000D7; 32 &#x000D7; 32 grid size and K = 8, we apply four possible combinations, such as ABAB, BABA, AABB, and BBAA. We train the model from the scratch. As shown in <xref ref-type="table" rid="T6">Table 6</xref>, the classification accuracy is fairly stable across the different combinations. The combination of ABAB has the highest accuracy and the lowest total log loss, with AABB coming in second. Although the two other combinations, BABA and BBAA, have lower accuracy, their overall performance is generally stable. The above result seems to indicate that, in line with He et al. (<xref ref-type="bibr" rid="B20">2016a</xref>), adding BN after addition forces skip connections to perturb the output, which is problematic. The main advantage of applying BN before addition here is that it speeds up training and allows a wider range of learning rates without sacrificing training convergence.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Different combinations of feature encoding units on ModelNet10.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>FEB unit</bold></th>
<th valign="top" align="center"><bold>Acc (%)</bold></th>
<th valign="top" align="center"><bold>Logloss</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>ABAB</bold></th>
<th valign="top" align="center"><bold>95.6</bold></th>
<th valign="top" align="center"><bold>2.22</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">BABA</td>
<td valign="top" align="center">93.8</td>
<td valign="top" align="center">2.38</td>
</tr> <tr>
<td valign="top" align="left">AABB</td>
<td valign="top" align="center">94.54</td>
<td valign="top" align="center">2.25</td>
</tr> <tr>
<td valign="top" align="left">BBAA</td>
<td valign="top" align="center">93.94</td>
<td valign="top" align="center">2.32</td>
</tr></tbody>
</table>
</table-wrap></sec>
<sec>
<title>4.6.3. Time complexity</title>
<p><xref ref-type="table" rid="T7">Table 7</xref> compares the average testing time for classification and segmentation with other similar methods. TensorFlow 1.1 is used to record forward time using Nvidia Geforce Titan GTX GPU. The proposed method requires less testing time than many other methods, such as (Leng et al., <xref ref-type="bibr" rid="B34">2016</xref>; Charles et al., <xref ref-type="bibr" rid="B7">2017</xref>; Huang et al., <xref ref-type="bibr" rid="B25">2018</xref>), DGCNN (Wang Y. et al., <xref ref-type="bibr" rid="B64">2019</xref>), SpecGCN (Wang et al., <xref ref-type="bibr" rid="B60">2018</xref>), and 3D-UNet (Cicek et al., <xref ref-type="bibr" rid="B11">2016</xref>), because of its strong data closeness and consistency. Because zeros are padded to empty voxel, the proposed voxelization and sampling approaches both include random memory accesses, which help to decrease unnecessary computation. As observed, using the same voxel resolution of 32<sup>3</sup>, the proposed improved fused feature residual network is faster than the 3DCNN (Leng et al., <xref ref-type="bibr" rid="B34">2016</xref>) method and still outperforms it in terms of mIoU, as shown in <xref ref-type="table" rid="T5">Table 5</xref>. Another advantage of this strategy is that the same number of points is kept in each grid cell while still being able to describe neighborhood information. Now lets analyze the approach to the PointNet&#x0002B;&#x0002B; (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>), set abstraction module. If we have a batch of 2,048 points with 64-channel characteristics, the technique can model the entire point cloud, but the SA module must aggressively downsample the input, resulting in information loss. The proposed method does not necessitate dynamic kernel computing, which is typically rather expensive. Even though RSNet (Huang et al., <xref ref-type="bibr" rid="B25">2018</xref>) outperformed ours in terms of Mean IoU by 0.7%, the proposed improved fused feature residual network is much faster and requires less memory consumption, as shown in <xref ref-type="table" rid="T7">Table 7</xref>.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Average testing time of our method with others on ModelNet40.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Classification (ms)</bold></th>
<th valign="top" align="left"><bold>Segmentation (ms)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">PointNet&#x0002B;&#x0002B; (Qi et al., <xref ref-type="bibr" rid="B46">2017</xref>)</td>
<td valign="top" align="left">163</td>
<td valign="top" align="left">-</td>
</tr> <tr>
<td valign="top" align="left">3DCNN (Leng et al., <xref ref-type="bibr" rid="B34">2016</xref>)</td>
<td valign="top" align="left">49</td>
<td valign="top" align="left">137</td>
</tr> <tr>
<td valign="top" align="left">SpecGCN (Wang et al., <xref ref-type="bibr" rid="B60">2018</xref>)</td>
<td valign="top" align="left">11254</td>
<td valign="top" align="left">-</td>
</tr> <tr>
<td valign="top" align="left">DGCNN (Wang Y. et al., <xref ref-type="bibr" rid="B64">2019</xref>)</td>
<td valign="top" align="left">52</td>
<td valign="top" align="left">87.8</td>
</tr> <tr>
<td valign="top" align="left">3D-UNet (Cicek et al., <xref ref-type="bibr" rid="B11">2016</xref>)</td>
<td valign="top" align="left">-</td>
<td valign="top" align="left">682.1</td>
</tr> <tr>
<td valign="top" align="left">RSNet (Huang et al., <xref ref-type="bibr" rid="B25">2018</xref>)</td>
<td valign="top" align="left">-</td>
<td valign="top" align="left">74.6</td>
</tr> <tr>
<td valign="top" align="left">(Ours)</td>
<td valign="top" align="left"><bold>28</bold></td>
<td valign="top" align="left"><bold>19</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values used to differentiate our results from the rest of the other methods.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.6.4. Effects of neighborhood query</title>
<p>In this section, we experiment with ball query and sift query, two other popular neighbor querying methods to sample local areas and experiment with general search radius. For all experiments, we use a 32 &#x000D7; 32 &#x000D7; 32 grid size with a K = 8 value on the ModelNet10 dataset. <xref ref-type="table" rid="T8">Table 8</xref> shows that KNN is more effective for our strategy. The sift query is the most inefficient method when compared with the KNN and ball query.</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Effects of neighborhood query on ModelNet10 classification.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left" colspan="2"><bold>Sift query</bold></th>
<th valign="top" align="left" colspan="2"><bold>Ball query</bold></th>
<th valign="top" align="left"><bold>KNN</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold><italic>r</italic> = 0.1</bold></th>
<th valign="top" align="left"><bold><italic>r</italic> = 0.2</bold></th>
<th valign="top" align="left"><bold><italic>r</italic> = 0.1</bold></th>
<th valign="top" align="left"><bold><italic>r</italic> = 0.2</bold></th>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">90.8%</td>
<td valign="top" align="left">91.0%</td>
<td valign="top" align="left">92.6%</td>
<td valign="top" align="left">93.0%</td>
<td valign="top" align="left">95.6%</td>
</tr></tbody>
</table>
</table-wrap></sec></sec></sec>
<sec id="s5">
<title>5. Conclusion and future work</title>
<p>In this study, we proposed the detail grid feature extraction (DGFE) module which is a highly efficient module. This module assists 3D convolutions in hierarchically capturing global information, reducing the grid size in each spatial dimension and managing overfitting by gradually lowering the spatial dimension of the representation, making it practical for high-resolution 3D objects. Furthermore, we design a feature encoding network that uses two different building blocks with layer skips containing batch normalization and non-linearity ReLU in between, resulting in fewer layers in the early training phase which helps speed learning and reduces the effect of gradients vanishing since there are few layers through which to propagate. The outputs of the two modules are fused in the feature fusion unit to produce a feature with improved contextual representation by utilizing both local and global shape structures. We built a network called improved fused feature residual network using the modules that have been proposed, which achieve a notable balance of accuracy and speed. In both ModelNet10 and ModelNet40 datasets, the proposed improved fused feature residual network offers a significant advantage over the bulk of voxel and point cloud-based approaches, as shown in <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>. Due to its scalability and efficiency, the proposed method can be used in extracting large-scale features of high-resolution inputs.</p>
<p>Although our method performs well with normal datasets, we note that when noise is added to the datasets, the performance drops, for example, when Gaussian noise is added to the 3D models, the performance decreases despite applying different parameters. In future, instead of directly sampling points, we will use sparse convolutions to convert them to a small number of voxels and sample non-empty voxels to ensure that precise point positions are retained.</p>
<p>In addition, numerous mechanisms for attention employed in transformer approaches are adaptable and offer a high potential for future advances. We think cutting-edge outcomes can be attained by extending generic point cloud processing innovation to transformer techniques. For instance, one possible option we are looking at is by swapping out the feature extraction module in our network design for one that is transformer/attention-based. Instead of just depending on transformers to extract features, we can conduct local feature extraction using non-transformer-based approaches and then couple it with a transformer for global feature interaction which will lead to the extraction of more fine grain features.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p></sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>AG: conceptualization of this study, methodology, writing&#x02014;original draft preparation, and software. CL: conceptualization, software, supervision, resources, project administration, and funding acquisition. HJ, YN, and MA: data curation, writing&#x02014;reviewing and editing, and software. HC: data curation, software, and supervision. All authors contributed to the article and approved the submitted version.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This study was supported by the Fujian Province University Key Lab for the Analysis and Application of Industry Big Data, Fujian Key Lab of Agriculture IOT Application, and IOT Application Engineering Research Center of Fujian Province Colleges and Universities.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>

<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arshad</surname> <given-names>S.</given-names></name> <name><surname>Shahzad</surname> <given-names>M.</given-names></name> <name><surname>Riaz</surname> <given-names>Q.</given-names></name> <name><surname>Fraz</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>DPRNet: deep 3D point based residual network for semantic segmentation and classification of 3D point clouds</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>68892</fpage>&#x02013;<lpage>68904</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2918862</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atzmon</surname> <given-names>M.</given-names></name> <name><surname>Maron</surname> <given-names>H.</given-names></name> <name><surname>Lipman</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Point convolutional neural networks by extension operators</article-title>. <source>ACM Trans. Graph.</source> <volume>37</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/3197517.3201301</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Latecki</surname> <given-names>L. J.</given-names></name></person-group> (<year>2016</year>).&#x0201C;GIFT: A real-time and scalable 3D shape search engine,&#x0201D; in <italic>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic> (Las Vegas, NV: IEEE), <fpage>5023</fpage>&#x02013;<lpage>5032</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.543</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bello</surname> <given-names>S. A.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Wambugu</surname> <given-names>N. M.</given-names></name> <name><surname>Adam</surname> <given-names>J. M.</given-names></name></person-group> (<year>2021</year>). <article-title>FFpointNet: local and global fused feature for 3D point clouds analysis</article-title>. <source>Neurocomputing</source> <volume>461</volume>, <fpage>55</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.07.044</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bello</surname> <given-names>Saifullahi, A.</given-names></name> <name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Adam</surname> <given-names>Jibril, M.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Review: deep learning on 3D point clouds</article-title>. <source>Remot. Sens.</source> <volume>12</volume>, <fpage>11</fpage>. <pub-id pub-id-type="doi">10.3390/rs12111729</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brock</surname> <given-names>A.</given-names></name> <name><surname>Lim</surname> <given-names>T.</given-names></name> <name><surname>Ritchie</surname> <given-names>J.</given-names></name> <name><surname>Weston</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>Generative and discriminative voxel modeling with convolutional neural networks</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1608.04236</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Charles</surname> <given-names>R.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Kaichun</surname> <given-names>M.</given-names></name> <name><surname>Guibas</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;PointNet: Deep learning on point sets for 3D classification and segmentation,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>77</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.16</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Jing</surname> <given-names>L.</given-names></name> <name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Tian</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Multimodal semi-supervised learning for 3D objects</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2110.11601</pub-id><pub-id pub-id-type="pmid">27609193</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chiotellis</surname> <given-names>I.</given-names></name> <name><surname>Triebel</surname> <given-names>R.</given-names></name> <name><surname>Windheuser</surname> <given-names>T.</given-names></name> <name><surname>Cremers</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Non-rigid 3D shape retrieval via large margin nearest neighbor embedding,&#x0201D;</article-title> in <source>European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Amsterdam</publisher-loc>). <pub-id pub-id-type="doi">10.1007/978-3-319-46475-6_21</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Choy</surname> <given-names>C.</given-names></name> <name><surname>Danfei</surname> <given-names>X.</given-names></name> <name><surname>JunYoung</surname> <given-names>G.</given-names></name> <name><surname>Kevin</surname> <given-names>C.</given-names></name> <name><surname>Savarese</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;3D-R2N2: A unified approach for single and multi-view 3D object reconstruction,&#x0201D;</article-title> in <source>European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Amsterdam</publisher-loc>).</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cicek</surname> <given-names>&#x000D6;.</given-names></name> <name><surname>Abdulkadir</surname> <given-names>A.</given-names></name> <name><surname>Lienkamp</surname> <given-names>S. S.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name> <name><surname>Ronneberger</surname> <given-names>O.</given-names></name></person-group> (<year>2016</year>). 3D U-Net: Learning dense volumetric segmentation from sparse annotation. <italic>ArXiv</italic>. <pub-id pub-id-type="doi">10.48550/arXiv.1606.06650</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dominguez</surname> <given-names>M.</given-names></name> <name><surname>Dhamdhere</surname> <given-names>R.</given-names></name> <name><surname>Petkar</surname> <given-names>A.</given-names></name> <name><surname>Jain</surname> <given-names>S.</given-names></name> <name><surname>Sah</surname> <given-names>S.</given-names></name> <name><surname>Ptucha</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;General-purpose deep point cloud feature extractor,&#x0201D;</article-title> in <source>2018 IEEE Winter Conference on Applications of Computer Vision (WACV)</source> (<publisher-loc>Lake Tahoe, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1972</fpage>&#x02013;<lpage>1981</lpage>. <pub-id pub-id-type="doi">10.1109/WACV.2018.00218</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eldar</surname> <given-names>Y.</given-names></name> <name><surname>Lindenbaum</surname> <given-names>M.</given-names></name> <name><surname>Porat</surname> <given-names>M.</given-names></name> <name><surname>Zeevi</surname> <given-names>Y.</given-names></name></person-group> (<year>1997</year>). <article-title>The farthest point strategy for progressive image sampling</article-title>. <source>IEEE Trans. Image Process.</source> <volume>9</volume>, <fpage>1305</fpage>&#x02013;<lpage>1315</lpage>. <pub-id pub-id-type="doi">10.1109/83.623193</pub-id><pub-id pub-id-type="pmid">18283019</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elhassan</surname> <given-names>M. A.</given-names></name> <name><surname>Huang</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Munea</surname> <given-names>T. L.</given-names></name></person-group> (<year>2021</year>). <article-title>DSANet: dilated spatial attention for real-time semantic segmentation in urban street scenes</article-title>. <source>Expert Syst. Appl.</source> <volume>183</volume>, <fpage>115090</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115090</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Erg&#x000FC;n</surname> <given-names>O.</given-names></name> <name><surname>Sahillioglu</surname> <given-names>Y.</given-names></name></person-group> (<year>2023</year>). <article-title>3D point cloud classification with ACGAN-3D and VACWGAN-GP</article-title>. <source>Turk. J. Electr. Eng. Comput. Sci.</source> <volume>31</volume>, <fpage>381</fpage>&#x02013;<lpage>395</lpage>. <pub-id pub-id-type="doi">10.55730/1300-0632.3990</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>M.</given-names></name> <name><surname>Ruan</surname> <given-names>N.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Deep neural network for 3D shape classification based on mesh feature</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>187040</fpage>. <pub-id pub-id-type="doi">10.3390/s22187040</pub-id><pub-id pub-id-type="pmid">36146387</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gezawa</surname> <given-names>A. S.</given-names></name> <name><surname>Bello</surname> <given-names>Z. A.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Yunqi</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>A voxelized point clouds representation for object classification and segmentation on 3D data</article-title>. <source>J. Supercomput.</source> <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1007/s11227-021-03899-x</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gezawa</surname> <given-names>Sulaiman, A.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Lei</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>A review on deep learning approaches for 3D data representations in retrieval and classifications</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>57566</fpage>&#x02013;<lpage>57593</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.2982196</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>Z.</given-names></name> <name><surname>Shang</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Vong</surname> <given-names>C.-M.</given-names></name> <name><surname>Liu</surname> <given-names>Y.-S.</given-names></name> <name><surname>Zwicker</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>SeqViews2SeqLabels: learning 3D global features via aggregating sequential views by RNN with attention</article-title>. <source>IEEE Trans. Image Process.</source> <volume>28</volume>, <fpage>658</fpage>&#x02013;<lpage>672</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2018.2868426</pub-id><pub-id pub-id-type="pmid">30183634</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016a</year>). <article-title>&#x0201C;Deep residual learning for image recognition,&#x0201D;</article-title> in <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id><pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016b</year>). <article-title>Identity mappings in deep residual networks</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1603.05027</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hua</surname> <given-names>B.-S.</given-names></name> <name><surname>Tran</surname> <given-names>M.-K.</given-names></name> <name><surname>Yueng</surname> <given-names>S.-K.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Pointwise convolutional neural networks,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>984</fpage>&#x02013;<lpage>993</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00109</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Tu</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Weight loss for point clouds classification</article-title>. <source>J. Phys.</source> <volume>1229</volume>, <fpage>e012045</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/1229/1/012045</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Densely connected convolutional networks,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2261</fpage>&#x02013;<lpage>2269</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Neumann</surname> <given-names>U.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Recurrent slice networks for 3D segmentation of point clouds,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2626</fpage>&#x02013;<lpage>2635</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00278</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Szegedy</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Batch normalization: accelerating deep network training by reducing internal covariate shift</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1502.03167</pub-id><pub-id pub-id-type="pmid">35496726</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kasaei</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>OrthographicNet: a deep learning approach for 3d object recognition in open-ended domains</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1902.03057</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>CoRR</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1412.6980</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Klokov</surname> <given-names>R.</given-names></name> <name><surname>Lempitsky</surname> <given-names>V.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Escape from cells: Deep Kd-networks for the recognition of 3D point cloud models,&#x0201D;</article-title> in <source>2017 IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>863</fpage>&#x02013;<lpage>872</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.99</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kohonen</surname> <given-names>T.</given-names></name></person-group> (<year>1998</year>). <article-title>The self-organizing map</article-title>. <source>Neurocomputing</source> <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1016/S0925-2312(98)00030-7</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuangen</surname> <given-names>Z.</given-names></name> <name><surname>Ming</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>de Silva</surname> <given-names>C. W.</given-names></name> <name><surname>Fu</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>Linked dynamic graph CNN: learning on point cloud via linking hierarchical features</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1904.10014</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Landrieu</surname> <given-names>L.</given-names></name> <name><surname>Simonovsky</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Large-scale point cloud semantic segmentation with superpoint graphs,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4558</fpage>&#x02013;<lpage>4567</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00479</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>T.</given-names></name> <name><surname>Duan</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;PointGrid: A deep network for 3D shape understanding,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>9204</fpage>&#x02013;<lpage>9214</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00959</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leng</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Xiong</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). <article-title>3D object understanding with 3D convolutional neural networks</article-title>. <source>Inf. Sci.</source> <volume>366</volume>, <fpage>188</fpage>&#x02013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2015.08.007</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Muller</surname> <given-names>M.</given-names></name> <name><surname>Qian</surname> <given-names>G.</given-names></name> <name><surname>Delgadillo</surname> <given-names>I. C.</given-names></name> <name><surname>Abualshour</surname> <given-names>A.</given-names></name> <name><surname>Thabet</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>DeepGCNs: making GCNs go as deep as CNNs</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>45</volume>, <fpage>6923</fpage>&#x02013;<lpage>6939</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2021.3074057</pub-id><pub-id pub-id-type="pmid">33872143</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>B. M.</given-names></name> <name><surname>Lee</surname> <given-names>G. H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;SO-Net: Self-organizing network for point cloud analysis,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>9397</fpage>&#x02013;<lpage>9406</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00979</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Bu</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>M.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Di</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;PointCNN: convolution on x-transformed points,&#x0201D;</article-title> in <source>Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS&#x00027;18)</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>828</fpage>&#x02013;<lpage>838</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Pirk</surname> <given-names>S.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Qi</surname> <given-names>C.</given-names></name> <name><surname>Guibas</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>FPNN: field probing neural networks for 3D data</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1605.06240</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Giles</surname> <given-names>L.</given-names></name> <name><surname>Ororbia</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Learning a hierarchical latent-variable model of 3D shapes,&#x0201D;</article-title> in <source>2018 International Conference on 3D Vision</source> (<publisher-loc>Verona</publisher-loc>), <fpage>542</fpage>&#x02013;<lpage>551</lpage>. <pub-id pub-id-type="doi">10.1109/3DV.2018.00068</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Lv</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>F.</given-names></name></person-group> (<year>2023</year>). <article-title>Point cloud classification using content-based transformer via clustering in feature space</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2303.04599</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>H.</given-names></name> <name><surname>Lee</surname> <given-names>S.-H.</given-names></name> <name><surname>Kwon</surname> <given-names>K.-R.</given-names></name></person-group> (<year>2021</year>). <article-title>A deep learning method for 3D object classification and retrieval using the global point signature plus and deep wide residual network</article-title>. <source>Sensors</source> <volume>21</volume>, <fpage>82644</fpage>. <pub-id pub-id-type="doi">10.3390/s21082644</pub-id><pub-id pub-id-type="pmid">33918845</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>D.</given-names></name> <name><surname>Xie</surname> <given-names>Q.</given-names></name> <name><surname>Gao</surname> <given-names>K.</given-names></name> <name><surname>Xu</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>3DCTN: 3D convolution-transformer network for point cloud classification</article-title>. <source>IEEE Trans. Intell. Transport. Syst.</source> <volume>23</volume>, <fpage>24854</fpage>&#x02013;<lpage>24865</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2022.3198836</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>C.</given-names></name> <name><surname>An</surname> <given-names>W.</given-names></name> <name><surname>Lei</surname> <given-names>Y.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;BV-CNNS: binary volumetric convolutional networks for 3D object recognition,&#x0201D;</article-title> in <source>British Machine Vision Conference 2017, BMVC 2017</source> (<publisher-loc>London</publisher-loc>: <publisher-name>BMVA Press</publisher-name>).</citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Maturana</surname> <given-names>D.</given-names></name> <name><surname>Scherer</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;VoxNet: A 3D Convolutional Neural Network for real-time object recognition,&#x0201D;</article-title> in <source>2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source> (<publisher-loc>Hamburg</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>922</fpage>&#x02013;<lpage>928</lpage>. <pub-id pub-id-type="doi">10.1109/IROS.2015.7353481</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nair</surname> <given-names>V.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Rectified linear units improve restricted Boltzmann machines,&#x0201D;</article-title> in <source>Proceedings of the 27th International Conference on International Conference on Machine Learning</source> (<publisher-loc>Madison, WI</publisher-loc>: <publisher-name>Omnipress</publisher-name>), <fpage>807</fpage>&#x02013;<lpage>814</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qi</surname> <given-names>C.</given-names></name> <name><surname>Yi</surname> <given-names>L.</given-names></name> <name><surname>Hao</surname> <given-names>S.</given-names></name> <name><surname>Guibas</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Pointnet<sup>&#x0002B;&#x0002B;</sup>: deep hierarchical feature learning on point sets in a metric space,&#x0201D;</article-title> in <source>Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS&#x00027;17)</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>5105</fpage>&#x02013;<lpage>5114</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qi</surname> <given-names>Z.</given-names></name> <name><surname>Dong</surname> <given-names>R.</given-names></name> <name><surname>Fan</surname> <given-names>G.</given-names></name> <name><surname>Ge</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ma</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Contrast with reconstruct: contrastive 3D representation learning guided by generative pretraining</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2302.02318</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Qiangeng</surname> <given-names>X.</given-names></name> <name><surname>Weiyue</surname> <given-names>W.</given-names></name> <name><surname>Duygu</surname> <given-names>C.</given-names></name> <name><surname>Mech</surname> <given-names>R.</given-names></name> <name><surname>Neumann</surname> <given-names>U.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;DISN: deep implicit surface network for high-quality single-view 3D reconstruction,&#x0201D;</article-title> in <source>Proceedings of the 33rd International Conference on Neural Information Processing Systems</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>492</fpage>&#x02013;<lpage>502</lpage>.</citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Riegler</surname> <given-names>G.</given-names></name> <name><surname>Ulusoy</surname> <given-names>A. O.</given-names></name> <name><surname>Geiger</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;OctNet: Learning deep 3D representations at high resolutions,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6620</fpage>&#x02013;<lpage>6629</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.701</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sfikas</surname> <given-names>K.</given-names></name> <name><surname>Theoharis</surname> <given-names>T.</given-names></name> <name><surname>Pratikakis</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Exploiting the PANORAMA representation for convolutional neural network classification and retrieval,&#x0201D;</article-title> in <source>Proceedings of the Workshop on 3D Object Retrieval (3Dor &#x00027;17)</source> (<publisher-loc>Goslar</publisher-loc>: <publisher-name>Eurographics Association</publisher-name>), 1-7. <pub-id pub-id-type="doi">10.2312/3dor.20171045</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>B.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>DeepPano: deep panoramic representation for 3-D shape recognition</article-title>. <source>IEEE Sign. Process. Lett.</source> <volume>22</volume>, <fpage>2339</fpage>&#x02013;<lpage>2343</lpage>. <pub-id pub-id-type="doi">10.1109/LSP.2015.2480802</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Simonovsky</surname> <given-names>M.</given-names></name> <name><surname>Komodakis</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Dynamic edge-conditioned filters in convolutional neural networks on graphs,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>29</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.11</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sinha</surname> <given-names>A.</given-names></name> <name><surname>Bai</surname> <given-names>J.</given-names></name> <name><surname>Ramani</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep learning 3D shape surfaces using geometry images,&#x0201D;</article-title> in <source>European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Amsterdam</publisher-loc>).</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Gao</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Shen</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>A novel point cloud encoding method based on local information for 3D classification and segmentation</article-title>. <source>Sensors</source> <volume>20</volume>, <fpage>92501</fpage>. <pub-id pub-id-type="doi">10.3390/s20092501</pub-id><pub-id pub-id-type="pmid">32354092</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Reed</surname> <given-names>S.</given-names></name> <name><surname>Anguelov</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>&#x0201C;Going deeper with convolutions,&#x0201D;</article-title> in <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298594</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>H.</given-names></name> <name><surname>Qi</surname> <given-names>C. R.</given-names></name> <name><surname>Deschaud</surname> <given-names>J.-E.</given-names></name> <name><surname>Marcotegui</surname> <given-names>B.</given-names></name> <name><surname>Goulette</surname> <given-names>F.</given-names></name> <name><surname>Guibas</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;KPConv: Flexible and deformable convolution for point clouds,&#x0201D;</article-title> in <source>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6410</fpage>&#x02013;<lpage>6419</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00651</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Song</surname> <given-names>W.</given-names></name> <name><surname>Sung</surname> <given-names>Y.</given-names></name> <name><surname>Woo</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>DGCB-Net: dynamic graph convolutional broad network for 3D object recognition in point cloud</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <fpage>66</fpage>. <pub-id pub-id-type="doi">10.3390/rs13010066</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varga</surname> <given-names>M.</given-names></name> <name><surname>Jadlovsk&#x000FD;</surname> <given-names>J.</given-names></name> <name><surname>Jadlovska</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Generative enhancement of 3D image classifiers</article-title>. <source>Appl. Sci.</source> <volume>2020</volume>, <fpage>10217433</fpage>. <pub-id pub-id-type="doi">10.3390/app10217433</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Cheng</surname> <given-names>M.</given-names></name> <name><surname>Sohel</surname> <given-names>F.</given-names></name> <name><surname>Bennamoun</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2019a</year>). <article-title>NormalNet: a voxel-based CNN for 3D object classification and retrieval</article-title>. <source>Neurocomputing</source> <volume>323</volume>, <fpage>139</fpage>&#x02013;<lpage>147</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2018.09.075</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Samari</surname> <given-names>B.</given-names></name> <name><surname>Siddiqi</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Local spectral graph convolution for point set feature learning,&#x0201D;</article-title> in <source>Computer Vision &#x02013; ECCV 2018: 15th European Conference, Munich, Germany, September 8&#x02013;14, 2018, Proceedings, Part IV</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>), <fpage>56</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01225-0_4</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>D. Z.</given-names></name> <name><surname>Posner</surname> <given-names>I.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Voting for voting in online point cloud object detection,&#x0201D;</article-title> in <source>Robotics: Science and Systems</source> (<publisher-loc>Rome</publisher-loc>), <fpage>10</fpage>&#x02013;<lpage>15607</lpage>.<pub-id pub-id-type="pmid">30040672</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>C.</given-names></name> <name><surname>Shi</surname> <given-names>S.</given-names></name> <name><surname>Lei</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>DSVT: dynamic sparse voxel transformer with rotated sets</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2301.06051</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Shan</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Graph attention convolution for point cloud semantic segmentation,&#x0201D;</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>10288</fpage>&#x02013;<lpage>10297</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01054</pub-id><pub-id pub-id-type="pmid">35286267</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Sarma</surname> <given-names>S. E.</given-names></name> <name><surname>Bronstein</surname> <given-names>M. M.</given-names></name> <name><surname>Solomon</surname> <given-names>J. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Dynamic graph cnn for learning on point clouds</article-title>. <source>ACM Trans. Graph.</source> <volume>38</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/3326362</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>X.</given-names></name> <name><surname>Yu</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;View-GCN: View-based graph convolutional network for 3D shape analysis,&#x0201D;</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1847</fpage>&#x02013;<lpage>1856</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00192</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Xue</surname> <given-names>T.</given-names></name> <name><surname>Freeman</surname> <given-names>B.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling,&#x0201D;</article-title> in <source>Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS&#x00027;16)</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>82</fpage>&#x02013;<lpage>90</lpage>.</citation>
</ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Qi</surname> <given-names>Z.</given-names></name> <name><surname>Fuxin</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;PointConv: Deep convolutional networks on 3D point clouds,&#x0201D;</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>9613</fpage>&#x02013;<lpage>9622</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00985</pub-id><pub-id pub-id-type="pmid">36272236</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Song</surname> <given-names>S.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Yu</surname> <given-names>F.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>&#x0201C;3D ShapeNets: A deep representation for volumetric shapes,&#x0201D;</article-title> in <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1912</fpage>&#x02013;<lpage>1920</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298801</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>B.</given-names></name> <name><surname>He</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>PVT-SSD: single-stage 3d object detector with point-voxel transformer</article-title>. <source>ArXiv</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2305.06621</pub-id></citation>
</ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Modeling point clouds with self-attention and gumbel subset sampling,&#x0201D;</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3318</fpage>&#x02013;<lpage>3327</lpage>, <pub-id pub-id-type="doi">10.1109/CVPR.2019.00344</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yavartanoo</surname> <given-names>M.</given-names></name> <name><surname>Hung</surname> <given-names>S.-H.</given-names></name> <name><surname>Neshatavar</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;PolyNet: Polynomial neural network for 3D shape recognition with polyshape representation,&#x0201D;</article-title> in <source>2021 International Conference on 3D Vision (3DV)</source> (<publisher-loc>London</publisher-loc>), 1014&#x02013;1023, <pub-id pub-id-type="doi">10.1109/3DV53792.2021.00109</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yi</surname> <given-names>L.</given-names></name> <name><surname>Kim</surname> <given-names>V. G.</given-names></name> <name><surname>Ceylan</surname> <given-names>D.</given-names></name> <name><surname>Shen</surname> <given-names>I.-C.</given-names></name> <name><surname>Yan</surname> <given-names>M.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>A scalable active framework for region annotation in 3D shape collections</article-title>. <source>ACM Trans. Graph.</source> <volume>35</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/2980179.2980238</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yifan</surname> <given-names>X.</given-names></name> <name><surname>Tianqi</surname> <given-names>F.</given-names></name> <name><surname>Mingye</surname> <given-names>X.</given-names></name> <name><surname>Long</surname> <given-names>Z.</given-names></name> <name><surname>Qiao</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;SpiderCNN: deep learning on point sets with parameterized convolutional filters,&#x0201D;</article-title> in <source>European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>).</citation>
</ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhijian</surname> <given-names>L.</given-names></name> <name><surname>Haotian</surname> <given-names>T.</given-names></name> <name><surname>Yujun</surname> <given-names>L.</given-names></name> <name><surname>Song</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Point-voxel CNN for efficient 3D deep learning,&#x0201D;</article-title> in <source>Proceedings of the 33rd International Conference on Neural Information Processing Systems</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>965</fpage>&#x02013;<lpage>975</lpage>.</citation>
</ref>
<ref id="B75">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Tuzel</surname> <given-names>O.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;VoxelNet: End-to-end learning for point cloud based 3D object detection,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4490</fpage>&#x02013;<lpage>4499</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00472</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>