<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">839273</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2022.839273</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Classification Model of Point Cloud Along Transmission Line Based on Group Normalization</article-title>
<alt-title alt-title-type="left-running-head">Yin et al.</alt-title>
<alt-title alt-title-type="right-running-head">Transmission Line Point Cloud Classification</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yin</surname>
<given-names>Zhimin</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1604947/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Ji</surname>
<given-names>Shichao</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Xuyong</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dai</surname>
<given-names>Jianhua</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yu</surname>
<given-names>Weiyong</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wu</surname>
<given-names>Song</given-names>
</name>
</contrib>
</contrib-group>
<aff>
<institution>State Grid Huzhou Power Supply Company</institution>, <addr-line>Hu&#x2019;zhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1442178/overview">Zaibin Jiao</ext-link>, Xi&#x2019;an Jiaotong University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1608794/overview">Xinyu Liu</ext-link>, Fuzhou University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1642083/overview">Keyu Wu</ext-link>, Institute for Infocomm Research (A&#x2217;STAR), Singapore</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Zhimin Yin, <email>hfhxbhjd123@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>839273</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>02</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Yin, Ji, Zhang, Dai, Yu and Wu.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yin, Ji, Zhang, Dai, Yu and Wu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>This article proposes a point cloud classification model based on group normalization to increase the classification accuracy when the computing power of the terminal device is limited. This model groups and normalizes the features of point cloud during inference and increases the classification accuracy when the computing power is limited. The group normalization first groups the features of point cloud by their channel, then computes their statistic metrics and normalizes them. Also, one-dimensional convolution layers are used to replace the fully connected layers to decrease the model parameters and keep the model&#x27;s performance when the computing power is limited. In the experiment, PointNet&#x2b;&#x2b; is used to pretrain on ModelNet40 and then fine-tune on the point cloud data of transmission lines. The result shows that the proposed method can effectively increase the classification accuracy and help the 3D modeling process of the transmission line.</p>
</abstract>
<kwd-group>
<kwd>group normalization</kwd>
<kwd>PointNet&#x2b;&#x2b;</kwd>
<kwd>ModelNet40</kwd>
<kwd>batch normalization</kwd>
<kwd>point cloud of transmission line</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Transmission lines are the main part of power transmission and play a key role in the process of power transmission. However, due to the rapid development of power grids in recent years and the increasing length of transmission lines, manual inspections have become more time-consuming and labor-intensive, and the use of unmanned aerial vehicle inspections has become a trend (<xref ref-type="bibr" rid="B18">Qin et al., 2018</xref>; <xref ref-type="bibr" rid="B27">Yao et al., 2021</xref>). Equipped with a laser radar, an unmanned aerial vehicle can scan along the transmission line to gather point clouds of the transmission line in three dimensions (<xref ref-type="bibr" rid="B21">Teng et al., 2017</xref>; <xref ref-type="bibr" rid="B11">Li et al., 2021</xref>). Based on the classified point cloud data, various algorithms, such as clearance anomaly detection (<xref ref-type="bibr" rid="B3">Chen et al., 2018</xref>), can be applied on the transmission lines. Therefore, how to classify point cloud is an important topic in point cloud processing and analysis (<xref ref-type="bibr" rid="B29">Zhang et al., 2016</xref>; <xref ref-type="bibr" rid="B10">Li et al., 2020</xref>). On the other hand, though the computing ability is limited on the unmanned aerial vehicle, fast and real-time detection is effective when the internet is bad in the backcountry (<xref ref-type="bibr" rid="B6">Huang et al., 2020</xref>).</p>
<p>
<xref ref-type="bibr" rid="B23">Wang et al. (2014)</xref> proposed a multiscale and hierarchical point cluster method to classify point clouds, which first resamples the point clouds into different scales, and the latent Dirichlet allocation is used to reconstruct the features before an AdaBoost classifier classifies them. There are some researchers who use conditional random fields (<xref ref-type="bibr" rid="B14">Niemeyer et al., 2012</xref>) to classify point clouds which have a good generalization performance. In addition, the natural exponential function threshold can also be used to split the point cloud into hierarchical clusters and then jointly use latent Dirichlet allocation and sparse coding to extract and encode the shape features of the multilevel point clusters (<xref ref-type="bibr" rid="B29">Zhang et al., 2016</xref>). Some other methods based on the connected component analysis and voxel can also be used to detect pylons and wires from point cloud data (<xref ref-type="bibr" rid="B1">Awrangjeb et al., 2017</xref>; <xref ref-type="bibr" rid="B13">Munir et al., 2019</xref>).</p>
<p>However, these methods require data formats like image grids or 3D voxels, which need data transformation from the original point clouds and thus lose some space information.</p>
<p>In recent years, deep learning has shown good results in many different tasks, such as image classification (<xref ref-type="bibr" rid="B5">He et al., 2016</xref>; <xref ref-type="bibr" rid="B20">Szegedy et al., 2017</xref>), natural language processing (<xref ref-type="bibr" rid="B22">Vaswani et al., 2017</xref>; <xref ref-type="bibr" rid="B4">Devlin et al., 2018</xref>), and automatic speech recognition (<xref ref-type="bibr" rid="B2">Chan et al., 2015</xref>; <xref ref-type="bibr" rid="B24">Watanabe et al., 2017</xref>). There are also some methods based on deep learning, such as PointNet and PointNet&#x2b;&#x2b;, that can be applied to point clouds on classification (<xref ref-type="bibr" rid="B16">Qi et al., 2017a</xref>; <xref ref-type="bibr" rid="B17">Qi et al., 2017b</xref>). These methods need no data preprocessing and can be applied to the original point clouds, which keep the original space information. Based on these works, there are some other works focusing on data diversity or space information (<xref ref-type="bibr" rid="B28">Zhang and Rabbat, 2018</xref>; <xref ref-type="bibr" rid="B9">Li et al., 2020</xref>). However, these methods do not consider the condition where the model will be loaded in a platform where the computing power is limited (<xref ref-type="bibr" rid="B30">Zhao et al., 2019</xref>), such as an unmanned aerial vehicle, which may cause precision loss. Similar work on spatial localization of insulators exists, which can detect insulators in real-time with an unmanned aerial vehicle (<xref ref-type="bibr" rid="B12">Ma et al., 2021</xref>). However, there are few similar works in transmission line point cloud classification.</p>
<p>This article proposes an improved network that can resist the precision loss when computing power is limited. The improved network uses group normalization to replace batch normalization, which will cause accuracy loss when the batch size is small. To decrease the model size and increase the inference speed, one-dimensional (1D) convolution layers are used instead of the fully connected layer. Transfer learning is also used to overcome the lack of transmission line point clouds. The model is first trained on ModelNet40 data set, and then fine-tuned on the transmission line point clouds data set. The results on ModelNet40 are compared to other networks. And ablation study on how the batch size affects the accuracy of the model is done to show the effect of the proposed method.</p>
<p>The main contributions of this article are as follows:<list list-type="simple">
<list-item>
<p>1) This article proposes an improved model based on PointNet&#x2b;&#x2b;, which performs better when the computing power is limited and batch size is small, such as an unmanned aerial vehicle that inspections along transmission lines.</p>
</list-item>
<list-item>
<p>2) The proposed model has a smaller model size and faster inference time while not losing much accuracy, which is suitable for real-time detection on an unmanned aerial vehicle.</p>
</list-item>
</list>
</p>
<p>This article is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> introduces the related works. <xref ref-type="sec" rid="s3">Section 3</xref> explains the proposed method. <xref ref-type="sec" rid="s4">Section 4</xref> is the experimental study and analysis. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the article.</p>
</sec>
<sec id="s2">
<title>2 Related Works</title>
<p>This article proposes a point cloud classification and recognition model based on group normalization (<xref ref-type="bibr" rid="B25">Wu and He, 2018</xref>). In the forward propagation of the model, point cloud features are grouped and normalized to ensure the performance of the model when the computing power is limited. In addition, 1D convolution is used to replace the fully connected layer used in the model for feature extraction and classification to accelerate model training and reduce the model size. The experiment uses the PointNet&#x2b;&#x2b; as the basic network for improvement. As there is less point cloud data on the transmission line, transfer learning is used for training. First, the model is trained on the public data set ModelNet40 (<xref ref-type="bibr" rid="B26">Wu et al., 2015</xref>), which focuses on point cloud classification and then uses the transmission line point cloud data collected by laser radar to fine-tune the network and realize the point cloud classification of the transmission line. The models used and the improvements made will be discussed later.</p>
<sec id="s2-1">
<title>2.1 Batch Normalization</title>
<p>It is an important assumption in machine learning that the training data and test data are independent and identically distributed. However, the distribution of data will shift in the process of neural network training process, causing the subsequent layers of the network to learn new distribution changes, and the training speed becomes slower and more difficult to converge than independent and identical distribution. Batch normalization (<xref ref-type="bibr" rid="B7">Ioffe and Szegedy, 2015</xref>; <xref ref-type="bibr" rid="B19">Santurkar et al., 2018</xref>) can be used to solve this problem.</p>
<p>For the input vector <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of dimension (N, C, H, W), batch normalization is performed on the dimension C of the batch size. The calculation process is as follows:<list list-type="simple">
<list-item>
<p>1) First the mean value of the input vector <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is calculated in the entire batch:</p>
</list-item>
</list>
<disp-formula id="e1">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mtext>B</mml:mtext>
</mml:msub>
<mml:mo>&#x003D;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>2) Then the variance <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is calculated:</p>
</list-item>
</list>
<disp-formula id="e2">
<mml:math id="m5">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>3) After obtaining the mean and variance, normalization is performed:</p>
</list-item>
</list>
<disp-formula id="e3">
<mml:math id="m6">
<mml:mrow>
<mml:msubsup>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:msqrt>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">&#x3f5;</mml:mi>
</mml:msqrt>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf4">
<mml:math id="m7">
<mml:mi mathvariant="italic">&#x3f5;</mml:mi>
</mml:math>
</inline-formula> is a very small number added for numerical stability, and <inline-formula id="inf5">
<mml:math id="m8">
<mml:mrow>
<mml:msubsup>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the result of <inline-formula id="inf6">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> after normalization.<list list-type="simple">
<list-item>
<p>4) Finally, the result of batch normalization <inline-formula id="inf7">
<mml:math id="m10">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> can be obtained after scaling of two trainable parameters:</p>
</list-item>
</list>
<disp-formula id="e4">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msubsup>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf8">
<mml:math id="m12">
<mml:mi>&#x3b3;</mml:mi>
</mml:math>
</inline-formula> and <inline-formula id="inf9">
<mml:math id="m13">
<mml:mi>&#x3b2;</mml:mi>
</mml:math>
</inline-formula> are trainable parameters.<list list-type="simple">
<list-item>
<p>5) For the PointNet&#x2b;&#x2b; model using batch normalization, for the input point cloud <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, its point cloud category <inline-formula id="inf11">
<mml:math id="m15">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> can be expressed as:</p>
</list-item>
</list>
<disp-formula id="e5">
<mml:math id="m16">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x220f;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:munderover>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>B</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-2">
<title>2.2 One-Dimensional Convolution</title>
<p>Convolutional layers are always used to extract features. Among the various types of convolutional layers, 2D convolutional layer is often used in image-related tasks, such as image classification and detection. 1D convolutional layer is often used in natural language process&#x2013;related tasks. Due to the high computing efficiency of convolutional operators, in many recent works, it is a trend that fully connected layers are replaced by several convolutional layers.</p>
</sec>
<sec id="s2-3">
<title>2.3 Transfer Learning</title>
<p>In the case of insufficient training data, there are several methods to improve accuracy during the training stage, such as data argumentation, meta learning, and transfer learning. Data argumentation can expand the data set by applying different transforms on the original data while meta learning focuses on learning the method to learn.</p>
<p>Transfer learning depends on the similarity between the source data and target data. When the target data are insufficient, the model on a similar data set with enough data is trained first, then the target data set is fine-tuned. This method always performs better than when only training the model on the target data set owing to the knowledge learned in the big data set. It is a common method in image tasks when data are limited and can also be applied to other domains.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Proposed Model</title>
<sec id="s3-1">
<title>3.1 Normalization Improvement</title>
<p>Suppose the input data dimension is (N, C, H, W), then batch normalization is performed on the channel of N. However, due to the capacity of the graphics cards, N is generally not very large to keep fast inference speed, which may result in poor batch normalization. On the other hand, group normalization is performed on the channel of C. Generally, C may reach a value of 128 or 256, or even greater within the network, so it will not adversely affect the result of normalization. Group normalization is performed on the channel dimension, and therefore group normalization can solve the problem of poor performance when batch normalization is used with small batches. This is the reason why this article uses group normalization in PointNet&#x2b;&#x2b; to replace the original batch normalization to improve its accuracy.</p>
<p>The calculation process of group normalization is as follows:<list list-type="simple">
<list-item>
<p>1) Different from batch normalization, group normalization first groups different channels and normalizes each group:</p>
</list-item>
</list>
<disp-formula id="e6">
<mml:math id="m17">
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x230a;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>The meaning of <inline-formula id="inf12">
<mml:math id="m18">
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>G</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x230a;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>G</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is to round the quotient of the number of channels and the number of groups of elements with indices k and i. If they are the same, then they are in the same group. The group size is a hyper-parameter and is generally set to 32.<list list-type="simple">
<list-item>
<p>2) After the grouping, the process of group normalization is similar to batch normalization, assuming that the input vector in the same group is:</p>
</list-item>
</list>
<disp-formula id="e7">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>3) First the mean value of the input vector in the entire group <inline-formula id="inf13">
<mml:math id="m20">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is calculated:</p>
</list-item>
</list>
<disp-formula id="e8">
<mml:math id="m21">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>4) Then the variance <inline-formula id="inf14">
<mml:math id="m22">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is calculated:</p>
</list-item>
</list>
<disp-formula id="e9">
<mml:math id="m23">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>5) After obtaining the mean and variance, normalization can be performed:</p>
</list-item>
</list>
<disp-formula id="e10">
<mml:math id="m24">
<mml:mrow>
<mml:msubsup>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:msqrt>
<mml:msup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">&#x3f5;</mml:mi>
</mml:msqrt>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>6) Finally, the result of batch normalization can be obtained after the scaling of two trainable parameters <inline-formula id="inf15">
<mml:math id="m25">
<mml:mi>&#x3b3;</mml:mi>
</mml:math>
</inline-formula> and <inline-formula id="inf16">
<mml:math id="m26">
<mml:mi>&#x3b2;</mml:mi>
</mml:math>
</inline-formula>:</p>
</list-item>
</list>
<disp-formula id="e11">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msubsup>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>7) For the PointNet&#x2b;&#x2b; model using batch normalization, for the input point cloud <inline-formula id="inf17">
<mml:math id="m28">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, its point cloud category <inline-formula id="inf18">
<mml:math id="m29">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> can be expressed as:</p>
</list-item>
</list>
<disp-formula id="e12">
<mml:math id="m30">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x220f;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:munderover>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>G</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
</sec>
<sec id="s3-2">
<title>3.2 Proposed Model Structure</title>
<p>The proposed model uses PointNet&#x2b;&#x2b; as the basic model to improve. Group normalization is used to replace batch normalization, and the fully connected layer is replaced with several 1D convolutional layers. The network structure is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, which contains two cascaded sampling modules with a feature extraction module, then the feature is fed to three cascaded 1D convolutional layers to get the final classification.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Diagram of the improved model.</p>
</caption>
<graphic xlink:href="fenrg-10-839273-g001.tif"/>
</fig>
<p>Suppose the <inline-formula id="inf19">
<mml:math id="m31">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf20">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mtext>i</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is a d dimension vector, and the feature extracted by the PointNet layer has a dimension of C. The input point cloud should be sampled first, and a subset of N1 points is sampled from the initial N points. The sampling method adopts the farthest point sampling method that the farthest point is selected each time to maximize the mutual distance between the points in the sampling subset. The sampling result obtained by this sampling method is easier to converge than the result obtained by random sampling. Then, for each point in the subset, K points are selected with the closest distance, and <inline-formula id="inf21">
<mml:math id="m33">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> neighboring points are extracted in total, and these points are sent to the PointNet layer to extract the resulting <inline-formula id="inf22">
<mml:math id="m34">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> features. After this, through a sampling and feature extraction module of the same structure, <inline-formula id="inf23">
<mml:math id="m35">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> features are extracted. Then, the extracted <inline-formula id="inf24">
<mml:math id="m36">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> features by the two cascading sampling and PointNet modules are sent to the final PointNet layer to extract the <inline-formula id="inf25">
<mml:math id="m37">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>C</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> features. The PointNet layer is the same as in PointNet&#x2b;&#x2b;, which includes multilayer perceptron and max pooling layers. Then, three cascading 1D convolution layers are used to get the final classification and the numbers of the features decreased after each 1D convolution layer. Assuming the final classification category as <inline-formula id="inf26">
<mml:math id="m38">
<mml:mrow>
<mml:mtext>cls</mml:mtext>
</mml:mrow>
</mml:math>
</inline-formula>, then <inline-formula id="inf27">
<mml:math id="m39">
<mml:mrow>
<mml:mtext>cls</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>&#xa0;conv</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mtext>d</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x220f;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:munderover>
<mml:mtext>pointnet </mml:mtext>
<mml:mrow>
<mml:mrow>
<mml:mtext>sample</mml:mtext>
</mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</p>
</sec>
</sec>
<sec id="s4">
<title>4 Computational Analysis</title>
<sec id="s4-1">
<title>4.1 Data Set Description and Model Pretraining</title>
<p>The deep learning model is a data-driven model. Because of the small amount of power equipment data, it is necessary to pretrain on a large public data set first, and then fine-tune on the collected power equipment data set. The public data set ModelNet40 is used to pretrain the model. ModelNet40 was released in 2015, containing 3D models of 40 categories (including computers, bottles, airplanes, etc.). <xref ref-type="fig" rid="F2">Figure 2</xref> shows some examples of ModelNet40. This data set is collected by the Princeton Vision and Robotics Laboratory. It is often used to evaluate point cloud deep learning models for semantic segmentation, instance segmentation, and classification.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Point cloud examples of ModelNet40.</p>
</caption>
<graphic xlink:href="fenrg-10-839273-g002.tif"/>
</fig>
<p>ModelNet40 contains a total of 12,311 point cloud models, each of which includes three models which have 512, 1024, or 2048 points. Among them, there are 9843 point cloud models in the training set, and the remaining 2468 point cloud models are divided into test sets. The improved PointNet&#x2b;&#x2b; model is pretrained using point clouds with a single model point cloud number of 1024.</p>
<p>The pretraining is based on the Ubuntu 18.04 platform, the graphics card model is 16&#xa0;GB Telsa T4, and the PyTorch (<xref ref-type="bibr" rid="B15">Paszke et al., 2019</xref>) version is 1.6.</p>
<p>The training parameters are configured as follows. The training epochs are set to 200, the batch size is 24, and the initial learning rate is 0.001. After every 20 rounds, the learning rate decays to 0.7 times. The optimizer is the Adam optimizer (<xref ref-type="bibr" rid="B8">Kingma and Ba, 2014</xref>), and the momentum parameter is set to 0.9, and the L2 decay is set to 0.0001.</p>
<p>After 200 epochs of training, the accuracy of the model on ModelNet40 reaches 92.9% at the instance level, and the average accuracy on the class is 90.7%. Subsequent experiments are fine-tuned on this pretrained model. The comparison of the results between our model and previous networks is shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Results comparison on ModelNet40.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Accuracy (%)</th>
<th align="center">Model size (MB)</th>
<th align="center">Inference time (ms)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">3D ShapeNet</td>
<td align="char" char=".">84.7</td>
<td align="center">&#x2014;</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">Point</td>
<td align="char" char=".">89.2</td>
<td align="char" char=".">11.2</td>
<td align="char" char=".">33</td>
</tr>
<tr>
<td align="left">PointNet&#x2b;&#x2b;</td>
<td align="char" char=".">
<bold>91.9</bold>
</td>
<td align="char" char=".">40.7</td>
<td align="char" char=".">312</td>
</tr>
<tr>
<td align="left">Ours</td>
<td align="char" char=".">90.7</td>
<td align="char" char=".">
<bold>7.8</bold>
</td>
<td align="char" char=".">
<bold>24</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values means best in the column.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-2">
<title>4.2 Transmission Line Point Cloud Data Set</title>
<p>The RS-Ruby laser radar of Sagitar can be used to collect the point cloud data of the transmission line. Using CloudCompare point cloud processing software, cloud points of the transmission line are collected by the laser radar and can be segmented and marked as different types. <xref ref-type="fig" rid="F3">Figure 3</xref> shows the point cloud distribution of the transmission line data set.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Point cloud of transmission lines.</p>
</caption>
<graphic xlink:href="fenrg-10-839273-g003.tif"/>
</fig>
<p>In this article, four commonly used categories of transmission lines: transmission lines, transmission towers, ground, and vegetation are segmented and labelled, and a labelled data set with 343 instances was obtained. <xref ref-type="fig" rid="F4">Figure 4</xref> shows some examples of the transmission line data set. The number of labelled instances of each type of equipment is shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Examples of the transmission line point cloud data set.</p>
</caption>
<graphic xlink:href="fenrg-10-839273-g004.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Distribution of transmission line point cloud data set.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Device type</th>
<th align="center">Number of instances</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Transmission lines</td>
<td align="char" char=".">97</td>
</tr>
<tr>
<td align="left">Pylons</td>
<td align="char" char=".">81</td>
</tr>
<tr>
<td align="left">Grounds</td>
<td align="char" char=".">87</td>
</tr>
<tr>
<td align="left">Vegetation</td>
<td align="char" char=".">78</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The transmission line point cloud data set is divided by random selection, 80% of which is used for training on the basis of the pretrained PointNet&#x2b;&#x2b;, and the rest is used for testing.</p>
</sec>
<sec id="s4-3">
<title>4.3 Fine-Tune Results</title>
<p>First, the proposed model with batch normalization is trained. The training parameter configuration is as follows: the training epochs are set to 100, the batch size is 24, the initial learning rate is 0.001, the initial 10 rounds learning rate linearly rises from 0 to 0.001, and after this, the learning rate decays to 0.99 times in each round. The optimizer uses the SGD optimizer, the momentum is set to 0.9, and the L2 decay is set to 0.0001. After 100 epochs of training, the accuracy of the model on the transmission line point cloud data set reached 94.9%.</p>
<p>In the case of limited computing power, the proposed model with group normalization was also trained to compare between two normalization methods. The training parameters are the same as before except the batch size is set to 2. After 100 rounds of training, the accuracy of the model on the transmission line point cloud data set reached 92.4%.</p>
<p>It can be seen from the results shown in <xref ref-type="table" rid="T3">Table 3</xref> that the recognition accuracy is also affected due to the limited computing power. Analyzing the model structure, it can be found that when the batch size is reduced, the number of participants in the batch normalization at the same time is decreased, which leads to the loss of training accuracy. Through group normalization, only in the normalization process, the advantage of being affected by the number of channels, not related to the batch size, improves the normalization effect of the model, thereby improving the performance of the model.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Classification accuracy when computing power is enough or not enough.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Accuracy (%)</th>
<th colspan="2" align="center">Batch size</th>
</tr>
<tr>
<th align="center">2</th>
<th align="center">24</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Ours (with BN)</td>
<td align="char" char=".">92.4</td>
<td align="char" char=".">94.9</td>
</tr>
<tr>
<td align="left">Ours (with GN)</td>
<td align="char" char=".">94.6</td>
<td align="char" char=".">95.1</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>BN, Batch Normalization; GN, Group Normalization</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>To show the effectiveness of the proposed model, we also tried PointNet and PointNet&#x2b;&#x2b; in our transmission lines data set when the batch size was 24, and the results are shown in <xref ref-type="table" rid="T4">Table 4</xref>. The results show that the proposed model can achieve a similar accuracy with PointNet&#x2b;&#x2b; while having a faster inference speed and smaller model size.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Results comparison on transmission line data set.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Accuracy (%)</th>
<th align="center">Model size (MB)</th>
<th align="center">Inference time (ms)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PointNet</td>
<td align="char" char=".">93.2</td>
<td align="char" char=".">11.2</td>
<td align="char" char=".">33</td>
</tr>
<tr>
<td align="left">PointNet&#x2b;&#x2b;</td>
<td align="char" char=".">
<bold>95.4</bold>
</td>
<td align="char" char=".">40.7</td>
<td align="char" char=".">312</td>
</tr>
<tr>
<td align="left">Ours (with BN)</td>
<td align="char" char=".">94.9</td>
<td align="char" char=".">7.8</td>
<td align="char" char=".">25</td>
</tr>
<tr>
<td align="left">Ours (with GN)</td>
<td align="char" char=".">95.1</td>
<td align="char" char=".">
<bold>7.8</bold>
</td>
<td align="char" char=".">
<bold>24</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values means best in the column.</p>
</fn>
<fn>
<p>BN, Batch Normalization; GN, Group Normalization</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-4">
<title>4.4 Ablation Study</title>
<p>Using the same training parameters as the previous training, the training batch size was only changed and the experiments were performed when the batch sizes were 2, 4, and 8; the experimental results are shown in <xref ref-type="table" rid="T5">Table 5</xref>. It can be seen from the table that when the batch size becomes smaller, the model accuracy of the batch normalization method decreases. As the batch size increases, the model accuracy also increases; while the accuracy with group normalization method is stable, leading to more suitable scenarios where the computing power of the device is limited. And even when the batch size is 2, the accuracy achieved can be 94.6, which is not much smaller than 95.4 that PointNet&#x2b;&#x2b; achieves when the batch size is 24.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Classification accuracy with different normalization methods and batch size.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Accuracy (%)</th>
<th colspan="3" align="center">Batch size</th>
</tr>
<tr>
<th align="center">2</th>
<th align="center">4</th>
<th align="center">8</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Ours (with BN)</td>
<td align="char" char=".">92.4</td>
<td align="char" char=".">93.7</td>
<td align="char" char=".">94.4</td>
</tr>
<tr>
<td align="left">Ours (with GN)</td>
<td align="char" char=".">94.6</td>
<td align="char" char=".">94.7</td>
<td align="char" char=".">94.7</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>BN, Batch Normalization; GN, Group Normalization</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>
<xref ref-type="fig" rid="F5">Figure 5</xref> shows the visualized results of the part of our transmission lines data sets. It can be seen that the transmission lines and pylons are classified well. And if more data are collected, the vegetation and grounds can also be classified better.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Visualized classification results. <bold>(A)</bold> Front view. <bold>(B)</bold> Vertical view.</p>
</caption>
<graphic xlink:href="fenrg-10-839273-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion</title>
<p>In this article, point cloud data of the transmission line are collected by laser radar and are labeled to four point cloud categories. Due to less cloud data of the transmission line, the transfer learning method is used during training. First, the ModelNet40 data set is used to pretrain the improved PointNet&#x2b;&#x2b;, and then the data of the transmission line is used to fine-tune the pretrained network to realize the classification of the point cloud of the transmission line. The experimental results show that when the batch size is small, the improved PointNet&#x2b;&#x2b; model proposed in this article can effectively classify the point cloud of the transmission line.</p>
<p>In the case of limited computing power and the batch size being small, the model classification accuracy will be impaired. This article proposes to use group normalization instead of batch normalization to improve the classification accuracy. Experiments show that group normalization improves the accuracy of the model when the batch is small, and it is not limited by the batch size during model training when compared with batch normalization. The point cloud classification and recognition model based on group normalization has an accuracy of 94.9% for the point cloud classification of the transmission lines, which can effectively classify the point cloud of the transmission line.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The data sets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/supplementary material.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>ZY and SW prepared the transmission lines data set. XZ and JD wrote the original draft. ZY and SJ contributed to the study method and result analysis. All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This research is supported by the Science and Technology Project of the State Grid Zhejiang Electric Power Company (5211UZ190057).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>Authors YZ, JS, ZX, DJ, YW, and WS were employed by the State Grid Huzhou Power Supply Company.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Awrangjeb</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jonas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>20172017</year>). &#x201c;<article-title>An Automatic Technique for Power Line Pylon Detection from point Cloud Data</article-title>,&#x201d; in <conf-name>International Conference on Digital Image Computing: Techniques and Applications (DICTA)</conf-name>, <conf-loc>Sydney, NSW, Australia</conf-loc>, <conf-date>29 Nov.-1 Dec. 2017</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/dicta.2017.8227407</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chan</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Jaitly</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>Q. V.</given-names>
</name>
<name>
<surname>Vinyals</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Listen, Attend and Spell</article-title>. <comment>arXiv preprint arXiv:1508.01211</comment>. </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Automatic Clearance Anomaly Detection for Transmission Line Corridors Utilizing UAV-Borne LIDAR Data</article-title>. <source>Remote Sensing</source> <volume>10</volume> (<issue>4</issue>), <fpage>613</fpage>. <pub-id pub-id-type="doi">10.3390/rs10040613</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Devlin</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>M. W.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Toutanova</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>. <comment>arXiv preprint arXiv:1810.04805</comment>. </citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep Residual Learning for Image Recognition</article-title>,&#x201d; in <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Las Vegas, NV, USA</conf-loc>, <conf-date>27-30 June 2016</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>770</fpage>&#x2013;<lpage>778</lpage>.<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Fast Reconstruction of 3D Point Cloud Model Using Visual SLAM on Embedded UAV Development Platform</article-title>. <source>Remote Sensing</source> <volume>12</volume> (<issue>20</issue>), <fpage>3308</fpage>. <pub-id pub-id-type="doi">10.3390/rs12203308</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ioffe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift</article-title>,&#x201d; in <conf-name>In International conference on machine learning</conf-name>, <conf-date>July 2015</conf-date> (<publisher-name>PMLR</publisher-name>), <fpage>448</fpage>&#x2013;<lpage>456</lpage>. </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Method for Stochastic Optimization</article-title>. <comment>arXiv preprint arXiv:1412.6980</comment>. </citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Heng</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>C. W.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Pointaugment: an Auto-Augmentation Framework for point Cloud Classification</article-title>,&#x201d; in <conf-name>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <fpage>6378</fpage>&#x2013;<lpage>6387</lpage>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.-D.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>G.-S.</given-names>
</name>
</person-group> (<year>2020a</year>). <article-title>A Geometry-Attentional Network for ALS point Cloud Classification</article-title>. <source>ISPRS J. Photogrammetry Remote Sensing</source> <volume>164</volume>, <fpage>26</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.03.016</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Unmanned Aerial Vehicle for Transmission Line Inspection: Status, Standardization, and Perspectives</article-title>. <source>Front. Energ. Res.</source> <volume>9</volume>, <fpage>336</fpage>. <pub-id pub-id-type="doi">10.3389/fenrg.2021.713634</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Chu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Real-time Detection and Spatial Localization of Insulators for UAV Inspection Based on Binocular Stereo Vision</article-title>. <source>Remote Sensing</source> <volume>13</volume> (<issue>2</issue>), <fpage>230</fpage>. <pub-id pub-id-type="doi">10.3390/rs13020230</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Munir</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Awrangjeb</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Stantic</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Islam</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Voxel-based Extraction of Individual Pylons and Wires from Lidar point Cloud Data</article-title>. <source>ISPRS Ann.</source> <volume>4</volume>, <fpage>91</fpage>&#x2013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.5194/isprs-annals-IV-4-W8-91-2019</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niemeyer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rottensteiner</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Soergel</surname>
<given-names>U.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Conditional Random Fields for Lidar Point Cloud Classification in Complex Urban Areas</article-title>. <source>ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci.</source> <volume>I-3</volume> (<issue>3</issue>), <fpage>263</fpage>&#x2013;<lpage>268</lpage>. <pub-id pub-id-type="doi">10.5194/isprsannals-i-3-263-2012</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Paszke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gross</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Massa</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lerer</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bradbury</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Pytorch: An Imperative Style, High-Performance Deep Learning Library</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>32</volume>, <fpage>8026</fpage>&#x2013;<lpage>8037</lpage>. <comment>arXiv:1912.01703</comment> </citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Qi</surname>
<given-names>C. R.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mo</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Guibas</surname>
<given-names>L. J.</given-names>
</name>
</person-group> (<year>2017a</year>). &#x201c;<article-title>Pointnet: Deep Learning on point Sets for 3d Classification and Segmentation</article-title>,&#x201d; in <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <fpage>652</fpage>&#x2013;<lpage>660</lpage>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qi</surname>
<given-names>C. R.</given-names>
</name>
<name>
<surname>Yi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Guibas</surname>
<given-names>L. J.</given-names>
</name>
</person-group> (<year>2017b</year>). <article-title>Pointnet&#x2b;&#x2b;: Deep Hierarchical Feature Learning on point Sets in a Metric Space</article-title>. <comment>arXiv preprint arXiv:1706.02413</comment>. </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qin</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Mei</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Novel Method of Autonomous Inspection for Transmission Line Based on cable Inspection Robot Lidar Data</article-title>. <source>Sensors</source> <volume>18</volume> (<issue>2</issue>), <fpage>596</fpage>. <pub-id pub-id-type="doi">10.3390/s18020596</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Santurkar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tsipras</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ilyas</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>M&#x105;dry</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>How Does Batch Normalization Help Optimization</article-title>,&#x201d; in <conf-name>In Proceedings of the 32nd international conference on neural information processing systems</conf-name>, <conf-date>December 2018</conf-date>. </citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ioffe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vanhoucke</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Alemi</surname>
<given-names>A. A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Inception-v4, Inception-Resnet and the Impact of Residual Connections on Learning</article-title>,&#x201d; in <conf-name>In Thirty-first AAAI conference on artificial intelligence</conf-name>, <conf-loc>San Francisco, CA</conf-loc>, <conf-date>February 4-9, 2017</conf-date> (<publisher-loc>San Francisco</publisher-loc>: <publisher-name>AAAI Press</publisher-name>). <pub-id pub-id-type="doi">10.5555/3298023.3298188</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teng</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C. R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>H. H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>F. R.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Mini-uav Lidar for Power Line Inspection</article-title>. <source>Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci.</source> <volume>XLII-2/W7</volume>, <fpage>297</fpage>&#x2013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.5194/isprs-archives-xlii-2-w7-297-2017</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shazeer</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Parmar</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Uszkoreit</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Attention Is All You Need</article-title>,&#x201d; in <conf-name>31st Conference on Neural Information Processing Systems (NIPS 2017)</conf-name>, <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>. </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Mathiopoulos</surname>
<given-names>P. T.</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Multiscale and Hierarchical Feature Extraction Method for Terrestrial Laser Scanning point Cloud Classification</article-title>. <source>IEEE Trans. Geosci. Remote Sensing</source> <volume>53</volume> (<issue>5</issue>), <fpage>2409</fpage>&#x2013;<lpage>2425</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2014.2359951</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Watanabe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hori</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hershey</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Hayashi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Hybrid CTC/attention Architecture for End-To-End Speech Recognition</article-title>. <source>IEEE J. Sel. Top. Signal. Process.</source> <volume>11</volume> (<issue>8</issue>), <fpage>1240</fpage>&#x2013;<lpage>1253</lpage>. <pub-id pub-id-type="doi">10.1109/jstsp.2017.2763455</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Group Normalization</article-title>.&#x201d; in <conf-name>Proceedings of the European conference on computer vision (ECCV)</conf-name>, <fpage>3</fpage>&#x2013;<lpage>19</lpage>. </citation>
</ref>
<ref id="B26">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Khosla</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). &#x201c;<article-title>A Deep Representation for Volumetric Shapes</article-title>,&#x201d; in <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Boston, MA, USA</conf-loc>, <conf-date>7-12 June 2015</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>1912</fpage>&#x2013;<lpage>1920</lpage>. </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jun-Hua</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhun</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>An-Min</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Biao</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Autonomous Control Method of Rotor UAVs for Power Inspection with Renewable Energy Based on Swarm Intelligence</article-title>. <source>Front. Energ. Res.</source> <volume>9</volume>, <fpage>229</fpage>. <pub-id pub-id-type="doi">10.3389/fenrg.2021.697054</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Rabbat</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>A Graph-Cnn for 3d point Cloud Classification</article-title>,&#x201d; in <conf-name>IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>, <conf-loc>Calgary, AB, Canada</conf-loc>, <conf-date>15-20 April 2018</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>6279</fpage>&#x2013;<lpage>6283</lpage>.<pub-id pub-id-type="doi">10.1109/ICASSP.2018.8462291</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Mathiopoulos</surname>
<given-names>P. T.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>A Multilevel point-cluster-based Discriminative Feature for ALS point Cloud Classification</article-title>. <source>IEEE Trans. Geosci. Remote Sensing</source> <volume>54</volume> (<issue>6</issue>), <fpage>3309</fpage>&#x2013;<lpage>3321</lpage>. <pub-id pub-id-type="doi">10.1109/tgrs.2016.2514508</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Als point Cloud Classification with Small Training Data Set Based on Transfer Learning</article-title>,&#x201d; in <conf-name>IEEE Geoscience and Remote Sensing Letters</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1406</fpage>&#x2013;<lpage>1410</lpage>. </citation>
</ref>
</ref-list>
</back>
</article>