<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2023.1243847</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Truck model recognition for an automatic overload detection system based on the improved MMAL-Net</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Sun</surname>
<given-names>Jiachen</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2351813/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Su</surname>
<given-names>Jin</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yan</surname>
<given-names>Zhenhao</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gao</surname>
<given-names>Zenggui</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2353989/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sun</surname>
<given-names>Yanning</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Liu</surname>
<given-names>Lilan</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Shanghai Key Laboratory of Intelligent Manufacturing and Robotics, School of Mechatronic Engineering and Automation, Shanghai University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>College of Information Engineering, Lanzhou University of Finance and Economics</institution>, <addr-line>Lanzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001"><p>Edited by: Qieshi Zhang, Chinese Academy of Sciences (CAS), China</p></fn>
<fn fn-type="edited-by" id="fn0002"><p>Reviewed by: Ziliang Ren, Dongguan University of Technology, China; Jing Ji, Xidian University, China; Ketsuseki Cyou, Waseda University, Japan</p></fn>
<corresp id="c001">&#x002A;Correspondence: Lilan Liu, <email>lancy@shu.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1243847</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Sun, Su, Yan, Gao, Sun and Liu.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Sun, Su, Yan, Gao, Sun and Liu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Efficient and reliable transportation of goods through trucks is crucial for road logistics. However, the overloading of trucks poses serious challenges to road infrastructure and traffic safety. Detecting and preventing truck overloading is of utmost importance for maintaining road conditions and ensuring the safety of both road users and goods transported. This paper introduces a novel method for detecting truck overloading. The method utilizes the improved MMAL-Net for truck model recognition. Vehicle identification involves using frontal and side truck images, while APPM is applied for local segmentation of the side image to recognize individual parts. The proposed method analyzes the captured images to precisely identify the models of trucks passing through automatic weighing stations on the highway. The improved MMAL-Net achieved an accuracy of 95.03% on the competitive benchmark dataset, Stanford Cars, demonstrating its superiority over other established methods. Furthermore, our method also demonstrated outstanding performance on a small-scale dataset. In our experimental evaluation, our method achieved a recognition accuracy of 85% when the training set consisted of 20 sets of photos, and it reached 100% as the training set gradually increased to 50 sets of samples. Through the integration of this recognition system with weight data obtained from weighing stations and license plates information, the method enables real-time assessment of truck overloading. The implementation of the proposed method is of vital importance for multiple aspects related to road traffic safety.</p>
</abstract>
<kwd-group>
<kwd>overload detection</kwd>
<kwd>truck model recognition</kwd>
<kwd>automatic weighing station</kwd>
<kwd>fine-grained visual categorization</kwd>
<kwd>MMAL-Net</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="1"/>
<equation-count count="5"/>
<ref-count count="37"/>
<page-count count="10"/>
<word-count count="6042"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Visual Neuroscience</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1.</label>
<title>Introduction</title>
<p>With the rapid development of the global economy and the acceleration of urbanization processes, highways play a crucial role in connecting different regions and cities. In the realm of road transportation, trucks serve as vital transportation tools, undertaking the task of transporting a substantial amount of goods. However, the issue of truck overloading has become one of the primary challenges in road traffic safety and road damage. Overloaded trucks exert significant pressure on road infrastructure, increasing the risk of traffic accidents and potentially leading to severe road collapse incidents. Therefore, the development of an accurate and efficient truckload monitoring method holds significant practical significance.</p>
<p>Traditional methods for truckload monitoring mainly rely on static weight measurement equipment such as weighbridges (<xref ref-type="bibr" rid="ref36">Zhou et al., 2005</xref>) and fixed scales. However, these devices have several limitations, including the need for trucks to stop for measurement and high time and labor costs. Additionally, static measurement methods cannot provide real-time monitoring and detection capabilities for violations, limiting their effectiveness in practical applications.</p>
<p>To address these challenges, a promising solution has emerged: utilizing camera images from highway weigh stations for truck model recognition and combining them with weighing information obtained from a dynamic weighing system. By leveraging truck photos captured near these stations and employing advanced image processing and pattern recognition techniques, truck models can be identified accurately. Regrettably, the current network architectures utilized for recognition in this context often exhibit a simplistic nature, leading to suboptimal accuracy in the identification process. As a consequence, determining whether a truck is overloaded becomes inaccurate. Moreover, the ability to perform real-time recognition using captured photographs poses an unresolved challenge that demands urgent attention.</p>
<p>This paper aims to propose a truck model recognition method based on highway automatic weighing station camera images, with the objective of accurately identifying truck models. Consequently, the maximum load capacity of the trucks is determined. Through the integration of license plates and weighing information, the system can accurately determine if a truck is carrying excessive load. By doing so, it can prevent the occurrence of misjudgments caused by incidents of license plates damage, which are likely to happen in schemes that rely solely on license plates recognition to obtain vehicle models. Through this method, precise truck information can be provided for freight management, transportation safety, and highway planning, promoting the development of the logistics industry and enhancing traffic safety.</p>
</sec>
<sec id="sec2">
<label>2.</label>
<title>Literature review</title>
<sec id="sec3">
<label>2.1.</label>
<title>FGVC</title>
<p>Fine-Grained Visual Categorization (FGVC) refers to the task of classifying objects into different subcategories or fine-grained classes within a broader category. In FGVC, the goal is to achieve detailed discrimination and classification among visually similar objects, such as different species of birds, breeds of dogs, or models of cars. This field of research focuses on developing computer vision algorithms and techniques to accurately recognize and classify objects at a fine-grained level, where subtle differences between subclasses need to be distinguished.</p>
<p>Fine-grained image classification, in contrast to conventional image classification tasks, encompasses a low signal-to-noise ratio, restricting the presence of highly discriminating information to minuscule local regions. Thus, the crux of achieving success in fine-grained image classification algorithms lies in the identification and efficient utilization of these valuable local region insights. Presently, most classification algorithms adhere to a common workflow: initial localization of the foreground object and its distinct local regions, followed by individual feature extraction from these regions. The processed features are subsequently utilized for classifier training and prediction purposes. To attain satisfactory classification results, numerous existing algorithms heavily depend on manual annotation information (<xref ref-type="bibr" rid="ref27">Wei et al., 2018</xref>), such as bounding boxes and part locations. The annotation frame aids in foreground object detection, effectively mitigating background noise interference. Local region positions serve to identify valuable regions or align perspectives, facilitating the extraction of local features. Nevertheless, the costly acquisition of manual annotation information severely limits the practicality of these classification algorithms. In recent years, an increasing number of studies have opted to exclude such labeling information, relying solely on labels to accomplish image classification tasks (<xref ref-type="bibr" rid="ref15">Lin et al., 2005</xref>; <xref ref-type="bibr" rid="ref33">Zhang et al., 2016</xref>), resulting in commendable outcomes.</p>
<p>In the research and development of FGVC, traditional classification algorithms based on handcrafted features were initially employed. These algorithms typically begin by extracting local features, such as Histogram of Oriented Gradients (HOG) (<xref ref-type="bibr" rid="ref4">Dalal and Triggs, 2005</xref>), from the images. Subsequently, an encoding model like Vector of Locally Aggregated Descriptors (VLAD) (<xref ref-type="bibr" rid="ref13">J&#x00E9;gou et al., 2010</xref>) is employed for feature encoding, resulting in the desired feature representation. However, the limited descriptive power of handcrafted features often leads to suboptimal classification performance. In the early stages of fine-grained visual categorization research, the representation capacity of features became a primary bottleneck hindering performance improvement.</p>
<p>In recent years, Convolutional Neural Network (CNN)-based methods for fine-grained image recognition (<xref ref-type="bibr" rid="ref26">Wang et al., 2017</xref>; <xref ref-type="bibr" rid="ref28">Xie et al., 2017</xref>) have significantly matured. <xref ref-type="bibr" rid="ref5">Donahue et al. (2014)</xref> conducted an analysis of a CNN model trained on the ImageNet dataset, revealing that the features extracted from the CNN possess more robust semantic characteristics and exhibit superior differentiation compared to artificial features. Building upon these findings, the researchers applied the convolution features to various domain-specific tasks, including fine-grained classification, resulting in improved classification performance. Nevertheless, the crucial components of the task tend to be subtle and were not adequately captured by conventional CNN approaches. Consequently, researchers have directed their attention toward internal enhancements within the framework. <xref ref-type="bibr" rid="ref31">Zhang et al. (2014)</xref> introduced the Part R-CNN algorithm, which leverages R-CNN (<xref ref-type="bibr" rid="ref9">Girshick et al., 2014</xref>) for image detection. This methodology aims to achieve precise localization of crucial components and enhance feature representation. <xref ref-type="bibr" rid="ref2">Branson et al. (2014)</xref> proposed the Pose Normalized Convolutional Neural Network (Pose Normalized CNN) algorithm. Their approach comprises several steps: localization detection is performed on local regions for each input image, followed by cropping the image based on the detected annotation boxes, extracting hierarchical local information, and conducting pose alignment. Subsequently, distinct layers of convolutional features are extracted for different body parts. Finally, these convolutional features are concatenated into a feature vector and utilized for SVM model training. These approaches have demonstrated robust feature representation capabilities and yielded promising results in fine-grained image recognition tasks.</p>
<p>Compared to regular classification tasks, acquiring fine-grained image databases poses greater challenges and requires stronger domain expertise for data collection and annotation. However, in recent years, there has been a significant increase in the availability of fine-grained image databases, which reflects the flourishing development trend and strong real-world demand in this field. Currently, commonly used fine-grained image databases include (1) CUB200-2011: It comprises a total of 11,788 bird images belonging to 200 different categories. This database provides rich manual annotations, including 15 local part locations, 312 binary attributes, 1 bounding box, and semantic segmentation images, (2) Stanford Dogs: This database offers a collection of 20,580 images featuring 120 different breeds of dogs. It provides only bounding box annotations, (3) Oxford Flowers: This database is divided into two scales, containing 17 and 102 categories of flowers, respectively. The 102-category database is more commonly used, with each category containing 40 to 258 images. In total, there are 8,189 images in this database, which provides only semantic segmentation images without any additional annotations, (4) Cars: This database provides a collection of 16,185 vehicle images belonging to 196 different categories, encompassing various brands, years, and models. Only bounding box annotations are provided, and (5) FGVC-Aircraft: This database consists of 10,200 images of 102 different aircraft categories, with each category containing 100 distinct photos. Only bounding box annotations are provided. In recent years, extensive research has been conducted on fine-grained image databases. DCL (<xref ref-type="bibr" rid="ref3">Chen et al., 2019</xref>) employed a deconstruction and reconstruction approach to learn semantic correlations among local regions in input images. API-Net (<xref ref-type="bibr" rid="ref37">Zhuang et al., 2020</xref>) progressively recognized pairs of fine-grained images through iterative interaction. GCP (<xref ref-type="bibr" rid="ref21">Song et al., 2022</xref>) introduced a dedicated network branch to magnify the importance of small eigenvalues. MSHQP (<xref ref-type="bibr" rid="ref23">Tan et al., 2022</xref>) effectively modeled intra and inter-layer feature interactions, integrating multi-layer features to enhance part responses. These methods primarily focus on locating and utilizing key regions for final recognition, yielding promising performance. However, they tend to overlook the potential contribution of complementary regions that can also play a positive role in the recognition process.</p>
</sec>
<sec id="sec4">
<label>2.2.</label>
<title>Vehicle recognition and classification</title>
<p>Vehicle recognition and classification are essential components of FGVC field. In the context of vehicles, this entails distinguishing between closely related classes such as different car models, brands, and types, where subtle visual differences in features become crucial for accurate classification. Currently, research on vehicle recognition and classification primarily centers around three main approaches: pattern recognition based on matching method, pattern recognition based on machine learning and pattern recognition based on deep learning.</p>
<p>The first approach involves the identification of vehicles through license plates and vehicle tag detection using a matching method. While the license plate number and label characteristics can directly identify the vehicle&#x2019;s brand and model (<xref ref-type="bibr" rid="ref17">Psyllos et al., 2010</xref>; <xref ref-type="bibr" rid="ref12">Huang et al., 2015</xref>), this method has a limitation: it does not encompass all the fine-grained features associated with the vehicle brand and model. Apart from the license plates and labels, vehicle lights and other textural information also bear the characteristics of the vehicle model. Relying solely on license plates and tags is insufficient. Additionally, the license plates of trucks are prone to being contaminated by dirt and dust, which leads to reduced visibility and clarity. In such scenarios, this method becomes ineffective.</p>
<p>The second approach involves using machine learning to classify vehicle brands and models. The traditional machine learning method comprises two steps: feature extraction and classifier classification. Fraz et al. proposed a method for recognizing vehicle brands and models based on a SIFT feature dictionary (<xref ref-type="bibr" rid="ref7">Fraz et al., 2014</xref>). In this method, SIFT features of pictures from the training set&#x2019;s vehicles were treated as &#x201C;words&#x201D; to create a dictionary of vehicle brands and models. However, this method necessitates extensive computation and takes a considerable amount of time to identify each image, making it unsuitable for real-time vehicle brand and model classification in practical scenarios. Abdul et al. proposed a method employing a cascade classifier (<xref ref-type="bibr" rid="ref20">Siddiqui et al., 2016</xref>). Initially, representative features were extracted from the samples instead of using all features. Subsequently, a cascade-based SVM classifier was employed, resulting in significant improvements in real-time recognition. <xref ref-type="bibr" rid="ref1">Biglari et al. (2019)</xref> introduced an algorithm based on the histogram of gradient directions feature and cascade classifier. Multiple vehicle brand models were trained first, followed by classification using a cascade SVM classifier, achieving an impressive classification accuracy of up to 96.78%. However, this method still requires hardware acceleration for real-time classification.</p>
<p>The third approach involves vehicle pattern recognition based on deep learning. Yang et al. proposed a method for recognizing vehicle brands and models based on the joint attributes of vehicles (<xref ref-type="bibr" rid="ref29">Yang et al., 2015</xref>). This method extracts vehicle features from multiple perspectives and angles, fuses the extracted features, and performs recognition. While this method is well-suited for recognizing vehicle brands and models in complex scenes, its real-time performance is compromised due to the abundance of features. Huang et al. suggested randomly discarding certain layers during the training of ResNet to obtain a convolutional neural network with random depth (<xref ref-type="bibr" rid="ref11">Huang et al., 2016</xref>), thereby addressing the issue of gradient vanishing caused by excessively deep networks. Fang et al. introduced a fine-grained method for recognizing vehicle brands and models (<xref ref-type="bibr" rid="ref6">Fang et al., 2016</xref>), utilizing a CNN model to extract local and overall features of vehicles, and combining them for classification. <xref ref-type="bibr" rid="ref24">Wang et al. (2020)</xref> proposed a method based on structural graph to learn discriminative representations for vehicle recognition. This approach first constructs a global structural graph from the features generated by a convolutional network. Then, it utilizes this structural graph as guidance to generate effective vehicle representations. <xref ref-type="bibr" rid="ref16">Mo et al. (2020)</xref> analyzed the relationship between the number and distribution of vehicle axles and the weight limit of trucks. They proposed a circular detection method based on an improved Hough and clustering algorithm to identify the axles of trucks. Presently, most studies on deep learning for vehicle brand recognition rely on a single convolutional neural network model. However, for the intricate task of truck brand classification, a single model falls short in achieving satisfactory classification accuracy. Consequently, integrating multiple convolutional neural network models to develop a fusion model suitable for truck brand classification becomes a problem that requires resolution in this study.</p>
</sec>
</sec>
<sec sec-type="methods" id="sec5">
<label>3.</label>
<title>Method</title>
<p>In typical scenarios, automatic weighing stations on highways are equipped with multiple cameras to capture frontal and side images of trucks. When utilizing these images for model recognition, the initial step involves utilizing the frontal image (front view) for identification. Analyzing the frontal image allows for the determination of the truck&#x2019;s model. Additionally, the side image is utilized to enhance accuracy in identifying the frontal view. The side image provides supplementary perspectives and details, thereby improving the accuracy of frontal view recognition. Moreover, the side image enables the segmentation of the truck into multiple parts, further refining model recognition precision. Through comprehensive analysis of both frontal and side images, we can achieve more accurate truck identification and conduct additional analysis based on its body features. Knowing the truck&#x2019;s model provides information regarding its rated load capacity. The weight measurement data obtained in the automatic weighing area enables straightforward determination of whether the truck is overloaded. Moreover, the inclusion of license plates information enables efficient monitoring and regulation by traffic authorities. <xref rid="fig1" ref-type="fig">Figure 1</xref> illustrates the process described above.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>The process of the overload detection.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g001.tif"/>
</fig>
<sec id="sec6">
<label>3.1.</label>
<title>The improved MMAL-Net</title>
<p>We improved MMAL-Net (<xref ref-type="bibr" rid="ref32">Zhang et al., 2021</xref>) and employed it for truck recognition and classification. In <xref rid="fig2" ref-type="fig">Figure 2</xref>, we illustrate the network architecture that was constructed during the training phase, consisting of three branches: frontal, side, and part branches. The frontal branch is responsible for recognizing and classifying frontal truck images, while the side branch receives side images and segments them into multiple parts using the Attention Part Proposal Module (APPM). The part branch, on the other hand, specializes in recognizing and classifying part images. All three branches utilize a ResNet-50 (<xref ref-type="bibr" rid="ref10">He et al., 2016</xref>) for feature extraction and employ a Fully Connected (FC) layer for classification, employing cross-entropy loss as the classification loss function.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>The improved MMAL-Net in the training phase.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g002.tif"/>
</fig>
<p>ResNet-50 is a CNN architecture that belongs to the ResNet family. The ResNet family of architectures was specifically developed to address the problem of vanishing gradients in deep neural networks. In ResNet-50, the numerical suffix &#x201C;50&#x201D; indicates that the network consists of a total of 50 layers, including convolutional layers, pooling layers, fully connected layers, and shortcut connections. The key innovation of ResNet lies in the introduction of residual or skip connections, which allow information to bypass certain layers. This enables the network to learn more effectively by facilitating the propagation of gradients during training and enabling the acquisition of deeper and more complex representations. These skip connections also mitigate the problem of degradation, wherein the accuracy of a deep network decreases as its depth increases, by facilitating the training of deeper networks. ResNet-50 has been widely utilized and has achieved significant success in various machine vision tasks, such as image classification, object detection, and image segmentation. It has proven to be a powerful architecture that has advanced the field of computer vision and deep learning.</p>
<p>Formulas 1, 2, and 3 represent the loss function of the three branches, respectively.</p>
<disp-formula id="E1">
<label>(1)</label>
<mml:math id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math></disp-formula>
<disp-formula id="E2">
<label>(2)</label>
<mml:math id="M2">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math></disp-formula>
<disp-formula id="E3">
<label>(3)</label>
<mml:math id="M3">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math></disp-formula>
<p>Where <inline-formula>
<mml:math id="M4">
<mml:mi>c</mml:mi>
</mml:math>
</inline-formula> represents the ground truth label of the input image, while <inline-formula>
<mml:math id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denote the category probabilities obtained from the last softmax layer outputs of the frontal and side branches, respectively. <inline-formula>
<mml:math id="M7">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> refers to the output of the softmax layer in the part branch that corresponds to the <inline-formula>
<mml:math id="M8">
<mml:mi>n</mml:mi>
</mml:math>
</inline-formula>th part image. <inline-formula>
<mml:math id="M9">
<mml:mi>N</mml:mi>
</mml:math>
</inline-formula> represents the total count of part images.</p>
<p>The total loss is defined as Formula 4:</p>
<disp-formula id="E4">
<label>(4)</label>
<mml:math id="M10">
<mml:mrow>
<mml:mtable equalrows="true" equalcolumns="true">
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math></disp-formula>
<p>The total loss is calculated as the cumulative sum of losses from the three branches, collaborating to enhance the model&#x2019;s performance during backpropagation. This enables the final converged model to generate classification predictions by considering both the global structural attributes of the object and its detailed features. During the testing phase, the part branch was excluded to minimize computational complexity, ensuring efficient prediction times for practical applications of our method.</p>
</sec>
<sec id="sec7">
<label>3.2.</label>
<title>APPM</title>
<p>By analyzing the activation map <inline-formula>
<mml:math id="M11">
<mml:mi>A</mml:mi>
</mml:math>
</inline-formula>, we observed that areas with high activation values corresponded to key parts, such as the front area of the truck. To identify these informative regions, we adopted a sliding window approach inspired by object detection techniques. This approach allowed us to extract part images from windows containing relevant information. Additionally, we employed a modified version of the traditional sliding window method using a fully convolutional network, similar to the approach used in Overfeat (<xref ref-type="bibr" rid="ref18">Sermanet et al., 2013</xref>). This method involved obtaining feature maps for different windows from the output feature map of the previous network branch. Subsequently, we aggregated the activation maps <inline-formula>
<mml:math id="M12">
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of each window along the channel dimension and computed their mean activation value <inline-formula>
<mml:math id="M13">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>a</mml:mi>
<mml:mo>&#x00AF;</mml:mo>
</mml:mover>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, as described in Formula 5. Here, <inline-formula>
<mml:math id="M14">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M15">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denote the height and width of a window&#x2019;s feature map, respectively. We then ranked the windows based on their <inline-formula>
<mml:math id="M16">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>a</mml:mi>
<mml:mo>&#x00AF;</mml:mo>
</mml:mover>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> values, with higher values indicating more informative regions, as illustrated in <xref rid="fig3" ref-type="fig">Figure 3</xref>.</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>The simple pipeline of the APPM. We use red, orange, yellow and green colors to indicate the order of windows&#x2019; <inline-formula>
<mml:math id="M19">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>a</mml:mi>
<mml:mo>&#x00AF;</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula></p>
</caption>
<graphic xlink:href="fnins-17-1243847-g003.tif"/>
</fig>
<disp-formula id="E5">
<label>(5)</label>
<mml:math id="M17">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mo>&#x00AF;</mml:mo>
</mml:mover>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="normal">x</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="normal">y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">x,y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
<mml:mo>&#x00D7;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mi mathvariant="normal">w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math></disp-formula>
<p>However, we cannot directly select the initial windows because they are often adjacent to the windows with the highest average activation values <inline-formula>
<mml:math id="M18">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>a</mml:mi>
<mml:mo>&#x00AF;</mml:mo>
</mml:mover>
<mml:mi>w</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and contain nearly identical parts. Nonetheless, our objective is to choose a diverse range of parts. To minimize redundancy in the regions, we employ Non-Maximum Suppression (NMS) to select a fixed number of windows as part images at different scales. The visualization of the module&#x2019;s output in <xref rid="fig4" ref-type="fig">Figure 4</xref> demonstrates that the proposed method effectively identifies distinct part regions with varying levels of importance. We utilize red, orange, yellow, and green rectangles to highlight the regions proposed by APPM that have the highest average activation values at various scales, with the red rectangle indicating the highest value. <xref rid="fig4" ref-type="fig">Figure 4</xref> illustrates that the proposed approach captures detailed information and exhibits a more logical ordering at the same scale, thus significantly enhancing the model&#x2019;s robustness to scale variations. Notably, the head region stands out as the most discriminative region for truck recognition.</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption>
<p>Visualization of part regions.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g004.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="results" id="sec8">
<label>4.</label>
<title>Results and discussion</title>
<p>To validate the advantages of the enhanced MMAL-Net, we conducted an evaluation of our method on the well-established and competitive benchmark dataset, Stanford Cars (<xref ref-type="bibr" rid="ref14">Krause et al., 2013</xref>).</p>
<p>In our experiments, we adopted a consistent preprocessing approach. Initially, we resized the images to dimensions of 512&#x2009;&#x00D7;&#x2009;512, serving as inputs for both the frontal and side branches. Additionally, all part images were uniformly resized to 256&#x2009;&#x00D7;&#x2009;256 for the part branch. To ensure efficient initialization, we pre-trained ResNet-50 on the widely used ImageNet dataset, allowing us to effectively obtain the network&#x2019;s initial weights. Throughout both the training and testing phases, we exclusively relied on image-level labels, refraining from employing any additional annotations. Our optimization process involved utilizing SGD with specific hyperparameters: a momentum value of 0.9 and a weight decay of 0.0001. To enhance training efficiency, we employed a mini-batch size of 6, utilizing a Tesla P100 GPU for computation. For fine-tuning the learning process, we set the initial learning rate to 0.001, which we later scaled down by a factor of 0.1 after 60 epochs. This step was instrumental in facilitating smoother convergence during training. We utilized PyTorch as the foundational framework.</p>
<p>In the experiments, we compared the proposed method to several baseline approaches and achieved competitive results, as shown in <xref rid="tab1" ref-type="table">Table 1</xref>. By comparison, we can observe that our method attains the highest accuracy 95.03%.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Comparison of different methods on the Stanford Cars dataset.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Methods</th>
<th align="center" valign="top">Backbone</th>
<th align="center" valign="top">Source</th>
<th align="center" valign="top">Accuracy (%)</th></tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">RA-CNN <xref ref-type="bibr" rid="ref8">Fu et al. (2017)</xref></td>
<td align="center" valign="top">VGGNet-19</td>
<td align="center" valign="top">CVPR&#x2019;2017</td>
<td align="left" valign="middle">92.5</td></tr>
<tr>
<td align="left" valign="middle">MA-CNN <xref ref-type="bibr" rid="ref34">Zheng et al. (2017)</xref></td>
<td align="center" valign="top">VGGNet-19</td>
<td align="center" valign="top">ICCV&#x2019;2017</td>
<td align="left" valign="middle">92.8</td></tr>
<tr>
<td align="left" valign="middle">NTS-Net <xref ref-type="bibr" rid="ref30">Yang et al. (2018)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">ECCV&#x2019;2018</td>
<td align="left" valign="middle">93.9</td></tr>
<tr>
<td align="left" valign="middle">MAMC <xref ref-type="bibr" rid="ref22">Sun et al. (2018)</xref></td>
<td align="center" valign="top">ResNet-101</td>
<td align="center" valign="top">ECCV&#x2019;2018</td>
<td align="left" valign="middle">93.0</td></tr>
<tr>
<td align="left" valign="middle">TASN <xref ref-type="bibr" rid="ref35">Zheng et al. (2019)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">CVPR&#x2019;2019</td>
<td align="left" valign="middle">93.8</td></tr>
<tr>
<td align="left" valign="middle">DCL <xref ref-type="bibr" rid="ref3">Chen et al. (2019)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">CVPR&#x2019;2019</td>
<td align="left" valign="middle">94.5</td></tr>
<tr>
<td align="left" valign="middle">API-Net <xref ref-type="bibr" rid="ref37">Zhuang et al. (2020)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">AAAI&#x2019;2020</td>
<td align="left" valign="middle">94.8</td></tr>
<tr>
<td align="left" valign="middle">DP-Net <xref ref-type="bibr" rid="ref25">Wang et al. (2021)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">AAAI&#x2019;2021</td>
<td align="left" valign="middle">94.8</td></tr>
<tr>
<td align="left" valign="middle">SAM <xref ref-type="bibr" rid="ref19">Shu et al. (2022)</xref></td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">ECCV&#x2019;2022</td>
<td align="left" valign="middle">94.18</td></tr>
<tr>
<td align="left" valign="middle">MSHQP <xref ref-type="bibr" rid="ref23">Tan et al. (2022)</xref></td>
<td align="center" valign="top">ResNet-152</td>
<td align="center" valign="top">TOMM&#x2019;2022</td>
<td align="left" valign="middle">94.9</td></tr>
<tr>
<td align="left" valign="middle">The Improved MMAL-Net</td>
<td align="center" valign="top">ResNet-50</td>
<td align="center" valign="top">This paper</td>
<td align="left" valign="middle">95.03</td></tr>
</tbody>
</table>
</table-wrap>
<p>In practice, there is a continuous emergence of new truck models. Given their recent introduction, it becomes challenging to obtain an adequate number of instances for constructing a comprehensive dataset. Hence, we utilized a customized dataset on a smaller scale to validate the applicability and effectiveness of our method. The personalized truck dataset includes four truck models: FAW J7, Shaanxi Delong X3000, Dongfeng Dorica D6, and JAC Junling V6. FAW J7 and Shaanxi Delong X3000 are heavy-duty trucks, whereas Dongfeng Dorica D6 and JAC Junling V6 are light-duty trucks. Each truck category is composed of 50 sets of training images and 20 sets of test images and each set comprises one frontal image and one side image. An example is depicted in <xref rid="fig5" ref-type="fig">Figure 5</xref>.</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption>
<p>The personalized truck dataset. <bold>(A&#x2013;D)</bold> represent FAW J7, Shaanxi Delong X3000, Dongfeng Dorica D6, and JAC Junling V6.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g005.tif"/>
</fig>
<p>Thereafter, the overall network structure with specific features was fine-tuned to achieve fine-grained recognition of multiple target models. Moreover, we evaluated the effect of varied training samples on the recognition performance of the intelligent identification model by testing the same dataset using different incremental levels of training data. This simulation emulated the impact of increasing the number of target truck images collected in actual scenarios on the enhancement of the recognition model&#x2019;s performance.</p>
<p>In our experiment, we selected sets of 20, 30, 40, and 50 images for each classifier as training datasets and used the same number of test set to compare the performance of API-Net, DP-Net, MSHQP and the improved MMAL-Net. It is worth mentioning that API-Net, DP-Net, and MSHQP were the top three performing methods in our experiments on the Stanford Cars dataset, excluding our proposed method. The results indicate that as the training data increases, the network&#x2019;s ability to identify and extract features from target trucks gradually improves, suggesting that larger datasets can effectively enhance the model&#x2019;s capability to extract potential features. The improved MMAL-Net exhibits comparable or superior performance to other methods across all numbers of training sets, demonstrating its superior ability to extract fine-grained features of target trucks (see <xref rid="fig6" ref-type="fig">Figure 6</xref>).</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption>
<p>The comparison result on different number of training sets.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g006.tif"/>
</fig>
<p>In our small-scale custom dataset, it is evident that the recognition accuracy reaches 85% when the training set consists of 20 sets of photos. This greatly addresses the practical issue of scarce images of a particular type of truck. The improved MMAL-Net demonstrated remarkable resilience to image quality and scene noise, as evidenced by its recognition accuracy of 100% when trained on a dataset comprising 50 sets of samples. This noteworthy achievement further supports the superior performance of the enhanced network.</p>
<p>Confusion matrices in <xref rid="fig7" ref-type="fig">Figure 7</xref> illustrate the test results of the improved MMAL-Net. At a training set size of 40 sets of images, the improved MMAL-Net had a single misclassification on the test set, misclassifying a Dongfeng Dorica D6 as a JAC Junling V6. However, at a training set size of 50 sets, all classifications were accurate. The improved MMAL-Net accurately classified heavy-duty trucks, avoiding misclassification as light-duty trucks. Similarly, it correctly identified light-duty trucks without misclassification as heavy-duty trucks. This is crucial because misidentifying an overloaded light-duty truck as a heavy-duty truck can result in undetected overweight issues, thus posing safety concerns.</p>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption>
<p>Confusion matrices of the improved MMAL-Net. <bold>(A&#x2013;D)</bold> represent 20, 30, 40 and 50 sets of images, respectively, which are used as the training dataset.</p>
</caption>
<graphic xlink:href="fnins-17-1243847-g007.tif"/>
</fig>
</sec>
<sec sec-type="conclusions" id="sec9">
<label>5.</label>
<title>Conclusion</title>
<p>This paper introduces a method for precise identification of truck models. In our experimental evaluation, this method achieved an accuracy of 95.03% on the competitive benchmark dataset, Stanford Cars. Furthermore, it achieved an accuracy of 100% on our custom truck dataset. When integrated with weighing and license plates systems, it can be applied in highway automatic weighing stations to determine if a truck is overloaded. By providing accurate truck information, this method contributes to freight management, transportation safety, and highway planning, thereby fostering the development of the logistics industry and improving traffic safety. However, the accuracy of truck model recognition may decrease in real-world scenarios due to the reduced data quality. Consequently, future research will focus on addressing this issue, with specific emphasis on long-distance shooting conditions.</p>
</sec>
<sec sec-type="data-availability" id="sec10">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="sec11">
<title>Author contributions</title>
<p>JiaS contributed to conception and design of the study. JiaS and JinS organized the methodology. JiaS and ZY performed the statistical analysis. JiaS wrote the first draft of the manuscript. JiaS, ZG, and YS contributed to the visualization of the results. JiaS and LL contributed to the supervision of the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="sec12">
<title>Funding</title>
<p>This research was supported by National Key R&#x0026;D Program of China (Grant no. 2021YFB3300503).</p>
</sec>
<sec sec-type="COI-statement" id="sec13">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Biglari</surname> <given-names>M.</given-names></name> <name><surname>Soleimani</surname> <given-names>A.</given-names></name> <name><surname>Hassanpour</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>A cascading scheme for speeding up multiple classifier systems</article-title>. <source>Pattern. Anal. Applic.</source> <volume>22</volume>, <fpage>375</fpage>&#x2013;<lpage>387</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10044-017-0637-4</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Branson</surname> <given-names>S.</given-names></name> <name><surname>Van Horn</surname> <given-names>G.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>Bird species categorization using pose normalized deep convolutional nets</article-title>. <source>arXiv preprint arXiv</source> <volume>2952</volume>:<fpage>1406</fpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1406.2952</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Bai</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Mei</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Destruction and construction learning for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name> (pp. <fpage>5157</fpage>&#x2013;<lpage>5166</lpage>).</citation></ref>
<ref id="ref4"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Dalal</surname> <given-names>N.</given-names></name> <name><surname>Triggs</surname> <given-names>B.</given-names></name></person-group> (<year>2005</year>). <article-title>Histograms of oriented gradients for human detection</article-title>. <conf-name>In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR&#x2019;05)</conf-name> (Vol. <volume>1</volume>, pp. <fpage>886</fpage>&#x2013;<lpage>893</lpage>). <publisher-name>Ieee</publisher-name>.</citation></ref>
<ref id="ref5"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Hoffman</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>N.</given-names></name> <name><surname>Tzeng</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2014</year>). &#x201C;<article-title>Decaf: a deep convolutional activation feature for generic visual recognition</article-title>&#x201D; in <source>International conference on machine learning</source> (<publisher-loc>Amsterdam</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>647</fpage>&#x2013;<lpage>655</lpage>.</citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Du</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Fine-grained vehicle model recognition using a coarse-to-fine convolutional neural network architecture</article-title>. <source>IEEE Trans. Intell. Transp. Syst.</source> <volume>18</volume>, <fpage>1782</fpage>&#x2013;<lpage>1792</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2016.2620495</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Fraz</surname> <given-names>M.</given-names></name> <name><surname>Edirisinghe</surname> <given-names>E. A.</given-names></name> <name><surname>Sarfraz</surname> <given-names>M. S.</given-names></name></person-group> (<year>2014</year>). <article-title>Mid-level-representation based lexicon for vehicle make and model recognition</article-title>. <conf-name>In 2014 22nd International Conference on Pattern Recognition</conf-name> (pp. <fpage>393</fpage>&#x2013;<lpage>398</lpage>). <publisher-name>IEEE</publisher-name>.</citation></ref>
<ref id="ref8"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>J.</given-names></name> <name><surname>Zheng</surname> <given-names>H.</given-names></name> <name><surname>Mei</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Look closer to see better: recurrent attention convolutional neural network for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name> (pp. <fpage>4438</fpage>&#x2013;<lpage>4446</lpage>).</citation></ref>
<ref id="ref9"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Malik</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>. <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name> (pp. <fpage>580</fpage>&#x2013;<lpage>587</lpage>).</citation></ref>
<ref id="ref10"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name> (pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>).</citation></ref>
<ref id="ref11"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Sedra</surname> <given-names>D.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep networks with stochastic depth</article-title>. <conf-name>In Computer Vision&#x2013;ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11&#x2013;14, 2016, Proceedings, Part IV</conf-name> <volume>14</volume> (pp. <fpage>646</fpage>&#x2013;<lpage>661</lpage>). <publisher-name>Springer International Publishing</publisher-name>. <publisher-loc>United States</publisher-loc></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Ding</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>Vehicle logo recognition system based on convolutional neural networks with a pretraining strategy</article-title>. <source>IEEE Trans. Intell. Transp. Syst.</source> <volume>16</volume>, <fpage>1951</fpage>&#x2013;<lpage>1960</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2014.2387069</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>J&#x00E9;gou</surname> <given-names>H.</given-names></name> <name><surname>Douze</surname> <given-names>M.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name> <name><surname>P&#x00E9;rez</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Aggregating local descriptors into a compact image representation</article-title>. <conf-name>In 2010 IEEE computer society conference on computer vision and pattern recognition</conf-name> (pp. <fpage>3304</fpage>&#x2013;<lpage>3311</lpage>). <publisher-name>IEEE</publisher-name>.</citation></ref>
<ref id="ref14"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Stark</surname> <given-names>M.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>3d object representations for fine-grained categorization</article-title>. <conf-name>In Proceedings of the IEEE international conference on computer vision workshops</conf-name> (pp. <fpage>554</fpage>&#x2013;<lpage>561</lpage>).</citation></ref>
<ref id="ref15"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>T. Y.</given-names></name> <name><surname>Roy Chowdhury</surname> <given-names>A.</given-names></name> <name><surname>Maji</surname> <given-names>S</given-names></name></person-group>. (<year>2005</year>) <article-title>Bilinear CNN models for fine-grained visual recognition</article-title>. <conf-name>In Proceedings of the 15th IEEE International Conference on Computer Vision (ICCV)</conf-name>. <publisher-name>Santiago, Chile</publisher-name>: <publisher-loc>IEEE</publisher-loc>, <fpage>1449</fpage>&#x2013;<lpage>1457</lpage>.</citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mo</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>Research on the method of determining highway truck load limit based on image processing</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>205477</fpage>&#x2013;<lpage>205486</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3037195</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Psyllos</surname> <given-names>A. P.</given-names></name> <name><surname>Anagnostopoulos</surname> <given-names>C. N. E.</given-names></name> <name><surname>Kayafas</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>Vehicle logo recognition using a sift-based enhanced matching scheme</article-title>. <source>IEEE Trans. Intell. Transp. Syst.</source> <volume>11</volume>, <fpage>322</fpage>&#x2013;<lpage>328</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2010.2042714</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Eigen</surname> <given-names>D.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Mathieu</surname> <given-names>M.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Overfeat: integrated recognition, localization and detection using convolutional networks</article-title>. <source>arXiv preprint arXiv</source> <volume>1312</volume>:<fpage>6229</fpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1312.6229</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Shu</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>Improving fine-grained visual recognition in low data regimes via self-boosting attention mechanism</article-title>. <conf-name>In European Conference on Computer Vision</conf-name> (pp. <fpage>449</fpage>&#x2013;<lpage>465</lpage>). <publisher-name>Cham</publisher-name>: <publisher-loc>Springer Nature Switzerland</publisher-loc>.</citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siddiqui</surname> <given-names>A. J.</given-names></name> <name><surname>Mammeri</surname> <given-names>A.</given-names></name> <name><surname>Boukerche</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Real-time vehicle make and model recognition based on a bag of SURF features</article-title>. <source>IEEE Trans. Intell. Transp. Syst.</source> <volume>17</volume>, <fpage>3205</fpage>&#x2013;<lpage>3219</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2016.2545640</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>On the eigenvalues of global covariance pooling for fine-grained visual recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>45</volume>, <fpage>1</fpage>&#x2013;<lpage>3566</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2022.3178802</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>M.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>F.</given-names></name> <name><surname>Ding</surname> <given-names>E.</given-names></name></person-group> (<year>2018</year>). <article-title>Multi-attention multi-class constraint for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the European conference on computer vision (ECCV)</conf-name> (pp. <fpage>805</fpage>&#x2013;<lpage>821</lpage>).</citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>M.</given-names></name> <name><surname>Yuan</surname> <given-names>F.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Gu</surname> <given-names>X.</given-names></name></person-group> (<year>2022</year>). <article-title>Fine-grained image classification via multi-scale selective hierarchical biquadratic pooling</article-title>. <source>ACM Trans. Multimedia Comp, Commun, Appl (TOMM)</source> <volume>18</volume>, <fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3492221</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Fu</surname> <given-names>H.</given-names></name> <name><surname>Ma</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Global structure graph guided fine-grained vehicle recognition</article-title>. <conf-name>In ICASSP 2020&#x2013;2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name> (pp. <fpage>1913</fpage>&#x2013;<lpage>1917</lpage>). <publisher-name>IEEE</publisher-name>.</citation></ref>
<ref id="ref25"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Ouyang</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>Dynamic position-aware network for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the AAAI Conference on Artificial Intelligence</conf-name> (Vol. <volume>35</volume>. <fpage>2791</fpage>&#x2013;<lpage>2799</lpage>).</citation></ref>
<ref id="ref26"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>P.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>G2DeNet: global Gaussian distribution embedding network and its application to visual recognition</article-title>. <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name> (pp. <fpage>2730</fpage>&#x2013;<lpage>2739</lpage>).</citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>X. S.</given-names></name> <name><surname>Xie</surname> <given-names>C. W.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>Mask-CNN: localizing parts and selecting descriptors for fine-grained bird species categorization</article-title>. <source>Pattern Recogn.</source> <volume>76</volume>, <fpage>704</fpage>&#x2013;<lpage>714</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2017.10.002</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>G. S.</given-names></name> <name><surname>Zhang</surname> <given-names>X. Y.</given-names></name> <name><surname>Yang</surname> <given-names>W.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Yan</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>C. L.</given-names></name></person-group> (<year>2017</year>). <article-title>LG-CNN: from local parts to global discrimination for fine-grained recognition</article-title>. <source>Pattern Recogn.</source> <volume>71</volume>, <fpage>118</fpage>&#x2013;<lpage>131</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2017.06.002</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Luo</surname> <given-names>P.</given-names></name> <name><surname>Change Loy</surname> <given-names>C.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>A large-scale car dataset for fine-grained categorization and verification</article-title>. <conf-name>In Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name> (pp. <fpage>3973</fpage>&#x2013;<lpage>3981</lpage>).</citation></ref>
<ref id="ref30"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Hu</surname> <given-names>Z.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>Learning to navigate for fine-grained classification</article-title>. <conf-name>In Proceedings of the European conference on computer vision (ECCV)</conf-name> (pp. <fpage>420</fpage>&#x2013;<lpage>435</lpage>).</citation></ref>
<ref id="ref31"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>N.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Part-based R-CNNs for fine-grained category detection</article-title>. <conf-name>In Computer Vision&#x2013;ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6&#x2013;12, 2014, Proceedings, Part I</conf-name> <volume>13</volume> (pp. <fpage>834</fpage>&#x2013;<lpage>849</lpage>). <publisher-name>Springer International Publishing</publisher-name>.</citation></ref>
<ref id="ref32"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-branch and multi-scale attention learning for fine-grained visual categorization</article-title>. <conf-name>In Multi Media Modeling: 27th International Conference, MMM 2021, Prague, Czech Republic, June 22&#x2013;24, 2021, Proceedings, Part I</conf-name> <volume>27</volume> (pp. <fpage>136</fpage>&#x2013;<lpage>147</lpage>). <publisher-name>Springer International Publishing</publisher-name>.</citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wei</surname> <given-names>X. S.</given-names></name> <name><surname>Wu</surname> <given-names>J. X.</given-names></name> <name><surname>Cai</surname> <given-names>J. F.</given-names></name> <name><surname>Lu</surname> <given-names>J. B.</given-names></name> <name><surname>Nguyen</surname> <given-names>V. A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Weakly supervised fine-grained categorization with part-based image representation</article-title>. <source>IEEE Trans. Image Process.</source> <volume>25</volume>, <fpage>1713</fpage>&#x2013;<lpage>1725</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2016.2531289</pub-id>, PMID: <pub-id pub-id-type="pmid">26890872</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>H.</given-names></name> <name><surname>Fu</surname> <given-names>J.</given-names></name> <name><surname>Mei</surname> <given-names>T.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Learning multi-attention convolutional neural network for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the IEEE international conference on computer vision</conf-name> (pp. <fpage>5209</fpage>&#x2013;<lpage>5217</lpage>).</citation></ref>
<ref id="ref35"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>H.</given-names></name> <name><surname>Fu</surname> <given-names>J.</given-names></name> <name><surname>Zha</surname> <given-names>Z. J.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Looking for the devil in the details: learning trilinear attention sampling network for fine-grained image recognition</article-title>. <conf-name>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name> (pp. <fpage>5012</fpage>&#x2013;<lpage>5021</lpage>).</citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Ou</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>A new kind of high durable traffic weighbridge based on fbg sensors</article-title>. <source>Proceedings of SPIE - The Int. Soc. Optical Eng.</source> <volume>5855</volume>, <fpage>735</fpage>&#x2013;<lpage>738</lpage>. doi: <pub-id pub-id-type="doi">10.1117/12.623436</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zhuang</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Qiao</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Learning attentive pairwise interaction for fine-grained classification</article-title>. <conf-name>In Proceedings of the AAAI conference on artificial intelligence</conf-name>. <volume>34</volume>, <fpage>13130</fpage>&#x2013;<lpage>13137</lpage>).</citation></ref>
</ref-list>
</back>
</article>