<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Imaging.</journal-id>
<journal-title>Frontiers in Imaging</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Imaging.</abbrev-journal-title>
<issn pub-type="epub">2813-3315</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fimag.2023.1271885</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Imaging</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Height reverse perspective transformation for crowd counting</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhao</surname> <given-names>Xiaomei</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2278113/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Honggang</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/2414457/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhao</surname> <given-names>Zhan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/2541346/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Shuo</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/2541531/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Shandong Key Laboratory of Intelligent Buildings Technology, School of Information and Electrical Engineering, Shandong Jianzhu University</institution>, <addr-line>Jinan</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Hong Fu, The Education University of Hong Kong, Hong Kong SAR, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jiyong Wang, Shanghai Pulse Medical Technology, Inc., China; Wanli Xie, Qufu Normal University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Xiaomei Zhao <email>zhaoxiaomei20&#x00040;sdjzu.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>2</volume>
<elocation-id>1271885</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>08</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Zhao, Li, Zhao and Li.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Zhao, Li, Zhao and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license></permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Crowd counting plays a critical role in the intelligent video surveillance of public areas. A significant challenge to this task is the perspective effect on human heads, which causes serious scale variations. Height reverse perspective transformation (HRPT) alleviates this problem by narrowing the height gap among human heads.</p></sec>
<sec>
<title>Methods</title>
<p>It employs depth maps to calculate the rescaling factors of image rows, and then it performs image transformation accordingly. HRPT enlarges small human heads in far areas to make them more noticeable and shrinks large human heads in closer areas to reduce redundant information. Then, convolutional neural networks can be used for crowd counting. Previous crowd-counting methods mainly solve the scale variation problem by designing specific networks, such as multi-scale or perspective-aware networks. These networks cannot be conveniently employed by other methods. In contrast, HRPT solves the scale variation problem through image transformation. It can be used as a preprocessing step and easily employed by other crowd-counting methods without changing their original structures.</p></sec>
<sec>
<title>Results and discussion</title>
<p>Experimental results show that HRPT successfully narrows the height gap among human heads and achieves state-of-the-art performance on a large crowd-counting RGB-D dataset.</p></sec></abstract>
<kwd-group>
<kwd>crowd counting</kwd>
<kwd>scale variation problem</kwd>
<kwd>perspective effect</kwd>
<kwd>height reverse perspective transformation</kwd>
<kwd>RGB-D image</kwd>
</kwd-group>
<contract-sponsor id="cn001">Natural Science Foundation of Shandong Province<named-content content-type="fundref-id">10.13039/501100007129</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="5"/>
<equation-count count="14"/>
<ref-count count="47"/>
<page-count count="15"/>
<word-count count="9531"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Image Retrieval</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Crowd-counting technology estimates the number of people in images, which is instrumental in intelligent video surveillance. It plays a vital role in safeguarding public safety by recognizing abnormal crowd gatherings automatically.</p>
<p>Over the past decades, countless researchers have focused on improving the accuracy of crowd-counting techniques (Sindagi and Patel, <xref ref-type="bibr" rid="B29">2018</xref>; Gao et al., <xref ref-type="bibr" rid="B3">2020</xref>; Fan et al., <xref ref-type="bibr" rid="B2">2022</xref>). One standard method is to count the number of human heads captured by surveillance cameras, since occlusions on human heads are less severe than the other parts of human bodies. However, the perspective effect can cause significant scale variations in human head size, posing a critical challenge to accurate counting. This study proposes a novel height reverse perspective transformation (HRPT) method to alleviate the scale variation problem. This technique creatively narrows the height gap among human heads, particularly enlarging small human heads in far areas and shrinking large human heads in closer areas. <xref ref-type="fig" rid="F1">Figure 1</xref> displays a group of crowd RGB images to demonstrate the effect of HRPT.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>A group of crowd RGB images to show the effect of HRPT: <bold>(A)</bold> original crowd RGB image; <bold>(B)</bold> crowd RGB image transformed by HRPT. The heads of three pedestrians are marked in <bold>(A, B)</bold>. Their head heights, measured in pixels, are shown on the right of the images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0001.tif"/>
</fig>
<p>Existing crowd-counting methods are categorized into detection-based, regression-based, and density map-based methods. Detection-based methods generally count people by detecting people or their heads (Idrees et al., <xref ref-type="bibr" rid="B8">2015</xref>; Stewart et al., <xref ref-type="bibr" rid="B30">2016</xref>; Liu Y. et al., <xref ref-type="bibr" rid="B23">2019</xref>), while regression-based methods extract image features from the whole crowd image and regress the number of people according to these features (Liu and Vasconcelos, <xref ref-type="bibr" rid="B18">2015</xref>; Wang et al., <xref ref-type="bibr" rid="B33">2015</xref>; Shang et al., <xref ref-type="bibr" rid="B27">2016</xref>). Density map-based methods count people by estimating the density map and summing the density over the whole image (Ma et al., <xref ref-type="bibr" rid="B24">2022</xref>; Wang et al., <xref ref-type="bibr" rid="B36">2022</xref>, <xref ref-type="bibr" rid="B34">2023</xref>; Yan et al., <xref ref-type="bibr" rid="B40">2022</xref>). Detection-based methods usually perform poorly when dealing with dense crowds far from the camera (Liu et al., <xref ref-type="bibr" rid="B19">2018</xref>; Xu et al., <xref ref-type="bibr" rid="B39">2019</xref>; Fan et al., <xref ref-type="bibr" rid="B2">2022</xref>), and regression-based methods ignore spatial information (Gao et al., <xref ref-type="bibr" rid="B3">2020</xref>). Therefore, density map-based methods are more popular than the other two types of methods.</p>
<p>Many density map-based methods solve the scale variation problem using multi-scale networks (Liu W. et al., <xref ref-type="bibr" rid="B21">2019</xref>; Ma et al., <xref ref-type="bibr" rid="B24">2022</xref>; Wang et al., <xref ref-type="bibr" rid="B36">2022</xref>, <xref ref-type="bibr" rid="B34">2023</xref>). However, these networks only consider a finite number of discrete scales, limiting their ability to handle scale variations (Yan et al., <xref ref-type="bibr" rid="B40">2022</xref>). Therefore, many researchers have focused on solving the scale variation problem by utilizing perspective-aware approaches (Yan et al., <xref ref-type="bibr" rid="B40">2022</xref>; Zhang and Li, <xref ref-type="bibr" rid="B43">2022</xref>). Perspective-aware approaches generally extract perspective information from RGB images (Yan et al., <xref ref-type="bibr" rid="B40">2022</xref>; Zhang and Li, <xref ref-type="bibr" rid="B43">2022</xref>). In contrast, a depth map can be directly used as perspective information, and it is more accurate than the perspective information extracted from an RGB image. Using a depth map to its maximum capacity can alleviate the scale variation problem caused by the perspective effect.</p>
<p>RGB-D cameras capture both RGB images and depth maps, which contain complementary information. Applying multiple complementary information is popular in many areas, such as automatic malfunction detection (Jing et al., <xref ref-type="bibr" rid="B11">2017</xref>) and automatic driving (Wang et al., <xref ref-type="bibr" rid="B37">2019</xref>; Zhuang et al., <xref ref-type="bibr" rid="B47">2021</xref>). This is because using multiple types of complementary information can improve the robustness and accuracy of automatic systems. At present, although the RGB-D camera is less popular than the RGB camera owing to its higher cost, as the economy develops, the RGB-D camera is expected to be used more widely.</p>
<p>This study proposes the HRPT method, developed based on RGB-D images, to solve the scale variation problem in the crowd-counting task. By narrowing the height gap among human heads according to the perspective information in the depth map, HRPT alleviates the scale variation problem, reduces redundant information in near areas, and makes small human heads in faraway areas more visible. As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, HRPT successfully narrows the height gap among human heads. It enlarges small human heads in far areas, making them more visible, and it shrinks large human heads in closer areas to reduce redundant information, making the counting network focus more attention on remote areas where the human heads are denser and harder to detect.</p>
<p>Previously developed crowd-counting methods often use specialized networks to tackle scale variation, employing multi-scale or perspective-aware strategies. Other methods must change their original network structures to employ these strategies. In contrast, the proposed HRPT can be used as a preprocessing step and easily integrated into existing methods without modifying their original network structures. This study combines HRPT with four well-performing crowd-counting methods: CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>), DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>), and CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>) to produce promising results.</p>
<p>In addition to narrowing height gap among human heads by HRPT, we also try to narrow the width gap among human heads. However, narrowing the width gap does not achieve promising results. Section 5.2.1 details the discussion. Furthermore, if we ignore the human height, the perspective effect on the 2D ground can be eliminated using a homograph (Hartley and Zisserman, <xref ref-type="bibr" rid="B5">2003</xref>). However, experimental results show that even though the homograph successfully eliminates the perspective effect on the 2D ground, it cannot achieve satisfactory results in alleviating the scale variations of human heads. Section 5.2.2 presents a more detailed discussion.</p>
<p>In summary, our main contributions are given below.</p>
<list list-type="order">
<list-item><p>HRPT is a novel technique to alleviate the scale variation problem in the crowd-counting task. It uses the perspective information in depth maps and creatively narrows the height gap among human heads via image transformation. After HRPT, small human heads in outlying areas are enlarged and become more apparent, and large human heads in closer areas are shrunken to reduce redundant information.</p></list-item>
<list-item><p>HRPT is integrated with the following well-performing crowd-counting methods: CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>), DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>), and CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>), and experimental results demonstrate that HRPT successfully improves crowd-counting performance. Our method (GL&#x0002B;HRPT) achieves state-of-the-art results and outperforms other methods that employ depth maps.</p></list-item>
<list-item><p>Another two image transformation methods (i.e. a shape-changing method that narrows the width gap among human heads and a geometric method that attempts to eliminate the perspective effect) are also experimented to solve the scale variation problem. Experimental results show that these two methods have poorer performance than the proposed HRPT.</p></list-item>
</list></sec>
<sec id="s2">
<title>2. Related work</title>
<p>The proposed method aims to alleviate the scale variation problem in crowd-counting tasks by making use of depth maps and image transformation. This problem has traditionally been tackled by employing multi-scale or perspective-aware crowd-counting methods. The following section first introduces these two types of previous methods, and then presents other related crowd-counting methods that also employ depth maps and image transformation.</p>
<sec>
<title>2.1. Multi-scale crowd-counting methods</title>
<p>Multi-scale crowd-counting methods solve the scale variation problem by employing multi-scale network structures. Wang et al. (<xref ref-type="bibr" rid="B36">2022</xref>) and Wang et al. (<xref ref-type="bibr" rid="B34">2023</xref>) built multi-scale networks by employing multiple branches with different convolutional dilation rates. Liu W. et al. (<xref ref-type="bibr" rid="B21">2019</xref>) built a multi-scale network by using multiple branches with different pooling sizes. Jiang et al. (<xref ref-type="bibr" rid="B9">2019</xref>) and Ma et al. (<xref ref-type="bibr" rid="B24">2022</xref>) built multi-scale networks by combining image features of multiple layers. Jiang et al. (<xref ref-type="bibr" rid="B10">2020</xref>) and Du et al. (<xref ref-type="bibr" rid="B1">2023</xref>) built multi-scale networks by combining the estimated crowd density maps of multiple scales. However, multi-scale methods only consider a finite number of discrete scales, limiting their potential to solve the scale variation problem (Yan et al., <xref ref-type="bibr" rid="B40">2022</xref>).</p></sec>
<sec>
<title>2.2. Perspective-aware crowd-counting methods</title>
<p>Many impressive perspective-aware methods have been proposed. For example, Zhang et al. (<xref ref-type="bibr" rid="B42">2015</xref>) estimated the number of pixels representing one meter and used this perspective information to normalize the density map. Yan et al. (<xref ref-type="bibr" rid="B40">2022</xref>) estimated the same perspective information as Zhang et al. (<xref ref-type="bibr" rid="B42">2015</xref>) and used it to select different dilation kernels. Zhang and Li (<xref ref-type="bibr" rid="B43">2022</xref>) embedded perspective information into a point-supervised network to better handle the scaling problem. Wan et al. (<xref ref-type="bibr" rid="B31">2021</xref>) used a perspective-guided cost function with a larger penalty to density far from the camera. Zhao et al. (<xref ref-type="bibr" rid="B46">2019</xref>) used the depth map predicted from RGB image as perspective information and embedded it into their density map prediction network. The abovementioned methods extract perspective information from RGB images, which is complicated and inaccurate. In contrast, depth maps can be directly used as accurate perspective information. In the following subsection, we introduce methods that employ depth maps.</p></sec>
<sec>
<title>2.3. Depth maps and crowd counting</title>
<p>Depth maps are the source of information, more accurate than the perspective information extracted from RGB images. Thanks to the current development of RGB-D cameras, several excellent crowd-counting methods have emerged to take advantage of depth maps. As density map-based crowd-counting methods have better performance in dealing with far-view areas, and as detection-based methods have better performance in dealing with near-view areas, Xu et al. (<xref ref-type="bibr" rid="B39">2019</xref>) used depth maps to segment RGB images into the far-view and near-view areas, and used density map-based and detection-based methods to deal with these two areas, respectively. However, the density map-based and detection-based methods employed in their framework do not use depth maps during counting. Liu et al. (<xref ref-type="bibr" rid="B20">2021</xref>, <xref ref-type="bibr" rid="B22">2023</xref>), Zhang et al. (<xref ref-type="bibr" rid="B44">2021</xref>), and Li et al. (<xref ref-type="bibr" rid="B13">2023</xref>) proposed cross-model frameworks to estimate density maps, fusing image features extracted from RGB images and depth maps to make use of the complementary information in these two kinds of images. They only used depth maps as input and did not explicitly utilize the perspective information contained in depth maps. Lian et al. (<xref ref-type="bibr" rid="B16">2019</xref>) used depth-adaptive Gaussian kernels and depth-aware anchors to improve crowd-counting and localization results, using the perspective information in depth maps to improve the quality of ground-truth density maps and human head anchors. Additionally, Lian et al. (<xref ref-type="bibr" rid="B15">2022</xref>) used depth-guided dynamic dilated convolution to further improve the method proposed by Lian et al. (<xref ref-type="bibr" rid="B16">2019</xref>). Compared with the above methods, the proposed HRPT utilizes the perspective information in depth maps more intuitively, narrowing the height gap among human heads through image transformation to alleviate the scale variation problem in crowd counting.</p></sec>
<sec>
<title>2.4. Image transformation and crowd counting</title>
<p>Yang et al. (<xref ref-type="bibr" rid="B41">2020</xref>) proposed a reverse perspective network to evaluate and correct the perspective distortion in crowd images. Both Yang et al.&#x00027;s (<xref ref-type="bibr" rid="B41">2020</xref>) method and our HRPT attempt to narrow the scale gap among human heads by image transformation. However, the two methods use different perspective information. Theirs uses the perspective information extracted from RGB images; our HRPT uses the perspective information in depth maps. The perspective information in depth maps is more accurate than that extracted from RGB images. Moreover, Yang et al. (<xref ref-type="bibr" rid="B41">2020</xref>) designed a specific network structure to estimate and correct perspective distortion. Other methods must change their original network structures to employ this approach. In contrast, they can easily employ our HRPT approach without changing any part of their original network structures.</p></sec></sec>
<sec sec-type="methods" id="s3">
<title>3. Methods</title>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> depicts the flowchart of our crowd-counting framework. As shown in this figure, the original RGB-D image is first sent into the proposed HRPT module. HRPT employs the perspective information in the depth map to narrow the height gap among human heads in the RGB image by image transformation. Afterward, the transformed RGB image is sent into a density map-based crowd-counting network to estimate the crowd density map, and the counting result is calculated by summing the density over the whole image. The proposed HRPT comprises two main steps: rescaling factor calculation and height transformation. In the following, we first introduce each step of HRPT in detail and then briefly introduce the crowd-counting networks employed in our framework.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Flowchart of our crowd-counting framework.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0002.tif"/>
</fig>
<sec>
<title>3.1. Rescaling factor calculation</title>
<p>The proposed HRPT is designed to narrow the height gap among the human heads in RGB images. To accomplish this goal, the rescaling factor should be inversely proportional to the height of the head in the original RGB image:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>s</italic> denotes the rescaling factor; <italic>h</italic> denotes the head height in the original RGB image; and <italic>a</italic><sub>1</sub> is a hyper-parameter that equals the rescaled head height.</p>
<p>According to Lian et al. (<xref ref-type="bibr" rid="B16">2019</xref>), the head height <italic>h</italic> is inversely proportional to the depth <italic>d</italic>. We formulate their relationship as follows:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>a</italic><sub>2</sub> is a parameter related to camera intrinsic parameters, such as focal length. After combining Equations (1, 2), we deduce the relationship between the rescaling factor <italic>s</italic> and depth <italic>d</italic> as follows:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:mi>d</mml:mi><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>As each pixel has a different depth value and the depth map always misses the depth value of some areas, such as the top-left area in the original depth map shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, transforming the crowd image according to the rescaling factors calculated by Equation (3) is hard. Fortunately, the depth usually complies with the following rule: pixels with smaller y-coordinates generally have higher depth values. Considering Equation (3) and the above rule, we conclude that it is possible to find the general relationship between the rescaling factor and y-coordinate. According to Rodriguez et al. (<xref ref-type="bibr" rid="B26">2011</xref>), under the assumption that people stand on the ground plane and the camera has no horizontal or in-plane rotation, the relationship between head height and y-coordinate is formulated as follows:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>y</italic><sub><italic>o</italic></sub> denotes the y-coordinate in the original RGB image; <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the y-coordinate of the horizon in the original RGB image; and <italic>a</italic><sub>3</sub> is a parameter related to camera extrinsic parameters, such as the camera height above the ground. We obtain the relationship between y-coordinate <italic>y</italic><sub><italic>o</italic></sub> and depth <italic>d</italic> by substituting Equation (4) into Equation (2):</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We set <inline-formula><mml:math id="M7"><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M8"><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> to simplify Equation (5). Then, Equation (5) is rewritten as follows:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>According to Equation (6), <inline-formula><mml:math id="M10"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula> is positively correlated with <italic>y</italic><sub><italic>o</italic></sub>. Next, we obtain the relationship between y-coordinate <italic>y</italic><sub><italic>o</italic></sub> and rescaling factor <italic>s</italic> by substituting Equation (6) into Equation (3):</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We compute rescaling factors using Equation (7), where <italic>a</italic><sub>1</sub> is a parameter whose value is determined by experience; <italic>a</italic><sub>2</sub> is related to intrinsic camera parameters; and <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are related to both extrinsic and intrinsic camera parameters. In general, camera intrinsic parameters can remain unchanged. The value of <italic>a</italic><sub>2</sub> is the same for images captured by cameras with identical intrinsic parameters, and it can be determined by fitting Equation (2). In contrast, it is difficult to keep the camera&#x00027;s extrinsic parameters constant. As a result, the values of <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> should be recalculated for each image by fitting Equation (6) based on its associated depth map.</p></sec>
<sec>
<title>3.2. Height transformation</title>
<p>Height transformation is implemented according to the rescaling factors calculated by Equation (7), showing that each image row has a specific value of rescaling factor. Thus, height transformation can be performed by adjusting the height of each image row according to its rescaling factor. <xref ref-type="fig" rid="F3">Figure 3</xref> displays the height transformation principle using a toy example, where <italic>I</italic><sub><italic>o</italic></sub> denotes the original RGB image; <italic>I</italic><sub><italic>t</italic></sub> denotes the RGB image after height transformation. The heights and y-coordinates of image rows are displayed in <italic>I</italic><sub><italic>o</italic></sub> and <italic>I</italic><sub><italic>t</italic></sub>. These heights are measured in pixels. As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, before height transformation, the height of each row is 1. After height transformation, the height of each row is multiplied by its corresponding rescaling factor. For example, in <italic>I</italic><sub><italic>o</italic></sub>, the height of the image row with y-coordinate <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is 1. After height transformation, the height of its corresponding row in <italic>I</italic><sub><italic>t</italic></sub> is changed to <inline-formula><mml:math id="M15"><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, and the y-coordinate is changed to <inline-formula><mml:math id="M16"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:munderover><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. <italic>s</italic>(<italic>y</italic><sub><italic>o</italic></sub>) denotes the rescaling factor of the image row with y-coordinate <italic>y</italic><sub><italic>o</italic></sub>. It is calculated by Equation (7). <inline-formula><mml:math id="M17"><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Height transformation process of a toy example. <italic>I</italic><sub><italic>o</italic></sub> denotes the original RGB image; <italic>H</italic><sub><italic>o</italic></sub> denotes the height of <italic>I</italic><sub><italic>o</italic></sub>; <italic>I</italic><sub><italic>t</italic></sub> denotes the RGB image after height transformation; <italic>s</italic>(<italic>y</italic><sub><italic>o</italic></sub>) denotes the rescaling factor calculated using Equation (7), and <inline-formula><mml:math id="M12"><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:math></inline-formula> <italic>s</italic>(<italic>H</italic><sub><italic>o</italic></sub>) &#x0003D; <italic>s</italic>(<italic>y</italic><sub><italic>o</italic></sub> &#x0003D; <italic>H</italic><sub><italic>o</italic></sub>). The height of <italic>I</italic><sub><italic>t</italic></sub> is <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0003.tif"/>
</fig>
<p>The toy example shown in <xref ref-type="fig" rid="F3">Figure 3</xref> depicts an ideal height transformation process. In the ideal process, the calculated heights and y-coordinates of image rows in <italic>I</italic><sub><italic>t</italic></sub> have a high probability of being decimals. However, in practice, they should be integers. Therefore, this ideal process cannot be implemented in practice. To solve this problem, we approximate this ideal height transformation process using some equations that can be easily implemented in practice.</p>
<p>As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the relationship between <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:munderover><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. We use integration to replace summing:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:msubsup></mml:mstyle><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Then, we substitute Equation (7) into Equation (8) and obtain the following expression:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M22"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:mi>l</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In Equation (9), <italic>a</italic><sub>4</sub>, <italic>b</italic><sub>4</sub>, and <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> are above 0. Equation (9) can be reformulated as:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The y-coordinates of image rows in the RGB image after height transformation are integers, <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where <italic>H</italic><sub><italic>t</italic></sub> denotes the height of the RGB image after height transformation. For a particular <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, we can use Equation (10) to calculate its corresponding <inline-formula><mml:math id="M27"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>. Let us use <italic>I</italic><sub><italic>o</italic></sub>(<inline-formula><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>) to denote the image row with y-coordinate <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> in <italic>I</italic><sub><italic>o</italic></sub>, and use <inline-formula><mml:math id="M30"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> to denote the image row with y-coordinate <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> in <italic>I</italic><sub><italic>t</italic></sub>. In our height transformation process, the pixels in <inline-formula><mml:math id="M32"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> are assigned to the pixels in <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. We use linear interpolation to calculate <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> since the calculated <inline-formula><mml:math id="M35"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> has a high probability of being a decimal:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M36"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>&#x0230A;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x0230B;</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>&#x02308;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x02309;</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>&#x02308;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x02309;</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>&#x0230A;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x0230B;</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M38"><mml:mrow><mml:mo>&#x0230A;</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>&#x0230B;</mml:mo></mml:mrow></mml:math></inline-formula> denotes the nearest integer lower than <inline-formula><mml:math id="M39"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>;</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M40"><mml:mrow><mml:mo>&#x02308;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow><mml:mo>&#x02309;</mml:mo></mml:mrow></mml:math></inline-formula> denotes the nearest integer higher than <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M42"><mml:mrow><mml:mo>&#x02308;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x02309;</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>&#x0230A;</mml:mo> <mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>o</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow> <mml:mo>&#x0230B;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x000A0;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>.</p></sec>
<sec>
<title>3.3. Crowd-counting networks</title>
<p>The proposed HRPT can be used as an image preprocessing step in crowd-counting networks. Then, this study selects four well-performing crowd-counting methods: CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>), DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>), and CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>) with their brief introductions presented below:</p>
<p>CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>) is a representative crowd-counting method based on density map estimation. It generates ground-truth density maps by blurring the dot annotations of human heads with Gaussian kernels. CSRNet uses VGG-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B28">2015</xref>) as its backbone and uses the L2 loss between predicted density map and ground-truth density map as its loss function. A popular tactic used in density map-based methodologies is to generate ground-truth density maps using Gaussian kernels. Their counting performance is strongly associated with the quality of generated ground-truth density maps (Ma et al., <xref ref-type="bibr" rid="B25">2019</xref>). However, a recent study indicates that using Gaussian kernels is detrimental to the generalization performance (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>). DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>) is proposed to solve the above problem by abandoning Gaussian kernels. It does not generate any ground-truth density maps in advance, and it uses optimally balanced transport to calculate the training loss between the predicted density map and dot annotations of human heads. GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>) offers a similar technique to DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>). Differently from DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>) uses unbalanced optimal transport, which preserves the predicted and annotated counts and generates pixel and point-wise loss.</p>
<p>The above three methods are based on Convolutional Neural Networks (CNNs). In recent years, transformer has been successfully used in many computer vision tasks and has achieved higher performance than CNN (Han et al., <xref ref-type="bibr" rid="B4">2023</xref>). Therefore, we also employ a transformer-based crowd counting method, CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>). CLTR takes image features extracted by CNN and trainable embeddings as input of a transformer-decoder. It directly predicts the localizations of human heads.</p></sec>
<sec>
<title>3.4. The processing steps of our method</title>
<p>The processing steps of our method are shown in the following <xref ref-type="table" rid="T6">Algorithm 1</xref>, where <italic>n</italic> denotes the <italic>n</italic><sup><italic>th</italic></sup> crowd image in the dataset; <italic>N</italic> denotes the total number of images in the dataset. The meanings of <italic>a</italic><sub>1</sub>, <italic>a</italic><sub>2</sub>, <italic>a</italic><sub>4</sub>, and <italic>b</italic><sub>4</sub> have been introduced in Section 3.1. The meanings of <inline-formula><mml:math id="M43"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M45"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M46"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, <italic>H</italic><sub><italic>t</italic></sub>, and <italic>I</italic><sub><italic>t</italic></sub> have been introduced in Section 3.2.</p>
<table-wrap position="float" id="T6">
<label>Algorithm 1</label>
<caption><p>Our proposed crowd-counting method.</p></caption> 
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td valign="top" align="left"><monospace><bold>Input:</bold> Crowd RGBD dataset, where images are captured by cameras with same intrinsic parameters.</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(1) Set <italic>a</italic><sub>1</sub> by experience.</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(2) Set <italic>a</italic><sub>2</sub> by fitting Equation (2) on the training subset. The value of <italic>a</italic><sub>2</sub> is the same for all images in this dataset. We need to use the depths in depth maps to fit Equation (2).</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(3) <bold>For</bold> <italic>n</italic><bold>&#x0003D;1 to</bold> <italic>N</italic> <bold>do</bold></monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(4) &#x000A0;&#x000A0;&#x000A0;&#x000A0;Set <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> by fitting Equation (6). The values of <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are recalculated for each image. We need to use the depths in depth maps to fit Equation (6).</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(5) &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>For</bold> <inline-formula><mml:math id="M47"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x02217;</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula><bold>&#x0003D;1 to</bold> <italic>H</italic><sub><italic>t</italic></sub> <bold>do</bold></monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(6) &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Calculate the corresponding <inline-formula><mml:math id="M48"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> of each <inline-formula><mml:math id="M49"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> according to Equation (10).</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(7) &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Calculate <inline-formula><mml:math id="M50"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> using linear interpolation according to Equation (11).</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(8) &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Assign <inline-formula><mml:math id="M51"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M52"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace>(9) &#x000A0;&#x000A0;&#x000A0;&#x000A0;Count the number of people by sending <italic>I</italic><sub><italic>t</italic></sub> to the crowd counting networks.</monospace></td></tr>
<tr>
<td valign="top" align="left"><monospace><bold>Output:</bold> The crowd-counting results.</monospace></td></tr>
</tbody>
</table>
</table-wrap>
<p>As shown in the above algorithm, <italic>a</italic><sub>2</sub> is set by fitting Equation (2), and <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are set by fitting Equation (6). We need to use the depths in depth maps to fit these two equations. Thus, the depth information is used in the above step (2) and step (4).</p></sec></sec>
<sec id="s4">
<title>4. Experiment</title>
<p>The proposed HRPT requires the perspective information in depth maps to accomplish image transformation. However, most public crowd-counting datasets only contain RGB images (Wang Q. et al., <xref ref-type="bibr" rid="B35">2020</xref>). Fortunately, Lian et al. (<xref ref-type="bibr" rid="B16">2019</xref>) released a large RGB-D crowd-counting dataset called ShanghaiTechRGBD in 2019, comprising 1193 training images and 1000 testing images. Most of our experiments are done on this RGBD dataset. Besides, our method can be extended to the RGB dataset by predicting the depth maps of the RGB images. We choose ShanghaiTech PartB dataset (Zhang et al., <xref ref-type="bibr" rid="B45">2016</xref>) to evaluate the performance of our method on the RGB dataset. Our experiments are implemented with Pytorch framework. We use one Nvidia RTX 2080ti GPU and one Intel Core i7 9700k CPU.</p>
<p>The image transformation and crowd-counting performances of our method are discussed in the subsequent subsections.</p>
<sec>
<title>4.1. Performance of image transformation</title>
<p>We use our experience to set <italic>a</italic><sub>1</sub> to 40. The counting performances with different values of <italic>a</italic><sub>1</sub> are shown in Section 4.2.4. Then, <italic>a</italic><sub>2</sub> is set to 350 by fitting Equation (2), and its value is the same for all images. Using the least-square algorithm, <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are set by fitting Equation (6). Their values differ based on different images. The distributions of <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are shown in Section 4.2.4. HRPT uses image processing to narrow the height gap among human heads. <xref ref-type="fig" rid="F4">Figure 4</xref> depicts three groups of crowd RGB images to demonstrate the effectiveness of HRPT. The heads of three pedestrians are marked in each RGB image. Their pixel-measured head heights are indicated on the left side of the images. HRPT stretches the top areas of RGB images and shrinks the bottom areas of RGB images, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. By comparing the head heights shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, we observe that HRPT successfully narrows the height gap among human heads. After HRPT, the head heights are approximated to the value of <italic>a</italic><sub>1</sub>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Three examples show the effectiveness of HRPT. The heads of three pedestrians are marked in each RGB image. Their head heights, measured in pixels, are shown on the left of the images. In this figure, <italic>H</italic><sub><italic>o</italic></sub> denotes the height of the original crowd RGB image; <italic>H</italic><sub><italic>t</italic></sub> denotes the height of the crowd RGB image transformed by HRPT; <italic>C</italic><sup><italic>gt</italic></sup> denotes the ground-truth number of people; <italic>C</italic><sup><italic>p</italic></sup> denotes the predicted number of people; and <italic>AE</italic> denotes the absolute error.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0004.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F5">Figure 5</xref>, we compare the head heights to demonstrate the effectiveness of HRPT. As shown in this figure, HRPT successfully enlarges the heights of small heads in far areas, reduces the heights of large heads in near areas, and narrows the height gap among human heads.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparisons of head heights to show the effectiveness of HRPT: <bold>(A)</bold> head heights in Example 1 of <xref ref-type="fig" rid="F4">Figure 4</xref>; <bold>(B)</bold> head heights in Example 2 of <xref ref-type="fig" rid="F4">Figure 4</xref>; and <bold>(C)</bold> head heights in Example 3 of <xref ref-type="fig" rid="F4">Figure 4</xref>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0005.tif"/>
</fig></sec>
<sec>
<title>4.2. Performance of crowd counting</title>
<sec>
<title>4.2.1. Training details and metrics</title>
<p>To verify the effectiveness of HRPT on crowd counting, we combine four well-performing crowd-counting neural networks, CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>), DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>), and CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>), with HRPT. During training of these four neural networks, Adam (Kingma and Ba, <xref ref-type="bibr" rid="B12">2014</xref>) is used as the optimizer, and the learning rate is set to 1 &#x000D7; 10<sup>&#x02212;5</sup>. The performance of different methods is evaluated using the mean absolute error (MAE) and mean square error (MSE), as follows:</p>
<disp-formula id="E13"><label>(12)</label><mml:math id="M53"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E14"><label>(13)</label><mml:math id="M54"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> is the number of testing images; <inline-formula><mml:math id="M55"><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M56"><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are the predicted and ground-truth number of people in <italic>i</italic><sup><italic>th</italic></sup> testing image, respectively.</p></sec>
<sec>
<title>4.2.2. Comparisons to the baselines</title>
<p>This study combines HRPT with four well-performing crowd-counting methods: CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>), DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>), GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>), and CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>). Therefore, we use CSRNet, DM-Count, GL, and CLTR as our baselines. Comparisons to these four baselines are shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Comparisons to the baselines on the ShanghaiTechRGBD dataset.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left" colspan="2"><bold>Methods</bold></th>
<th valign="top" align="left"><bold>Architecture</bold></th>
<th valign="top" align="center" colspan="2"><bold>Speed (ms)</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497;color:#ffffff">
<td valign="top" align="left" colspan="2"></td>
<td/>
<td valign="top" align="center"><bold>Preprocessing (CPU)</bold></td>
<td valign="top" align="center"><bold>Neural network (GPU)</bold></td>
<td/>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="4">Baselines</td>
<td valign="top" align="left">CLTR (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>)</td>
<td valign="top" align="left">Transformer</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">249</td>
<td valign="top" align="center">5.29</td>
<td valign="top" align="center">7.70</td>
</tr>
 <tr>
<td valign="top" align="left">CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>)</td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">164</td>
<td valign="top" align="center">5.11</td>
<td valign="top" align="center">9.99</td>
</tr>
 <tr>
<td valign="top" align="left">DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>)</td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">149</td>
<td valign="top" align="center">4.00</td>
<td valign="top" align="center">5.95</td>
</tr>
 <tr>
<td valign="top" align="left">GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>)</td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">152</td>
<td valign="top" align="center">3.96</td>
<td valign="top" align="center">5.97</td>
</tr> <tr>
<td valign="top" align="left" rowspan="4"><bold>Ours</bold></td>
<td valign="top" align="left"><bold>CLTR&#x0002B;HRPT</bold></td>
<td valign="top" align="left">Transformer</td>
<td valign="top" align="center">268</td>
<td valign="top" align="center">173</td>
<td valign="top" align="center">4.61</td>
<td valign="top" align="center">6.71</td>
</tr>
 <tr>
<td valign="top" align="left"><bold>CSRNet</bold> <bold>&#x0002B;</bold> <bold>HRPT</bold></td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">268</td>
<td valign="top" align="center">124</td>
<td valign="top" align="center">3.76</td>
<td valign="top" align="center">5.65</td>
</tr>
 <tr>
<td valign="top" align="left"><bold>DM-Count</bold> <bold>&#x0002B;</bold> <bold>HRPT</bold></td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">268</td>
<td valign="top" align="center">108</td>
<td valign="top" align="center">3.78</td>
<td valign="top" align="center">5.59</td>
</tr>
 <tr>
<td valign="top" align="left"><bold>GL</bold> <bold>&#x0002B;</bold> <bold>HRPT</bold></td>
<td valign="top" align="left">CNN</td>
<td valign="top" align="center">268</td>
<td valign="top" align="center">111</td>
<td valign="top" align="left"><bold>3.70</bold></td>
<td valign="top" align="left"><bold>5.42</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Scores marked in bold indicate the best results on the corresponding metric.</p>
</table-wrap-foot>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="T1">Table 1</xref>, CLTR is built based on transformer (Liang et al., <xref ref-type="bibr" rid="B17">2022</xref>), while CSRNet, DM-Count, and GL are built based on CNN. Many studies have proved that transformer has higher performance than other types of networks, such as CNN (Han et al., <xref ref-type="bibr" rid="B4">2023</xref>). However, the transformer-based crowd-counting method CLTR has poorer performance than the other three CNN-based methods in our experiments. This is because, transformer models are more sensitive to the hyper-parameters for training, such as batch size, and are huger and more computationally expensive (Han et al., <xref ref-type="bibr" rid="B4">2023</xref>). However, we only have one Nvidia RTX 2080ti GPU. To train CLTR on our device, we set a much smaller batch size than the authors of CLTR. As a result, the maximum performance of CLTR is not achieved. Even so, our experimental results still demonstrate that HRPT can improve the crowd-counting performance of CLTR.</p>
<p><xref ref-type="table" rid="T1">Table 1</xref> also demonstrates that HRPT offers a more significant improvement to CSRNet than to DM-Count and GL. This is because CSRNet uses Gaussian kernels to generate ground-truth density maps. Its performance is highly dependent on the quality of the generated ground-truth density maps, which is reduced due to the scale variations of human heads (Ma et al., <xref ref-type="bibr" rid="B25">2019</xref>). The above quality reduction is narrowed since HRPT can narrow the height gap among human heads. In contrast, DM-Count and GL do not use Gaussian kernels to generate ground-truth density maps. Their sensitivities to the scale variations of human heads are smaller than those of CSRNet. Therefore, HRPT offers a more considerable improvement to CSRNet than to DM-Count and GL. In addition, because GL uses unbalanced optimal transport, which preserves the predicted and annotated counts, it has better performance than CSRNet and DM-Count. Therefore, GL&#x0002B;HRPT achieves the best crowd counting performance.</p>
<p>Besides the evaluation scores of MAE and MSE, the processing speeds are also shown in <xref ref-type="table" rid="T1">Table 1</xref>. The baselines do not employ the preprocessing step. Therefore, they spend 0 ms in the preprocessing step. Our methods use HRPT as the preprocessing step. HRPT spends 268 ms for each image, which is a considerable amount of time. This is because our current version of HRPT is operated in CPU. In the future, we will put the for-loop in step (5) of Algorithm 1 in GPU to increase the processing speed. In addition, as shown in <xref ref-type="table" rid="T1">Table 1</xref>, after HRPT, the processing speeds of neural networks are faster than the baselines. This is because HRPT effectively reduces much redundancy information in the near areas. Image examples in <xref ref-type="fig" rid="F4">Figure 4</xref> show that, in the original crowd RGB images, the near areas contain much fewer people but take up much more image spaces. In contrast, in the crowd RGB images transformed by HRPT, the near areas are shrunken by a large margin. The average image size is reduced by HRPT.</p>
<p><xref ref-type="fig" rid="F4">Figure 4</xref> also depicts the density map estimation results of GL and GL&#x0002B;HRPT, displaying the effectiveness of HRPT on crowd counting qualitatively. By comparing these density map estimation results, we can see that HRPT distributes human heads more evenly. In the density map estimation results of Example 1, a couple of corresponding regions located at the top of images are marked. After comparing these two regions, it becomes clear that without HRPT, it is hard to distinguish between human heads that are far from the camera. However, with HRPT, it is much easier to identify these heads based on the estimated density map. Although these two regions have different heights, they correspond to the same area in the actual scene.</p>
<p>As GL&#x0002B;HRPT achieves the best performance, in the following, we focus on the experiments of GL&#x0002B;HRPT, and use <bold>ours</bold> to denote GL&#x0002B;HRPT.</p></sec>
<sec>
<title>4.2.3. Comparisons to other methods</title>
<p>In this subsection, we compare our method with other crowd-counting methods that also use depth maps. Evaluation results of different methods on the ShanghaiTechRGBD dataset are shown in the following <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparisons to other crowd-counting methods that also use depth maps on the ShanghaiTechRGBD dataset.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>Year</bold></th>
<th valign="top" align="left"><bold>Backbones</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="left"><bold>MSE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">RDNet (Lian et al., <xref ref-type="bibr" rid="B16">2019</xref>)</td>
<td valign="top" align="center">2019</td>
<td valign="top" align="left">ResNet-101 &#x0002B; VGG-16</td>
<td valign="top" align="center">4.96</td>
<td valign="top" align="center">7.22</td>
</tr> <tr>
<td valign="top" align="left">CSRNet &#x0002B; IADM (Liu et al., <xref ref-type="bibr" rid="B20">2021</xref>)</td>
<td valign="top" align="center">2021</td>
<td valign="top" align="left">VGG-16 &#x000D7; 3</td>
<td valign="top" align="center">4.38</td>
<td valign="top" align="center">7.06</td>
</tr> <tr>
<td valign="top" align="left">DPDNet (Lian et al., <xref ref-type="bibr" rid="B15">2022</xref>)</td>
<td valign="top" align="center">2022</td>
<td valign="top" align="left">ResNet-101 &#x0002B; VGG-16</td>
<td valign="top" align="center">4.23</td>
<td valign="top" align="center">6.75</td>
</tr> <tr>
<td valign="top" align="left">Cross-model (Zhang et al., <xref ref-type="bibr" rid="B44">2021</xref>)</td>
<td valign="top" align="center">2021</td>
<td valign="top" align="left">VGG-16 &#x000D7; 2</td>
<td valign="top" align="center">3.76</td>
<td valign="top" align="center">5.46</td>
</tr> <tr>
<td valign="top" align="left">Li et al. (<xref ref-type="bibr" rid="B13">2023</xref>)</td>
<td valign="top" align="center">2023</td>
<td valign="top" align="left">VGG-16 &#x000D7; 2</td>
<td valign="top" align="center">4.03</td>
<td valign="top" align="center">5.81</td>
</tr> <tr>
<td valign="top" align="left">CCANet (Liu et al., <xref ref-type="bibr" rid="B22">2023</xref>)</td>
<td valign="top" align="center">2023</td>
<td valign="top" align="left">VGG-16 &#x0002B; Designed</td>
<td valign="top" align="center">3.78</td>
<td valign="top" align="center">5.56</td>
</tr> <tr>
<td valign="top" align="left"><bold>Ours</bold></td>
<td valign="top" align="center">2023</td>
<td valign="top" align="left">VGG-19</td>
<td valign="top" align="left"><bold>3.70</bold></td>
<td valign="top" align="left"><bold>5.42</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Scores marked in bold indicate the best results on the corresponding metric.</p>
</table-wrap-foot>
</table-wrap>
<p>In addition to the evaluation results of different methods, <xref ref-type="table" rid="T2">Table 2</xref> also displays their type of backbones. In <xref ref-type="table" rid="T2">Table 2</xref>, &#x0201C;Designed&#x0201D; denotes that the authors designed the network backbone; &#x0201C;VGG-16 &#x000D7; 3&#x0201D; depicts that the network contains three branches whose backbones are VGG-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B28">2015</xref>), as do &#x0201C;VGG-16 &#x000D7; 2&#x0201D;; &#x0201C;ResNet-101 &#x0002B; VGG-16&#x0201D; means that the network contains two branches whose backbones are ResNet-101 (He et al., <xref ref-type="bibr" rid="B6">2016</xref>) and VGG-16. Similarly, &#x0201C;VGG-16 &#x0002B; Designed&#x0201D; means that the network contains two branches whose backbones are VGG-16 and Designed.</p>
<p>As shown in <xref ref-type="table" rid="T2">Table 2</xref>, our method outperforms other methods that also employ depth maps. This demonstrate that, although employing depth maps improves the crowd-counting results, different methods have different performances. Compared with other methods, the proposed HRPT more efficiently uses depth maps and improves the crowd-counting performance. Moreover, other depth map methods use complex network backbones because they employ additional network branches to deal with depth maps. In contrast, our method uses depth maps in the HRPT module, serving as a preprocessing step of an excellent crowd-counting network. Therefore, our method can use depth maps without changing any part of the original network structure. Thus, the backbone used in our method is much simpler than those used in other methods, which also use depth maps.</p></sec>
<sec>
<title>4.2.4. Study on image transformation parameters</title>
<p><italic>a</italic><sub>1</sub> is a hyper-parameter of HRPT. As shown in Equation (1), it represents the head height after HRPT. To study its effectiveness on crowd-counting performance, we change its value to 10, 20, 30, and 40, and then evaluate the corresponding crowd-counting results. Evaluation results with different values of <italic>a</italic><sub>1</sub> are shown in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Evaluation results with different values of <italic>a</italic><sub>1</sub> on the ShanghaiTechRGBD dataset.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>a<sub>1</sub></bold></th>
<th valign="top" align="center" colspan="2"><bold>Speed (ms)</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497;color:#ffffff">
<td/>
<td valign="top" align="left"><bold>Preprocessing (CPU)</bold></td>
<td valign="top" align="left"><bold>Neural network (GPU)</bold></td>
<td/>
<td/>
</tr> <tr>
<td valign="top" align="left">10</td>
<td valign="top" align="center">147</td>
<td valign="top" align="center">29</td>
<td valign="top" align="center">4.43</td>
<td valign="top" align="center">6.63</td>
</tr> <tr>
<td valign="top" align="left">20</td>
<td valign="top" align="center">166</td>
<td valign="top" align="center">55</td>
<td valign="top" align="center">3.74</td>
<td valign="top" align="center">5.51</td>
</tr> <tr>
<td valign="top" align="left">30</td>
<td valign="top" align="center">188</td>
<td valign="top" align="center">81</td>
<td valign="top" align="center">3.70</td>
<td valign="top" align="center">5.60</td>
</tr> <tr>
<td valign="top" align="left">40</td>
<td valign="top" align="center">268</td>
<td valign="top" align="center">111</td>
<td valign="top" align="center"><bold>3.70</bold></td>
<td valign="top" align="center"><bold>5.42</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Scores marked in bold indicate the best results on the corresponding metric.</p>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T3">Table 3</xref> shows that with the increase of <italic>a</italic><sub>1</sub>, the crowd-counting performance improves. <italic>a</italic><sub>1</sub> represents the head heights after HRPT. The smaller the value of <italic>a</italic><sub>1</sub>, the more image details are lost by HRPT, which is detrimental to crowd counting. When we set <italic>a</italic><sub>1</sub> to 50, some images in our experimental dataset will become very large and cannot be processed by the crowd-counting network on our device. Therefore, we finally set <italic>a</italic><sub>1</sub> to 40.</p>
<p>In addition, the processing speeds in <xref ref-type="table" rid="T3">Table 3</xref> show that, with the increase of <italic>a</italic><sub>1</sub>, the processing speed becomes slower and slower. This is because, the larger the value of <italic>a</italic><sub>1</sub>, the larger the images outputted by HRPT. Then, the preprocessing step and neural network need to spend more time to deal with these images.</p>
<p><italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> are set by fitting Equation (6). Their values differ based on different images. Their distributions on the ShanghaiTechRGBD dataset are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. As shown in this figure, in both the training and testing datasets, the maximum number of <italic>a</italic><sub>4</sub> falls into (3, 4) &#x000D7; 10<sup>&#x02212;4</sup>, and the maximum number of <italic>b</italic><sub>4</sub> falls into [2.5, 4.0] &#x000D7; 10<sup>&#x02212;2</sup>. In addition, the distribution of <italic>a</italic><sub>4</sub> on the training subset is similar to the distribution of <italic>a</italic><sub>4</sub> on the testing subset; the distribution of <italic>b</italic><sub>4</sub> on the training subset is similar to the distribution of <italic>b</italic><sub>4</sub> on the testing subset.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The histograms of <italic>a</italic><sub>4</sub> and <italic>b</italic><sub>4</sub> on ShanghaiTechRGBD dataset: <bold>(A)</bold> the histogram of <italic>a</italic><sub>4</sub> on the training subset; <bold>(B)</bold> the histogram of <italic>b</italic><sub>4</sub> on the training subset; <bold>(C)</bold> the histogram of <italic>a</italic><sub>4</sub> on the testing subset; <bold>(D)</bold> the histogram of <italic>b</italic><sub>4</sub> on the testing subset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0006.tif"/>
</fig></sec>
<sec>
<title>4.2.5. Evaluation on RGB dataset</title>
<p>Our method can be extended to the RGB dataset by predicting the depth maps of the RGB images (Lian et al., <xref ref-type="bibr" rid="B15">2022</xref>). HRPT estimates the relationship between the y-coordinate and rescaling factor. It is built under the assumption that in each image, people stand on the same ground plane and the camera has no horizontal rotation. Therefore, HRPT suits images captured from flat areas with surveillance views, and it requires the horizontal lines of captured images to be parallel to the image rows. In addition, HRPT narrows the head height gap by stretching the far areas. If some rows in the image are close to or above the horizontal lines, the depths of these rows are infinite. According to Equation (3), the rescaling factors of these rows are also infinite. Thus, HRPT does not suit images that contain image rows close to or above the horizontal lines. Images in the ShanghaiTech PartB dataset (Zhang et al., <xref ref-type="bibr" rid="B45">2016</xref>) satisfy the above requirements. Therefore, we choose ShanghaiTech PartB to evaluate the performance of our method on the RGB dataset. Evaluation results of our method and many other well-performing methods are shown in <xref ref-type="table" rid="T4">Table 4</xref>. As shown in this table, our method achieves high performance.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Evaluation results of different methods on the ShanghaiTech PartB dataset.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">MCNN (Zhang et al., <xref ref-type="bibr" rid="B45">2016</xref>)</td>
<td valign="top" align="center">26.4</td>
<td valign="top" align="center">41.3</td>
</tr> <tr>
<td valign="top" align="left">DecideNet (Liu et al., <xref ref-type="bibr" rid="B19">2018</xref>)</td>
<td valign="top" align="center">20.8</td>
<td valign="top" align="center">29.4</td>
</tr> <tr>
<td valign="top" align="left">CSRNet (Li et al., <xref ref-type="bibr" rid="B14">2018</xref>)</td>
<td valign="top" align="center">10.6</td>
<td valign="top" align="center">16.0</td>
</tr> <tr>
<td valign="top" align="left">RDNet (Lian et al., <xref ref-type="bibr" rid="B16">2019</xref>)</td>
<td valign="top" align="center">8.8</td>
<td valign="top" align="center">12.9</td>
</tr> <tr>
<td valign="top" align="left">DM-Count (Wang B. et al., <xref ref-type="bibr" rid="B32">2020</xref>)</td>
<td valign="top" align="center">7.4</td>
<td valign="top" align="center">11.8</td>
</tr> <tr>
<td valign="top" align="left">GL (Wan et al., <xref ref-type="bibr" rid="B31">2021</xref>)</td>
<td valign="top" align="center">7.3</td>
<td valign="top" align="center">11.7</td>
</tr> <tr>
<td valign="top" align="left">Cross-model (Zhang et al., <xref ref-type="bibr" rid="B44">2021</xref>)</td>
<td valign="top" align="center">8.3</td>
<td valign="top" align="center">12.9</td>
</tr> <tr>
<td valign="top" align="left">CCANet (Liu et al., <xref ref-type="bibr" rid="B22">2023</xref>)</td>
<td valign="top" align="center">8.1</td>
<td valign="top" align="center">13.5</td>
</tr> <tr>
<td valign="top" align="left">DPDNet (Lian et al., <xref ref-type="bibr" rid="B15">2022</xref>)</td>
<td valign="top" align="center">7.9</td>
<td valign="top" align="center">12.4</td>
</tr> <tr>
<td valign="top" align="left">AutoScale (Xu et al., <xref ref-type="bibr" rid="B38">2022</xref>)</td>
<td valign="top" align="left"><bold>6.8</bold></td>
<td valign="top" align="center">11.3</td>
</tr> <tr>
<td valign="top" align="left"><bold>Ours</bold></td>
<td valign="top" align="left"><bold>6.8</bold></td>
<td valign="top" align="left"><bold>11.2</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Scores marked in bold indicate the best results on the corresponding metric.</p>
</table-wrap-foot>
</table-wrap></sec></sec></sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>In this section, we first analyze why the proposed HRPT can improve the crowd-counting performance from another point of view. Afterward, we discuss another two image transformation methods, the drawbacks of our methods and our future work.</p>
<sec>
<title>5.1. Analysis</title>
<p>The proposed HRPT is designed to improve the crowd-counting performance by alleviating the scale variation problem. In the following, we analyze why the proposed HRPT can improve the crowd-counting performance from another point of view.</p>
<p>First, we analyze the effect of HRPT on the bottom areas of crowd RGB images. Human heads in the bottom areas are large and sparse before HRPT, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. Larger human heads contain much more detailed information. However, current crowd-counting methods do not require so much detailed information to pick out human heads. Therefore, the bottom areas of crowd RGB images contain much redundant information. Moreover, large and sparse human heads occupy too much image space, making counting networks spend too much energy on those &#x0201C;easy samples.&#x0201D; After HRPT, the bottom areas are shrunken to shorten the heights of human heads. Shrinking the bottom areas reduces redundant information and compels counting networks to pay more attention to the top areas with more &#x0201C;hard samples.&#x0201D;</p>
<p>Next, we analyze the effect of HRPT on the top areas of crowd RGB images. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, before HRPT, human heads in the top areas are tiny and dense. Information about these small human heads may be lost when extracting high-level image features through crowd-counting networks. After HRPT, the top areas are stretched to enlarge the heights of human heads. Even though human heads in these top areas become thinner, they can still be identified as human heads. The possibility of losing their information while extracting high-level image features becomes much smaller. Therefore, the proposed HRPT can improve crowd-counting performance.</p></sec>
<sec>
<title>5.2. Other two image transformation methods</title>
<sec>
<title>5.2.1. Shape-changing method</title>
<p>Section 4 reports the experimental results demonstrating that narrowing the height gap among human heads helps improve crowd-counting performance. What about narrowing the width gap among human heads? If we attempt to narrow the width gap by changing each row&#x00027;s width according to the rescaling factor calculated from Equation (7), the shape of the RGB image will be changed from rectangle to trapezoid. Moreover, human heads in the top-right and top-left regions will be seriously deformed and lose their essential characteristics. To avoid this severe deformation, we narrow the width gap among human heads by changing rectangle-shaped images into fan-shaped images. This shape-changing method naturally enlarges the widths of human heads in the top regions and shortens the widths in the bottom regions. <xref ref-type="fig" rid="F7">Figure 7</xref> displays a group of crowd RGB images to show the effect of narrowing the width gap among human heads through shape changing.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>A group of images to show the effect of narrowing the width gap among human heads by changing a rectangle-shaped image into a fan-shaped image: <bold>(A)</bold> before shape changing; <bold>(B)</bold> after shape changing. In this figure, the height gap among human heads has been narrowed by HRPT.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0007.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, this shape-changing method is operated by changing the shape of each image row from straight to semi-circle, and the width gap among human heads is successfully narrowed after shape changing. However, this shape-changing method has a shortcoming: it adds additional rotation. We combine this shape-changing method with GL and GL &#x0002B; HRPT to evaluate its effect on crowd-counting performance. The evaluation results are shown in <xref ref-type="table" rid="T5">Table 5</xref>.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Evaluation results of shaping changing method on ShanghaiTechRGBD dataset.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">GL</td>
<td valign="top" align="center">3.96</td>
<td valign="top" align="center">5.97</td>
</tr> <tr>
<td valign="top" align="left">GL &#x0002B; Shape-Changing</td>
<td valign="top" align="center">4.36</td>
<td valign="top" align="center">6.43</td>
</tr> <tr>
<td valign="top" align="left">GL &#x0002B; HRPT</td>
<td valign="top" align="left"><bold>3.70</bold></td>
<td valign="top" align="left"><bold>5.42</bold></td>
</tr> <tr>
<td valign="top" align="left">GL &#x0002B; HRPT &#x0002B; Shape-Changing</td>
<td valign="top" align="center">4.27</td>
<td valign="top" align="center">6.30</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Scores marked in bold indicate the best results on the corresponding metric.</p>
</table-wrap-foot>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="T5">Table 5</xref>, after we combine this shape-changing method with GL, MAE and MSE rise to 4.36 and 6.43, respectively; after we combine this shape-changing method with GL &#x0002B; HRPT, MAE and MSE rise to 4.27 and 6.30, respectively. These experimental results demonstrate that this shape changing negatively affects crowd counting, implying that the crowd-counting network is not robust enough to handle the additional rotation.</p></sec>
<sec>
<title>5.2.2. Geometric method</title>
<p>The perspective effect mainly causes the scale variation problem in crowd images. Accurately eliminating the perspective effect on the 3D world is very hard in surveillance scenes. In contrast, homographs can quickly eliminate the perspective effect on the 2D ground (Hartley and Zisserman, <xref ref-type="bibr" rid="B5">2003</xref>). Here, we test whether the homograph can narrow the scale gap (both height and width gap) among human heads by eliminating the perspective effect on the 2D ground. <xref ref-type="fig" rid="F8">Figure 8</xref> displays a group of crowd RGB images to show the effect of the homograph.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>A group of crowd RGB images to show the effect of eliminating the perspective effect by homograph: <bold>(A)</bold> the original crowd RGB image; <bold>(B)</bold> the crowd RGB image transformed by homograph.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimag-02-1271885-g0008.tif"/>
</fig>
<p>We zoom in on two areas of <xref ref-type="fig" rid="F8">Figure 8B</xref> to display the effect of the homograph. We can observe the shape of floor tiles in these two areas. As shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, the shape of the floor tiles is trapezoid before the homograph, and these tiles have different sizes. After homograph, the shape of the floor tiles returns to square, and these floor tiles have the same size. This result demonstrates that the homograph can eliminate the perspective effect on the 2D ground and narrow the scale gap (both height gap and width gap) among floor tiles. However, as shown in <xref ref-type="fig" rid="F8">Figure 8B</xref>, homograph cannot narrow the scale gap among human heads. After the homograph, the human heads far from the camera are stretched too much, and those near the camera are shrunken too much. Additionally, human heads at the top-right and top-left regions are seriously deformed and lose their essential characteristics. Severe deformation is harmful to crowd counting. The homograph transformation result shown in <xref ref-type="fig" rid="F8">Figure 8B</xref> is similar to a frame of crowd video rectified by estimated scene geometry in Rodriguez et al. (<xref ref-type="bibr" rid="B26">2011</xref>).</p></sec></sec>
<sec>
<title>5.3. Drawbacks and future works</title>
<p>This study proposes HRPT to narrow the height gap among human heads. HRPT narrows the head height gap by stretching the far areas. If some rows in the image are close to or above the horizontal lines, the depths of these rows are infinite. According to Equation (3), the rescaling factors of these rows are also infinite. Thus, we cannot transform them by our HRPT. Furthermore, HRPT is built under the assumption that in each image, people stand on the same ground plane and the camera has no horizontal rotation. Therefore, it only suits images captured from flat areas with surveillance views, and it requires the horizontal lines of captured images to be parallel to the image rows. The above requirements limit the wide usage of HRPT. In the future, we plan to address the above first problem by a segmentation method to remove image rows above the horizontal lines. Moreover, we plan to address the above second problem by building a more advanced image transformation model that suits more scenarios and does not require the horizontal lines of captured images to be parallel to the image rows.</p>
<p>In addition, as mentioned above, we do not achieve satisfactory results when narrowing the width gap among human heads. In the future, we will improve crowd-counting performance by proposing a more efficient image transformation method to narrow the width gap among human heads and an efficient crowd-counting network to handle the additional rotation, as shown in <xref ref-type="fig" rid="F7">Figure 7B</xref>.</p>
<p>As shown in Section 4.2.5, our method can be extended to the RGB dataset by predicting the depth maps of the RGB images. However, owing to the limitation of monocular vision, single image-based depth prediction methods only provide relative depth information (Lian et al., <xref ref-type="bibr" rid="B15">2022</xref>), which limits their depth prediction accuracy. Previous research has demonstrated that embedding focal length can overcome the problem of single image-based depth prediction and acquire accurate depths (He et al., <xref ref-type="bibr" rid="B7">2018</xref>). In the future, we will build a crowd-counting dataset with known focal lengths, then use these focal lengths to predict accurate depths, and finally associate these accurate depth prediction results with our method to further improve the crowd-counting accuracy on RGB images.</p>
<p>Besides, as shown in <xref ref-type="table" rid="T1">Table 1</xref>, HRPT spends 268 ms for each image, which is a considerable amount of time. Our current version of HRPT is operated in CPU. In the future, we will put the for-loop in step (5) of Algorithm 1 in GPU to increase the processing speed of HRPT.</p></sec></sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusions</title>
<p>This study uses HRPT to alleviate the scale variation problem in crowd-counting tasks. HRPT creatively narrows the height gap among human heads by using the perspective information contained in the depth map. Moreover, it enlarges small human heads in far areas to make them more visible, and it shrinks large human heads in closer areas to reduce redundant information. Other excellent crowd-counting methods can easily employ HRPT as a preprocessing step. Experimental results show that our method achieves high crowd-counting performance.</p></sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://github.com/svip-lab/RGBD-counting">https://github.com/svip-lab/RGBD-counting</ext-link>.</p></sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>XZ: Funding acquisition, Methodology, Project administration, Software, Writing&#x02014;original draft. HL: Methodology, Software, Writing&#x02014;original draft. ZZ: Formal analysis, Writing&#x02014;review and editing. SL: Formal analysis, Writing&#x02014;review and editing.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by the Natural Science Foundation of Shandong Province (grant ZR2021QF094) and the Youth Innovation Team Technology Project of Higher School in Shandong Province (2022KJ204).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>Z.</given-names></name> <name><surname>Shi</surname> <given-names>M.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Zafeiriou</surname> <given-names>S.</given-names></name></person-group> (<year>2023</year>). <article-title>Redesigning multi-scale neural network for crowd counting</article-title>. <source>IEEE T Image Proc</source>. <volume>32</volume>, <fpage>3664</fpage>&#x02013;<lpage>3678</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2023.3289290</pub-id><pub-id pub-id-type="pmid">37384475</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>A survey of crowd counting and density estimation based on convolutional neural network</article-title>. <source>Neurocomputing</source> <volume>472</volume>, <fpage>224</fpage>&#x02013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.02.103</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>G.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>CNN-based density estimation and crowd counting: a survey</article-title>. <source>arXiv preprint</source> arXiv:2003.12783. <pub-id pub-id-type="doi">10.48550/arXiv.2003.12783</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>A survey on vision transformer</article-title>. <source>IEEE T Pattern Anal</source>. <volume>45</volume>, <fpage>87</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2022.3152247</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hartley</surname> <given-names>R.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <source>Multiple View Geometry in Computer Vision</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep residual learning for image recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>Z.</given-names></name></person-group> (<year>2018</year>). <article-title>Learning depth from single images with deep neural network embedding focal length</article-title>. <source>IEEE T Image Process</source>. <volume>27</volume>, <fpage>4676</fpage>&#x02013;<lpage>4689</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2018.2832296</pub-id><pub-id pub-id-type="pmid">29994526</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Idrees</surname> <given-names>H.</given-names></name> <name><surname>Soomro</surname> <given-names>K.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Detecting humans in dense crowds using locally-consistent scale prior and global occlusion reasoning</article-title>. <source>IEEE T Pattern Anal</source>. <volume>37</volume>, <fpage>1986</fpage>&#x02013;<lpage>1998</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2396051</pub-id><pub-id pub-id-type="pmid">26340254</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Zhen</surname> <given-names>X.</given-names></name> <name><surname>Cao</surname> <given-names>X.</given-names></name> <name><surname>Doermann</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Crowd counting and density estimation by trellis encoder-decoder networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6126</fpage>&#x02013;<lpage>6135</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Lv</surname> <given-names>P.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Learning multi-level density maps for crowd counting</article-title>. <source>IEEE T Neur. Net. Lear</source>. <volume>31</volume>, <fpage>2705</fpage>&#x02013;<lpage>2715</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2019.2933920</pub-id><pub-id pub-id-type="pmid">31562106</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jing</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>An adaptive multi-sensor data fusion method based on deep convolutional neural networks for fault diagnosis of planetary gearbox</article-title>. <source>Sensors</source> <volume>17</volume>, <fpage>414</fpage>. <pub-id pub-id-type="doi">10.3390/s17020414</pub-id><pub-id pub-id-type="pmid">28230767</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>arXiv preprint</source> arXiv:1412.6980. <pub-id pub-id-type="doi">10.48550/arXiv.1412.6980</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Kong</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>RGB-D crowd counting with cross-modal cycle-attention fusion and fine-coarse supervision</article-title>. <source>IEEE T Ind. Inform</source>. <volume>19</volume>, <fpage>306</fpage>&#x02013;<lpage>316</lpage>. <pub-id pub-id-type="doi">10.1109/TII.2022.3171352</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;CSRNet: dilated convolutional neural networks for understanding the highly congested scenes,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1091</fpage>&#x02013;<lpage>1100</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lian</surname> <given-names>D.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>Locating and counting heads in crowds with a depth prior</article-title>. <source>IEEE T Pattern Anal</source>. <volume>44</volume>, <fpage>9056</fpage>&#x02013;<lpage>9072</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2021.3124956</pub-id><pub-id pub-id-type="pmid">34735337</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lian</surname> <given-names>D.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Zheng</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Density map regression guided detection network for RGB-D crowd counting and localization,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1821</fpage>&#x02013;<lpage>1830</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>D.</given-names></name> <name><surname>Xu</surname> <given-names>W.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;An end-to-end transformer model for crowd localization,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>38</fpage>&#x02013;<lpage>54</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Bayesian model adaptation for crowd counts,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Santiago</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4175</fpage>&#x02013;<lpage>4183</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Gao</surname> <given-names>C.</given-names></name> <name><surname>Meng</surname> <given-names>D.</given-names></name> <name><surname>Hauptmann</surname> <given-names>A. G.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Decidenet: counting varying density crowds through attention guided detection and density estimation,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5197</fpage>&#x02013;<lpage>5206</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Lin</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Cross-modal collaborative representation learning and a large-scale RGBT benchmark for crowd counting,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4821</fpage>&#x02013;<lpage>4831</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Salzmann</surname> <given-names>M.</given-names></name> <name><surname>Fua</surname> <given-names>P.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Context-aware crowd counting,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5094</fpage>&#x02013;<lpage>5103</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>G.</given-names></name> <name><surname>Shi</surname> <given-names>B.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name></person-group> (<year>2023</year>). <article-title>CCANet: a collaborative cross-modal attention network for RGB-D crowd counting</article-title>. <source>IEEE T Multimedia</source> <volume>25</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2023.3262978</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Shi</surname> <given-names>M.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Point in, box out: beyond counting persons in crowds,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6462</fpage>&#x02013;<lpage>6471</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Y.</given-names></name> <name><surname>Sanchez</surname> <given-names>V.</given-names></name> <name><surname>Guha</surname> <given-names>T.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Fusioncount: efficient crowd counting via multiscale feature fusion,&#x0201D;</article-title> in <source>Proceedings of the International Conference on Image Processing</source> (<publisher-loc>Bordeaux</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3256</fpage>&#x02013;<lpage>3260</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Z.</given-names></name> <name><surname>Wei</surname> <given-names>X.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Gong</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Bayesian loss for crowd count estimation with point supervision,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6141</fpage>&#x02013;<lpage>6150</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>M.</given-names></name> <name><surname>Laptev</surname> <given-names>I.</given-names></name> <name><surname>Sivic</surname> <given-names>J.</given-names></name> <name><surname>Audibert</surname> <given-names>J.-Y.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Density-aware person detection and tracking in crowds,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Barcelona</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2423</fpage>&#x02013;<lpage>2430</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shang</surname> <given-names>C.</given-names></name> <name><surname>Ai</surname> <given-names>H.</given-names></name> <name><surname>Bai</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;End-to-end crowd counting via joint learning local and global count,&#x0201D;</article-title> in <source>Proceedings of the International Conference on Image Processing</source> (<publisher-loc>Phoenix, AZ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1215</fpage>&#x02013;<lpage>1219</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv preprint</source> arXiv:1409.1556. <pub-id pub-id-type="doi">10.48550/arXiv.1409.1556</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sindagi</surname> <given-names>V. A.</given-names></name> <name><surname>Patel</surname> <given-names>V. M.</given-names></name></person-group> (<year>2018</year>). <article-title>A survey of recent advances in CNN-based single image crowd counting and density estimation</article-title>. <source>Pattern Recogn Lett</source>. <volume>107</volume>, <fpage>3</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2017.07.007</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stewart</surname> <given-names>R.</given-names></name> <name><surname>Andriluka</surname> <given-names>M.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;End-to-end people detection in crowded scenes,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2325</fpage>&#x02013;<lpage>2333</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wan</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Chan</surname> <given-names>A. B.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;A generalized loss function for crowd counting and localization,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1974</fpage>&#x02013;<lpage>1983</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00201</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Samaras</surname> <given-names>D.</given-names></name> <name><surname>Nguyen</surname> <given-names>M. H.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Distribution matching for crowd counting,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Beijing</publisher-loc>).</citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Cao</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Deep people counting in extremely dense crowds,&#x0201D;</article-title> in <source>MM 2015 - Proceedings of the 2015 ACM Multimedia Conference</source>, <fpage>1299</fpage>&#x02013;<lpage>1302</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Cai</surname> <given-names>H.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Gong</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>STNet: scale tree network with multi-level auxiliator for crowd counting</article-title>. <source>IEEE T Multimedia</source>. <volume>25</volume>, <fpage>2074</fpage>&#x02013;<lpage>2084</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2022.3142398</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>NWPU-crowd: a large-scale benchmark for crowd counting and localization</article-title>. <source>IEEE T Patt. Anal</source>. <volume>43</volume>, <fpage>2141</fpage>&#x02013;<lpage>2149</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2020.3013269</pub-id><pub-id pub-id-type="pmid">32750840</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Pyramid-dilated deep convolutional neural network for crowd counting</article-title>. <source>Appl. Intell</source>. <volume>52</volume>, <fpage>1825</fpage>&#x02013;<lpage>1837</lpage>. <pub-id pub-id-type="doi">10.1007/s10489-021-02537-6</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Niu</surname> <given-names>Q.</given-names></name></person-group> (<year>2019</year>). <article-title>Multi-sensor fusion in automated driving: a survey</article-title>. <source>IEEE Access</source>. <volume>8</volume>, <fpage>2847</fpage>&#x02013;<lpage>2868</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2962554</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Liang</surname> <given-names>D.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Zhan</surname> <given-names>W.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Autoscale: learning to scale for crowd counting</article-title>. <source>Int. J. Comput. Vision</source>. <volume>130</volume>, <fpage>405</fpage>&#x02013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-021-01542-z</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Ge</surname> <given-names>Z.</given-names></name> <name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Cui</surname> <given-names>G.</given-names></name> <name><surname>Lv</surname> <given-names>P.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Depth information guided crowd counting for complex crowd scenes</article-title>. <source>Pattern Recogn. Lett</source>. <volume>125</volume>, <fpage>563</fpage>&#x02013;<lpage>569</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2019.02.026</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Crowd counting via perspective-guided fractional-dilation convolution</article-title>. <source>IEEE T Multimedia</source>. <volume>24</volume>, <fpage>2633</fpage>&#x02013;<lpage>2647</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2021.3086709</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Su</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>Q.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Reverse perspective network for perspective-aware object counting,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4373</fpage>&#x02013;<lpage>4382</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Cross-scene crowd counting via deep convolutional neural networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>433</fpage>&#x02013;<lpage>841</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Perspective-guided point supervision network for crowd counting,&#x0201D;</article-title> in <source>2022 International Conference on High Performance Big Data and Intelligent Systems, HDIS 2022</source> (<publisher-loc>Tianjin</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>212</fpage>&#x02013;<lpage>217</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Kong</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>A cross-modal fusion based approach with scale-aware deep representation for RGB-D crowd counting and density estimation</article-title>. <source>Exp. Syst. Appl</source>. <volume>180</volume>, <fpage>115071</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115071</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Single-image crowd counting via multi-column convolutional neural network,&#x0201D;</article-title> in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>589</fpage>&#x02013;<lpage>597</lpage>.<pub-id pub-id-type="pmid">33315562</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Porikli</surname> <given-names>F.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Scale-aware crowd counting via depth-embedded convolutional neural networks</article-title>. <source>IEEE T Circ. Syst. Vid</source>. <volume>30</volume>, <fpage>3651</fpage>&#x02013;<lpage>3662</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2019.2943010</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhuang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Jia</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Tan</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Perception-aware multi-sensor fusion for 3D lidar semantic segmentation,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>16260</fpage>&#x02013;<lpage>16270</lpage>.</citation>
</ref>
</ref-list> 
</back>
</article> 