<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2022.731816</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Background-Aware Domain Adaptation for Plant Counting</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Shi</surname> <given-names>Min</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1639767/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Xing-Yi</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1376958/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Lu</surname> <given-names>Hao</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1149459/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Cao</surname> <given-names>Zhi-Guo</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/769885/overview"/>
</contrib>
</contrib-group>
<aff><institution>Key Laboratory of Image Processing and Intelligent Control, Ministry of Education, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology</institution>, <addr-line>Wuhan</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Kioumars Ghamkhar, AgResearch Ltd, New Zealand</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Mohsen Yoosefzadeh Najafabadi, University of Guelph, Canada; Youshan Zhang, Cornell University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Zhi-Guo Cao <email>zgcao&#x00040;hust.edu.cn</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Technical Advances in Plant Science, a section of the journal Frontiers in Plant Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>02</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>731816</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>01</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Shi, Li, Lu and Cao.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Shi, Li, Lu and Cao</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Deep learning-based object counting models have recently been considered preferable choices for plant counting. However, the performance of these data-driven methods would probably deteriorate when a discrepancy exists between the training and testing data. Such a discrepancy is also known as the domain gap. One way to mitigate the performance drop is to use unlabeled data sampled from the testing environment to correct the model behavior. This problem setting is also called unsupervised domain adaptation (UDA). Despite UDA has been a long-standing topic in machine learning society, UDA methods are less studied for plant counting. In this paper, we first evaluate some frequently-used UDA methods on the plant counting task, including feature-level and image-level methods. By analyzing the failure patterns of these methods, we propose a novel background-aware domain adaptation (BADA) module to address the drawbacks. We show that BADA can easily fit into object counting models to improve the cross-domain plant counting performance, especially on background areas. Benefiting from learning where to count, background counting errors are reduced. We also show that BADA can work with adversarial training strategies to further enhance the robustness of counting models against the domain gap. We evaluated our method on 7 different domain adaptation settings, including different camera views, cultivars, locations, and image acquisition devices. Results demonstrate that our method achieved the lowest Mean Absolute Error on 6 out of the 7 settings. The usefulness of BADA is also supported by controlled ablation studies and visualizations.</p></abstract>
<kwd-group>
<kwd>plant counting</kwd>
<kwd>maize tassels</kwd>
<kwd>rice plants</kwd>
<kwd>domain adaptation</kwd>
<kwd>adversarial training</kwd>
<kwd>local count models</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="11"/>
<table-count count="5"/>
<equation-count count="15"/>
<ref-count count="36"/>
<page-count count="16"/>
<word-count count="8716"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Estimating the number of plants accurately and efficiently is an important task in agriculture breeding and plant phenotyping. Counting plants (Liu et al., <xref ref-type="bibr" rid="B14">2020</xref>) or their flowers (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>) and fruits (Bargoti and Underwood, <xref ref-type="bibr" rid="B3">2017</xref>) can help farmers to monitor the status of crops and estimate yield. Recently, deep learning-based object counting models (Zhang et al., <xref ref-type="bibr" rid="B35">2015</xref>), which directly infer object counts from a single image, can be a promising choice for plant counting. Thanks to the strong representation ability of convolutional neural networks (CNNs), these methods can achieve high accuracy on standard plant counting datasets (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>; David et al., <xref ref-type="bibr" rid="B5">2020</xref>). It seems that the applications of counting models are around the corner. However, a vital problem has been neglected: the training data can be significantly different from the scenes where the counting models are deployed. Such a difference is given as a scientific term <italic>domain gap</italic>. In plant counting, various factors can contribute to domain gaps, e.g., different camera views, cultivars, object sizes or background. The performance of a counting model trained on one domain (source domain) usually deteriorates when tested on another domain (target domain) due to the domain gap. A straight-forward solution is to annotate additional data, while the consumption of time and labor is expensive. Naturally, one comes to the thought whether the unlabeled data in the target domain can be used to correct the model performance as much as possible. This problem setting is called unsupervised domain adaptation (UDA).</p>
<p>UDA is a long-standing topic in machine learning. A large number of task-specific UDA methods have been proposed for tasks such as semantic segmentation (Vu et al., <xref ref-type="bibr" rid="B26">2019</xref>), image classification (Ganin and Lempitsky, <xref ref-type="bibr" rid="B8">2015</xref>), and object detection (D&#x00027;Innocente et al., <xref ref-type="bibr" rid="B7">2020</xref>; Xu et al., <xref ref-type="bibr" rid="B31">2020</xref>). By contrast, UDA for object counting, especially for plant counting, has been less studied. To our knowledge, existing UDA methods (Giuffrida et al., <xref ref-type="bibr" rid="B10">2019</xref>; Ayalew et al., <xref ref-type="bibr" rid="B2">2020</xref>) applied to plant counting are often direct adoptions of generic UDA ideas without considering the particularities of domain gaps in plant counting. In fact, different from crowd counting or car counting, domain gaps in plant counting are much more diverse. The shapes of plants can change with time, cultivars and their growth environment; plants in different locations show different appearances; different image acquisition devices and viewpoints also intensify the domain gap. Considering that camera views and image perspectives are less diverse than those in crowd counting datasets, these factors make the domain adaptation for plant counting tricky. Some typical causes of domain gaps are shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Typical causes of domain gaps in plant counting. <bold>(A)</bold> Different camera views. <bold>(B)</bold> Scale variations in different locations. <bold>(C)</bold> Different appearances due to different growth stages.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0001.tif"/>
</fig>
<p>In this work, we first evaluated some frequently-used UDA methods in the context of plant counting and analyzed the weaknesses that these methods expose. In particular, we found that the counting models produce large errors on background areas that show similar appearances with the plants, e.g., similar colors or textures. Targeting these weaknesses, we propose a novel background-aware domain adaptation (BADA) module. This module can fit into existing plant counting models to enhance their cross-domain performance. Specifically, BADA is implemented as a parallel branch in the CNN model. This branch aims to segment areas which potentially contain counting objects, i.e., the foreground. The predicted foregrounds are merged into the feature maps as useful cues. In this way the network learns where to count. We also found that adding only a background-aware branch was insufficient to yield satisfactory cross-domain performance. Hence, two additional domain discriminators were connected to the input feature maps and the output foreground masks. We use adversarial training strategy to jointly optimize the discriminators and other parts of the model, facilitating to extract domain-invariant features and refine the predicted foreground masks.</p>
<p>We evaluated our method on three public datasets: MTC (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>), RPC (Liu et al., <xref ref-type="bibr" rid="B14">2020</xref>) and MTC-UAV (Lu et al., <xref ref-type="bibr" rid="B17">2021</xref>), including 7 different and representative domain adaptation settings close to real applications. We split data into different domains by different cultivars, locations, and image acquisition devices. It is worth noticing that one of our settings was to train a model using images captured by phenopoles and to test the model on images captured by UAVs. The results showed that, comparing with directly applying generic UDA ideas, our method achieved better cross-domain performance. We also verified each module of our method via ablation study. Moreover, the visualizations further show that our method significantly improves the performance on background areas.</p>
<p>Our contributions have two folds:</p>
<list list-type="bullet">
<list-item><p>We present a thorough evaluation of some frequently-used UDA methods under several plant counting tasks and analyze their weaknesses;</p></list-item>
<list-item><p>We propose a novel background-aware UDA module, which can easily fit into existing object counting models to prompt cross-domain performance.</p></list-item>
</list>
</sec>
<sec id="s2">
<title>2. Related work</title>
<p>In this section, we briefly review the applications of machine learning in plant science. Then we focus on the object counting methods and the unsupervised domain adaptation (UDA) methods in open literature.</p>
<p><bold>Machine Learning</bold>. Machine learning is a useful tool for plant science, which can model the relationships and patterns between targets and factors given a set of data. It is widely used in many none-destructive phenotyping tasks, e.g., field estimation (Yoosefzadeh-Najafabadi et al., <xref ref-type="bibr" rid="B34">2021</xref>) and plant identification (Tsaftaris et al., <xref ref-type="bibr" rid="B24">2016</xref>). A dominating trend in machine learning is deep learning, as deep learning models can learn to extract robust features and complete the tasks in a end-to-end manner. Deep learning-based methods have shown great advantages in different tasks of plant phenomics, e.g., plant counting (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>), detection (Bargoti and Underwood, <xref ref-type="bibr" rid="B3">2017</xref>; Madec et al., <xref ref-type="bibr" rid="B19">2019</xref>), segmentation (Tsaftaris et al., <xref ref-type="bibr" rid="B24">2016</xref>), and classification (Lu et al., <xref ref-type="bibr" rid="B15">2017a</xref>). For in-field plant counting tasks (from RGB images), deep learning-based methods show great robustness against different illuminations, scales and complex backgrounds (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>). The release of datasets (David et al., <xref ref-type="bibr" rid="B5">2020</xref>; Lu et al., <xref ref-type="bibr" rid="B17">2021</xref>) also accelerates the development of deep learning-based plant counting methods. Therefore, the deep learning has become the default choice for in-field plant counting.</p>
<p><bold>Object counting</bold>. Plant counting is a subset of object counting. Object counting aims to inference the number of target objects in the input images. Current cutting-edge object counting methods (Lempitsky and Zisserman, <xref ref-type="bibr" rid="B11">2010</xref>; Zhang et al., <xref ref-type="bibr" rid="B35">2015</xref>; Arteta et al., <xref ref-type="bibr" rid="B1">2016</xref>; Onoro-Rubio and L&#x000F3;pez-Sastre, <xref ref-type="bibr" rid="B21">2016</xref>; Li et al., <xref ref-type="bibr" rid="B12">2018</xref>; Ma et al., <xref ref-type="bibr" rid="B18">2019</xref>; Xiong et al., <xref ref-type="bibr" rid="B30">2019b</xref>; Wang et al., <xref ref-type="bibr" rid="B27">2020</xref>) utilize the power of deep learning and formulate the object counting problem as a regression task. A fully-convolutional neural network is trained to predict density maps (Lempitsky and Zisserman, <xref ref-type="bibr" rid="B11">2010</xref>) for target objects, where the value of each pixel denotes the local counting value. The integral of the density map is equal to the total number of objects. Inspired by the success of these methods in crowd counting, a constellation of methods (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>; Xiong et al., <xref ref-type="bibr" rid="B29">2019a</xref>; Liu et al., <xref ref-type="bibr" rid="B14">2020</xref>) and datasets (David et al., <xref ref-type="bibr" rid="B5">2020</xref>; Lu et al., <xref ref-type="bibr" rid="B17">2021</xref>) are proposed for plant counting. However, existing plant counting methods neglect the influence of domain gap, which is common in real applications.</p>
<p><bold>Unsupervised domain adaptation</bold>. The harm of domain gaps is common for data-driven methods (Ganin and Lempitsky, <xref ref-type="bibr" rid="B8">2015</xref>; Vu et al., <xref ref-type="bibr" rid="B26">2019</xref>). Therefore, UDA has been a long-standing topic in deep learning society, where unlabeled data collected in the target domain are utilized to prompt the model performance on the target domain. Ben-David et al. (<xref ref-type="bibr" rid="B4">2010</xref>) theoretically prove that domain adaptation can be achieved by narrowing the domain gap. One can achieve this from the feature level, or, more directly, from the image level. The feature-level methods (Ganin and Lempitsky, <xref ref-type="bibr" rid="B8">2015</xref>; Tzeng et al., <xref ref-type="bibr" rid="B25">2017</xref>) align the feature to be domain-invariant. And the image-level methods (Zhu et al., <xref ref-type="bibr" rid="B36">2017</xref>; Wang et al., <xref ref-type="bibr" rid="B28">2019</xref>; Yang and Soatto, <xref ref-type="bibr" rid="B33">2020</xref>; Yang et al., <xref ref-type="bibr" rid="B32">2020</xref>) manipulate the styles of images, e.g., hues, illuminations, textures to make the images in two different domains closer. Some of the UDA methods are proposed to address the domain gap for plant counting (Giuffrida et al., <xref ref-type="bibr" rid="B10">2019</xref>; Ayalew et al., <xref ref-type="bibr" rid="B2">2020</xref>). However, existing UDA methods for plant counting directly adopt the generic feature-level UDA methods. This motivates us to test different UDA methods under the context of plant counting.</p>
</sec>
<sec sec-type="materials and methods" id="s3">
<title>3. Materials and Methods</title>
<sec>
<title>3.1. Plant Counting Datasets</title>
<p>We evaluated the performance of UDA on three public plant counting datasets: Maize Tassel Counting (MTC) dataset (Lu et al., <xref ref-type="bibr" rid="B16">2017b</xref>), Rice Plant Counting (RPC) dataset (Liu et al., <xref ref-type="bibr" rid="B14">2020</xref>) and Maize Tassel Counting UAV (MTC-UAV) (Lu et al., <xref ref-type="bibr" rid="B17">2021</xref>) dataset. Here, we briefly introduce the statistics and characteristics of these datasets.</p>
<sec>
<title>3.1.1. The MTC Dataset</title>
<p>The MTC dataset contains 361 images of maize fields. Each center of maize tassel is manually annotated with a dot. The samples were collected from 4 different places in China, including 6 different maize cultivars. We split the dataset into 6 domains according to cultivars. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, domain gaps not only reflect in the different shapes of maize tassels, but also reflect in different backgrounds, illuminations and camera views.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Samples in MTC datasets. <bold>(A)</bold> Images captured at different locations. Camera views, backgrounds and illuminations are different. <bold>(B&#x02013;G)</bold> Maize tassels of different cultivars.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.1.2. The RPC Dataset</title>
<p>The RPC dataset contains 382 images of rice seedlings captured in Jiangxi, China and Guangxi, China. The rice seedlings are manually annotated with dots. We split the dataset into 2 domains according to locations. For samples from Guangxi, the images were captured shortly after the rice seedlings were transplanted, while most of the rice seedlings in Jiangxi had been growing for some time. Thus, rice seedlings in Guangxi were much smaller and with less occlusions. On the contrary, rice seedlings in Jiangxi had grown more leaves and block each other. Besides, the hues and camera views are very different, images from Guangxi show dimmer illuminations and hues. We show some typical samples in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Samples from RPC datasets. <bold>(A)</bold> Images captured in Guangxi. <bold>(B)</bold> Images captured in Jiangxi.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0003.tif"/>
</fig>
</sec>
<sec>
<title>3.1.3. The MTC-UAV Dataset</title>
<p>The MTC-UAV dataset is very different from the two aforementioned plant counting datasets, as the samples were captured by an unmanned aircraft vehicle (UAV). The UAV took 306 pictures of an experimental field which covered around 1 ha. Images were captured at the height of 12.5 m, and the focal length of the camera was 28 mm. Thus, the ground sampling resolution is about 0.3 cm/pixel.</p>
<p>This dataset was adopted to evaluate the UDA performance between different image acquisition devices. This setup is challenging as camera views, perspectives, and object scales in images captured by a UAV are significantly different from those of the images captured by phenopoles.</p>
</sec>
</sec>
<sec>
<title>3.2. Background-Aware Domain Adaptation</title>
<p>Assume that we have two domains of data under different distributions: labeled data from the source domain and unlabeled data from the target domain. Labeled data from source domain can be denoted by <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>Y</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, where <bold>X</bold><sub><italic>s</italic></sub> denotes the images in the source domain and <bold>Y</bold><sub><italic>s</italic></sub> stores the point annotations for each image. Unlabeled data is denoted by <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. UDA for plant counting aims at jointly utilizing <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to prompt counting performance on the target domain.</p>
<p>We verified our BADA module on a popular and straight-forward object counting method CSRNet (Li et al., <xref ref-type="bibr" rid="B12">2018</xref>). For convenience, we first define the variables in <xref ref-type="table" rid="T1">Table 1</xref> and the I/O of each module in <xref ref-type="table" rid="T2">Table 2</xref>, where <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula><mml:math id="M6"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M10"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are parameterized by &#x003B8;<sub><italic>E</italic></sub>, &#x003B8;<sub><italic>D</italic></sub>, &#x003B8;<sub><italic>S</italic></sub>, &#x003B8;<sub><italic>C</italic></sub>, &#x003B8;<sub><italic>F</italic></sub> and &#x003B8;<sub><italic>M</italic></sub>, respectively. [<italic>M</italic><sub><italic>s</italic></sub>, <italic>M</italic><sub><italic>c</italic></sub>] denotes the channel-wise concatenation of <italic>M</italic><sub><italic>s</italic></sub> and <italic>M</italic><sub><italic>c</italic></sub>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Definition of variables.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Variable</bold></th>
<th valign="top" align="left"><bold>Symbol</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Input image</td>
<td valign="top" align="left"><italic>I</italic></td>
</tr>
<tr>
<td valign="top" align="left">Source image</td>
<td valign="top" align="left"><italic>I</italic><sub><italic>s</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Target image</td>
<td valign="top" align="left"><italic>I</italic><sub><italic>t</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Basic feature maps</td>
<td valign="top" align="left"><italic>M</italic><sub><italic>f</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Counting feature maps</td>
<td valign="top" align="left"><italic>M</italic><sub><italic>c</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Estimated foreground mask</td>
<td valign="top" align="left"><italic>M</italic><sub><italic>s</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Estimated local count map</td>
<td valign="top" align="left"><italic>C</italic><sub><italic>est</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Domain class map for feature map</td>
<td valign="top" align="left"><italic>C</italic><sub><italic>f</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left">Domain class map for foreground mask</td>
<td valign="top" align="left"><italic>C</italic><sub><italic>m</italic></sub></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>I/O for each module.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Module</bold></th>
<th valign="top" align="left"><bold>Symbol</bold></th>
<th valign="top" align="left"><bold>I/O function</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Feature extractor</td>
<td valign="top" align="left"><inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M12"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Counting feature decoder</td>
<td valign="top" align="left"><inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M14"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Segmentation branch</td>
<td valign="top" align="left"><inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Local count regressor</td>
<td valign="top" align="left"><inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M18"><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Feature discriminator</td>
<td valign="top" align="left"><inline-formula><mml:math id="M19"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M20"><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Foreground mask discriminator</td>
<td valign="top" align="left"><inline-formula><mml:math id="M21"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the input of the whole model is an RGB image. The image is first processed by the feature encoder <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to obtain feature maps <italic>M</italic><sub><italic>f</italic></sub>. Then, the extracted feature maps <italic>M</italic><sub><italic>f</italic></sub> are sent to the counting feature decoder <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the segmentation branch <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> further refines the feature maps to generate the counting feature maps <italic>M</italic><sub><italic>c</italic></sub>. And the segmentation branch segments the regions which potentially contain the counting objects, i.e., the foreground mask. The foreground mask <italic>M</italic><sub><italic>s</italic></sub> is then concatenated with <italic>M</italic><sub><italic>c</italic></sub> to form the input of local count regressor <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M28"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> outputs the local count map for the input image.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>The overview of our method. The BADA module works as a parallel branch in the CNN model. Two discriminators are connected to the input and output of BADA model and are imposed with adversarial training strategies.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0004.tif"/>
</fig>
<p>To extract domain-invariant features, we applied two domain discriminators, including a feature discriminator <inline-formula><mml:math id="M29"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and a mask discriminator <inline-formula><mml:math id="M30"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The discriminators are fully-convolutional, which receive the feature map <italic>M</italic><sub><italic>f</italic></sub> and the foreground mask <italic>M</italic><sub><italic>s</italic></sub> as inputs, and output domain class maps. The adversarial training strategy was imposed on the discriminators. Segmentation branch <inline-formula><mml:math id="M31"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, feature discriminator <inline-formula><mml:math id="M32"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and mask discriminator <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> together constitute the BADA module.</p>
<p>To train the network, we jointly optimized three loss functions: counting loss, segmentation loss and the adversarial training loss.</p>
<sec>
<title>3.2.1. Feature Encoder</title>
<p>We adopted part of the VGG16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B23">2014</xref>) network as the feature encoder. As shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, the feature encoder includes 3 stride-2 max pooling layers. Given an image of size <italic>H</italic> &#x000D7; <italic>W</italic>, the feature encoder outputs features maps <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mn>512</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>. At the beginning of the training process, the feature encoder was initialized by parameters pretrained on ImageNet (Deng et al., <xref ref-type="bibr" rid="B6">2009</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The architecture of our model. <bold>(A)</bold> The architecture of the feature encoder. <bold>(B)</bold> The architecture of the multi-branch decoder and the local count regressor. <bold>(C,D)</bold> The architecture of the feature discriminator and the foreground mask discriminator.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0005.tif"/>
</fig>
</sec>
<sec>
<title>3.2.2. Multi-Branch Decoder</title>
<p>The multi-branch decoder consists of a counting feature decoder <inline-formula><mml:math id="M35"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and a segmentation branch <inline-formula><mml:math id="M36"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M38"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> share almost the same network architecture. Both the two branches replace standard convolutions with dilated convolutions, which can enlarge the receptive fields without introducing extra parameters.</p>
<p>As shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, <italic>M</italic><sub><italic>f</italic></sub> is sent to <inline-formula><mml:math id="M39"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M40"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M41"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> outputs feature maps <italic>M</italic><sub><italic>c</italic></sub> with 64 channels. The last <monospace>softmax</monospace> layer of the <inline-formula><mml:math id="M42"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> outputs a 2-channel segmentation map, where each pixel can be viewed as a 2-d vector. The second element refers to the probability of a pixel being the foreground. We denote the second channel as the segmentation mask <italic>M</italic><sub><italic>s</italic></sub>. <italic>M</italic><sub><italic>s</italic></sub> and <italic>M</italic><sub><italic>c</italic></sub> are concatenated to form a feature map with 65 channels as the output of multi-branch decoder.</p>
</sec>
<sec>
<title>3.2.3. Local Count Regressor</title>
<p>Most object counting methods are based on density map regression, which predicts the counting value pixel by pixel. However, this paradigm is not robust to shape variations of non-rigid objects in plant counting, e.g., maize tassels or rice seedlings. In plant counting, shape and appearance of an object often change with different growth stages and cultivars. Density map-based methods tend to generate responses at every pixel that shares similar patterns with the counting objects. Thus, per-pixel density estimation often leads to accumulated error when summing the density map. To alleviate this, we followed Lu et al. (<xref ref-type="bibr" rid="B16">2017b</xref>) to estimate patch-wise counting values. As shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, the local count regressor <inline-formula><mml:math id="M43"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> includes two average pooling layers with the stride of 2 and 4, respectively. Thus, given an image of size <italic>H</italic> &#x000D7; <italic>W</italic>, the spatial resolution of the estimated local count map <italic>C</italic><sub><italic>est</italic></sub> is <inline-formula><mml:math id="M44"><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>. Each element in <italic>C</italic><sub><italic>est</italic></sub> denotes the number of counting objects in a 64 &#x000D7; 64 patch of the input image.</p>
</sec>
<sec>
<title>3.2.4. Discriminator</title>
<p>Adding a segmentation branch can guide the network to learn where to count (Lu et al., <xref ref-type="bibr" rid="B17">2021</xref>; Modolo et al., <xref ref-type="bibr" rid="B20">2021</xref>). Nevertheless, under the cross-domain setting, the segmentation branch also suffers from the domain gap. The foreground masks may also contain some false positives. Thus, we added two domain discriminators: a feature discriminator <inline-formula><mml:math id="M45"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and a foreground mask discriminator <inline-formula><mml:math id="M46"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M47"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> aims to force <italic>M</italic><sub><italic>f</italic></sub> extracted by <inline-formula><mml:math id="M48"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to be domain-invariant. <inline-formula><mml:math id="M49"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can help the segmentation branch to predict foreground mask <italic>M</italic><sub><italic>s</italic></sub> with reasonable shapes and high accuracy on both the source domain and target domain. This is motivated by the observation that the shape of the foreground mask is irregular and scattered when directly applying the model on target domain without discriminators. Readers can refer to section 4.4.2 for detailed visualizations.</p>
<p>The architectures of <inline-formula><mml:math id="M50"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M51"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. The <monospace>softmax</monospace> layer outputs a domain class map <italic>C</italic><sub><italic>f</italic></sub> (<italic>C</italic><sub><italic>m</italic></sub>). Each element in <italic>C</italic><sub><italic>f</italic></sub> (<italic>C</italic><sub><italic>m</italic></sub>) can be viewed as a 2-d vector, and the first element in the vector denotes the probability of the corresponding 4 &#x000D7; 4 patch in <italic>M</italic><sub><italic>f</italic></sub> (<italic>M</italic><sub><italic>s</italic></sub>) being the target domain. Similarly, the second dimension denotes the probability being the source domain.</p>
<p>To train the discriminator, we adopted the adversarial training strategy (Ganin and Lempitsky, <xref ref-type="bibr" rid="B8">2015</xref>). While discriminators <inline-formula><mml:math id="M52"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M53"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> learn to classify <italic>M</italic><sub><italic>f</italic></sub> and <italic>M</italic><sub><italic>s</italic></sub> into source and target domains, <inline-formula><mml:math id="M54"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M55"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> attempt to confuse the discriminators by generating domain-invariant <italic>M</italic><sub><italic>f</italic></sub> and <italic>M</italic><sub><italic>s</italic></sub>. This can be achieved by adding a gradient reversal layer (Ganin and Lempitsky, <xref ref-type="bibr" rid="B8">2015</xref>) before the input layers of the two discriminators. During forward propagation, the gradient reversal layer passes the input to the next layer with no change, but reverses the sign of the gradient during back propagation. The operation rule of the gradient reversal layer can be defined by</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M56"><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mtext>x</mml:mtext><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mtext>x</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mtext>x</mml:mtext></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where x denotes the input of the gradient reverse layer, and <bold>I</bold> denotes the identity matrix. &#x003BB; is a pre-defined parameter which adjusts the attenuation ratio when propagating the gradients back. This is useful as the adversarial training could interfere with the main task (counting) at the beginning of the training process. We will discuss the updating strategy of &#x003BB; in section 3.2.6.</p>
</sec>
<sec>
<title>3.2.5. Loss Function</title>
<p>1) Counting loss</p>
<p>The counting loss <inline-formula><mml:math id="M57"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is used to measure the differences between estimated local count maps <italic>C</italic><sub><italic>est</italic></sub> and the ground truth local count maps <italic>C</italic><sub><italic>gt</italic></sub>. One can obtain <italic>C</italic><sub><italic>gt</italic></sub> from the ground truth density map <italic>D</italic><sub><italic>gt</italic></sub>. Supposing the image <italic>I</italic><sub><italic>i</italic></sub> have <italic>n</italic> annotated points <italic>P</italic> &#x02208; &#x0211D;<sup><italic>n</italic>&#x000D7;2</sup> and the corresponding density map can be defined by</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M58"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M59"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> denotes a 2-d Gaussian kernel with the mean <italic>P</italic><sub><italic>k</italic></sub> and the variance &#x003C3;<sup>2</sup>. Then the ground truth local count map <italic>C</italic><sub><italic>gt</italic></sub> can be obtained by</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M60"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002A;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>&#x0002A;<bold>1</bold><sub><italic>h</italic> &#x000D7; <italic>w</italic></sub> denotes the convolution operation using a <italic>h</italic> &#x000D7; <italic>w</italic> matrix with all ones as kernel. The horizontal and vertical strides are <italic>h</italic> and <italic>w</italic>, respectively. In our method, we set <italic>h</italic> &#x0003D; 64 and <italic>w</italic> &#x0003D; 64. Then, we define the counting loss by:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M61"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> &#x0003D; <italic>H</italic> &#x000B7; <italic>W</italic>, i.e., the number of pixels in the local count map.</p>
<p>2) Segmentation loss</p>
<p>Akin to semantic segmentation (Lin et al., <xref ref-type="bibr" rid="B13">2017</xref>), the foreground segmentation can be viewed as a 2-class semantic segmentation task, and can be supervised by the cross-entropy loss. However, pixel-wise foreground labels are not available in plant counting datasets. Thus, we generated pseudo foreground masks <italic>S</italic><sub><italic>gt</italic></sub> from ground truth density maps. <italic>S</italic><sub><italic>gt</italic></sub> is obtained by</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M62"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02265;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>t</italic><sub><italic>c</italic></sub> is a pre-defined threshold. For different datasets, <italic>t</italic><sub><italic>c</italic></sub> can be adjusted conditioned on the empirical estimate of object size to make sure that every counting object can be fully covered by the foreground mask.</p>
<p>The standard cross-entropy loss was adopted as the segmentation loss. Given the estimated foreground mask <italic>M</italic><sub><italic>s</italic></sub> and the ground truth <italic>S</italic><sub><italic>gt</italic></sub>, the segmentation loss <inline-formula><mml:math id="M63"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be formulated by</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M64"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> &#x0003D; <italic>H</italic> &#x000B7; <italic>W</italic>, i.e., the number of pixels in the foreground mask.</p>
<p>3) Loss for adversarial training</p>
<p>The adversarial training loss function <inline-formula><mml:math id="M65"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> supervises the training of domain discriminators. We labeled the source domain as 1 and the target domain as 0. Then, the ground truth domain class map <italic>A</italic><sub><italic>gt</italic></sub> can be obtained by</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M66"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>1</mml:mn></mml:mstyle><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>0</mml:mn></mml:mstyle><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <bold>1</bold> and <bold>0</bold> denote matrices filled with ones and zeros. <italic>I</italic> &#x02208; <italic>I</italic><sub><italic>t</italic></sub> denotes the image <italic>I</italic> belongs to the source domain, and <italic>I</italic> &#x02208; <italic>I</italic><sub><italic>t</italic></sub> means <italic>I</italic> comes from the target domain.</p>
<p>Let the second channel of <italic>C</italic><sub><italic>f</italic></sub> and <italic>C</italic><sub><italic>m</italic></sub> be <italic>A</italic><sub><italic>est</italic></sub>, i.e., the probability that the feature maps (foreground masks) are from the source domain. Then, the adversarial training loss <inline-formula><mml:math id="M67"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is defined by</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M68"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x02112;</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mstyle><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> &#x0003D; <italic>H</italic> &#x000B7; <italic>W</italic>.</p>
</sec>
<sec>
<title>3.2.6. Implementation Details</title>
<p>1) Training details</p>
<p>We used <monospace>PyTorch</monospace> (Paszke et al., <xref ref-type="bibr" rid="B22">2019</xref>) to train and evaluate our model. Stochastic gradient descent (SGD) was adopted as the optimizer. We trained the datasets for 500 epochs. The initial learning rate was set to 0.01, and at the 250th and the 400th epoch, the learning rate decayed by 10 times.</p>
<p>As the resolution of samples was high, images were resized during training and evaluation. For data augmentation, 512 &#x000D7; 512 patches were randomly cropped from resized images, and then the cropped images were flipped along horizontal directions randomly.</p>
<p>2) Parameters update</p>
<p>Here we specify the parameter updating strategy during training. At each epoch, <inline-formula><mml:math id="M69"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M70"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> were jointly optimized while <inline-formula><mml:math id="M71"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> was optimized separately. The detailed updating strategy is defined in <xref ref-type="table" rid="T6">Algorithm 1</xref>.</p>
<table-wrap position="float" id="T6">
<label>Algorithm 1</label>
<caption><p>Parameters updating strategy</p></caption>
<graphic xlink:href="fpls-13-731816-i0001.tif"/>
</table-wrap>
<p>At the beginning of each epoch, &#x003BB; of the gradient reverse layer was updated by</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M76"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003BB;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>p</italic> denotes the ratio of the current epoch to total epochs. And &#x003B3; denotes a pre-defined parameter that controls the speed when &#x003BB; ascends. As the training proceeds, &#x003BB; increases from 0 to 1.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments</title>
<p>Here we report the experiments results. We first evaluated multiple UDA methods on 7 different domain adaptation settings. The results were compared with our method. The efficiency of each module in our method was verified via ablation study. We also conducted visualizations to show the qualitative results of our method. First, we introduce the evaluation metrics.</p>
<sec>
<title>4.1. Evaluation Metrics</title>
<p>We used mean absolute error (MAE) and root mean square error (MSE) as the main evaluation metrics, which can be defined by:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M77"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E11"><label>(11)</label><mml:math id="M78"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> denotes the number of samples on the test set. <inline-formula><mml:math id="M79"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> and <italic>y</italic><sub><italic>n</italic></sub> denote the estimated count and the ground truth count of the <italic>n</italic><sup><italic>th</italic></sup> sample.</p>
<p>To measure the ratio of counting error to the total count of each sample, we used mean absolute percentage error (MAPE), which can be calculated by:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M80"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mn>100</mml:mn><mml:mi>%</mml:mi><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In addition, we measured the correlation between estimated counts and annotations by <italic>R</italic><sup>2</sup>:</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M81"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:mover accent='true'><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover></mml:mrow></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:msup><mml:mo stretchy='false'>]</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy='false'>(</mml:mo></mml:mstyle><mml:msub><mml:mover accent='true'><mml:mi>y</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>We also noticed that, the false positive responses in the estimated density maps may compensate for errors from missing targets. This indicated that MAE may not fully reflect the real performance of counting models. Therefore, we designed a decoupled MAE where errors on target areas and background areas are calculated independently and then summed up, instead of directly comparing the total counts. For example, if the model wrongly predicts density responses on background and omits some targets. The density responses on background will not compensate for the error on real targets when calculating metrics. To be specific, DMAE is defined as follows,</p>
<disp-formula id="E14"><label>(14)</label><mml:math id="M82"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>&#x00177;<sub><italic>b,n</italic></sub> and &#x00177;<sub><italic>f,n</italic></sub> denote the estimated object count in the background areas and target areas. Similarly, <italic>y</italic><sub><italic>b,n</italic></sub> and <italic>y</italic><sub><italic>f,n</italic></sub> denote the ground truth count in the background areas and target areas. To obtain &#x00177;<sub><italic>b,n</italic></sub> and &#x00177;<sub><italic>f,n</italic></sub>, we used the same pesudo segmentation mask <italic>S</italic><sub><italic>gt</italic></sub> mentioned in section 3.2.5 to divide the image into background areas <italic>B</italic> and target areas. This process can be defined as follows,</p>
<disp-formula id="E15"><label>(15)</label><mml:math id="M83"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>B</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x00177;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02209;</mml:mo><mml:mi>B</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>y</italic><sub><italic>b,n</italic></sub> and <italic>y</italic><sub><italic>f,n</italic></sub> can be obtained likewise.</p>
</sec>
<sec>
<title>4.2. Experimental Settings</title>
<p>Here we specify the experimental settings, including the split of source domain and the introduction of other tested algorithm.</p>
<sec>
<title>4.2.1. The MTC Dataset</title>
<p>We split the MTC dataset according to cultivars. As Zhengdan No.958 contains more samples while samples of other cultivars are much fewer, we used Zhengdan No.958 as the source domain, and the other 5 cultivars as the target domains. Accordingly, there were 5 different adaptation pairs for MTC dataset. For convenience, we named the adaptation pairs by abbreviation, e.g., adaption from Zhengdan No.958 to Jundan No.20 was marked as Z&#x02192;Jun. The abbreviations for cultivars Zhengdan No.958, Jundan No.20, Wuyue No.3, Jidan No.32, Tianlong No.9 and Nongda No.108 were Z, Jun, W, Ji, T and N, respectively.</p>
</sec>
<sec>
<title>4.2.2. The RPC Dataset</title>
<p>We split the RPC dataset into two domains according to different locations. Since only 62 images were captured from Guangxi, we adapted the model from Jiangxi to Guangxi, marking this setting as J&#x02192;G.</p>
</sec>
<sec>
<title>4.2.3. The MTC-UAV Dataset</title>
<p>The MTC-UAV dataset and the MTC dataset shared the same counting object. We used data from MTC dataset as source domain and data from MTC-UAV dataset as target domain to constitute a domain adaptation setting.</p>
</sec>
</sec>
<sec>
<title>4.3. Comparison With Other Methods</title>
<p>As UDA for plant counting has seldom been studied, we first evaluated some frequently-used UDA methods. We trained these methods on the plant counting datasets using official implementations (CSRNet, FDA, PCEDA) when available. If no codes are released, we implement the method according to the their papers (CSRNet_DA, MFA).</p>
<sec>
<title>4.3.1. Baseline Approaches</title>
<p>1) CSRNet</p>
<p>CSRNet (Li et al., <xref ref-type="bibr" rid="B12">2018</xref>) is a generic object counting method with simple network architecture and competitive performance. For a fair comparison, all the UDA methods compared were based on CSRNet. We trained the counting model with only source data and directly evaluated the model on the target domain.</p>
<p>2) CSRNet_DA</p>
<p>CSRNet_DA refers to a na&#x000EF;ve upgrade of CSRNet. We added a discriminator for CSRNet and applied adversarial training strategy discussed in section 3.2.5. The discriminator receives the features extracted by decoder as input and outputs domain class maps.</p>
<p>3) Multi-level feature-aware domain adaptation</p>
<p>Multi-level Feature Aware (MFA) domain adaption is a feature-level UDA method purposed by Gao et al. (<xref ref-type="bibr" rid="B9">2021</xref>). Multi-level refers to a setup where the adversarial training is conducted on 2 intermediate feature maps and the estimated density maps. Specifically, two discriminators are connected to the output of VGG16 backbone and the output of the decoder.</p>
<p>4) PCEDA</p>
<p>PCEDA is an image-level unsupervised domain adaptation method based on Cycle GAN framework (Zhu et al., <xref ref-type="bibr" rid="B36">2017</xref>). Most image-level domain adaptation methods are designed for adaptation between synthetic data and real-world data. Since evident and unified style differences exist between computer-rendered images and real-world images, directly applying GAN to transfer images between two real-world domains could produce many artifacts. To alleviate this, we used PCEDA (Yang et al., <xref ref-type="bibr" rid="B32">2020</xref>), which preserves the high-frequency details of the source images, to evaluate the GAN-based UDA method.</p>
<p>PCEDA adds a phase consistency constraint between the original images and the transferred images. Fourier transform of an image consists of phase and amplitude, and the phase contains the semantic information (edges, textures) of the image. The phase consistency requires the phases of the original and transferred images to be close. Thus, instead of manipulating the shapes or textures, the generator tends to transfer the illuminations, hues or colors to target domain. For different domain adaptation setups, we used the official implementation to transfer source images to the target domain, and used the transferred images to train CSRNet, and directly evaluated the model on target data.</p>
<p>5) Fourier domain adaptation</p>
<p>Fourier Domain Adaptation (FDA) (Yang and Soatto, <xref ref-type="bibr" rid="B33">2020</xref>) is image-level UDA method which does not need to train a complex GAN. The transfer process is achieved by swapping low frequency spectrums of two images. This simple procedure can achieve comparable performance on UDA semantic segmentation benchmarks against GAN-based methods.</p>
</sec>
<sec>
<title>4.3.2. Comparison on the MTC Dataset</title>
<p><xref ref-type="table" rid="T3">Table 3</xref> presents the quantitative comparison of aforementioned methods on 5 different domain adaptation settings of the MTC dataset. Comparing with the non-adaptation method CSRNet, all UDA methods more or less reduced the MAE, MSE as well as the MAPE. We also noticed that, even with comparable MAE (Z&#x02192;W), the UDA methods can improve the DMAE by a large margin, indicating that UDA methods can also generate more correct density maps. Then we focused on the comparison between different UDA methods. Averaging the performance of five settings, the proposed method obtained the best MAE, MSE, MAPE and DMAE. Comparing with the second best, our method brought a relative improvement 42% on the DMAE. For different domain adaptation settings, our method obtained the best MAE except for Z&#x02192;W. It can also be observed that our method was more stable under different settings. On a difficult setting Z&#x02192;T, BADA reduced the MAE and DMAE by 43% and 56% comparing with the second best method. Domain gap under Z&#x02192;T is dramatic due to different viewpoints, illuminations and background elements. We believe results under Z&#x02192;T setup can better reflect the adaptation effectiveness of UDA methods. The visualizations of different methods on MTC dataset are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Quantitative comparisons on MTC dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Settings</bold></th>
<th valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Z</bold>&#x02192;<bold>Jun</bold></th>
<th valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Z</bold>&#x02192;<bold>W</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>methods</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
<th valign="top" align="center"><bold>MAPE</bold></th>
<th valign="top" align="center"><bold>DMAE</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
<th valign="top" align="center"><bold>MAPE</bold></th>
<th valign="top" align="center"><bold>DMAE</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CSRNet</td>
<td valign="top" align="center">5.89</td>
<td valign="top" align="center">8.22</td>
<td valign="top" align="center">77.5%</td>
<td valign="top" align="center">6.52</td>
<td valign="top" align="center">0.9682</td>
<td valign="top" align="center">3.56</td>
<td valign="top" align="center">4.56</td>
<td valign="top" align="center">5.9%</td>
<td valign="top" align="center">45.23</td>
<td valign="top" align="center">0.9106</td>
</tr>
<tr>
<td valign="top" align="left">CSRNet_DA</td>
<td valign="top" align="center">2.38</td>
<td valign="top" align="center">3.33</td>
<td valign="top" align="center">13.7%</td>
<td valign="top" align="center">7.56</td>
<td valign="top" align="center">0.9864</td>
<td valign="top" align="center">5.46</td>
<td valign="top" align="center">7.41</td>
<td valign="top" align="center">9.9%</td>
<td valign="top" align="center">15.77</td>
<td valign="top" align="center">0.8173</td>
</tr>
<tr>
<td valign="top" align="left">PCEDA</td>
<td valign="top" align="center">4.53</td>
<td valign="top" align="center">6.58</td>
<td valign="top" align="center">41.3%</td>
<td valign="top" align="center">10.68</td>
<td valign="top" align="center">0.9369</td>
<td valign="top" align="center">3.47</td>
<td valign="top" align="center">2.46</td>
<td valign="top" align="center">5.7%</td>
<td valign="top" align="center">29.86</td>
<td valign="top" align="center">0.9388</td>
</tr>
<tr>
<td valign="top" align="left">FDA</td>
<td valign="top" align="center">4.92</td>
<td valign="top" align="center">6.44</td>
<td valign="top" align="center">57.0%</td>
<td valign="top" align="center">6.23</td>
<td valign="top" align="center">0.9853</td>
<td valign="top" align="center"><bold>2.48</bold></td>
<td valign="top" align="center"><bold>2.99</bold></td>
<td valign="top" align="center"><bold>4.3%</bold></td>
<td valign="top" align="center"><bold>5.92</bold></td>
<td valign="top" align="center"><bold>0.9573</bold></td>
</tr>
<tr>
<td valign="top" align="left">MFA</td>
<td valign="top" align="center">4.02</td>
<td valign="top" align="center">5.65</td>
<td valign="top" align="center">37.1%</td>
<td valign="top" align="center">6.11</td>
<td valign="top" align="center">0.9655</td>
<td valign="top" align="center">3.76</td>
<td valign="top" align="center">4.72</td>
<td valign="top" align="center">6.6%</td>
<td valign="top" align="center">9.17</td>
<td valign="top" align="center">0.9463</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>1.92</bold></td>
<td valign="top" align="center"><bold>2.83</bold></td>
<td valign="top" align="center"><bold>10.1%</bold></td>
<td valign="top" align="center"><bold>3.78</bold></td>
<td valign="top" align="center"><bold>0.9884</bold></td>
<td valign="top" align="center">3.83</td>
<td valign="top" align="center">5.13</td>
<td valign="top" align="center">6.7%</td>
<td valign="top" align="center">7.98</td>
<td valign="top" align="center">0.9041</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left"><bold>Settings</bold></td>
<td valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Z</bold>&#x02192;<bold>Ji</bold></td>
<td valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Z</bold>&#x02192;<bold>T</bold></td>
</tr>
<tr>
<td valign="top" align="left"><bold>methods</bold></td>
<td valign="top" align="center"><bold>MAE</bold></td>
<td valign="top" align="center"><bold>MSE</bold></td>
<td valign="top" align="center"><bold>MAPE</bold></td>
<td valign="top" align="center"><bold>DMAE</bold></td>
<td valign="top" align="center"><italic>R</italic><sup>2</sup></td>
<td valign="top" align="center"><bold>MAE</bold></td>
<td valign="top" align="center"><bold>MSE</bold></td>
<td valign="top" align="center"><bold>MAPE</bold></td>
<td valign="top" align="center"><bold>DMAE</bold></td>
<td valign="top" align="center"><italic>R</italic><sup>2</sup></td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">CSRNet</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">1.16</td>
<td valign="top" align="center">10.3%</td>
<td valign="top" align="center">14.2</td>
<td valign="top" align="center">0.9776</td>
<td valign="top" align="center">15.76</td>
<td valign="top" align="center">19.42</td>
<td valign="top" align="center">134.9%</td>
<td valign="top" align="center">35.87</td>
<td valign="top" align="center">0.9039</td>
</tr>
<tr>
<td valign="top" align="left">CSRNet_DA</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">10.9%</td>
<td valign="top" align="center">5.11</td>
<td valign="top" align="center">0.9869</td>
<td valign="top" align="center">12.38</td>
<td valign="top" align="center">15.14</td>
<td valign="top" align="center">102.9%</td>
<td valign="top" align="center">26.32</td>
<td valign="top" align="center">0.9275</td>
</tr>
<tr>
<td valign="top" align="left">PCEDA</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">1.27</td>
<td valign="top" align="center">12.8%</td>
<td valign="top" align="center">6.76</td>
<td valign="top" align="center">0.9752</td>
<td valign="top" align="center">16.61</td>
<td valign="top" align="center">23.09</td>
<td valign="top" align="center">116.9%</td>
<td valign="top" align="center">39.83</td>
<td valign="top" align="center">0.6549</td>
</tr>
<tr>
<td valign="top" align="left">FDA</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center"><bold>9.1%</bold></td>
<td valign="top" align="center">1.82</td>
<td valign="top" align="center">0.9856</td>
<td valign="top" align="center">12.29</td>
<td valign="top" align="center">16.14</td>
<td valign="top" align="center">138.0%</td>
<td valign="top" align="center">28.35</td>
<td valign="top" align="center"><bold>0.9312</bold></td>
</tr>
<tr>
<td valign="top" align="left">MFA</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">1.09</td>
<td valign="top" align="center">14.2%</td>
<td valign="top" align="center"><bold>1.81</bold></td>
<td valign="top" align="center">0.9762</td>
<td valign="top" align="center">13.77</td>
<td valign="top" align="center">16.81</td>
<td valign="top" align="center">94.8%</td>
<td valign="top" align="center">33.93</td>
<td valign="top" align="center">0.8567</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>0.50</bold></td>
<td valign="top" align="center"><bold>0.69</bold></td>
<td valign="top" align="center">9.5%</td>
<td valign="top" align="center">2.92</td>
<td valign="top" align="center"><bold>0.9939</bold></td>
<td valign="top" align="center"><bold>6.96</bold></td>
<td valign="top" align="center"><bold>9.33</bold></td>
<td valign="top" align="center"><bold>31.9%</bold></td>
<td valign="top" align="center"><bold>11.61</bold></td>
<td valign="top" align="center">0.9115</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left"><bold>Settings</bold></td>
<td valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Z</bold>&#x02192;<bold>N</bold></td>
<td valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>Avg</bold>.</td>
</tr>
<tr>
<td valign="top" align="left"><bold>methods</bold></td>
<td valign="top" align="center"><bold>MAE</bold></td>
<td valign="top" align="center"><bold>MSE</bold></td>
<td valign="top" align="center"><bold>MAPE</bold></td>
<td valign="top" align="center"><bold>DMAE</bold></td>
<td valign="top" align="center"><italic>R</italic><sup>2</sup></td>
<td valign="top" align="center"><bold>MAE</bold></td>
<td valign="top" align="center"><bold>MSE</bold></td>
<td valign="top" align="center"><bold>MAPE</bold></td>
<td valign="top" align="center"><bold>DMAE</bold></td>
<td valign="top" align="center"><italic>R</italic><sup>2</sup></td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">CSRNet</td>
<td valign="top" align="center">2.59</td>
<td valign="top" align="center">3.57</td>
<td valign="top" align="center">11.2%</td>
<td valign="top" align="center">24.92</td>
<td valign="top" align="center"><bold>0.9891</bold></td>
<td valign="top" align="center">5.74</td>
<td valign="top" align="center">7.39</td>
<td valign="top" align="center">48.0%</td>
<td valign="top" align="center">12.01</td>
<td valign="top" align="center">0.9499</td>
</tr>
<tr>
<td valign="top" align="left">CSRNet_DA</td>
<td valign="top" align="center">1.88</td>
<td valign="top" align="center">2.43</td>
<td valign="top" align="center">9.2%</td>
<td valign="top" align="center">5.83</td>
<td valign="top" align="center">0.9864</td>
<td valign="top" align="center">4.56</td>
<td valign="top" align="center">5.83</td>
<td valign="top" align="center">29.3%</td>
<td valign="top" align="center">5.232</td>
<td valign="top" align="center">0.9409</td>
</tr>
<tr>
<td valign="top" align="left">PCEDA</td>
<td valign="top" align="center">2.45</td>
<td valign="top" align="center">3.18</td>
<td valign="top" align="center">18.7%</td>
<td valign="top" align="center">15.23</td>
<td valign="top" align="center">0.9765</td>
<td valign="top" align="center">5.61</td>
<td valign="top" align="center">7.72</td>
<td valign="top" align="center">39.1%</td>
<td valign="top" align="center">9.454</td>
<td valign="top" align="center">0.8965</td>
</tr>
<tr>
<td valign="top" align="left">FDA</td>
<td valign="top" align="center">1.84</td>
<td valign="top" align="center">2.44</td>
<td valign="top" align="center">9.9%</td>
<td valign="top" align="center">3.92</td>
<td valign="top" align="center">0.9850</td>
<td valign="top" align="center">4.44</td>
<td valign="top" align="center">5.79</td>
<td valign="top" align="center">43.7%</td>
<td valign="top" align="center">4.316</td>
<td valign="top" align="center"><bold>0.9689</bold></td>
</tr>
<tr>
<td valign="top" align="left">MFA</td>
<td valign="top" align="center">2.04</td>
<td valign="top" align="center">2.87</td>
<td valign="top" align="center">7.6%</td>
<td valign="top" align="center">3.47</td>
<td valign="top" align="center">0.9706</td>
<td valign="top" align="center">4.88</td>
<td valign="top" align="center">6.22</td>
<td valign="top" align="center">32.1%</td>
<td valign="top" align="center">5.148</td>
<td valign="top" align="center">0.9431</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>1.54</bold></td>
<td valign="top" align="center"><bold>2.06</bold></td>
<td valign="top" align="center"><bold>7.0%</bold></td>
<td valign="top" align="center"><bold>3.41</bold></td>
<td valign="top" align="center">0.9846</td>
<td valign="top" align="center"><bold>2.95</bold></td>
<td valign="top" align="center"><bold>4.01</bold></td>
<td valign="top" align="center"><bold>13.0%</bold></td>
<td valign="top" align="center"><bold>2.49</bold></td>
<td valign="top" align="center">0.9565</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The best performance is in boldface</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Qualitative comparisons on MTC dataset. From top to bottom alternating: RGB image, ground truth density map, density maps (count maps) estimated by CSRNet, CSRNet_DA, FDA, MFA, PCEDA, and our method. Numbers in the upper left corner of estimated density maps (count maps) represent the ground-truth or predicted counting value (rounded).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0006.tif"/>
</fig>
</sec>
<sec>
<title>4.3.3. Comparisons on RPC Dataset</title>
<p>The experiments on RPC dataset further demonstrated the effectiveness of our method. As shown in <xref ref-type="table" rid="T4">Table 4</xref>, our method achieved the lowest MAE, MSE, MAPE, and DMAE. Most of the methods underestimated the number of rice seedlings, mainly because the rice seedlings in the target domain are smaller than those in the source domain due to different growth stages. The other methods only generated responses for rice seedlings with more leaves and larger scales. In contrast, our method attained the accurate prediction results. The visualizations on RPC dataset are illustrated in <xref ref-type="fig" rid="F7">Figure 7</xref>. For results on the RSC dataset, the DMAE were very close to the MAE, as the targets appeared densely throughout the images.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Quantitative comparisons under J&#x02192;G setup (RPC dataset) and MTC&#x02192;MTC-UAV setup.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>J</bold>&#x02192;<bold>G setup</bold></th>
<th valign="top" align="center" colspan="5" style="border-bottom: thin solid #000000;"><bold>MTC</bold>&#x02192;<bold>MTC-UAV setup</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
<th valign="top" align="center"><bold>MAPE</bold></th>
<th valign="top" align="center"><bold>DMAE</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
<th valign="top" align="center"><bold>MAE</bold></th>
<th valign="top" align="center"><bold>MSE</bold></th>
<th valign="top" align="center"><bold>MAPE</bold></th>
<th valign="top" align="center"><bold>DMAE</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CSRNet</td>
<td valign="top" align="center">310.09</td>
<td valign="top" align="center">326.50</td>
<td valign="top" align="center">38.16%</td>
<td valign="top" align="center">310.44</td>
<td valign="top" align="center">0.1467</td>
<td valign="top" align="center">54.27</td>
<td valign="top" align="center">73.58</td>
<td valign="top" align="center">32.86%</td>
<td valign="top" align="center">98.14</td>
<td valign="top" align="center">0.6425</td>
</tr>
<tr>
<td valign="top" align="left">CSRNet_DA</td>
<td valign="top" align="center">209.56</td>
<td valign="top" align="center">282.93</td>
<td valign="top" align="center">26.46%</td>
<td valign="top" align="center">209.76</td>
<td valign="top" align="center"><bold>0.2599</bold></td>
<td valign="top" align="center">43.17</td>
<td valign="top" align="center">60.61</td>
<td valign="top" align="center">25.77%</td>
<td valign="top" align="center">61.52</td>
<td valign="top" align="center">0.7252</td>
</tr>
<tr>
<td valign="top" align="left">PCEDA</td>
<td valign="top" align="center">152.41</td>
<td valign="top" align="center">203.14</td>
<td valign="top" align="center">18.95%</td>
<td valign="top" align="center">152.55</td>
<td valign="top" align="center">0.1926</td>
<td valign="top" align="center">57.56</td>
<td valign="top" align="center">75.37</td>
<td valign="top" align="center">33.90%</td>
<td valign="top" align="center">69.10</td>
<td valign="top" align="center">0.6442</td>
</tr>
<tr>
<td valign="top" align="left">FDA</td>
<td valign="top" align="center">356.28</td>
<td valign="top" align="center">370.63</td>
<td valign="top" align="center">44.06%</td>
<td valign="top" align="center">356.54</td>
<td valign="top" align="center">0.2433</td>
<td valign="top" align="center">36.78</td>
<td valign="top" align="center">51.06</td>
<td valign="top" align="center"><bold>21.83%</bold></td>
<td valign="top" align="center">102.31</td>
<td valign="top" align="center">0.8247</td>
</tr>
<tr>
<td valign="top" align="left">MFA</td>
<td valign="top" align="center">243.05</td>
<td valign="top" align="center">283.61</td>
<td valign="top" align="center">30.48%</td>
<td valign="top" align="center">244.14</td>
<td valign="top" align="center">0.1593</td>
<td valign="top" align="center">46.56</td>
<td valign="top" align="center">65.91</td>
<td valign="top" align="center">29.11%</td>
<td valign="top" align="center">62.42</td>
<td valign="top" align="center">0.6755</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>111.17</bold></td>
<td valign="top" align="center"><bold>161.46</bold></td>
<td valign="top" align="center"><bold>14.54%</bold></td>
<td valign="top" align="center"><bold>117.27</bold></td>
<td valign="top" align="center">0.2057</td>
<td valign="top" align="center"><bold>35.88</bold></td>
<td valign="top" align="center"><bold>47.41</bold></td>
<td valign="top" align="center">23.99%</td>
<td valign="top" align="center"><bold>60.04</bold></td>
<td valign="top" align="center"><bold>0.8655</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The best performance is in boldface</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Visualizations on RPC dataset. From top to bottom alternating: RGB image, ground truth density map, density maps (count maps) predicted by CSRNet, CSRNet DA, FDA, MFA, PCEDA, and our method. Numbers on the upper left corner of the estimated density maps (count maps) represent the ground-truth or predicted counting value (rounded).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0007.tif"/>
</fig>
</sec>
<sec>
<title>4.3.4. Comparisons on MTC-UAV Dataset</title>
<p>Nowadays, UAVs have become useful image acquisition devices for agriculture. In practice, a model trained with images collected by phenopoles may be tested on images collected by UAVs. We adapted the model from MTC dataset to MTC-UAV dataset under this setting. As shown in <xref ref-type="table" rid="T4">Table 4</xref>, our method surpassed others in all metrics except for MAPE. The visualizations on MTC-UAV dataset is illustrated in <xref ref-type="fig" rid="F8">Figure 8</xref>.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Visualizations on MTC-UAV dataset. From top to bottom alternating: RGB image, ground truth density maps, density maps (count maps) predicted by CSRNet, CSRNet DA, FDA, MFA, PCEDA, and our method. Numbers on the upper left corner represent the ground-truth or predicted counting value (rounded).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0008.tif"/>
</fig>
</sec>
</sec>
<sec>
<title>4.4. Ablation Study</title>
<p>First, we compared two different regression paradigms: local count regression and density map regression. Then we demonstrated the effectiveness of the feature discriminator and foreground mask discriminator in the proposed BADA module.</p>
<sec>
<title>4.4.1. Local Count Regression</title>
<p>We found that local count regression were more robust than density map regression for cross-domain settings. To verify this, we replaced the local count regressor of the original BADANet with a local count regressor without any downsampling operations. The local count regressor consisted of a series of convolution layers and directly predicted the density maps. The training strategy was kept the same. As shown in <xref ref-type="table" rid="T5">Table 5</xref>, on all settings of the MTC dataset, local count regression obtained better results than the density map regression.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Ablation study on regression targets and discriminator configurations (MAE).</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center" colspan="2"><bold>Settings</bold></th>
<th valign="top" align="left"><bold>Z&#x02192;Jun</bold></th>
<th valign="top" align="left"><bold>Z&#x02192;W</bold></th>
<th valign="top" align="left"><bold>Z&#x02192;Ji</bold></th>
<th valign="top" align="left"><bold>Z&#x02192;T</bold></th>
<th valign="top" align="left"><bold>Z&#x02192;N</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Regression<break/> Target</td>
<td valign="top" align="left">Density map regression</td>
<td valign="top" align="left">2.37</td>
<td valign="top" align="left">4.76</td>
<td valign="top" align="left">1.31</td>
<td valign="top" align="left">7.01</td>
<td valign="top" align="left">1.87</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Local count regression</td>
<td valign="top" align="left"><bold>1.92</bold></td>
<td valign="top" align="left"><bold>3.83</bold></td>
<td valign="top" align="left"><bold>0.50</bold></td>
<td valign="top" align="left"><bold>6.96</bold></td>
<td valign="top" align="left"><bold>1.54</bold></td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Discriminators</td>
<td valign="top" align="left">None</td>
<td valign="top" align="left">2.15</td>
<td valign="top" align="left"><bold>2.37</bold></td>
<td valign="top" align="left">0.85</td>
<td valign="top" align="left">13.37</td>
<td valign="top" align="left">1.51</td>
</tr>
<tr>
<td/>
<td valign="top" align="left"><inline-formula><mml:math id="M84"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left"><bold>1.88</bold></td>
<td valign="top" align="left">2.65</td>
<td valign="top" align="left">1.06</td>
<td valign="top" align="left">8.64</td>
<td valign="top" align="left">1.48</td>
</tr>
<tr>
<td/>
<td valign="top" align="left"><inline-formula><mml:math id="M85"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left">2.11</td>
<td valign="top" align="left">3.48</td>
<td valign="top" align="left">0.79</td>
<td valign="top" align="left">11.02</td>
<td valign="top" align="left"><bold>1.41</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left"><inline-formula><mml:math id="M86"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left">1.92</td>
<td valign="top" align="left">3.83</td>
<td valign="top" align="left"><bold>0.50</bold></td>
<td valign="top" align="left"><bold>6.96</bold></td>
<td valign="top" align="left">1.54</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The best performance is in boldface</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.4.2. Discriminators</title>
<p>The domain discriminators were imposed at the input and output of BADA module. The feature discriminator <inline-formula><mml:math id="M87"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can help the CNNs extract domain-invariant feature maps. And the mask discriminator <inline-formula><mml:math id="M88"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can help refine the predicted foreground masks. As shown in <xref ref-type="table" rid="T5">Table 5</xref>, the combination of <inline-formula><mml:math id="M89"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M90"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can achieve lowest MAEs on 2 different settings, and the performance was more stable than only applying one or none of the discriminator. Although on some settings, the full method slightly fell behind the other versions. We believe this was because the domain gaps in these settings were not obvious, as the MAEs were already relatively low when no domain adaptation modules were attached. Under such circumstances, the adversarial training strategy might hurt the training process.</p>
<p>To understand the effectiveness of discriminators more intuitively, we show the visualizations of methods with/without discriminators in <xref ref-type="fig" rid="F9">Figure 9</xref>. With foreground mask discriminator <inline-formula><mml:math id="M91"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the network was more confident about the segmentation results and produced less error. The shapes of foreground masks were regular and neat. By contrast, when no discriminators were attached, the shapes of foreground masks were irregular and scattered. Besides, more backgrounds were mistaken for foregrounds, which provided incorrect target distribution information for the network. For scenes on row 3, 4, and 5 of <xref ref-type="fig" rid="F9">Figure 9</xref>, although the estimated foreground masks were correct, non-adversarial method produced more errors.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Visualizations of model with and without discriminators. From left to right are input images, ground truth density maps, estimated foreground mask without discriminators, estimated local count maps without discriminators, estimated foreground masks with discriminators, estimated local count maps with discriminators. The foreground masks have been binarized. The white numbers on the corners of ground truth density maps and estimated local count maps denote the ground truth counts and the inferred counts, respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0009.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>Here we conclude all the tested methods and discuss their advantages and drawbacks. For all the tested domain adaptation settings, we find that UDA methods more or less improve the cross-domain performance, which demonstrates the necessity and effectiveness of domain adaptation. According to the proposed metric DMAE, UDA method can also help the model predict more precise density maps. Among all the UDA methods, the proposed BADA module is more stable and obtains the best MAE and DMAE on 5 out of the 7 domain adaptation settings, which demonstrates its effectiveness.</p>
<p>For feature-level domain adaptation methods (CSRNet_DA, MFA and our method), the results show that adversarial training can help aligning the features for different domains in plant counting datasets. Compared with CSRNet_DA, MFA aligns features at different scales with multiple discriminators. MFA showed marginal improvement on MTC datasets, while significant improvements on the setting MTC&#x02192;MTC-UAV were obtained. This indicates that multi-scale adversarial training is more suitable when objects in different domains are with different scales. This also inspires us that the proposed BADA module can be further improved with multi-scale adaptation strategy. Our methods aligns the features as well as the foreground segmentation results. The visualizations show that, our method can better distinguish the targets and other background elements and generate more precise density maps comparing with other UDA methods. Therefore, the overall MAE and DMAE can be effectively reduced.</p>
<p>For image-level domain adaptaion methods (FDA and PCEDA), domain adaptation is achieved by aligning the image styles. We visualize the transferred images in <xref ref-type="fig" rid="F10">Figure 10</xref>. Although these methods fail to modify the core difference like camera views, target scales and appearances, some global style like illuminations, hues and textures can be transferred between source and target domain. However, the transferred images showed some artifacts. For example, some blue and red shadows can be observed in the transferred images from the source domain of J&#x02192;G settings. The PCEDA model recognized the texture of blue and red poles in the target domain while incorrectly added it on irrelevant objects like plants. We also noticed that the better quality of transferred images may not guarantee better cross-dataset counting performance. FDA can better boost the cross-domain performance on MTC dataset, while the quality of style transfer was inferior to PCEDA. However, when failure cases occur, the image-based UDA method will significantly harm the cross-domain performance. As shown in the third row of <xref ref-type="fig" rid="F10">Figure 10</xref>. FDA generated wrong hues and colors for the source domain, which led to performance drop on setting J&#x02192;G in <xref ref-type="table" rid="T4">Table 4</xref>. While the experimental results showed that these methods can improve the performance, we were suspicious whether the boost came from the reduction of image-level domain gaps, or from data augmentation. As style-transfer can be viewed as a data augmentation method which will change the hues, contrasts or illuminations of the original images. To validate this, we also conducted an experiment where we randomly swap the low frequency spectrums of two source domain images (identical to FDA) on the MTC dataset, and obtained almost the same performance improvement.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Visualizations of style transferred images with different image-level UDA methods. From top to bottom alternating: source domain image, transferred images by PCEDA and transferred images with FDA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-731816-g0010.tif"/>
</fig>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusion</title>
<p>In this paper, we investigate the influence of domain gap for deep learning-based plant counting method and show how to alleviate the influence with unsupervised domain adaptation methods. We evaluated the performance of several popular UDA methods. We found that these methods only prompted limited cross-domain performance due to the characteristics of domain gaps in plant counting. Particularly, the counting models produced large errors on background areas. To address this, we purpose a flexible background-aware domain adaptation module, which can easily fit into existing object counting methods and enhance the cross-domain performance. We evaluated our methods under 7 different domain adaptation settings. The results showed that our method can obtain better cross-domain accuracy than existing UDA methods on plant counting task.</p>
<p>Nowadays, despite the rapid development of deep learning-based plant counting methods, the scale and diversity of plant counting datasets are still limited. When applying data-driven plant counting methods on new scenes, it is necessary to consider the hazard of domain gaps. We hope our work can help more researchers and practitioners noticing this issue and bring more solutions for UDA in plant counting. In the future, we will investigate how to extract more generic features for plant counting.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>MS proposed the idea of BADA module, implemented the algorithm in <monospace>PyTorch</monospace>, conducted the experiments, analyzed the results, drafted, and revised the manuscript. X-YL helped draft the manuscript and organized part of the figures and tables. HL helped refine the idea, organized part of the experiments, and revised the manuscript. Z-GC provided the funding and supervised the study. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This work was supported in part by the National Natural Science Foundation of China under grant no. 61876211 and in part by the Chinese Fundamental Research Funds for the Central Universities under grant no. 2021XXJS095.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Arteta</surname> <given-names>C.</given-names></name> <name><surname>Lempitsky</surname> <given-names>V.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Counting in the wild</article-title>, in <source>Proceedings of European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Amsterdam</publisher-loc>), <fpage>483</fpage>&#x02013;<lpage>498</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ayalew</surname> <given-names>T. W.</given-names></name> <name><surname>Ubbens</surname> <given-names>J. R.</given-names></name> <name><surname>Stavness</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>Unsupervised domain adaptation for plant organ counting</article-title>, in <source>Proceedings of European Conference on Computer Vision (ECCV)</source>, <fpage>330</fpage>&#x02013;<lpage>346</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bargoti</surname> <given-names>S.</given-names></name> <name><surname>Underwood</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep fruit detection in orchards</article-title>, in <source>2017 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3626</fpage>&#x02013;<lpage>3633</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ben-David</surname> <given-names>S.</given-names></name> <name><surname>Blitzer</surname> <given-names>J.</given-names></name> <name><surname>Crammer</surname> <given-names>K.</given-names></name> <name><surname>Kulesza</surname> <given-names>A.</given-names></name> <name><surname>Pereira</surname> <given-names>F.</given-names></name> <name><surname>Vaughan</surname> <given-names>J. W.</given-names></name></person-group> (<year>2010</year>). <article-title>A theory of learning from different domains</article-title>. <source>Mach. Learn</source>. <volume>79</volume>, <fpage>151</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-009-5152-4</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>David</surname> <given-names>E.</given-names></name> <name><surname>Madec</surname> <given-names>S.</given-names></name> <name><surname>Sadeghi-Tehran</surname> <given-names>P.</given-names></name> <name><surname>Aasen</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Global wheat head detection (gwhd) dataset: a large and diverse dataset of high resolution rgb labelled images to develop and benchmark wheat head detection methods</article-title>. <source>Plant Phenomics</source> <volume>2020</volume>:<fpage>3521852</fpage>. <pub-id pub-id-type="doi">10.34133/2020/3521852</pub-id><pub-id pub-id-type="pmid">33313551</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>ImageNet: a large-scale hierarchical image database</article-title>, in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>248</fpage>&#x02013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>D&#x00027;Innocente</surname> <given-names>A.</given-names></name> <name><surname>Borlino</surname> <given-names>F. C.</given-names></name> <name><surname>Bucci</surname> <given-names>S.</given-names></name> <name><surname>Caputo</surname> <given-names>B.</given-names></name> <name><surname>Tommasi</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>One-shot unsupervised cross-domain detection</article-title>, in <source>Proceedings of European Conference on Computer Vision (ECCV)</source>, eds A. Vedaldi, H. Bischof, T. Brox, and J.-M. <volume>Frahm</volume>, <fpage>732</fpage>&#x02013;<lpage>748</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Ganin</surname> <given-names>Y.</given-names></name> <name><surname>Lempitsky</surname> <given-names>V.</given-names></name></person-group> (<year>2015</year>). <article-title>Unsupervised domain adaptation by backpropagation</article-title>, in <source>Proceedings of International Conference on Machine Learning (ICML), volume 37 of Proceedings of Machine Learning Research</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>1180</fpage>&#x02013;<lpage>1189</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://proceedings.mlr.press/v37/ganin15.html">http://proceedings.mlr.press/v37/ganin15.html</ext-link></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name></person-group> (<year>2021</year>). <article-title>Feature-aware adaptation and density alignment for crowd counting in video surveillance</article-title>. <source>IEEE Trans. Cybern</source>. <volume>51</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2020.3034316</pub-id><pub-id pub-id-type="pmid">33259318</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Giuffrida</surname> <given-names>M. V.</given-names></name> <name><surname>Dobrescu</surname> <given-names>A.</given-names></name> <name><surname>Doerner</surname> <given-names>P.</given-names></name> <name><surname>Tsaftaris</surname> <given-names>S. A.</given-names></name></person-group> (<year>2019</year>). <article-title>Leaf counting without annotations using adversarial unsupervised domain adaptation</article-title>, in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2590</fpage>&#x02013;<lpage>2599</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Lempitsky</surname> <given-names>V.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Learning</article-title>, in <source>Proceedings of Advances in Neural Information Processing Systems (NeurIPS), Vol. 23</source>, <fpage>1324</fpage>&#x02013;<lpage>1332</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf">https://proceedings.neurips.cc/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf</ext-link>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Csrnet: dilated convolutional neural networks for understanding the highly congested scenes</article-title>, in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1091</fpage>&#x02013;<lpage>1100</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>G.</given-names></name> <name><surname>Milan</surname> <given-names>A.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Reid</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). <article-title>Refinenet: multi-path refinement networks for high-resolution semantic segmentation</article-title>, in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1925</fpage>&#x02013;<lpage>1934</lpage>. <pub-id pub-id-type="pmid">30668461</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>High-throughput rice density estimation from transplantation to tillering stages using deep networks</article-title>. <source>Plant Phenomics</source> <volume>2020</volume>:<fpage>1375957</fpage>. <pub-id pub-id-type="doi">10.34133/2020/1375957</pub-id><pub-id pub-id-type="pmid">33313541</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name> <name><surname>Fang</surname> <given-names>Z.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name></person-group> (<year>2017a</year>). <article-title>Towards fine-grained maize tassel flowering status recognition: dataset, theory and practice</article-title>. <source>Appl. Soft. Comput</source>. <volume>56</volume>, <fpage>34</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2017.02.026</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name> <name><surname>Zhuang</surname> <given-names>B.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2017b</year>). <article-title>TasselNet: counting maize tassels in the wild via local counts regression network</article-title>. <source>Plant Methods</source> <volume>13</volume>:<fpage>79</fpage>. <pub-id pub-id-type="doi">10.1186/s13007-017-0224-0</pub-id><pub-id pub-id-type="pmid">29118821</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Y. N.</given-names></name> <name><surname>Zhao</surname> <given-names>X. M.</given-names></name> <name><surname>Wang</surname> <given-names>X. Q.</given-names></name> <name><surname>Cao</surname> <given-names>Z. G.</given-names></name></person-group> (<year>2021</year>). <article-title>Tasselnetv3: explainable plant counting with guided upsampling and background suppression</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>60</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2021.3058962</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Z.</given-names></name> <name><surname>Wei</surname> <given-names>X.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Gong</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Bayesian loss for crowd count estimation with point supervision</article-title>, in <source>Proceedings of IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>6142</fpage>&#x02013;<lpage>6151</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Madec</surname> <given-names>S.</given-names></name> <name><surname>Jin</surname> <given-names>X.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>de Solan</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Duyme</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Ear density estimation from high resolution rgb imagery using deep learning technique</article-title>. <source>Agric. Forest Meteorol</source>. <volume>264</volume>, <fpage>225</fpage>&#x02013;<lpage>234</lpage>. <pub-id pub-id-type="doi">10.1016/j.agrformet.2018.10.013</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Modolo</surname> <given-names>D.</given-names></name> <name><surname>Shuai</surname> <given-names>B.</given-names></name> <name><surname>Varior</surname> <given-names>R. R.</given-names></name> <name><surname>Tighe</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Understanding the impact of mistakes on background regions in crowd counting</article-title>, in <source>Proceedings of Winter Conference on Applications of Computer Vision (WACV)</source>, <fpage>1650</fpage>&#x02013;<lpage>1659</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Onoro-Rubio</surname> <given-names>D.</given-names></name> <name><surname>L&#x000F3;pez-Sastre</surname> <given-names>R. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Towards perspective-free object counting with deep learning</article-title>, in <source>Proceedings of European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Amsterdam</publisher-loc>), <fpage>615</fpage>&#x02013;<lpage>629</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Paszke</surname> <given-names>A.</given-names></name> <name><surname>Gross</surname> <given-names>S.</given-names></name> <name><surname>Massa</surname> <given-names>F.</given-names></name> <name><surname>Lerer</surname> <given-names>A.</given-names></name> <name><surname>Bradbury</surname> <given-names>J.</given-names></name> <name><surname>Chanan</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>PyTorch: an imperative style, high-performance deep learning library</article-title>, in <source>Proceedings of Advances in Neural Information Processing Systems (NeurIPS)</source> (<publisher-name>Curran Associates, Inc.</publisher-name>), <fpage>8026</fpage>&#x02013;<lpage>8037</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf">https://proceedings.neurips.cc/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf</ext-link></citation>
</ref>
<ref id="B23">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <source>Very deep convolutional networks for large-scale image recognition. arXiv preprint <italic>arXiv:1409.1556</italic></source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1409.1556.pdf">https://arxiv.org/pdf/1409.1556.pdf</ext-link>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tsaftaris</surname> <given-names>S. A.</given-names></name> <name><surname>Minervini</surname> <given-names>M.</given-names></name> <name><surname>Scharr</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>Machine learning for plant phenotyping needs image processing</article-title>. <source>Trends Plant Sci</source>. <volume>21</volume>, <fpage>989</fpage>&#x02013;<lpage>991</lpage>. <pub-id pub-id-type="doi">10.1016/j.tplants.2016.10.002</pub-id><pub-id pub-id-type="pmid">27810146</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tzeng</surname> <given-names>E.</given-names></name> <name><surname>Hoffman</surname> <given-names>J.</given-names></name> <name><surname>Saenko</surname> <given-names>K.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Adversarial discriminative domain adaptation</article-title>, in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2962</fpage>&#x02013;<lpage>2971</lpage>. <pub-id pub-id-type="pmid">32635540</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vu</surname> <given-names>T.-H.</given-names></name> <name><surname>Jain</surname> <given-names>H.</given-names></name> <name><surname>Bucher</surname> <given-names>M.</given-names></name> <name><surname>Cord</surname> <given-names>M.</given-names></name> <name><surname>P&#x000E9;rez</surname> <given-names>P.</given-names></name></person-group> (<year>2019</year>). <article-title>Advent: adversarial entropy minimization for domain adaptation in semantic segmentation</article-title>, in <source>Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>2512</fpage>&#x02013;<lpage>2521</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Samaras</surname> <given-names>D.</given-names></name> <name><surname>Nguyen</surname> <given-names>M. H.</given-names></name></person-group> (<year>2020</year>). <article-title>Distribution matching for crowd counting</article-title>, in <source>Proceedings of Advances in Neural Information Processing Systems (NeurIPS), Vol. 33</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://papers.nips.cc/paper/2020/file/11d558033a1016fcc82560c65cca5f-Paper.pdf">https://papers.nips.cc/paper/2020/file/11d558033a1016fcc82560c65cca5f-Paper.pdf</ext-link></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Learning from synthetic data for crowd counting in the wild</article-title>, in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8190</fpage>&#x02013;<lpage>8199</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiong</surname> <given-names>H.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Madec</surname> <given-names>S.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2019a</year>). <article-title>Tasselnetv2: in-field counting of wheat spikes with context-augmented local regression networks</article-title>. <source>Plant Methods</source> <volume>15</volume>:<fpage>150</fpage>. <pub-id pub-id-type="doi">10.1186/s13007-019-0537-2</pub-id><pub-id pub-id-type="pmid">31857821</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiong</surname> <given-names>H.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2019b</year>). <article-title>From open set to closed set: Counting objects by spatial divide-and-conquer</article-title>, in <source>Proceedings of IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>8362</fpage>&#x02013;<lpage>8371</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name> <name><surname>Tian</surname> <given-names>Q.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Cross-domain detection via graph-induced prototype alignment</article-title>, in <source>Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>12352</fpage>&#x02013;<lpage>12361</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Lao</surname> <given-names>D.</given-names></name> <name><surname>Sundaramoorthi</surname> <given-names>G.</given-names></name> <name><surname>Soatto</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Phase consistent ecological domain adaptation</article-title>, in <source>Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>9008</fpage>&#x02013;<lpage>9017</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Soatto</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Fda: fourier domain adaptation for semantic segmentation</article-title>, in <source>Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>4084</fpage>&#x02013;<lpage>4094</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoosefzadeh-Najafabadi</surname> <given-names>M.</given-names></name> <name><surname>Tulpan</surname> <given-names>D.</given-names></name> <name><surname>Eskandari</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Using hybrid artificial intelligence and evolutionary optimization algorithms for estimating soybean yield and fresh biomass using hyperspectral vegetation indices</article-title>. <source>Remote Sens</source>. <volume>13</volume>, <fpage>2555</fpage>&#x02013;<lpage>2575</lpage>. <pub-id pub-id-type="doi">10.3390/rs13132555</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>Cross-scene crowd counting via deep convolutional neural networks</article-title>, in <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>833</fpage>&#x02013;<lpage>841</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Park</surname> <given-names>T.</given-names></name> <name><surname>Isola</surname> <given-names>P.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Unpaired image-to-image translation using cycle-consistent adversarial networks</article-title>, in <source>2017 IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2242</fpage>&#x02013;<lpage>2251</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>