<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2025.1664718</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>CGA-ASNet: an RGB-D amodal segmentation network for restoring occluded tomato regions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Zhaoyang</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/3146072/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yin</surname>
<given-names>Yong</given-names>
</name>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Xing</surname>
<given-names>Zhihong</given-names>
</name>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Deng</surname>
<given-names>Hanbing</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2713923/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>College of Information and Electrical Engineering, Shenyang Agricultural University</institution>, <addr-line>Shenyang</addr-line>,&#xa0;<country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/172404/overview">Alejandro Isabel Luna-Maldonado</ext-link>, Autonomous University of Nuevo Le&#xf3;n, Mexico</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1838456/overview">Xing Sheng</ext-link>, Shandong Normal University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3135773/overview">Mohana Saranya S</ext-link>, Kongu Engineering College, India</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Hanbing Deng, <email xlink:href="mailto:denghanbing@syau.edu.cn">denghanbing@syau.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>09</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1664718</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>07</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>08</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Li, Yin, Xing and Deng.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Li, Yin, Xing and Deng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Obtaining the complete morphology of tomato fruits under non-destructive conditions is essential for phenotype research, yet fruit occlusions often hinder deep learning-based image segmentation methods from capturing the true shape of occluded regions. This limitation reduces prediction accuracy and adversely impacts phenotype data acquisition. To overcome this challenge, we propose CGA-ASNet, an RGB-D amodal segmentation network incorporating a Contextual and Global Attention (CGA) module. A synthetic tomato dataset (Tomato-sim) was constructed using NVIDIA Isaac Sim&#x2019;s Replicator Composer (ISRC) to realistically simulate tomato morphology and greenhouse environments, and the network was trained on this dataset. To evaluate generalization, CGA-ASNet was tested on both the synthetic and a separate real-world dataset. While no explicit domain adaptation techniques were adopted, diverse lighting conditions (strong, normal, and weak illumination) were simulated to implicitly reduce the domain gap, and a mean coordinate fusion algorithm was introduced to improve annotation completeness in real-world occlusion scenarios. By leveraging contextual information among feature input keys for self-attention learning, capturing global information, and expanding the receptive field, CGA-ASNet enhanced representation capacity, semantic understanding, and localization accuracy. Experimental results demonstrated that CGA-ASNet achieved an F@0.75 score of 94.2 and a mean Intersection over Union (mIoU) of 82.4% in greenhouse amodal segmentation tasks. These findings indicate that training with well-designed synthetic datasets can effectively support accurate occlusion-aware segmentation in real environments, providing a practical solution for tomato phenotyping in greenhouse conditions.</p>
</abstract>
<kwd-group>
<kwd>amodal segmentation</kwd>
<kwd>occlusion-aware segmentation</kwd>
<kwd>RGB-D image segmentation</kwd>
<kwd>plant phenotyping</kwd>
<kwd>tomato</kwd>
<kwd>smart agriculture</kwd>
</kwd-group>
<contract-num rid="cn001">2022&#x5e74;YFD2002303-01</contract-num>
<contract-sponsor id="cn001">Ministry of Science and Technology of the People's Republic of China<named-content content-type="fundref-id">10.13039/501100002855</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Department of Education of Liaoning Province<named-content content-type="fundref-id">10.13039/501100007620</named-content>
</contract-sponsor>
<counts>
<fig-count count="16"/>
<table-count count="9"/>
<equation-count count="20"/>
<ref-count count="41"/>
<page-count count="17"/>
<word-count count="8931"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Technical Advances in Plant Science</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Tomatoes are among the most widely cultivated vegetables globally, with countries such as the United States, China, and Japan extensively utilizing greenhouse cultivation methods. In recent years, the area dedicated to greenhouse tomato farming has steadily expanded. However, despite these advancements in controlled-environment agriculture, tomato harvesting remains largely dependent on manual labor, which is not only labor-intensive but also inefficient (<xref ref-type="bibr" rid="B3">C&#xe1;mara-Zapata et&#xa0;al., 2019</xref>). To address these challenges, automated growth monitoring systems and intelligent harvesting machines are gradually emerging as key solutions in modern agriculture. These technologies are increasingly being adopted to mitigate labor shortages in regions that heavily rely on manual harvesting. However, fruit morphology remains indispensable in processes such as biological control and biomass detection.</p>
<p>Accurate information on fruit morphology is crucial for multiple aspects of agricultural management. It not only aids in determining the growth status of plants but also supports precision fertilization and irrigation decisions. Furthermore, changes in fruit morphology can serve as early indicators of pests and diseases, enabling farmers to detect issues early and take appropriate action. Automated detection systems, through continuous monitoring of plant morphology, offer higher precision and real-time feedback, thereby reducing dependency on manual labor and improving both crop yield and quality. With the rapid advancements in machine vision and deep learning, many computer vision tasks&#x2014;such as image recognition (<xref ref-type="bibr" rid="B14">He et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B33">Szegedy et&#xa0;al., 2016</xref>), object detection (<xref ref-type="bibr" rid="B12">Girshick, 2015</xref>; <xref ref-type="bibr" rid="B26">Redmon and Farhadi, 2018</xref>; <xref ref-type="bibr" rid="B27">Ren et&#xa0;al., 2015</xref>);, and semantic segmentation (<xref ref-type="bibr" rid="B23">Long et&#xa0;al., 2015</xref>)&#x2014;have enabled precise localization and shape determination of fruits based on their appearance. These techniques have been widely applied in disease detection (<xref ref-type="bibr" rid="B7">Dhaka et&#xa0;al., 2021</xref>), maturity assessment, and growth monitoring. However, these tasks typically rely on each pixel in the image corresponding to a single label. In occlusion scenarios, models can only process visible portions, leaving occluded areas unaddressed or inadequately evaluated.</p>
<p>Traditional computer vision algorithms, including edge-based segmentation methods (<xref ref-type="bibr" rid="B31">Sheng et&#xa0;al., 2023</xref>), struggle to handle occlusions, particularly in agriculture, where fruits grow in random positions and complex lighting conditions further complicate scene interpretation. While edge-based approaches have shown effectiveness in fruit segmentation under certain conditions, current technologies face challenges in effectively dealing with occluded fruits, resulting in lower recognition accuracy. Current technologies face challenges in effectively dealing with occluded fruits, resulting in lower recognition accuracy. This issue presents a significant barrier to the implementation of automation in agriculture, particularly in automated harvesting and growth monitoring systems, where the presence of occlusions severely impacts recognition accuracy and operational efficiency. To effectively address this problem, the occluded portions of the fruit must be accurately reconstructed.</p>
<p>Amodal segmentation aims to infer and complete the occluded portions of objects by providing their full masks, as illustrated in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. Several recent studies have explored occlusion-aware perception and shape reconstruction techniques to improve fruit detection in complex agricultural environments. An occluder&#x2013;occludee relational network (O2RNet) was proposed to explicitly model spatial interactions between overlapping objects and achieved state-of-the-art performance in clustered apple detection (<xref ref-type="bibr" rid="B4">Chu et&#xa0;al., 2023</xref>). A zero-shot Sim2Real reinforcement learning strategy was introduced to manipulate deformable plants and reveal hidden fruits, achieving 86.7% success without real-world fine-tuning (<xref ref-type="bibr" rid="B32">Subedi et&#xa0;al., 2025</xref>). Furthermore, a safe leaf manipulation method was proposed to improve pose and shape estimation accuracy by uncovering occluded fruits (<xref ref-type="bibr" rid="B38">Yao et al., 2025</xref>).</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>
<bold>(A)</bold> Two ripe tomatoes on a plant with green leaves. <bold>(B)</bold> An amodal mask of the same tomatoes, highlighting their form against the background.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g001.tif">
<alt-text content-type="machine-generated">Panel A shows two ripe tomatoes on a plant with green leaves. Panel B displays an amodal mask of the same tomatoes, highlighting their form against the background.</alt-text>
</graphic>
</fig>
<p>While these techniques have shown promising performance in orchard and open-field conditions, greenhouse environments present a distinct set of challenges that remain underexplored. Greenhouse-grown crops, such as tomatoes, are typically cultivated in densely packed rows with limited spacing, resulting in more frequent intra-class occlusions. The constrained physical layout, along with complex lighting and structural occluders (e.g., stems, trellises, or support wires), imposes high demands on vision-based perception systems. These conditions significantly degrade the accuracy and reliability of fruit detection and localization, which are critical for robotic harvesting and automated yield estimation.</p>
<p>Therefore, it is essential to develop occlusion-resilient perception methods tailored to greenhouse-specific scenarios. Amodal segmentation, which infers complete object masks including invisible parts, offers a promising solution to this challenge. In this study, a deep learning-based amodal segmentation method is proposed for greenhouse tomatoes, targeting the reconstruction of occluded regions to support robust visual perception and task execution in controlled-environment agriculture.</p>
<p>This task offers significant benefits to various downstream applications. For instance, in 3D reconstruction (<xref ref-type="bibr" rid="B30">Seitz et&#xa0;al., 2006</xref>), having complete shape information is crucial for generating more accurate 3D models, especially when objects can only be observed from limited viewpoints. In cases where the view is restricted or objects are partially occluded, understanding the full structure of the object helps enhance the model&#x2019;s realism and reconstruction accuracy. For video segmentation tasks, objects in videos are often partially obscured by other elements, and having complete shape information aids in maintaining object consistency across frames, thereby improving segmentation quality and precision. In dynamic scenes, the continuity of object shapes significantly reduces errors caused by occlusion or movement. Additionally, in agricultural machine vision systems, particularly in controlled-environment agriculture, perceiving the full structure of occluded objects is essential for navigation and task execution. The complexity of greenhouse environments, where objects such as plants or machinery frequently cause occlusions, makes global shape perception vital for optimal path planning, obstacle avoidance, and harvesting strategy refinement. Accurate object perception allows the system to reliably assess fruit ripeness and determine the optimal harvesting time, thus improving harvesting efficiency, reducing manual intervention, and ultimately lowering labor costs. Traditional segmentation algorithms, such as thresholding, region growing, and edge-based methods, are primarily designed for semantic segmentation, which involves dividing images into predefined categories. While these methods can achieve reasonable results in simple scenarios, they often struggle in complex agricultural environments where multiple instances of the same class are present and occlusions are common. Some recent works have demonstrated instance-level segmentation capabilities using point cloud data (<xref ref-type="bibr" rid="B15">Jiang et&#xa0;al., 2025</xref>). Nevertheless, these methods still face challenges when dealing with heavily occluded scenes or when full object masks, including invisible regions, are required. As a result, more researchers are applying deep learning techniques to amodal segmentation tasks. Leveraging convolutional neural network (CNN) and other advanced architectures such as U-Net (<xref ref-type="bibr" rid="B28">Ronneberger et&#xa0;al., 2015</xref>) and Mask R-CNN (<xref ref-type="bibr" rid="B13">He et&#xa0;al., 2017</xref>), amodal segmentation not only performs semantic segmentation but also enables precise instance-level segmentation and even part-level segmentation within images.</p>
<p>The earliest work on amodal segmentation can be traced back to the research by (<xref ref-type="bibr" rid="B18">Li and Malik, 2016</xref>), where they synthesized images to create the first amodal instance segmentation dataset and trained and tested their proposed model, the Amodal Segmentation Network (ASN). To further validate the effectiveness of the amodal segmentation task, (<xref ref-type="bibr" rid="B41">Zhu et&#xa0;al., 2017</xref>) conducted additional studies. They invited multiple annotators to label the same image with amodal annotations, and the results showed a high level of agreement among annotators regarding regions and edges, demonstrating the task&#x2019;s clear operability. They also provided amodal annotations for 5000 images from the COCO dataset, known as the COCOA dataset. Building on this, they proposed the ExpandMask network, where the input consisted of image patches and visible mask predictions, and the output was the occluded part of the target object. (<xref ref-type="bibr" rid="B9">Follmann et al., 2019</xref>) further improved upon Mask R-CNN, introducing a dedicated module for amodal mask segmentation called ORCNN (Occlusion Region Convolutional Neural Network). They also compiled and organized two amodal segmentation datasets, D2SA and COCOA. Subsequently, (<xref ref-type="bibr" rid="B2">Blok et&#xa0;al., 2021</xref>) and (<xref ref-type="bibr" rid="B11">Gen&#xe9;-Mola et&#xa0;al., 2024</xref>) applied ORCNN to broccoli and apple datasets, achieving promising results. Their experiments demonstrated that their models outperformed other methods on these datasets, further validating their effectiveness and superiority.</p>
<p>This deep learning-based approach to amodal segmentation has significantly improved the perception of occluded objects, offering more precise solutions for scene understanding and complex visual tasks. However, most of these studies have been tested on public datasets or applied to agricultural datasets in a very limited capacity. In the agricultural domain, amodal segmentation faces several critical limitations. For instance, the lack of large-scale, high-quality training datasets and issues of domain mismatch often result in poor Sim-to-real (<xref ref-type="bibr" rid="B39">Zhao et&#xa0;al., 2020</xref>) transfer. In real-world greenhouse environments, images typically contain numerous instances of the same class that are occluded by one another, making amodal segmentation tasks for such occluded objects far more challenging. Although existing computational models can perform well when trained on large-scale datasets under supervised learning conditions, their performance is often significantly restricted when applied to complex greenhouse scenarios, where large datasets are scarce. Particularly in unstructured agricultural environments, frequent changes in lighting conditions, the visual similarity between crops and weeds, and the unpredictability of weather add significant complexity to the model&#x2019;s ability to process such scenes. Moreover, due to the diversity of crop species, the complexity of background environments, and the difficulties associated with data collection, large-scale deep learning datasets in agriculture are relatively rare. As a result, the training and evaluation of algorithms often rely on small datasets collected by researchers, which may not adequately represent the complexities of real-world situations. In dense tomato crops, for example, occlusions between similar objects are frequent, and manually annotating such complex scenes in real datasets is both costly and prone to human bias and inaccuracies. Therefore, constructing a high-quality synthetic dataset is a more suitable solution to address this problem. Synthetic data can provide precise ground-truth annotations and allow for variable control to simulate different occlusion and lighting conditions, thereby offering models more diverse and comprehensive training data.</p>
<p>The use of synthetic datasets effectively compensates for the challenges in real data collection in agricultural greenhouse environments and provides more consistent training and testing conditions in Sim-to-real transfer scenarios. This approach enables models to better generalize in complex greenhouse settings, improving the accuracy and robustness of amodal perception tasks. As a result, it offers more reliable technical support for automated detection, disease recognition, and fruit harvesting in agriculture. In some weakly supervised learning studies (<xref ref-type="bibr" rid="B5">Cinbis et&#xa0;al., 2016</xref>), ground-truth labels are derived from self-generated annotations. For example, in (<xref ref-type="bibr" rid="B37">Yang et&#xa0;al., 2024</xref>), self-supervised learning is used to train deep learning models for target segmentation, with self-generated labels acting as ground truth. These models are then evaluated and tested in experimental environments. However, self-generated labels are based on model predictions of object shapes, which may differ from the actual shape of the target. To bridge this gap, researchers have made various attempts. A notable effort is the tomato dataset created by (<xref ref-type="bibr" rid="B40">Zhou et&#xa0;al., 2021</xref>). They proposed a synthetic dataset method by simulating tomato growth environments using software, followed by rendering tomato images and generating segmentation labels. However, their dataset only annotated the visible parts of the instances, without addressing the occluded parts. In amodal instance segmentation, occluded regions must be annotated in alignment with the ground truth, though the ground truth itself may sometimes be inaccurate. Our new dataset offers valuable solutions for addressing the challenge of obtaining ground-truth labels for occluded tomatoes in greenhouse environments. With this dataset, models can more effectively handle occluded instances, leading to enhanced precision and robustness in machine learning models for agricultural applications.</p>
<p>To tackle the problem of occluded tomato segmentation in greenhouse scenarios, this study proposes a deep learning-based amodal segmentation method focused on reconstructing the occluded parts of tomatoes. Starting from the requirements of greenhouse vision systems, this research assists in detecting and locating grasp points. A synthetic dataset, Tomato-sim, was constructed in a virtual environment to meet the data needs of current vision systems. The model was then trained and tested using an RGB-D amodal instance segmentation network embedded with the CGA module. Finally, the model was validated on a tomato occlusion test dataset from real greenhouse scenes. This method aims to improve amodal segmentation performance for occluded tomatoes in complex greenhouse environments by leveraging synthetic datasets and the CGA module.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Materials and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Data acquisition</title>
<p>In this study, 2,000 images containing both RGB and depth information were generated across 40 different scenario conditions, using Blender Proc (<xref ref-type="bibr" rid="B6">Denninger et&#xa0;al., 2019</xref>) for photorealistic rendering. In the Blender software, various tomato models of different shapes and colors, along with their branches and leaves, were constructed. Between 1 to 10 tomatoes were randomly placed in the scene, and images were captured by randomly setting the camera pose. This approach enabled the acquisition of ground-truth RGB images for each instance or frame directly from the computer-rendered 3D scenes.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Camera sampling and lighting condition settings</title>
<p>To capture images of tomatoes under various occlusion conditions, we constructed synthetic greenhouse scenes using pre-designed 3D models of tomato fruits and branches. These objects were randomly placed in 3D space with varying x, y, z coordinates and orientations. Camera viewpoints were uniformly sampled from two concentric hemispheres centered on the tomato plant, ensuring sufficient angular diversity for simulating occlusions (as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>).</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Diagram illustrating two setups. <bold>(A)</bold> Data collection platform with cameras positioned around tomatoes on a rectangular surface, showing dimensions width (w) and length (l) in a 3D space with axes x, y, and z. <bold>(B)</bold> Shooting distance with cameras viewing tomatoes at different ranges, labeled r_view_lower and r_view_upper.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g002.tif">
<alt-text content-type="machine-generated">Diagram illustrating two setups. Panel A depicts a data collection platform with cameras positioned around tomatoes on a rectangular surface, showing dimensions width (w) and length (l) in a 3D space with axes x, y, and z. Panel B shows the shooting distance with cameras viewing tomatoes at different ranges, labeled r_view_lower and r_view_upper.</alt-text>
</graphic>
</fig>
<p>The sampling range was controlled by two parameters, l and w, where l&#x2208;[1,2] and w&#x2208;[2,4] meters. The inner and outer radius bounds for the viewpoint sampling were defined as shown in <xref ref-type="disp-formula" rid="eq1">Equations 1</xref> and <xref ref-type="disp-formula" rid="eq2">2</xref>.</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>w</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext>max</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>w</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>u</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mn>1.7</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>w</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In greenhouse environments, variations in lighting can significantly affect the visual appearance, color distribution, and surface texture of objects, which in turn influence the performance of image-based object detection and segmentation algorithms. For instance, under strong lighting conditions, the increased illumination intensity enhances contrast within the image, making object edges appear sharper and more distinct, thereby facilitating foreground-background separation and improving segmentation accuracy. In contrast, under weak lighting, the reduction in contrast leads to less pronounced object boundaries, resulting in blurred edges and a higher risk of segmentation failure or misclassification.</p>
<p>Compared to single-modality RGB data, the use of RGB-D inputs provides richer multi-source information. In particular, depth data remains relatively invariant to changes in lighting conditions and shadows, offering more stable structural cues for object localization and shape estimation. This property enables the model to maintain reliable performance even in complex lighting environments, where RGB images alone may suffer from intensity distortion or loss of detail due to overexposure, underexposure, or shadowing effects.</p>
<p>To simulate diverse lighting conditions that realistically reflect the variability found in greenhouse settings, we introduced randomized lighting during synthetic data generation. Specifically, three distinct illumination scenarios were designed: strong, normal, and weak lighting, corresponding to different levels of intensity and contrast observed in real greenhouses. For each sampled camera viewpoint, between 0 and 2 spherical light sources were randomly added to the scene to emulate these lighting conditions (as shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>, in pink). These light sources were placed using a strategy consistent with the camera viewpoint sampling, namely within the same concentric hemispherical region centered on the target object. The spatial bounds for light source placement were defined relative to the camera&#x2019;s upper view radius.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Sampling of lighting conditions.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g003.tif">
<alt-text content-type="machine-generated">Diagram depicting a semicircular setup with three red tomatoes at the center. Three sensors or devices are positioned around the arc, labeled as &#x201c;Light.&#x201d; Two horizontal arrows extend from the tomatoes to the arc's edges, labeled &#x201c;rlight_lower&#x201d; and &#x201c;rlight_upper,&#x201d; indicating distances.</alt-text>
</graphic>
</fig>
<p>The sampling constraints for the lower and upper radii of the lighting hemisphere, r<sub>light_lower</sub> and r<sub>light_upper</sub> are defined by <xref ref-type="disp-formula" rid="eq3">Equations 3</xref>, <xref ref-type="disp-formula" rid="eq4">4</xref>:</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>u</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>0.1</mml:mn>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>u</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>This setup enabled us to simulate soft shadows, directional lighting, and realistic greenhouse illumination by rendering synthetic images under three distinct lighting conditions-strong, normal, and weak-reflecting the typical variability observed in natural greenhouse environments. Each illumination condition was rendered using physically-based materials with enabled shadow casting and reflection, producing realistic phenomena such as soft shadows, directional highlights, and illumination gradients. Example outputs, including RGB and corresponding depth images under different lighting levels, are shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>
<bold>(A)</bold> RGB image of tomato plants under strong illuminance. <bold>(B)</bold> RGB image of tomato plants under normal illuminance. <bold>(C)</bold> RGB image of tomato plants under low illuminance. Below each panel, the corresponding depth image shows varying shading levels, with <bold>(C)</bold> being the darkest and <bold>(A)</bold> the lightest. The depth values correspond to true distances ranging from 0.25 to 5.46 m.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g004.tif">
<alt-text content-type="machine-generated">RGB images of tomato plants under strong, normal, and low illuminance are labeled A, B, and C. Below, corresponding depth images show varying shading levels, with C being the darkest and A the lightest. The depth values correspond to true distances ranging from 0.25 to 5.46 m.</alt-text>
</graphic>
</fig>
<p>Although synthetic data cannot fully replicate real-world conditions, our dataset design incorporates variability in both lighting and viewpoints to minimize the domain gap. The model was trained solely on synthetic RGB-D images, and its generalization capability was evaluated through inference on both synthetic and real-world test sets.</p>
<p>In this study, we partitioned the dataset into training and testing sets at a ratio of 8:2. <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref> presents the distribution of RGB images in the training set, while <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> shows the distribution in the testing set. Since the position of each tomato on the plant affects both light intensity and the degree of occlusion, we categorized the images into three levels based on occlusion rate: 0&#x2013;10% (low occlusion), 10&#x2013;30% (moderate occlusion), and 30&#x2013;100% (high occlusion). Here, 0% indicates complete visibility, while 100% represents total occlusion.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>The distribution of RGB image data in the training set of tomato-sim.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Occlusion rate(%)</th>
<th valign="middle" align="center">Low illuminance</th>
<th valign="middle" align="center">Normal illuminance</th>
<th valign="middle" align="center">Strong illuminance</th>
<th valign="middle" align="center">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">[0,10]</td>
<td valign="middle" align="center">25/84</td>
<td valign="middle" align="center">50/158</td>
<td valign="middle" align="center">25/78</td>
<td valign="middle" align="center">100/320</td>
</tr>
<tr>
<td valign="middle" align="center">[10,30]</td>
<td valign="middle" align="center">50/220</td>
<td valign="middle" align="center">125/582</td>
<td valign="middle" align="center">50/228</td>
<td valign="middle" align="center">225/1030</td>
</tr>
<tr>
<td valign="middle" align="center">[30,100]</td>
<td valign="middle" align="center">125/694</td>
<td valign="middle" align="center">425/2155</td>
<td valign="middle" align="center">125704</td>
<td valign="middle" align="center">625/3553</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">200/998</td>
<td valign="middle" align="center">600/2895</td>
<td valign="middle" align="center">200/1010</td>
<td valign="middle" align="center">1000/4903</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>The distribution of RGB image data in the test set of tomato-sim.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Occlusion rate(%)</th>
<th valign="middle" align="center">Low illuminance</th>
<th valign="middle" align="center">Normal illuminance</th>
<th valign="middle" align="center">Strong illuminance</th>
<th valign="middle" align="center">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">[0,10]</td>
<td valign="middle" align="center">5/18</td>
<td valign="middle" align="center">12/38</td>
<td valign="middle" align="center">5/16</td>
<td valign="middle" align="center">22/72</td>
</tr>
<tr>
<td valign="middle" align="center">[10,30]</td>
<td valign="middle" align="center">10/48</td>
<td valign="middle" align="center">24/95</td>
<td valign="middle" align="center">10/45</td>
<td valign="middle" align="center">44/188</td>
</tr>
<tr>
<td valign="middle" align="center">[30,100]</td>
<td valign="middle" align="center">25/102</td>
<td valign="middle" align="center">84/342</td>
<td valign="middle" align="center">25/110</td>
<td valign="middle" align="center">134/554</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">40/168</td>
<td valign="middle" align="center">120/475</td>
<td valign="middle" align="center">40/171</td>
<td valign="middle" align="center">200/814</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Columns and rows contain image categories (number of images/number of instances). Each row corresponds to a different level of occlusion in the tomatoes.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>These thresholds were determined based on both the natural clustering of occlusion levels observed in our manually annotated real-world greenhouse images and commonly adopted practices in agricultural vision research (e.g., <xref ref-type="bibr" rid="B37">Yang et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B17">Li et&#xa0;al., 2022</xref>). For consistency, the same occlusion thresholds were also applied to the synthetic dataset. The same grouping standard is used consistently in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> for data organization and performance evaluation.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>The acquisition of occlusion masks</title>
<p>The synthetic 3D scene-generated dataset offers a high degree of annotation flexibility, providing amodal instance masks, complete appearances, occlusion order, and layer order for all objects in the scene. For each view, the system captures RGB and depth images of the desktop scene, utilizing the built-in instance segmentation feature of NVIDIA&#x2019;s Isaac Sim Replicator Composer to obtain instance segmentation masks for the entire scene. Subsequently, amodal and modal masks for each object are extracted from the instance segmentation masks. The occlusion mask and occlusion rate of each object are then calculated. The occlusion mask is obtained by subtracting the modal mask from the amodal mask, as illustrated in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref> and formulated in <xref ref-type="disp-formula" rid="eq5">Equation 5</xref>.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>
<bold>(A)</bold> Amodal mask shown as a full white circle on a black background. <bold>(B)</bold> Visible mask shown as a partial white crescent. <bold>(C)</bold> Occlusion mask shown as a smaller white shape.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g005.tif">
<alt-text content-type="machine-generated">Three panels labeled A, B, and C show different mask types on a black background. Panel A displays a full white circle labeled &#x201c;Amodal mask.&#x201d; Panel B shows a partial white crescent labeled &#x201c;Visible mask.&#x201d; Panel C features a smaller white shape labeled &#x201c;Occlusion mask."</alt-text>
</graphic>
</fig>
<disp-formula id="eq5">
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mi>V</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The occlusion rate is calculated by dividing the number of pixels in the occlusion mask by the number of pixels in the amodal mask. If an object&#x2019;s occlusion rate equals 1, it means the object is completely occluded from the viewpoint, and the annotation for that object is not saved for that view. In such cases, the object&#x2019;s visibility is disabled to capture the mask for the next object.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Greenhouse tomato dataset under real scenarios</title>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Collection equipment</title>
<p>The greenhouse tomato dataset used in this study was entirely collected using the Azure Kinect DK depth camera. During image acquisition, the Azure Kinect depth camera was utilized to capture RGB-D images, with the color camera set to a resolution of 1920&#xd7;1080 pixels at 30 frames per second (fps) and the depth camera set to a resolution of 640&#xd7;576 pixels at 30 fps.</p>
<p>Azure Kinect, developed by Microsoft, is a depth camera capable of simultaneously capturing both RGB and depth data. It features a high-resolution and high-sensitivity lens, capable of capturing high-quality depth information within a range of 0 to 10 meters. The depth camera of Azure Kinect uses time-of-flight (ToF) technology, which projects modulated light in the near-infrared spectrum onto the scene and records the time it takes for the light to travel from the camera to the scene and back. This travel time, along with the speed of light, is used to calculate depth values for different positions in the scene, generating a depth map. To ensure the generalizability of the data, the tomato plants were randomly photographed from multiple angles and positions under different lighting conditions within the greenhouse. Each image set includes RGB and corresponding depth images. The captured RGB and depth images were registered, ensuring that the pixels in the RGB image corresponded to the distance-representing pixels in the depth image. Finally, the images were cropped to 640&#xd7;480 pixels for both RGB and depth.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Greenhouse data acquisition</title>
<p>
<xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref> presents the data distribution of the real-world test set, where the intensity of light and the level of occlusion vary across different positions on the tomato plants. Based on the degree of occlusion, the images are categorized into three levels: 0-10%, 10-30%, and 30-100%, with 0% indicating no occlusion and 100% indicating full occlusion.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Construction of tomato datasets under real greenhouse scenarios.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Occlusion rate(%)</th>
<th valign="middle" align="center">Low illuminance</th>
<th valign="middle" align="center">Normal illuminance</th>
<th valign="middle" align="center">Strong illuminance</th>
<th valign="middle" align="center">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">[0,10]</td>
<td valign="middle" align="center">5/10</td>
<td valign="middle" align="center">12/25</td>
<td valign="middle" align="center">5/11</td>
<td valign="middle" align="center">22/46</td>
</tr>
<tr>
<td valign="middle" align="center">[10,30]</td>
<td valign="middle" align="center">8/18</td>
<td valign="middle" align="center">24/65</td>
<td valign="middle" align="center">8/16</td>
<td valign="middle" align="center">40/99</td>
</tr>
<tr>
<td valign="middle" align="center">[30,100]</td>
<td valign="middle" align="center">27/72</td>
<td valign="middle" align="center">84/272</td>
<td valign="middle" align="center">27/91</td>
<td valign="middle" align="center">138/435</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">40/100</td>
<td valign="middle" align="center">120/362</td>
<td valign="middle" align="center">40/118</td>
<td valign="middle" align="center">200/580</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Tomato images captured under different lighting conditions in real greenhouse scenarios, and their corresponding depth images are also collected, as shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>
<bold>(A)</bold> Brightly lit tomato in RGB image. <bold>(B)</bold> Normally lit tomatoes in RGB image. <bold>(C)</bold> Dimly lit tomatoes in RGB image. The bottom row shows the corresponding depth images under strong, normal, and low illuminance. Labels indicate lighting conditions. The depth values correspond to true distances ranging from 0.25 to 5.46 m.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g006.tif">
<alt-text content-type="machine-generated">Three panels showing tomatoes in different lighting conditions. Top row (RGB): A) Brightly lit tomato, B) Normally lit tomatoes, C) Dimly lit tomatoes. Bottom row (Depth): Corresponding depth images under strong, normal, and low illuminance. Labels indicate lighting  conditions.The depth values correspond to true distances ranging from 0.25 to 5.46 m.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Data annotation</title>
<p>Unlike other image segmentation tasks, instance segmentation requires pixel-level masks for visible objects, while amodal segmentation not only needs visible object masks but also integrates semantic labels for both visible and occluded parts of the scene. After mean cloning and fusion (<xref ref-type="bibr" rid="B8">Farbman et&#xa0;al., 2009</xref>), the dataset easily captures more semantic information about the target images. To segment the occluded areas, the combined mask of visible and invisible regions after image fusion is subtracted from the visible mask before fusion, as illustrated in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>
<bold>(A)</bold> Ripe red tomato outlined in blue for cropping, with a smaller green tomato nearby. <bold>(B)</bold> Target image with two ripe red tomatoes on a vine. <bold>(C)</bold> Final result, where the tomato from <bold>(A)</bold> has been cropped and pasted onto <bold>(B)</bold> to create an occlusion effect.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g007.tif">
<alt-text content-type="machine-generated">Panel A shows a ripe red tomato outlined in blue for cropping, with a smaller green tomato nearby. Panel B displays the target image with two ripe red tomatoes on a vine. Panel C depicts the final result, where the tomato from Panel A has been cropped and pasted onto Panel B to create an occlusion effect.</alt-text>
</graphic>
</fig>
<p>In this study, the LabelMe tool (<xref ref-type="bibr" rid="B29">Russell et&#xa0;al., 2008</xref>) was used to annotate each region hierarchically, and 200 images with ground-truth amodal masks were selected as the test set. Annotating an entire image takes approximately 5 minutes, with each instance requiring around 0.5 minutes on average. Compared to the efficient construction of synthetic datasets, manual annotation in real-world scenes is time-consuming, highlighting the advantage of synthetic datasets in improving data annotation efficiency.</p>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>RGB-D-based amodal instance segmentation method for tomatoes</title>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>RGB-D-based greenhouse tomato amodal segmentation model</title>
<p>The CGA-ASNet architecture, as illustrated in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8</bold>
</xref>, consists of two main components: the feature extraction network and the segmentation prediction network. The feature extraction and fusion network first extracts RGB features and depth features separately from the input RGB and Depth images. The depth features from the C3, C4, and C5 layers of the CGA-50 Backbone are concatenated with the corresponding RGB features from the same layers. A 1&#xd7;1 convolution is applied to fuse the RGB-D features, reducing the channel dimensions. This fusion forms an RGB-D feature pyramid, which is then passed through the Region Proposal Network (RPN) and RoIAlign (Region of Interest Align) layers to generate the RGB-D features. These multi-dimensional feature maps are then fed into the segmentation prediction network for segmentation tasks. The model incorporates the CGA module, based on the Unseen Object Amodal Instance Segmentation (UOAIS) architecture. The CGA module is composed of the CFT module (proposed in this study) and the GAM module (<xref ref-type="bibr" rid="B22">Liu Y. et&#xa0;al., 2021</xref>). The improved model is highly adaptable to the constructed synthetic dataset, ensuring both high accuracy and enhanced training and inference speed, even with smaller datasets. The model effectively handles variations in lighting conditions and tomato color changes in greenhouse environments. Additionally, the introduction of a shape convolution module strengthens the model&#x2019;s perception of tomato shape and position, reducing the impact of occlusions caused by branches and leaves. The model&#x2019;s loss function is defined as shown in <xref ref-type="disp-formula" rid="eq6">Equation 6</xref>.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>Overall architecture diagram of the CGA-ASNet model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g008.tif">
<alt-text content-type="machine-generated">Diagram illustrating a neural network model for tomato image analysis. The Feature Extraction Network processes RGB and Depth Inputs of tomatoes using CGA-backbones. Fusion FPN+RPN layers combine features, feeding into the Segmentation Prediction Network. This network uses RoI Align and Hierarchical Fusion for feature extraction, segmentation, and mask predictions. Outputs include visible, amodal, and occlusion masks, class and bounding box predictions, and a processed tomato image. Labels and legend explain components like CGA, FC (Fully Connect), and masks for segmentation.</alt-text>
</graphic>
</fig>
<disp-formula id="eq6">
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mrow>
<mml:mtext>loss</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mrow>
<mml:mtext>cls</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mrow>
<mml:mtext>box</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mrow>
<mml:msub>
<mml:mtext>O</mml:mtext>
<mml:mrow>
<mml:mtext>cls</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>L</mml:mtext>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>rpn</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>cls</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mtext>L</mml:mtext>
<mml:mo>_</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>rpn</mml:mtext>
<mml:mo>_</mml:mo>
<mml:mtext>loc</mml:mtext>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Among them, the abbreviation L<sub>cls</sub> refers to the loss of classes, L<sub>box</sub> refers to the loss of bounding boxes, L<sub>V</sub> refers to the loss of non-modal mask losses, and L<sub>Ocls</sub> refers to the loss of occlusion classification.</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>CFT attention</title>
<p>To enhance the global modeling capability of ResNet50, we propose a Contextual Features Transformer (CFT) module, the structure of which is illustrated in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9</bold>
</xref>. This module replaces the standard 3&#xd7;3 convolution in the residual block. Unlike conventional self-attention mechanisms that compute attention weights based on dot-product similarity, the proposed CFT module leverages learnable convolutions to generate attention scores. This design integrates the inductive bias of convolution with the long-range dependency modeling strength of attention, effectively avoiding the scale sensitivity issues of dot-product attention while providing more stable and spatially aware representations. Moreover, it introduces only minimal computational overhead, making it particularly suitable for dense prediction tasks.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>Structure of CFT self-attention module. Here, the * is not alone. When combined with conv, it forms conv(*), representing a 1x1 convolution.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g009.tif">
<alt-text content-type="machine-generated">Flowchart illustrating a convolutional neural network (CNN) process for identifying tomato features. The image shows a photograph of tomatoes connected to various processing stages labeled as queries, keys, and values. These stages include convolutional operations and outputs labeled as static, dynamic, and weights, resulting in a final output of &#x201c;Tomato Feathers (F)."</alt-text>
</graphic>
</fig>
<p>The proposed Convolutional Feature Transformer (CFT) module is designed to simultaneously model local details and global dependencies by leveraging a convolution-based attention mechanism in place of traditional dot-product attention. Specifically, the input feature map <italic>X</italic> is first passed through a 3&#xd7;3 convolution to extract the key feature <italic>K</italic>, preserving spatial context. A second 3&#xd7;3 convolution is applied to further enhance local information within the key representation. In parallel, the query <italic>Q</italic> and value<italic>V</italic> features are obtained from <italic>X</italic> using two separate 1&#xd7;1 convolutions for dimensionality reduction while preserving feature structure.</p>
<p>After computing these features, spatial dependencies are modeled by concatenating the key and query features along the channel dimension. This combined representation is passed through two 1&#xd7;1 convolutions to produce the spatial interaction logits &#x3b8;, which are normalized by a Softmax function to yield the attention weight matrix <italic>W</italic>. The matrix <italic>W</italic> is then applied to the value feature <italic>V</italic> via weighted summation. To further incorporate global context, the value feature is enhanced using a dilated convolution layer <italic>M</italic>, which increases the receptive field without reducing resolution. The result is fused with the original key <italic>K</italic> using a final 1&#xd7;1 convolution, producing the final output <italic>Y</italic>, denoted as the tomato feature <italic>F</italic> for subsequent segmentation tasks.</p>
<p>The core idea of CFT is to replace the conventional dot-product attention with convolution-based attention, allowing the network to better integrate spatial inductive bias and global dependencies. The complete formulation is given as shown in <xref ref-type="disp-formula" rid="eq7">Equations 7</xref>&#x2013;<xref ref-type="disp-formula" rid="eq9">9</xref>.</p>
<disp-formula id="eq7">
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mtext>&#x3b8;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>Conv</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>K</mml:mtext>
<mml:mo>&#x2295;</mml:mo>
<mml:mtext>Q</mml:mtext>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq8">
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:mtext>W</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>Softmax</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>&#x3b8;</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq9">
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mtext>Y</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>Conv</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>K</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>M</mml:mtext>
<mml:mo>&#x2297;</mml:mo>
<mml:mtext>V</mml:mtext>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where Conv(*) denotes a 1&#xd7;1 convolution; &#x2295; represents concatenation; &#x2297; denotes matrix multiplication.</p>
</sec>
<sec id="s2_3_3">
<label>2.3.3</label>
<title>GAM attention</title>
<p>The Global Attention Mechanism (GAM) enhances feature representations by applying attention along both channel and spatial dimensions. It consists of two independent submodules: the Channel Attention Module (CAM) and the Spatial Attention Module (SAM). The overall architecture is illustrated in <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref>.</p>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>Global attention mechanism.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g010.tif">
<alt-text content-type="machine-generated">Diagram illustrating an attention mechanism. Input features F1 are processed through a Channel Attention Module and a Spatial Attention Module, represented by diagrams labeled Mc and Ms. The outputs are combined and directed to produce Output features F3, depicted as three squares with varying intensities of blue circles.</alt-text>
</graphic>
</fig>
<p>As illustrated in <xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref>, CAM first applies global average pooling and global max pooling across spatial dimensions of the input feature map <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:mtext>F</mml:mtext>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mtext>R</mml:mtext>
<mml:mrow>
<mml:mtext>C&#xd7;H&#xd7;W</mml:mtext>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, resulting in two descriptors of size R<sup>C</sup>. These descriptors are then passed through a shared two-layer MLP, where the first layer reduces the dimension by a ratio r, and the second layer restores it to C. After element-wise summation and a sigmoid activation, the resulting attention map M<sub>c</sub> is used to reweight the input feature map channel-wise.</p>
<fig id="f11" position="float">
<label>Figure&#xa0;11</label>
<caption>
<p>Channel attention module.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g011.tif">
<alt-text content-type="machine-generated">Flowchart of a neural network process, starting with &#x201c;Input features F1&#x201d; represented by stacked blue squares. It proceeds through a &#x201c;Permutation&#x201d; box, reordering dimensions from C&#xd7;W&#xd7;H to W&#xd7;H&#xd7;C. Then, it passes through an &#x201c;MLP&#x201d; with overlapping triangles, followed by a &#x201c;Reverse Permutation.&#x201d; Finally, it goes through a &#x201c;Sigmoid&#x201d; function, resulting in the output labeled Mc(F&#x2081;), shown as overlaid blurry rectangles.</alt-text>
</graphic>
</fig>
<p>SAM further refines the output from CAM by emphasizing important spatial locations. As shown in <xref ref-type="fig" rid="f12">
<bold>Figure&#xa0;12</bold>
</xref>, it applies average pooling and max pooling across channels, producing two <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> feature maps, which are concatenated and passed through a 7&#xd7;7 convolution followed by a sigmoid activation to generate the spatial attention map M<sub>s</sub>. This map is multiplied element-wise with the input to produce the final attention-weighted output.</p>
<fig id="f12" position="float">
<label>Figure&#xa0;12</label>
<caption>
<p>Spatial attention module.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g012.tif">
<alt-text content-type="machine-generated">Diagram showing a sequence of operations on input features F2 with dimensions C&#xd7;H&#xd7;W. A 7&#xd7;7 convolution reduces dimensions to C/r&#xd7;H&#xd7;W. Another 7&#xd7;7 convolution returns dimensions to C&#xd7;H&#xd7;W, followed by a sigmoid function, resulting in Ms(F2).</alt-text>
</graphic>
</fig>
</sec>
<sec id="s2_3_4">
<label>2.3.4</label>
<title>Segmentation prediction network</title>
<p>The segmentation prediction network is composed of four main branches: the Bounding Box Prediction Branch, the Visible Mask Prediction Branch, the Amodal Mask Prediction Branch, and the Occlusion Classification Prediction Branch. The Bounding Box Prediction Branch takes the 7&#xd7;7 feature map output from the RPN (Region Proposal Network) and passes it through two fully connected layers to predict the bounding box B and class C. The feature map is then upsampled to a 14&#xd7;14 feature map to provide bounding box features for the subsequent branches, ensuring that instance masks are segmented within the predicted bounding box. The Visible Mask Prediction Branch, the Amodal Mask Prediction Branch, and the Occlusion Classification Prediction Branch utilize the 14&#xd7;14 feature map from the RPN, along with features fused from the previous branches, to predict the visible mask V, the amodal mask A, and the occlusion classification O, respectively. The mathematical formulations for each branch are expressed in <xref ref-type="disp-formula" rid="eq10">Equations 10</xref>&#x2013;<xref ref-type="disp-formula" rid="eq13">13</xref>.</p>
<disp-formula id="eq10">
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>h</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>B</mml:mtext>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mrow>
<mml:mtext>RoI</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq11">
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>h</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>B</mml:mtext>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mrow>
<mml:mtext>RoI</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq12">
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>O</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>h</mml:mtext>
<mml:mtext>O</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>B</mml:mtext>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mrow>
<mml:mtext>RoI</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq13">
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:mtext>V</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>A</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>O</mml:mtext>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext>P</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>V</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>P</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>A</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>P</mml:mtext>
<mml:mtext>O</mml:mtext>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mtext>O</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the segmentation prediction network, F<sub>B</sub>, F<sub>RoI</sub>, F<sub>V</sub>, F<sub>A</sub>, and F<sub>O</sub> represent the bounding box feature, the RoI feature, the visible mask feature, the amodal mask feature, and the occlusion mask feature, respectively. The hierarchical fusion modules hv, h<sub>A</sub> and ho correspond to the visible mask, amodal mask, and occlusion classification branches. Specifically, the hierarchical fusion module integrates each input feature and reduces the channel dimensions through three 3&#xd7;3 convolution layers to decrease the parameter count. These are then fed into another set of three 3&#xd7;3 convolution layers to generate the task-specific features for each branch. The prediction layers P<sub>V</sub>, P<sub>A</sub> and P<sub>o</sub> are responsible for predicting the visible mask, amodal mask, and occlusion classification, respectively. P<sub>V</sub> and P<sub>A</sub> use 2&#xd7;2 deconvolutions and a fully connected layer, while P<sub>o</sub> consists of a fully connected layer to output the final results.</p>
</sec>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results and analysis</title>
<sec id="s3_1">
<label>3.1</label>
<title>Training and parameter setting</title>
<p>This study aims to address the lack of real RGB-D datasets by applying deep learning models, specifically focusing on tomatoes. Through software, synthetic RGB-D images simulating occluded tomatoes in a greenhouse environment are generated to build a diverse and high-quality dataset. The convolutional neural network (CNN) extracts and integrates features from both RGB and depth images using feature extraction algorithms. Multiple detection branches are employed to predict the visible part masks and the contours of the occluded parts of the objects. A hierarchical occlusion modeling mechanism is applied to improve the accuracy of amodal segmentation for tomatoes.</p>
<p>During model training, comparisons between different datasets (synthetic and real-world datasets) are conducted for both training and testing. To ensure fairness in the training and testing process, all tasks are performed on the same hardware platform. The experimental platform consists of a Dell Precision 7920 with 64GB RAM, a 2.1GHz CPU with 16 cores, and an NVIDIA A6000 GPU with 48GB GDDR6 VRAM and 10,752 CUDA cores. Initial training parameters are listed in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Training process related parameters.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Parameter name</th>
<th valign="middle" align="center">Parameter values</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Image Size</td>
<td valign="middle" align="center">640&#xd7;480</td>
</tr>
<tr>
<td valign="middle" align="center">Batch Size of Images</td>
<td valign="middle" align="center">2</td>
</tr>
<tr>
<td valign="middle" align="center">Initial Learning Rate</td>
<td valign="middle" align="center">0.00125</td>
</tr>
<tr>
<td valign="middle" align="center">Maximum Number of Iterations</td>
<td valign="middle" align="center">90000</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Evaluation metrics for tomato amodal segmentation quality</title>
<p>We adopt several evaluation metrics to quantitatively assess the instance-level segmentation performance, including Precision, Recall, F1-Score, F@75 (<xref ref-type="bibr" rid="B24">Ochs et&#xa0;al., 2013</xref>), and mean Intersection over Union (mIoU), as defined in <xref ref-type="disp-formula" rid="eq14">Equations 14</xref>&#x2013;<xref ref-type="disp-formula" rid="eq18">18</xref>. Precision measures the proportion of correctly predicted positive instances among all predicted positives, while Recall measures the proportion of correctly predicted positive instances among all actual positives. The F1-Score is the harmonic mean of Precision and Recall, providing a balanced evaluation of model performance.</p>
<p>F@.75 is an instance-level metric based on the F1-score, representing the proportion of ground-truth instances that are successfully matched with predicted instances having an F1-score no less than 0.75. The pairwise F1-scores between predicted and ground-truth instances are computed, and the optimal one-to-one assignment is determined using the Hungarian algorithm. Finally, mean Intersection over Union (mIoU) is used to evaluate segmentation quality across all classes.</p>
<disp-formula id="eq14">
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:mtext>Precision</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext>TP</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>FP</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq15">
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mtext>Recall</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext>TP</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>FN</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq16">
<label>(16)</label>
<mml:math display="block" id="M16">
<mml:mrow>
<mml:mtext>F</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>Score</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mtext>Precision</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext>Recall</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>Precision</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>Recall</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq17">
<label>(17)</label>
<mml:math display="block" id="M17">
<mml:mrow>
<mml:mtext>F</mml:mtext>
<mml:mo>@</mml:mo>
<mml:mn>.75</mml:mn>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>j</mml:mtext>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mtext>M</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>F</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>j</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0.75</mml:mn>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mtext>N</mml:mtext>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq18">
<label>(18)</label>
<mml:math display="block" id="M18">
<mml:mrow>
<mml:mtext>mIoU=</mml:mtext>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mn>k+1</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mtext>&#xa0;</mml:mtext>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mtext>k</mml:mtext>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:mtext>TP</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>FN+FP+TP</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where TP is the model correctly predicts positive instances; FP is model incorrectly predicts positive instances; FN is the model incorrectly predicts negative instances; FP is the model correctly predicts negative instances. F<sub>i,j</sub> is the F1-score between predicted instance i and ground-truth instance j, M is the optimal one-to-one matching obtained via the Hungarian algorithm, and N is the total number of ground-truth instances.</p>
<p>In order to compare our method with other existing methods, we adopted the AP (Average Precision) and mAP (mean Average Precision) as evaluation metrics, which are commonly used for amodal segmentation tasks (<xref ref-type="bibr" rid="B16">Ke et&#xa0;al., 2021</xref>), as defined in <xref ref-type="disp-formula" rid="eq19">Equations 19</xref>, <xref ref-type="disp-formula" rid="eq20">20</xref>.</p>
<disp-formula id="eq19">
<label>(19)</label>
<mml:math display="block" id="M19">
<mml:mrow>
<mml:mtext>AP</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x222b;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mtext>P</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>r</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mtext>dr</mml:mtext>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq20">
<label>(20)</label>
<mml:math display="block" id="M20">
<mml:mrow>
<mml:mtext>mAP</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>k</mml:mtext>
</mml:mfrac>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>k</mml:mtext>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mtext>AP</mml:mtext>
</mml:mrow>
<mml:mtext>i</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Analysis of test results with different backbone networks</title>
<p>To investigate the impact of different feature extraction backbone networks on the performance of the CGA-ASNet model, we conducted a series of controlled experiments using six different backbone architectures: ResNet50, ResNet101 (<xref ref-type="bibr" rid="B14">He et&#xa0;al., 2016</xref>), ResNeXt50, ResNeXt101 (<xref ref-type="bibr" rid="B36">Xie et&#xa0;al., 2017</xref>), ConvNeXt-Tiny (<xref ref-type="bibr" rid="B21">Liu et&#xa0;al., 2022</xref>), and Swin-Tiny (<xref ref-type="bibr" rid="B20">Liu Z. et&#xa0;al., 2021</xref>). All experiments were performed under identical training and testing conditions, with RGB-D as the input modality and only the backbone network varied.</p>
<p>As shown in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>, ResNet50 consistently outperformed the other backbone networks in both amodal mask prediction and occlusion segmentation. Specifically, it achieved the highest amodal F@.75 score of 92.0 and a mean Intersection-over-Union (mIoU) of 81.4%. Although newer backbone architectures such as ConvNeXt-Tiny and Swin-Tiny showed competitive results, they did not surpass the performance of ResNet50 in our task setting. This suggests that ResNet50 remains a strong and stable backbone choice for occlusion-aware segmentation tasks, particularly in our CGA-ASNet framework.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Comparison results of different backbone networks.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Backbone</th>
<th valign="middle" colspan="2" align="center">Amodal</th>
</tr>
<tr>
<th valign="middle" align="center">F@.75</th>
<th valign="middle" align="center">mIoU(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">ResNet101</td>
<td valign="middle" align="center">81.9</td>
<td valign="middle" align="center">76.7</td>
</tr>
<tr>
<td valign="middle" align="center">ResNext50</td>
<td valign="middle" align="center">87.0</td>
<td valign="middle" align="center">75.7</td>
</tr>
<tr>
<td valign="middle" align="center">ResNext101</td>
<td valign="middle" align="center">87.7</td>
<td valign="middle" align="center">79.1</td>
</tr>
<tr>
<td valign="middle" align="center">ConvNext_Tiny</td>
<td valign="middle" align="center">88.4</td>
<td valign="middle" align="center">78.9</td>
</tr>
<tr>
<td valign="middle" align="center">Swin_Tiny</td>
<td valign="middle" align="center">89.3</td>
<td valign="middle" align="center">79.4</td>
</tr>
<tr>
<td valign="middle" align="center">ResNet50</td>
<td valign="middle" align="center">
<bold>92.0</bold>
</td>
<td valign="middle" align="center">
<bold>81.4</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bolded part is the most effective part in the backbone network and thus is supported.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Ablation study</title>
<p>Ablation experiments, commonly used to assess the influence of different components in a model, are an effective method for exploring the contributions of each module and gaining a deeper understanding of the model&#x2019;s behavior. As such, ablation experiments play a crucial role in the design of neural network structures. To verify the effectiveness of the CGA module, this study designed a series of ablation experiments. We used ResNet50 as the backbone network with R-50.pkl serving as the initial weight baseline. The experiments were divided into three parts: first, the CFT self-attention module and GAM attention module were individually embedded for testing; finally, both CFT and GAM were combined and embedded into the network for comparison to evaluate their specific contributions to improving network performance.</p>
<p>As shown in <xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>, the first row presents results from the baseline model without any modifications, achieving an F@.75 score of 92.0 and a mIoU of 81.4% for amodal masks. In the second experiment, where the CFT module was added, the F@.75 score increased to 93.5 and the mIoU to 82.6%, representing improvements of 1.5 and 1.2%, respectively. The third experiment introduced the GAM module, which raised the F@.75 score to 94.2, an increase of 2.2, and the mIoU to 82.4%, a 1% improvement. Finally, the model with the combined CFT and GAM modules, forming the CGA module, achieved an F@.75 score of 94.2 and a mIoU of 82.4%. These results demonstrate that the CGA module effectively captures more semantic information from tomatoes, significantly enhancing the segmentation performance.</p>
<table-wrap id="T6" position="float">
<label>Table&#xa0;6</label>
<caption>
<p>Ablation study.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Method</th>
<th valign="middle" colspan="2" align="center">Amodal</th>
</tr>
<tr>
<th valign="middle" align="center">F@.75</th>
<th valign="middle" align="center">mIoU(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Baseline</td>
<td valign="middle" align="center">92.0</td>
<td valign="middle" align="center">81.4</td>
</tr>
<tr>
<td valign="middle" align="center">Baseline+cft</td>
<td valign="middle" align="center">93.5</td>
<td valign="middle" align="center">82.6</td>
</tr>
<tr>
<td valign="middle" align="center">Baseline+gam</td>
<td valign="middle" align="center">94.2</td>
<td valign="middle" align="center">82.4</td>
</tr>
<tr>
<td valign="middle" align="center">Baseline+CGA</td>
<td valign="middle" align="center">94.2</td>
<td valign="middle" align="center">83.3</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Amodal segmentation results on test images with different degrees of occlusion</title>
<p>To evaluate the robustness of the improved amodal segmentation network, this study compared the baseline model with the CGA-embedded segmentation model across three subsets with occlusion levels greater than 0-10%, 10-30%, and 30-100%, using identical parameters. The results are shown in <xref ref-type="table" rid="T7">
<bold>Table&#xa0;7</bold>
</xref> and <xref ref-type="table" rid="T8">
<bold>Table&#xa0;8</bold>
</xref>. When the occlusion rate was below 10%, CGA-ASNet achieved an F@.75 score of 98.4 and a mIoU of 86.8%, both higher than the baseline model. For occlusion levels between 10% and 30%, and those above 30%, CGA-ASNet also outperformed the baseline model by 1.4 and 2.6, respectively.</p>
<table-wrap id="T7" position="float">
<label>Table&#xa0;7</label>
<caption>
<p>Baseline prediction results.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Occlusion Rate(%)</th>
<th valign="middle" colspan="2" align="center">Amodal</th>
</tr>
<tr>
<th valign="middle" align="center">F@.75</th>
<th valign="middle" align="center">mIoU(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">[0,10]</td>
<td valign="middle" align="center">98.2</td>
<td valign="middle" align="center">85.3</td>
</tr>
<tr>
<td valign="middle" align="center">[10,30]</td>
<td valign="middle" align="center">93.1</td>
<td valign="middle" align="center">82.8</td>
</tr>
<tr>
<td valign="middle" align="center">[30,100]</td>
<td valign="middle" align="center">86.7</td>
<td valign="middle" align="center">78.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T8" position="float">
<label>Table&#xa0;8</label>
<caption>
<p>CGA-ASNet prediction results.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Occlusion rate(%)</th>
<th valign="middle" colspan="2" align="center">Amodal</th>
</tr>
<tr>
<th valign="middle" align="center">F@.75</th>
<th valign="middle" align="center">mIoU(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">[0,10]</td>
<td valign="middle" align="center">98.4</td>
<td valign="middle" align="center">86.8</td>
</tr>
<tr>
<td valign="middle" align="center">[10,30]</td>
<td valign="middle" align="center">94.5</td>
<td valign="middle" align="center">83.4</td>
</tr>
<tr>
<td valign="middle" align="center">[30,100]</td>
<td valign="middle" align="center">89.3</td>
<td valign="middle" align="center">79.8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results indicate that, while segmentation performance declines as occlusion increases, CGA-ASNet consistently handles severe occlusions better than the baseline. As shown in <xref ref-type="fig" rid="f13">
<bold>Figure&#xa0;13</bold>
</xref>, when multiple tomatoes are stacked, the baseline model without the CGA module exhibited jagged contours in its predictions of occluded tomatoes, whereas our model generated smoother and more natural predictions. This demonstrates that the CGA module significantly enhances the model&#x2019;s ability to perceive and predict the edge shapes of segmented objects, improving overall prediction accuracy.</p>
<fig id="f13" position="float">
<label>Figure&#xa0;13</label>
<caption>
<p>
<bold>(A)</bold> Tomatoes with GT overlays in blue and orange. <bold>(B)</bold> Baseline result with mainly orange and green overlays. <bold>(C)</bold> CGA-ASNet result with accurate red and green overlays, showing segmentation improvements.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g013.tif">
<alt-text content-type="machine-generated">Three images labeled A, B, and C show tomatoes on a plant with different color overlays. A uses GT with blue and orange overlays, B is labeled Baseline with mainly orange and green overlays, and C shows CGA-ASNet with accurate red and green overlays indicating segmentation improvements.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Comparison of test results from different models</title>
<p>During the experimental design phase, we reviewed several recent representative amodal segmentation models, including pix2gestalt (<xref ref-type="bibr" rid="B25">Ozguroglu et&#xa0;al., 2024</xref>), AISDiff (<xref ref-type="bibr" rid="B34">Tran et&#xa0;al., 2024</xref>), and BLADE (<xref ref-type="bibr" rid="B19">Liu et&#xa0;al., 2024</xref>), etc. However, most models only support feature extraction of the RGB channels. These models cannot provide the feature support for image segmentation based on depth information. If the RGBD four-channel data is compressed into three channels for feature extraction, the obtained features cannot accurately represent the pixel semantics of the original image. To ensure reproducibility and fair comparison, we selected a group of well-established and publicly available models as baselines for evaluation.</p>
<p>In this study, the Tomato-sim dataset was trained on state-of-the-art (SOTA) models, including BC-net, AISFormer (<xref ref-type="bibr" rid="B35">Tran et&#xa0;al., 2022</xref>), ORCNN (<xref ref-type="bibr" rid="B10">Gen&#xe9;-Mola et&#xa0;al., 2023</xref>), and Uoais-net (<xref ref-type="bibr" rid="B1">Back et&#xa0;al., 2022</xref>), using identical parameters to compare different training data (Tomato-sim and real datasets). The models were tested on datasets constructed through mean clone fusion in both synthetic and real greenhouse scenarios. <xref ref-type="table" rid="T9">
<bold>Table&#xa0;9</bold>
</xref> presents the prediction results of different models.</p>
<table-wrap id="T9" position="float">
<label>Table&#xa0;9</label>
<caption>
<p>Comparison of predictions from different models.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Method</th>
<th valign="middle" align="center">Eval</th>
<th valign="middle" align="center">AP50(%)</th>
<th valign="middle" align="center">AP75(%)</th>
<th valign="middle" align="center">mAP(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="2" align="center">BC-net</td>
<td valign="middle" align="center">Tomato-sim</td>
<td valign="middle" align="center">87.4</td>
<td valign="middle" align="center">78.7</td>
<td valign="middle" align="center">70.2</td>
</tr>
<tr>
<td valign="middle" align="center">real</td>
<td valign="middle" align="center">85.7</td>
<td valign="middle" align="center">75.5</td>
<td valign="middle" align="center">66.7</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="center">AISFormer</td>
<td valign="middle" align="center">Tomato-sim</td>
<td valign="middle" align="center">92.3</td>
<td valign="middle" align="center">86.1</td>
<td valign="middle" align="center">74.4</td>
</tr>
<tr>
<td valign="middle" align="center">real</td>
<td valign="middle" align="center">89.9</td>
<td valign="middle" align="center">85.7</td>
<td valign="middle" align="center">72.5</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="center">ORCNN</td>
<td valign="middle" align="center">Tomato-sim</td>
<td valign="middle" align="center">73.3</td>
<td valign="middle" align="center">63.4</td>
<td valign="middle" align="center">55.7</td>
</tr>
<tr>
<td valign="middle" align="center">real</td>
<td valign="middle" align="center">72.3</td>
<td valign="middle" align="center">58.3</td>
<td valign="middle" align="center">52.4</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="center">Uoais-net</td>
<td valign="middle" align="center">Tomato-sim</td>
<td valign="middle" align="center">92.9</td>
<td valign="middle" align="center">82.3</td>
<td valign="middle" align="center">74.5</td>
</tr>
<tr>
<td valign="middle" align="center">real</td>
<td valign="middle" align="center">89.6</td>
<td valign="middle" align="center">78.3</td>
<td valign="middle" align="center">73.1</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="center">CGA-ASNet</td>
<td valign="middle" align="center">Tomato-sim</td>
<td valign="middle" align="center">94.3</td>
<td valign="middle" align="center">83.6</td>
<td valign="middle" align="center">78.3</td>
</tr>
<tr>
<td valign="middle" align="center">real</td>
<td valign="middle" align="center">93.1</td>
<td valign="middle" align="center">78.4</td>
<td valign="middle" align="center">75.0</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="f14">
<bold>Figure&#xa0;14</bold>
</xref> shows the performance of these segmentation models in the amodal segmentation task. From image (2), it can be observed that in the complex stacking scenario of tomatoes, our model exhibited strong robustness. Images (1) to (3) show prediction results from real greenhouse environments, while Images (4) to (6) display performance in virtual scenes. Although all models performed well in the virtual scenario, our model demonstrated the best segmentation ability, especially in handling complex occlusion and multi-layer stacking, achieving significantly higher segmentation accuracy compared to other models.</p>
<fig id="f14" position="float">
<label>Figure&#xa0;14</label>
<caption>
<p>Amodal segmentation results of different models.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g014.tif">
<alt-text content-type="machine-generated">Six rows of images show different object detection results on fruit clusters, each using a different method labeled GT, BC-net, AISFormer, ORCNN, Uoais-net, and CGA-ASNet. Each row represents a different set of fruit clusters under various lighting and backgrounds, highlighting the differences in segmentation quality and accuracy across the methods.</alt-text>
</graphic>
</fig>
<p>Furthermore, CGA-ASNet was evaluated in a real greenhouse environment to validate its practical applicability. As shown in <xref ref-type="fig" rid="f15">
<bold>Figure&#xa0;15</bold>
</xref>, we selected the best- and worst-performing baseline models&#x2014;ORCNN and AISFormer&#x2014;for direct comparison with our method. Most results demonstrate that our model produces high-quality amodal mask predictions, with natural and consistent mask distributions across the entire ROI. In contrast, both ORCNN and AISFormer exhibit varying degrees of segmentation incompleteness or inaccuracies. Our model achieves better overall shape recovery and boundary alignment, highlighting its superior performance under real-world conditions.</p>
<fig id="f15" position="float">
<label>Figure&#xa0;15</label>
<caption>
<p>
<bold>(A&#x2013;D)</bold> Tomatoes on vines with varying color changes under different algorithms. The bottom row shows ground truth (GT), ORCNN, AISFormer, and our method, with tomatoes highlighted using bounding boxes to indicate varying ripeness and detection accuracy. Each method depicts different levels of detail and color fidelity.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g015.tif">
<alt-text content-type="machine-generated">Comparison of image processing methods for tomato detection. The top row (labeled A-D) shows tomatoes on vines with varying color changes under different algorithms. The bottom row (labeled GT, ORCNN, AISFormer, Ours) shows tomatoes highlighted with bounding boxes, displaying varying levels of ripeness and detection accuracy. Each method depicts different levels of detail and color fidelity in processing tomato images.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Generalization evaluation on PApple_RGB-D-size dataset</title>
<p>To further assess the generalization capability of CGA-ASNet, we conducted cross-domain experiments on the PApple_RGB-D-Size dataset (<xref ref-type="bibr" rid="B10">Gen&#xe9;-Mola et&#xa0;al., 2023</xref>), which contains RGB-D images of apples under different illumination and occlusion conditions. This dataset significantly differs from the training domain in both fruit category, color distribution, and geometric structure, making it a suitable benchmark for evaluating robustness.</p>
<p>Without any additional fine-tuning, CGA-ASNet achieved an AP50 of 89.2%, AP75 of 76.1%, and a mean Average Precision (mAP) of 73.4%, demonstrating strong generalization ability and transferability across domains. These results suggest that the model can effectively learn domain-invariant features and accurately infer the complete shape of occluded objects even under unfamiliar visual and structural conditions. In addition, <xref ref-type="fig" rid="f16">
<bold>Figure&#xa0;16</bold>
</xref> illustrates representative qualitative results. Despite the domain shift, CGA-ASNet is able to predict coherent amodal masks and successfully complete severely occluded fruit regions.</p>
<fig id="f16" position="float">
<label>Figure&#xa0;16</label>
<caption>
<p>
<bold>(A)</bold> Apples with shadows under varied lighting. <bold>(B)</bold> Apples with color changes under uneven illumination. <bold>(C)</bold> Apples in brighter light showing increased brightness. <bold>(D)</bold> Apples in high illumination with strong contrast.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1664718-g016.tif">
<alt-text content-type="machine-generated">Panels labeled A to D show apple trees with different lighting conditions. Panels A and B depict apples with shadows and varied lighting causing color changes. Panels C and D feature apples in brighter light, highlighting contrasts in brightness and illumination.</alt-text>
</graphic>
</fig>
</sec>
</sec>
<sec id="s4" sec-type="conclusion">
<label>4</label>
<title>Conclusion</title>
<p>In the greenhouse environment, in order to ensure the accuracy of the non-destructive phenotype detection of tomato fruits, we constructed a virtual dataset of tomato fruits (Tomato-sim). This dataset simulated the shading conditions that occur during the actual growth of tomatoes. Additionally, for this dataset, we built an RGB-D image non-modal segmentation model based on the CGA module. We used the virtual data to train the model and then tested the model on the real data set. The following are some conclusions drawn based on the experimental results of this research work:</p>
<list list-type="order">
<list-item>
<p>The synthetic dataset used for amodal tomato segmentation, Tomato-sim, achieved an average precision of 78.3%, closely matching the 75.0% precision obtained from real data testing. This demonstrates that synthetic data can effectively compensate for the limitations of real data collection, especially in complex agricultural scenarios, by providing flexible and diverse training conditions that handle scene complexity and object occlusion.</p>
</list-item>
<list-item>
<p>The CGA module designed in this study effectively captures the semantic information of tomatoes, particularly excelling in handling occluded regions. Compared to the baseline model, the CGA module improved the Mean Intersection over Union (mIoU) by 1.9% when dealing with occluded areas, significantly enhancing segmentation accuracy and robustness. This result further validates the CGA module&#x2019;s segmentation capabilities in complex scenes, enabling better extraction of complete semantic information for partially occluded objects.</p>
</list-item>
</list>
<p>Experiments demonstrated that the CGA-ASNet model performed exceptionally well on the synthetic dataset and could effectively generalize to real greenhouse scenarios. Additionally, we tested the model on the PApple_RGB-D-Size dataset and observed similar generalization capabilities, indicating that the method is well-suited for amodal segmentation tasks involving approximately round crops like apples. The model showcased high accuracy and stability, suggesting that this approach is not limited to tomatoes but can be extended to other crops with similar shapes.</p>
<p>This study demonstrates that the combination of synthetic datasets and deep learning techniques provides an efficient and cost-effective solution for target segmentation in agricultural scenarios. In the future, with the expansion of dataset size and further model optimizations, the integration of synthetic and real-world data will further enhance the model&#x2019;s generalization capabilities, providing robust technical support for tasks such as automated crop harvesting and crop monitoring. This also highlights the significant potential of synthetic data in agricultural vision tasks.</p>
<p>Despite these promising results, this study still has some limitations. First, the current method primarily focuses on crops with relatively round shapes, and its effectiveness on more complex or irregularly shaped crops remains to be validated. Second, although synthetic data improves robustness, the domain gap between synthetic and real-world data may still limit generalization in more diverse or unconstrained environments. In future work, we plan to extend our dataset to include various crop types and environmental settings, explore domain adaptation techniques, and further enhance the model&#x2019;s architecture to support broader applications in agricultural perception.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>ZL: Conceptualization, Methodology, Project administration, Writing &#x2013; original draft, Writing &#x2013; review &amp; editing. YY: Data curation, Formal Analysis, Visualization, Writing &#x2013; review &amp; editing. ZX: Data curation, Methodology, Software, Validation, Writing &#x2013; review &amp; editing. HD: Funding acquisition, Writing &#x2013; review &amp; editing.</p>
</sec>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research and/or publication of this article. This research was funded by Supported by Sub-project of National Key R&amp;D Plan (Grant No. 2022YFD2002303-01 and No.2024YFD1501205-01); Liaoning Province Innovation Capability Enhancement Joint Fund Project (Grant No. JYTMS20231303).</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>The authors thank the editor and reviewers for providing helpful suggestions for improving the quality of this manuscript.</p>
</ack>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="ai-statement">
<title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Back</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Noh</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Kang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Bak</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). &#x201c;<article-title>Unseen object amodal instance segmentation via hierarchical occlusion modeling</article-title>,&#x201d; in <source>2022 International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Philadelphia, PA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICRA46639.2022.9811646</pub-id>
</citation></ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blok</surname> <given-names>P. M.</given-names>
</name>
<name>
<surname>van Henten</surname> <given-names>E. J.</given-names>
</name>
<name>
<surname>van Evert</surname> <given-names>F. K.</given-names>
</name>
<name>
<surname>Kootstra</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Image-based size estimation of broccoli heads under varying degrees of occlusion</article-title>. <source>Biosyst. Eng.</source> <volume>208</volume>, <fpage>213</fpage>&#x2013;<lpage>233</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.biosystemseng.2021.06.001</pub-id>
</citation></ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>C&#xe1;mara-Zapata</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Brotons-Mart&#xed;nez</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Sim&#xf3;n-Grao</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Martinez-Nicol&#xe1;s</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Garc&#xed;a-S&#xe1;nchez</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Cost&#x2013;benefit analysis of tomato in soilless culture systems with saline water under greenhouse conditions</article-title>. <source>J. Sci. Food Agric.</source> <volume>99</volume>, <fpage>5842</fpage>&#x2013;<lpage>5851</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/jsfa.9857</pub-id>, PMID: <pub-id pub-id-type="pmid">31206706</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Lammers</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>O2RNet: Occluder-occludee relational network for robust apple detection in clustered orchard environments</article-title>. <source>Smart Agric. Technol.</source> <volume>5</volume>, <elocation-id>100284</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.atech.2023.100284</pub-id>
</citation></ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cinbis</surname> <given-names>R. G.</given-names>
</name>
<name>
<surname>Verbeek</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Schmid</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Weakly supervised object localization with multi-fold multiple instance learning</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>39</volume>, <fpage>189</fpage>&#x2013;<lpage>203</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2535231</pub-id>, PMID: <pub-id pub-id-type="pmid">26930676</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Denninger</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Sundermeyer</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Winkelbauer</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zidan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Olefir</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Elbadrawy</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Blenderproc</article-title>. <source>arXiv</source>. <volume>arXiv</volume>:<elocation-id>1911.01911</elocation-id> doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1911.01911</pub-id>. arXiv:1911.0191.</citation></ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dhaka</surname> <given-names>V. S.</given-names>
</name>
<name>
<surname>Meena</surname> <given-names>S. V.</given-names>
</name>
<name>
<surname>Rani</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Sinwar</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Ijaz</surname> <given-names>M. F.</given-names>
</name>
<name>
<surname>Wo&#x17a;niak</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A survey of deep convolutional neural networks applied for prediction of plant leaf diseases</article-title>. <source>Sensors</source> <volume>21</volume>, <elocation-id>4749</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/s21144749</pub-id>, PMID: <pub-id pub-id-type="pmid">34300489</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Farbman</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Hoffer</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lipman</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Cohen-Or</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Lischinski</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Coordinates for instant image cloning</article-title>. <source>ACM Trans. Graphics (TOG), (New York, NY, USA) </source> <volume>28</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1145/1531326.1531373</pub-id>
</citation></ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Follmann</surname> <given-names>P.</given-names>
</name>
<name>
<surname>K&#xf6;nig</surname> <given-names>R.</given-names>
</name>
<name>
<surname>H&#xe4;rtinger</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Klostermann</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Learning to See the Invisible: End-to-End Trainable Amodal Instance Segmentation</article-title>.&#x201d; in <conf-name>2019 IEEE Winter Conference on Applications of Computer Vision (WACV)</conf-name>, (<conf-loc>Waikoloa, HI, USA</conf-loc>), <fpage>1328</fpage>&#x2013;<lpage>1336</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/WACV.2019.00146</pub-id>
</citation></ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gen&#xe9;-Mola</surname>
</name>
<name>
<surname>Ferrer-Ferrer</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Gregorio</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Blok</surname> <given-names>P. M.</given-names>
</name>
<name>
<surname>Hemming</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Morros</surname> <given-names>J. R.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Looking behind occlusions: a study on amodal segmentation for robust on-tree apple fruit size estimation</article-title>. <source>Comput. Electron. Agric.</source> <volume>209</volume>, <elocation-id>107854</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.COMPAG.2023.107854</pub-id>
</citation></ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gen&#xe9;-Mola</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ferrer-Ferrer</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Hemming</surname> <given-names>J.</given-names>
</name>
<name>
<surname>van Dalfsen</surname> <given-names>P.</given-names>
</name>
<name>
<surname>de Hoog</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Sanz-Cortiella</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>AmodalAppleSize_RGB-D dataset: RGB-D images of apple trees annotated with modal and amodal segmentation masks for fruit detection, visibility and size estimation</article-title>. <source>Data Brief</source> <volume>52</volume>, <elocation-id>110000</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.dib.2023.110000</pub-id>, PMID: <pub-id pub-id-type="pmid">38274155</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Fast R-CNN</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>. (<publisher-loc>Santiago, Chile</publisher-loc>), <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV.2015.169</pub-id>
</citation></ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Gkioxari</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Mask R-CNN</article-title>,&#x201d; in <conf-name>2017 IEEE International Conference on Computer Vision (ICCV)</conf-name>, (<conf-loc>Venice, Italy</conf-loc>), <fpage>2980</fpage>&#x2013;<lpage>2988</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV.2017.322</pub-id>
</citation></ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep residual learning for image recognition</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<conf-loc>Las Vegas, NV, USA</conf-loc>), pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
</citation></ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Apple tree architectural trait phenotyping with organ-level instance segmentation from point cloud</article-title>. <source>Comput. Electron. Agric.</source> <volume>229</volume>, <elocation-id>109708</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2024.109708</pub-id>
</citation></ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ke</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tai</surname> <given-names>Y. W.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>C. K.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Deep occlusion-aware instance segmentation with overlapping bilayers</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name> (<publisher-loc>Nashville, TN, USA</publisher-loc>), <fpage>4019</fpage>&#x2013;<lpage>4028</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00401</pub-id>
</citation></ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Qiu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Occluded apple fruit detection and localization with a frustum-based point-cloud-processing approach for robotic harvesting</article-title>. <source>Remote Sens.</source> <volume>14</volume>, <elocation-id>482</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14030482</pub-id>
</citation></ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Malik</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Amodal instance segmentation</article-title>,&#x201d; in <source>Computer Vision  - ECCV 2016. ECCV 2016. Lecture Notes in Computer Science</source>, vol <volume>9906</volume>. eds. <person-group person-group-type="editor">
<name>
<surname>Leibe</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Matas</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sebe</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Welling</surname> <given-names>M.</given-names>
</name>
</person-group>. (<publisher-loc>Cham Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-319-46475-6_42</pub-id>
</citation></ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2024</year>). &#x201c;<article-title>BLADE: Box-level supervised amodal segmentation through directed expansion</article-title>,&#x201d; in <conf-name>Proceedings of the AAAI conference on artificial intelligence</conf-name>. (<conf-loc>Vancouver, Canada</conf-loc>), <volume>38</volume> (<issue>4</issue>), <fpage>3846</fpage>&#x2013;<lpage>3854</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1609/aaai.v38i4.28176</pub-id>
</citation></ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Swin transformer: Hierarchical vision transformer using shifted windows</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF international conference on computer vision</conf-name>. (<conf-loc>Montreal, Canada</conf-loc>), <fpage>10012</fpage>&#x2013;<lpage>10022</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2103.14030</pub-id>
</citation></ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Mao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>C. Y.</given-names>
</name>
<name>
<surname>Feichtenhofer</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>A convnet for the 2020s</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<conf-loc>New Orleans, LA, USA</conf-loc>),  <fpage>11976</fpage>&#x2013;<lpage>11986</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2201.03545</pub-id>
</citation></ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shao</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Hoffmann</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Global attention mechanism: Retain information to enhance channel-spatial interactions</article-title>. (<publisher-loc>Dresden, Germany</publisher-loc>: <publisher-name>Helmholtz-Zentrum Dresden-Rossendorf</publisher-name>). <source>arXiv</source>. arXiv:2112.05561. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2112.05561</pub-id>
</citation></ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Long</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Shelhamer</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Fully convolutional networks for semantic segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<conf-loc>Boston, MA, USA</conf-loc>), <fpage>3431</fpage>&#x2013;<lpage>3440</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2572683</pub-id>, PMID: <pub-id pub-id-type="pmid">27244717</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ochs</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Malik</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Brox</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Segmentation of moving objects by long term video analysis</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>36</volume>, <fpage>1187</fpage>&#x2013;<lpage>1200</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2013.242</pub-id>, PMID: <pub-id pub-id-type="pmid">26353280</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ozguroglu</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sur&#xed;s</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Dave</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Tokmakov</surname> <given-names>P.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). &#x201c;<article-title>pix2gestalt: Amodal segmentation by synthesizing wholes</article-title>,&#x201d; in <source>CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Los Alamitos, CA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>), <fpage>3931</fpage>&#x2013;<lpage>3940</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2401.14398</pub-id>
</citation></ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Redmon</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Farhadi</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Yolov3: An incremental improvement</article-title>. <source>arXiv</source>. arXiv:1804.02767. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1804.02767</pub-id>
</citation></ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname> <given-names>S.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Faster r-cnn: Towards real-time object detection with region proposal networks</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>28</volume>, <fpage>91</fpage>&#x2013;<lpage>99</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id>, PMID: <pub-id pub-id-type="pmid">27295650</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ronneberger</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Fischer</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Brox</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>,&#x201d; in <conf-name>International Conference on Medical image computing and computer-assisted intervention. </conf-name>. (<publisher-loc>Cham Switzerland</publisher-loc>: <publisher-name>Springer international publishing</publisher-name>), <fpage>234</fpage>&#x2013;<lpage>241</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id>
</citation></ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Russell</surname> <given-names>B. C.</given-names>
</name>
<name>
<surname>Torralba</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Murphy</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Freeman</surname> <given-names>W. T.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>LabelMe: a database and web-based tool for image annotation</article-title>. <source>Int. J. Comput. Vision</source> <volume>77</volume>, <fpage>157</fpage>&#x2013;<lpage>173</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11263-007-0090-8</pub-id>
</citation></ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Seitz</surname> <given-names>S. M.</given-names>
</name>
<name>
<surname>Curless</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Diebel</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Scharstein</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Szeliski</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>A comparison and evaluation of multi-view stereo reconstruction algorithms</article-title>,&#x201d; in <source>2006 IEEE computer society conference on computer vision and pattern recognition (CVPR'06)</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>) Vol. <volume>1</volume>, pp. <fpage>519</fpage>&#x2013;<lpage>528</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2006.19</pub-id>
</citation></ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sheng</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Kang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lyu</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>An edge-guided method to fruit segmentation in complex environments</article-title>. <source>Comput. Electron. Agric.</source> <volume>208</volume>, <elocation-id>107788</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2023.107788</pub-id>
</citation></ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Subedi</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>H. J.</given-names>
</name>
<name>
<surname>Jha</surname> <given-names>D. K.</given-names>
</name>
<name>
<surname>Sarkar</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Find the fruit: designing a zero-shot sim2Real deep RL planner for occlusion aware plant manipulation</article-title>. <source>arXiv</source>. arXiv:2505.16547. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2505.16547</pub-id>.</citation></ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Szegedy</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Vanhoucke</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Ioffe</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Shlens</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wojna</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Rethinking the inception architecture for computer vision</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<conf-loc>Las Vegas, NV, USA</conf-loc>), <fpage>2818</fpage>&#x2013;<lpage>2826</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2016.308</pub-id>
</citation></ref>
<ref id="B34">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tran</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Vo</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Nguyen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Le</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2024</year>). &#x201c;<article-title>Amodal instance segmentation with diffusion shape prior estimation</article-title>,&#x201d; in <conf-name>Proceedings of the Asian Conference on Computer Vision</conf-name>. (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1181</fpage>&#x2013;<lpage>1196</lpage>.  doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2409.18256</pub-id>
</citation></ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tran</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Vo</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yamazaki</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fernandes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Kidd</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Le</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Aisformer: Amodal instance segmentation with transformer</article-title>. <source>arXiv</source>. arXiv:2210.06323. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2210.06323</pub-id>
</citation></ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Tu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Aggregated residual transformations for deep neural networks</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<conf-loc>Honolulu, HI, USA</conf-loc>), <fpage>1492</fpage>&#x2013;<lpage>1500</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1611.05431</pub-id>
</citation></ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Miao</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Application of amodal segmentation for shape reconstruction and occlusion recovery in occluded tomatoes</article-title>. <source>Front. Plant Sci.</source> <volume>15</volume>, <elocation-id>1376138</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2024.1376138</pub-id>, PMID: <pub-id pub-id-type="pmid">38938637</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yao</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Bennewitz</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Hauser</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2025</year>). &#x201c;<article-title>Safe leaf manipulation for accurateshape and pose estimation of occluded fruits</article-title>.&#x201d; in <conf-name>2025 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>, (<conf-loc>Atlanta, GA, USA</conf-loc>), <fpage>16795</fpage>&#x2013;<lpage>16802</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICRA55743.2025.11128788</pub-id>
</citation></ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Queralta</surname> <given-names>J. P.</given-names>
</name>
<name>
<surname>Westerlund</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Sim-to-real transfer in deep reinforcement learning for robotics: a survey</article-title>,&#x201d; in <source>2020 IEEE symposium series on computational intelligence (SSCI)</source> (<publisher-loc>Canberra, ACT, Australia</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>737</fpage>&#x2013;<lpage>744</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/SSCI47803.2020.9308468</pub-id>
</citation></ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>L. L.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W. X.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>Y. W.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yi</surname> <given-names>Z. Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Tomato dataset for agricultural scene visual-parsing tasks</article-title>. <source>J. Agric. Big Data</source> <volume>3</volume> (<issue>4</issue>), <fpage>70</fpage>&#x2013;<lpage>76</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.19788/j.issn.2096-6369.210408</pub-id>
</citation></ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Metaxas</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Semantic amodal segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<conf-loc>Honolulu, HI, USA</conf-loc>), (pp. <fpage>1464</fpage>&#x2013;<lpage>1472</lpage>). doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2017.320</pub-id>
</citation></ref>
</ref-list>
</back>
</article>