<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2025.1656381</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>YOLOv8-LBP: multi-scale attention enhanced YOLOv8 for ripe tomato detection and harvesting keypoint localization</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yong</surname>
<given-names>Wang</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2648924/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shunfa</surname>
<given-names>Xu</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2926311/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Konghao</surname>
<given-names>Cheng</given-names>
</name>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>School of Computer Science and Technology, Zhejiang University of Technology</institution>, <addr-line>Hangzhou</addr-line>,&#xa0;<country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/55589/overview">Milind B. Ratnaparkhe</ext-link>, ICAR Indian Institute of Soybean Research, India</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2850948/overview">Yibin Tian</ext-link>, Shenzhen University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2996782/overview">Robson de Sousa Nascimento</ext-link>, Federal University of Para&#xed;ba, Brazil</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3141279/overview">Shanglei Chai</ext-link>, Shenzhen University, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Wang Yong, <email xlink:href="mailto:wy@zjut.edu.cn">wy@zjut.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>09</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1656381</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>06</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>08</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Yong, Shunfa and Konghao.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Yong, Shunfa and Konghao</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>In the process of target detection for tomato harvesting robots, there are two primary challenges. First, most existing tomato harvesting robots are limited to fruit detection and recognition, lacking the capability to locate harvesting keypoints. As a result, they cannot be directly applied to the harvesting of ripe tomatoes. Second, variations in lighting conditions in natural environments, occlusions between tomatoes, and missegmentation caused by similar fruit colors often lead to keypoint localization errors during harvesting. To address these issues, we propose YOLOv8-LBP, an enhanced model based on YOLOv8-Pose, designed for both ripe tomato recognition and harvesting keypoint detection. Specifically, we introduce a Large Separable Kernel Attention (LSKA) module into the backbone network, which effectively decomposes large kernel convolutions to extract target feature matrices more efficiently, enhancing the model&#x2019;s adaptability and accuracy for multi-scale objects. Secondly, the weighted bidirectional feature pyramid network (BiFPN) introduces additional weights to learn the importance of different input features. Through top-down and bottom-up bidirectional paths, the model repeatedly fuses multi-scale features, thereby enhancing its ability to detect objects at multiple scales. Ablation experiments demonstrate that, on our self-constructed ripe tomato dataset, the YOLOv8-LBP model achieves improvements of 4.5% in Precision (<italic>P</italic>), 1.1% in <italic>mAP</italic>
<sub>50</sub>, 2.8% in <italic>mAP</italic>
<sub>50&#x2212;95</sub>, and 3.3% in <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic> compared to the baseline. When compared with the state-of-the-art YOLOv12-Pose, YOLOv8-LBP shows respective improvements of 5.7%, 0.5%, 3.5%, and 4.9% in the same metrics. While maintaining the improvement in model accuracy, our method introduces only a small computational overhead, with the number of parameters increasing from 3.08M to 3.175M, GFLOPs rising by 0.1, and the inference speed improving from 96.15 FPS to 99.01 FPS. This computational cost is reasonable and acceptable. Overall, the proposed YOLOv8-LBP model demonstrates significant advantages in recognizing ripe tomatoes and detecting harvesting keypoints under complex scenarios, offering a solid theoretical foundation for the advancement of robotic harvesting technologies.</p>
</abstract>
<kwd-group>
<kwd>ripe tomato recognition</kwd>
<kwd>keypoint prediction</kwd>
<kwd>YOLOv8-Pose</kwd>
<kwd>attention mechanism</kwd>
<kwd>multi-scale feature fusion</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="6"/>
<equation-count count="12"/>
<ref-count count="31"/>
<page-count count="15"/>
<word-count count="6424"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Technical Advances in Plant Science</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Smart agriculture is an emerging trend in the modernization of agricultural practices, playing an indispensable role in the high-quality development of agriculture in China. The vigorous development of smart agriculture has become a key driving force in achieving rural revitalization <xref ref-type="bibr" rid="B26">Yu and Haiyan (2025)</xref>. As an important tool of agricultural intelligence, harvesting robots are capable of performing efficient and precise picking operations, significantly reducing labor costs while improving production efficiency <xref ref-type="bibr" rid="B28">Zhang et&#xa0;al. (2023)</xref>.</p>
<p>At the core of harvesting robots lies the target detection algorithm, which provides accurate target localization guidance for the robot&#x2019;s hydraulic control system <xref ref-type="bibr" rid="B29">Zhang et&#xa0;al. (2024b)</xref>. <xref ref-type="bibr" rid="B1">Cai et&#xa0;al. (2024)</xref> proposed a cherry tomato detection method based on multimodal perception and an improved YOLOv7-tiny network. By utilizing RGB-D image inputs and introducing a &#x201c;Classness&#x201d; prediction mechanism along with a hybrid non-maximum suppression strategy, the method enhances detection accuracy and real-time performance. Experimental results demonstrate that the method achieves a high picking success rate and exhibits promising potential for practical applications in greenhouse environments. <xref ref-type="bibr" rid="B12">Liang et&#xa0;al. (2025)</xref> developed an improved CTDA model based on YOLOv8 by redesigning the backbone network and incorporating SoftPool and an attention-driven dynamic detection head to improve small object feature extraction and multi-scale feature fusion efficiency. Under complex harvesting conditions, the model achieved a detection accuracy of 94.3% and a <italic>mAP</italic>
<sub>50&#x2212;95</sub> of 76.5%, with a detection speed of 154.1 FPS and a model size of only 6.7 MB, demonstrating excellent real-time performance and robustness. <xref ref-type="bibr" rid="B24">Wu et&#xa0;al. (2025)</xref> introduced a lightweight detection network, LEFF-YOLO, based on an improved YOLOv8 framework. They incorporated the VanillaNet module into the backbone to reduce computational cost and enhanced feature fusion capability using the SiMAM attention mechanism and the SiMAMC2f module. Furthermore, a novel bounding box loss function was designed to improve model stability, resulting in improvements of 2.7% and 2.9% in <italic>mAP</italic>
<sub>50</sub> and <italic>mAP</italic>
<sub>50&#x2212;95</sub> respectively, with a 32% reduction in parameters and a 10.5% decrease in computational load, meeting the real-time and accuracy requirements of harvesting robots. <xref ref-type="bibr" rid="B14">Liu et&#xa0;al. (2024)</xref> proposed a &#x201c;coarse detection&#x2013;fine segmentation&#x201d; method, Y-HRNet, tailored for greenhouse environments. The approach first employs YOLOv7 for lightweight multi-class tomato detection, and then utilizes a Y-HRNet network enhanced with ECA and DR-ASPP modules for pixel-level segmentation of green, turning, ripe, and fully ripe tomatoes. This method achieved a mean IoU of 84.69% and an accuracy of 94.39% under complex backgrounds, with an average processing time of 0.35 seconds, providing strong support for tomato maturity grading and harvesting management. Zhang et&#xa0;al. proposed a tomato detection algorithm for harvesting robots based on Shufflenetv2-YOLOv5, utilizing Shufflenetv2 as the backbone to reduce computational cost, incorporating the ECA attention mechanism, and replacing the activation function with FReLU to enhance detection precision and model robustness <xref ref-type="bibr" rid="B27">Zhang et&#xa0;al. (2024a)</xref>. However, the above approaches primarily focus on fruit-level object detection, without addressing keypoint localization for harvesting. As a result, although these methods can effectively detect fruits, they lack the capability for precise picking point localization, making them difficult to directly apply in real-world tomato harvesting operations.</p>
<p>In the field of tomato picking point localization, Liu Rong et&#xa0;al. achieved automatic harvesting by detecting tomato branches, including coarse localization of fruits and pixel-level segmentation of multiple pedicels <xref ref-type="bibr" rid="B19">Rong et&#xa0;al. (2021)</xref>. However, their method exhibited a significant angular error of up to 5&#xb0;. <xref ref-type="bibr" rid="B16">Luo et&#xa0;al. (2018)</xref> proposed a grape stem cutting point detection method based on geometric modeling and contour analysis, which integrates K-means clustering segmentation with a geometric constraint strategy. In complex vineyard environments, this approach achieved an average recognition accuracy of 88.33% and a double-cluster stem cutting point detection success rate of 81.66%, providing a feasible solution for the application of harvesting robots in operations involving overlapping grape clusters. Sa et&#xa0;al. utilized HSV color space and geometric features (such as curvature differences between stems and fruit surfaces), combined with a support vector machine (SVM) to extract features and predict the location of sweet pepper pedicels <xref ref-type="bibr" rid="B20">Sa et&#xa0;al. (2017)</xref>.</p>
<p>These methods mainly rely on geometric features for predicting harvesting points. However, in practical operations, factors such as ambient lighting, fruit occlusion, and leaf interference can significantly impact accuracy, leading to misaligned or missed picking points, thereby increasing the risk of failed or inaccurate harvesting and limiting their application in complex environments.</p>
<p>Keypoint detection was originally applied in human pose estimation. For instance, JEONG et&#xa0;al. used the OpenPose algorithm to infer smoking behavior through skeleton graphs constructed from human keypoints <xref ref-type="bibr" rid="B6">Feng et&#xa0;al. (2024)</xref>. Today, many studies have extended this technique to industrial and agricultural domains. For example, Taehyeong Kim et&#xa0;al. used OpenPose as the backbone to estimate the posture of tomato fruits <xref ref-type="bibr" rid="B9">Kim et&#xa0;al. (2023)</xref>. Wu et&#xa0;al. proposed a stem localization method for grapes, based on their top-down growth characteristics, where grape clusters were detected using object detection and the stem positions were then accurately located through keypoint detection <xref ref-type="bibr" rid="B25">Wu et&#xa0;al. (2023)</xref>. <xref ref-type="bibr" rid="B8">Kang et&#xa0;al. (2025)</xref> focused on greenhouse-grown melons and employed Region of Interest (RoI) to coarsely localize ripe fruits, followed by Human Pose Estimation (HPE) to estimate the posture of melon-fruit-peduncle pairs within the RoI. Their approach achieves high performance with improved real-time efficiency, even under limited training data.</p>
<p>However, most of these methods were developed under controlled conditions such as greenhouses, relying on stable crop postures and standardized cultivation patterns. In contrast, tomatoes grown in natural environments may exhibit a wide range of unpredictable postures (e.g., hanging, inclined, or covered) due to wind, gravity, and occlusion by leaves and branches. These factors significantly limit the generalizability and robustness of existing approaches, making them difficult to apply effectively in real-world natural field conditions where tomato postures are highly variable and uncertain.</p>
<p>Debapriya Maji was among the first to unify keypoint detection and object detection in the context of human pose estimation <xref ref-type="bibr" rid="B17">Maji et&#xa0;al. (2022)</xref>. Zhang et&#xa0;al. built upon YOLO-Pose by proposing an algorithm that fuses pose estimation with object detection, achieving an accuracy of 90.14% in dense crowd scenes&#x2014;21.56% higher than conventional YOLO-Pose <xref ref-type="bibr" rid="B30">Zhang and Xue (2024)</xref>. Qin et&#xa0;al. developed PW-YOLO-Pose, which effectively addressed challenges in power operation scenarios, such as complex backgrounds, occlusions, small targets, and extreme viewpoints, thereby reducing keypoint loss and false detections <xref ref-type="bibr" rid="B21">Su et&#xa0;al. (2024)</xref>. Liu et&#xa0;al. developed an improved YOLOv8-Pose model by introducing a Slim-neck module and a CBAM attention mechanism module to recognize red-ripe strawberries and detect peduncle keypoints. The improved YOLOv8-Pose achieved a <italic>mAP<sub>kp</sub>
</italic> of 97.91% <xref ref-type="bibr" rid="B18">Mochen et&#xa0;al. (2023)</xref>.</p>
<p>Building on these insights, this paper improves upon human pose estimation algorithms for predicting keypoints related to the harvesting of ripe tomatoes. By analyzing the geometric configuration and growth posture of tomatoes, we aim to accurately predict keypoints at the fruit pedicel and picking region, thus offering technical support for robotic tomato harvesting. The main contributions of this study are as follows:</p>
<list list-type="order">
<list-item>
<p>An attention mechanism (LSKA) <xref ref-type="bibr" rid="B11">Lau et&#xa0;al. (2024)</xref> is incorporated into the backbone network to effectively decompose large convolutional kernels, extract richer target feature matrices, and enhance the network&#x2019;s feature extraction capability.</p>
</list-item>
<list-item>
<p>A Bidirectional Feature Pyramid Network (BiFPN) <xref ref-type="bibr" rid="B13">Lin et&#xa0;al. (2017)</xref> is introduced in the neck of the model to facilitate adaptive weighting during training and enable bidirectional information flow between high- and low-level features.</p>
</list-item>
<list-item>
<p>The YOLOv8-LBP model, when applied to our self-constructed ripe tomato dataset, achieves performance gains of +4.5% in <italic>P</italic>, +1.1% in <italic>mAP</italic>
<sub>50</sub>, +2.8% in <italic>mAP</italic>
<sub>50&#x2212;95</sub>, and +3.3% in <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic>, demonstrating the effectiveness of the proposed improvements.</p>
</list-item>
</list>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Material and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Material</title>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Self-constructed ripe tomato dataset</title>
<p>The dataset used in this study is derived from TomatoDiverse: An Open Dataset for Industrial Tomato Detection in Complex Natural Environments <xref ref-type="bibr" rid="B23">Tao (2024)</xref>. The images were captured at tomato production bases located in Bayingolin Mongol Autonomous Prefecture, Xinjiang Uygur Autonomous Region, and Jingxian County, China. From this dataset, a total of 236 high-resolution images (4000&#xd7;3000 pixels) depicting the growth posture of ripe tomatoes were selected.</p>
<p>These images reflect a variety of challenges present in complex natural environments, including variations in lighting conditions, different shooting distances, and occlusions. To enhance data diversity, the dataset was augmented using several techniques: random rotation (with angles randomly sampled within the range from -30 to +30) <xref ref-type="bibr" rid="B3">Cubuk et&#xa0;al. (2019)</xref>, horizontal flipping <xref ref-type="bibr" rid="B10">Krizhevsky et&#xa0;al. (2012)</xref>, and Gaussian blurring (with Gaussian kernel sizes randomly selected from [21, 31, 41, 51]) <xref ref-type="bibr" rid="B31">Zhang et&#xa0;al. (2019)</xref>. After augmentation, the dataset was expanded to a total of 780 images. <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref> illustrates examples of the augmented images.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Ripe tomato images after data augmentation.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g001.tif">
<alt-text content-type="machine-generated">Images of tomatoes illustrate various transformations. The top row shows original images. The second row applies Gaussian blur. The third row displays horizontal flips, and the fourth row shows random rotations. Each row contains three images of tomatoes displaying these transformations.</alt-text>
</graphic>
</fig>
<p>The criteria for determining red-ripe <xref ref-type="bibr" rid="B2">Choi et&#xa0;al. (1995)</xref> tomatoes in the dataset are as follows: 1. Visual Inspection: A tomato is labeled as ripe when the red area on its surface reaches or exceeds 90%. 2. Quantitative HSV-Based Determination: For samples where manual judgment is challenging, the Labelme tool is used to segment the tomato region, and the average values of Hue, Saturation, and Value (Brightness) are calculated in the HSV color space. A tomato is considered red-ripe if the following conditions are met: Hue (H): Within the range of 0&#xa0;&#x2013;&#xa0;11.2 or 144&#xa0;&#x2013;&#xa0;180 (where 11.2 is derived from dividing the visual threshold 10 by 0.9, and 144 is calculated similarly);Saturation (S): Between 100 and 255;Value (V): Between 100 and 255.</p>
<p>
<xref ref-type="bibr" rid="B18">Mochen et&#xa0;al. (2023)</xref> adopted two keypoints (the peduncle and the picking point) to annotate and detect strawberries grown in an elevated cultivation system. In this cultivation mode, strawberry fruits hang naturally and maintain a stable posture, making keypoint annotation relatively easy. In contrast, tomatoes in natural growing environments may either hang down or lie horizontally on branches and leaves, exhibiting a wide variety of postures. To enable more effective localization of picking points, we introduce three keypoints in this study: P1 (picking point), P2 (peduncle location), and P3 (fruit bottom). P2 and P3 together provide information about the size and orientation of the fruit, which is critical for evaluating the feasibility of grasping. P1 is located approximately 1&#xa0;&#x2013;&#xa0;2 cm above P2 along the direction of the peduncle and represents the ideal cutting point for the harvesting robot. This design ensures accurate picking without damaging the fruit itself and minimizes the risk of mistakenly cutting non-target parts such as branches or leaves.</p>
<p>Keypoint annotations for ripe tomatoes were performed using the Labelme software, with the tomato fruit and its picking points defined as annotation targets. Each ripe tomato was annotated using its minimum bounding rectangle, with the label assigned as &#x201c;tomato&#x201d;. Three keypoints were defined: P2, located at the junction of the calyx and the pedicel; P1, positioned 1&#xa0;&#x2013;&#xa0;2 cm upward along the pedicel from P2; and P3, located at the bottom end of the tomato. An example of the annotated image is shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. If a keypoint is partially occluded, the suffix &#x2018;hide&#x2019; is appended to its name; if a keypoint is invisible or absent, the suffix &#x2018;missing&#x2019; is used instead. All annotations are stored in standard JSON format, including the image path, image dimensions (width, height, number of channels), as well as the bounding box of the tomato and the coordinates of each keypoint. The JSON annotation files were converted into TXT format for training purposes. The dataset was split into training and testing sets at a ratio of 8:2, resulting in 624 images for training and 156 images for testing. The overall image distribution conforms to a 4:1 ratio between the training set and the testing set. The visibility statistics of the annotated keypoints are summarized in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Annotation of red-ripe tomatoes.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g002.tif">
<alt-text content-type="machine-generated">Two ripe red tomatoes are growing on a vine among green leaves. The image has bounding boxes around the tomatoes, indicating they are the focus of analysis or identification.</alt-text>
</graphic>
</fig>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Instance in dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Dataset</th>
<th valign="middle" align="center">Image</th>
<th valign="middle" align="center">Tomato</th>
<th valign="middle" align="center">P0</th>
<th valign="middle" align="center">P1</th>
<th valign="middle" align="center">P2</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Train</td>
<td valign="middle" align="center">624</td>
<td valign="middle" align="center">1934</td>
<td valign="middle" align="center">890</td>
<td valign="middle" align="center">800</td>
<td valign="middle" align="center">4112</td>
</tr>
<tr>
<td valign="middle" align="center">Val</td>
<td valign="middle" align="center">156</td>
<td valign="middle" align="center">451</td>
<td valign="middle" align="center">273</td>
<td valign="middle" align="center">174</td>
<td valign="middle" align="center">906</td>
</tr>
<tr>
<td valign="middle" align="center">All</td>
<td valign="middle" align="center">780</td>
<td valign="middle" align="center">2385</td>
<td valign="middle" align="center">1163</td>
<td valign="middle" align="center">974</td>
<td valign="middle" align="center">5018</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Methods</title>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>YOLOv8-Pose</title>
<p>YOLOv8-Pose is based on the YOLOv8 neural network and introduces a keypoint localization module for human pose estimation on top of object detection. It is primarily used for human detection and pose estimation. Both YOLOv8 and YOLOv5 models were developed by the same author. The network consists of five main components: the input, Backbone, Neck, Head, and output. The Backbone continues the structure of previous versions, where an input image of 640&#xd7;640 is processed to generate feature maps of sizes 80, 40, and 20, corresponding to downscaling factors of 8, 16, and 32, respectively. The C2f module is extensively used, which, compared to the C2 module in YOLOv5, introduces deeper convolutional feature fusion, providing richer gradient information. The Neck follows the structure of YOLOv5, utilizing a path aggregation network to enhance feature fusion. It achieves this by upsampling low-level features and fusing them with high-level features, as well as downsampling high-level features and merging them with low-level features. This improves the network&#x2019;s ability to handle objects at different scales. The output layer includes both classification and detection. The coupled head is replaced with a decoupled head, where classification and regression tasks are handled by separate branches. This separation improves efficiency and accuracy when handling different tasks.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Improved YOLOv8 - pose key - point recognition method for picking tomato fruits in the red - ripe stage</title>
<p>YOLOv8-Pose is a keypoint prediction model based on human pose estimation. Compared to human pose estimation, the number of keypoints, target categories, and feature characteristics of ripe tomatoes differ significantly. Therefore, the original model is not directly applicable to the recognition of ripe tomatoes and the detection of picking keypoints, requiring improvements to YOLOv8-Pose. To address challenges such as target occlusion, leaf occlusion, and mis-segmentation caused by the similar color of ripe tomatoes, this study designs and introduces a Large Separable Kernel Attention (LSKA) module to enhance the network&#x2019;s ability to distinguish target features <xref ref-type="bibr" rid="B4">Deng et&#xa0;al. (2024)</xref>. Additionally, a weighted Bi-directional Feature Pyramid Network (BiFPN) is used to replace the FPN structure in the Neck, allowing the network to learn the importance of different input features through learnable weights, thereby improving feature fusion capabilities. The overall framework of the improved YOLOv8-Pose algorithm is shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Structure of YOLOv8-LBP.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g003.tif">
<alt-text content-type="machine-generated">Diagram of a neural network architecture. It consists of three main sections: Backbone, Neck, and Head. The Backbone includes layers such as Conv, C2f, SPPF, and LSKA. The Neck features BiFPN, upsampling, and additional Conv layers. The Head completes with Conv2d layers leading to Bounding Box Loss and Classification Loss outputs. The flow indicates the data processing path through these components.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>BiFPN module</title>
<p>In object detection and keypoint estimation tasks, effectively acquiring and processing multi-scale feature information remains a major challenge. The traditional Feature Pyramid Network (FPN) (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4a</bold>
</xref>) <xref ref-type="bibr" rid="B15">Liu et&#xa0;al. (2018)</xref> aggregates multi-scale features in a top-down manner, which is often constrained by unidirectional information flow. To address this, the Path Aggregation Network (PANet) <xref ref-type="bibr" rid="B22">Tan et&#xa0;al. (2020)</xref> introduces an additional bottom-up path aggregation, as illustrated in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4b</bold>
</xref>. In the Neck of YOLOv8-Pose, the FPN structure from YOLOv5 is still adopted. High-level feature information is fused through the FPN+PAN structure, which improves the efficiency of feature transmission. However, this also increases computational complexity. Particularly when dealing with high-resolution inputs, it may result in high computational costs and reduced real-time performance. Moreover, the fixed architecture lacks adaptive adjustment capabilities for different tasks and datasets. To address the aforementioned issues such as unidirectional information flow, high accuracy at the cost of numerous parameters and heavy computation, as well as information loss and redundancy caused by simple feature concatenation, this paper proposes replacing the original FPN+PAN structure in YOLOv8-Pose with a BiFPN. Based on PANet and NAS-FPN (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4c</bold>
</xref>), BiFPN optimizes multi-scale feature fusion strategies. Its architecture is illustrated in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4d</bold>
</xref>.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Feature network.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g004.tif">
<alt-text content-type="machine-generated">Diagram comparing four neural network architectures: FPN, PANet, NAS-FPN, and BiFPN. Each architecture shows layers labeled P3 to P7 with different colored nodes and connecting arrows, illustrating distinct connection patterns.</alt-text>
</graphic>
</fig>
<p>Different feature input resolutions contribute differently to the final feature network output. Therefore, the network needs to learn the corresponding weights to balance feature importance. As shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4d</bold>
</xref>, feature maps from different levels are input into the BiFPN_Concat layer (taking BiFPN_Concat2 as an example). Given two feature maps <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>C</italic>
<sub>0</sub> and <italic>C</italic>
<sub>1</sub> denote the number of input channels, H and W represent the height and width of the feature maps respectively, the weights in the BiFPN_Concat2 module are initialized as w=[1,1], assigning equal initial importance to each input feature map. During each forward propagation step, weights are normalized to ensure stable training. To prevent division by zero, a small constant <italic>&#x3f5;</italic> =&#xa0;0.0001 is used. The formula for weight normalization is as follows (<xref ref-type="disp-formula" rid="eq1">Equation 1</xref>):</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In BiFPN_Concat2, the number of input feature maps is n = 2. After normalization, the weights are represented as <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> =&#xa0;[<inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>], where <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> + <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> = 1. Each input feature map is multiplied by its corresponding normalized weight: <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:msubsup>
<mml:mi>F</mml:mi>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:msubsup>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> During training, the model optimizes w based on loss feedback, which in turn updates weight values dynamically. Finally, BiFPN_Concat2 concatenates the weighted feature maps along the channel dimension to obtain the output feature map (<xref ref-type="disp-formula" rid="eq2">Equation 2</xref>):</p>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>F</mml:mi>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>At this stage, the output feature map is represented as <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, successfully fusing feature information from <italic>F</italic>
<sub>0</sub> and <italic>F</italic>
<sub>1</sub>.</p>
</sec>
<sec id="s2_2_4">
<label>2.2.4</label>
<title>LSKA attention mechanism</title>
<p>To address the challenges of occlusion among small targets and mis-segmentation caused by similar colors between ripe tomato fruits, this paper designs and incorporates a Large Separable Kernel Attention (LSKA) module to enhance the network&#x2019;s ability to distinguish target features <xref ref-type="bibr" rid="B4">Deng et&#xa0;al. (2024)</xref>. By enlarging the receptive field of the convolutional kernels, the module focuses more on the overall shape characteristics of the target, thereby reducing the interference from background textures and color similarity in the detection results <xref ref-type="bibr" rid="B5">Ding et&#xa0;al. (2022)</xref>. As the kernel size increases, LSKA can effectively capture broader contextual information, thus improving the recognition performance of small targets in complex scenes. However, directly increasing the kernel size also significantly raises the computational cost and memory usage, posing challenges to the model&#x2019;s efficiency. LSKA decomposes a <italic>k</italic> &#xd7; <italic>k</italic> depthwise convolution (DW-Conv) into two sequential 1D DW-Convs with kernel sizes of 1&#xd7;(2<italic>d</italic> &#x2212;1) and (2<italic>d</italic> &#x2212;1)&#xd7;1, as well as two sequential 1D depthwise-dilated convolutions (DW-D-Conv) with kernel sizes of <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. The outputs are then passed through a 1&#xd7;1 convolution. The depthwise-dilated convolutions are responsible for capturing global spatial information from the depthwise convolution output. This design achieves an effective receptive field equivalent to that of a large kernel while avoiding the quadratic increase in computational cost typically caused by using large kernels in depthwise convolutions alone. Given an input feature map <italic>F</italic>
<sub>0</sub> &#x2208; <italic>R</italic>
<sup>C&#xd7;<italic>H</italic>&#xd7;W</sup>, where C is the number of input channels and H and W represent the height and width of the feature map, the output of LSKA is computed as follows (<xref ref-type="disp-formula" rid="eq3">Equations 3</xref>&#x2013;<xref ref-type="disp-formula" rid="eq6">6</xref>):</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>C</mml:mi>
</mml:msubsup>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>C</mml:mi>
</mml:msubsup>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>C</mml:mi>
</mml:msubsup>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>C</mml:mi>
</mml:msubsup>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mover accent="true">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq5">
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msup>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq6">
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>&#x2297;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where <inline-formula>
<mml:math display="inline" id="im13">
<mml:mi>d</mml:mi>
</mml:math>
</inline-formula> is the dilation rate and <inline-formula>
<mml:math display="inline" id="im14">
<mml:mi>k</mml:mi>
</mml:math>
</inline-formula> represents the maximum receptive field. <inline-formula>
<mml:math display="inline" id="im15">
<mml:mo>*</mml:mo>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im16">
<mml:mo>&#x2297;</mml:mo>
</mml:math>
</inline-formula> denote convolution and Hadamard product (element-wise multiplication), respectively. <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the output of two cascaded 1D depthwise convolutions (DW-Conv) with kernel sizes of <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, respectively. <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the output of two cascaded 1D depthwise dilated convolutions (DW-D-Conv) with kernel sizes of <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>d</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, respectively (where <inline-formula>
<mml:math display="inline" id="im23">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes the floor function). <inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:msup>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the output of applying a 1 &#xd7; 1 convolution to <inline-formula>
<mml:math display="inline" id="im25">
<mml:mrow>
<mml:msup>
<mml:mi>Z</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the result of the Hadamard product between the feature map <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:msup>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and the input feature map <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>
</p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Model training</title>
<sec id="s3_1">
<label>3.1</label>
<title>Training environment and parameter settings</title>
<p>The main hardware configuration for training and testing is shown in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>. [] The training epoch is set to 100, batch size is 32, input image size is 640&#xd7;640 pixels, the initial learning rate is set to 0.001, momentum factor is 0.937, and other parameters are kept at default.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Experimental environment.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Configuration</th>
<th valign="middle" align="center">Parameter</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">CPU (Central Processing Unit)</td>
<td valign="middle" align="center">Intel (R) Xeon (R) Gold 6240 CPU @ 2.60GHz</td>
</tr>
<tr>
<td valign="middle" align="center">GPU (Graphics Processing Unit)</td>
<td valign="middle" align="center">NVIDIA GeForce RTX 2080Ti</td>
</tr>
<tr>
<td valign="middle" align="center">Video Memory Capacity/MB</td>
<td valign="middle" align="center">11264</td>
</tr>
<tr>
<td valign="middle" align="center">Training Environment</td>
<td valign="middle" align="center">CUDA 11.3, cuDNN 8.2</td>
</tr>
<tr>
<td valign="middle" align="center">Operating System</td>
<td valign="middle" align="center">Ubuntu 20.04 LTS</td>
</tr>
<tr>
<td valign="middle" align="center">Development Environment</td>
<td valign="middle" align="center">python 3.11.5, Pytorch 2.1.0</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Evaluation metrics</title>
<p>To validate the effectiveness of the proposed model in detecting tomatoes, commonly used quantitative metrics in the field of object detection were employed, including precision (<italic>P</italic>), mean average precision at Intersection over Union (IoU) threshold 50% (<italic>mAP</italic>
<sub>50</sub>), and mean average precision across IoU thresholds from 50% to 95% (<italic>mAP</italic>
<sub>50&#x2212;95</sub>), to comparatively analyze the tomato fruit detection performance. Specifically, <italic>mAP</italic>
<sub>50</sub> refers to the mean average precision calculated at an Intersection over Union (IoU) threshold of 50%, while <italic>mAP</italic>
<sub>50&#x2212;95</sub> represents the mean average precision calculated across IoU thresholds ranging from 50% to 95%, averaged over multiple thresholds. The detailed formulas are as follows (<xref ref-type="disp-formula" rid="eq7">Equation 7</xref>):</p>
<disp-formula id="eq7">
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the formula: <italic>TP</italic> is the number of true positive detections (correctly detected boxes). <italic>FP</italic> is the number of false positive detections (incorrectly detected boxes). <italic>FN</italic> is the number of false negative detections (missed targets). <italic>AP</italic> is the Average Precision for a specific class, and <italic>mAP</italic> is the mean Average Precision across all detection classes. The specific calculation formulas are as follows (<xref ref-type="disp-formula" rid="eq8">Equations 8</xref>, <xref ref-type="disp-formula" rid="eq9">9</xref>):</p>
<disp-formula id="eq8">
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mi>&#x222b;</mml:mi>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
</mml:mstyle>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq9">
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mi>A</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the formula: <italic>N</italic> is the number of object categories to be detected. In this experiment, only the red-ripe tomato category needs to be detected, so <italic>N</italic> =&#xa0;1. The target keypoint similarity <italic>OKS<sub>p</sub>
</italic> is obtained through detection. From <italic>OKS<sub>p</sub>
</italic> <xref ref-type="bibr" rid="B17">Maji et&#xa0;al. (2022)</xref>, the Average Precision <italic>AP</italic> can be derived, and from AP, the <italic>mAP<sub>kp</sub>
</italic> is calculated as the evaluation metric for keypoints. The formula for calculating <italic>OKS<sub>p</sub>
</italic> is (<xref ref-type="disp-formula" rid="eq10">Equation 10</xref>):</p>
<disp-formula id="eq10">
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:mi>K</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mi>exp</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mi>s</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:msubsup>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>)</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&gt;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the formula:</p>
<p>&#x2022; <inline-formula>
<mml:math display="inline" id="im29">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the Euclidean distance between the predicted and ground truth values of the i-th keypoint of the p-th target.</p>
<p>&#x2022; <inline-formula>
<mml:math display="inline" id="im30">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>is the visibility index of the i-th keypoint of the p-th target.</p>
<p>&#x2022; <italic>v<sub>i</sub>
</italic> is the area of the bounding box of the p-th target.</p>
<p>&#x2022; <inline-formula>
<mml:math display="inline" id="im31">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the standard deviation between the label and the actual value of the i-th keypoint.</p>
<p>The formula for calculating <italic>AP</italic> is (<xref ref-type="disp-formula" rid="eq11">Equation 11</xref>):</p>
<disp-formula id="eq11">
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>m</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>p</mml:mi>
</mml:munder>
<mml:mi>&#x3b2;</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>m</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>p</mml:mi>
</mml:munder>
<mml:mn>1</mml:mn>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the formula (<xref ref-type="disp-formula" rid="eq12">Equation 12</xref>):</p>
<disp-formula id="eq12">
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>O</mml:mi>
<mml:mi>K</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:mi>K</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>O</mml:mi>
<mml:mi>K</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&gt;</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In the formula, <italic>T</italic> represents the threshold, ranging from 0.5 to 0.95.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Model results and analysis</title>
<p>To comprehensively evaluate the overall performance of the improved YOLOv8-Pose network model, joint training and validation were conducted on a custom ripened-tomato dataset.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Experimental results on the custom dataset</title>
<p>The LSKA attention mechanism decomposes large convolutional kernels to achieve a larger receptive field, enhancing the model&#x2019;s feature extraction capability. We set kernel sizes of 7, 11, 23, 35, 41, and 53, and trained the YOLOv8-LBP model on a self-collected red-ripe tomato dataset. When the kernel size was 23, the model achieved the highest Precision of 0.944; at kernel size 7, the <italic>mAP</italic>
<sub>50</sub> peaked at 0.926 with the smallest parameter count of 3.164M; kernel size 41 yielded the best <italic>mAP</italic>
<sub>50&#x2212;95</sub> and keypoint <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic> scores of 0.71 and 0.785, outperforming other sizes; while kernel size 53 resulted in the highest FPS of 107.53. Considering the balance between detection performance and computational efficiency, we selected the LSKA attention mechanism with kernel size 41 for subsequent experiments. Details are shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Impact of different LSKA kernel sizes on YOLOv8-LBP detection performance and efficiency.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Kernel size</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">mAP<sub>50</sub>
</th>
<th valign="middle" align="center">mAP<sub>50&#x2212;95</sub>
</th>
<th valign="middle" align="center">mAP<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic>
</th>
<th valign="middle" align="center">Params (M)</th>
<th valign="middle" align="center">FPS</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">7</td>
<td valign="middle" align="center">0.924</td>
<td valign="middle" align="center">0.926</td>
<td valign="middle" align="center">0.689</td>
<td valign="middle" align="center">0.763</td>
<td valign="middle" align="center">3.164</td>
<td valign="middle" align="center">97.01</td>
</tr>
<tr>
<td valign="middle" align="center">11</td>
<td valign="middle" align="center">0.918</td>
<td valign="middle" align="center">0.904</td>
<td valign="middle" align="center">0.690</td>
<td valign="middle" align="center">0.742</td>
<td valign="middle" align="center">3.169</td>
<td valign="middle" align="center">98.04</td>
</tr>
<tr>
<td valign="middle" align="center">23</td>
<td valign="middle" align="center">0.944</td>
<td valign="middle" align="center">0.907</td>
<td valign="middle" align="center">0.678</td>
<td valign="middle" align="center">0.762</td>
<td valign="middle" align="center">3.172</td>
<td valign="middle" align="center">90.10</td>
</tr>
<tr>
<td valign="middle" align="center">35</td>
<td valign="middle" align="center">0.927</td>
<td valign="middle" align="center">0.903</td>
<td valign="middle" align="center">0.679</td>
<td valign="middle" align="center">0.749</td>
<td valign="middle" align="center">3.174</td>
<td valign="middle" align="center">85.47</td>
</tr>
<tr>
<td valign="middle" align="center">41</td>
<td valign="middle" align="center">0.927</td>
<td valign="middle" align="center">0.917</td>
<td valign="middle" align="center">0.710</td>
<td valign="middle" align="center">0.785</td>
<td valign="middle" align="center">3.175</td>
<td valign="middle" align="center">99.01</td>
</tr>
<tr>
<td valign="middle" align="center">53</td>
<td valign="middle" align="center">0.925</td>
<td valign="middle" align="center">0.915</td>
<td valign="middle" align="center">0.689</td>
<td valign="middle" align="center">0.759</td>
<td valign="middle" align="center">3.177</td>
<td valign="middle" align="center">107.53</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Experimental results on the self-constructed ripe tomato dataset</title>
<p>In the experiments on object recognition and keypoint prediction, we evaluated the performance of several models including YOLOv5-Pose, YOLOv8-Pose, YOLOv11-Pose, YOLOv12-Pose, YOLOv8-Pose-ss <xref ref-type="bibr" rid="B7">Guo et&#xa0;al. (2024)</xref>, YOLOv8-Pose-cv <xref ref-type="bibr" rid="B18">Mochen et&#xa0;al. (2023)</xref>, and YOLOv8-LBP on a self-collected red-ripe tomato dataset. The results are presented in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>. Comparative analysis shows that YOLOv8-LBP achieved 92.7%, 91.7%, 71.%, and 78.5% on Precision, <italic>mAP</italic>
<sub>50</sub>, <italic>mAP</italic>
<sub>50&#x2212;95</sub>, and <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic> metrics, respectively. Although its Precision is slightly lower than YOLOv8-Pose-cv&#x2019;s 0.945, it surpasses all other models in the remaining metrics, demonstrating its outstanding performance in red-ripe tomato detection and harvesting keypoint localization tasks. <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref> illustrate the experimental performance and comparative results of different models across various metrics.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Results of different models on self-made red-ripe stage tomato dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">model</th>
<th valign="middle" align="center">
<italic>P</italic>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50</sub>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50-95</sub>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic>
</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">YOLOv5-Pose</td>
<td valign="middle" align="center">0.874</td>
<td valign="middle" align="center">0.904</td>
<td valign="middle" align="center">0.675</td>
<td valign="middle" align="center">0.722</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv8-Pose</td>
<td valign="middle" align="center">0.882</td>
<td valign="middle" align="center">0.906</td>
<td valign="middle" align="center">0.682</td>
<td valign="middle" align="center">0.752</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv11-Pose</td>
<td valign="middle" align="center">0.863</td>
<td valign="middle" align="center">0.906</td>
<td valign="middle" align="center">0.684</td>
<td valign="middle" align="center">0.736</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv12-Pose</td>
<td valign="middle" align="center">0.87</td>
<td valign="middle" align="center">0.912</td>
<td valign="middle" align="center">0.675</td>
<td valign="middle" align="center">0.736</td>
</tr>
<tr>
<td valign="middle" align="center">yolov8-pose-ss</td>
<td valign="middle" align="center">0.88</td>
<td valign="middle" align="center">0.894</td>
<td valign="middle" align="center">0.647</td>
<td valign="middle" align="center">0.663</td>
</tr>
<tr>
<td valign="middle" align="center">yolov8-pose-cv</td>
<td valign="middle" align="center">0.945</td>
<td valign="middle" align="center">0.897</td>
<td valign="middle" align="center">0.687</td>
<td valign="middle" align="center">0.76</td>
</tr>
<tr>
<td valign="middle" align="center">Ours</td>
<td valign="middle" align="center">0.927</td>
<td valign="middle" align="center">0.917</td>
<td valign="middle" align="center">0.71</td>
<td valign="middle" align="center">0.785</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Results of different models on self-made red-ripe stage tomato dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g005.tif">
<alt-text content-type="machine-generated">Line graph comparing seven models based on four metrics: P, mAP@50, mAP@50-95, and mAP-kp@50-95. The models are yolov-pose, yolov8-pose, yolov11-pose, yolov12-pose, yolov8-pose-ss, yolov8-pose-cv, and Ours. Each metric is represented by different colored lines with distinct markers. The graph shows varying performance across models, with a notable peak at &#x201c;Ours&#x201d; for all metrics.</alt-text>
</graphic>
</fig>
<p>The experimental results comparing YOLOv5-Pose, YOLOv8-Pose, YOLOv11-Pose, YOLOv12-Pose, YOLOv8-Pose-ss, YOLOv8-Pose-cv, and YOLOv8-LBP models for ripe tomato recognition and picking keypoint detection are shown in <xref ref-type="fig" rid="f6">
<bold>Figures&#xa0;6</bold>
</xref>, <xref ref-type="fig" rid="f7">
<bold>7</bold>
</xref>. From column (i), it can be seen that models b and c have certain deviations in keypoint prediction, while models d, e, f, and g have lower tomato recognition confidence compared to model h; Column (ii) shows that under leaf occlusion, models c, d, e, and f exhibit multiple slight deviations in keypoint prediction, models b and g show larger deviations, while model f, despite some minor deviations, achieves better overall keypoint prediction accuracy than other models; Column (iii) indicates that models c, d, e, and g have false positives in tomato recognition, models c, d, e, f, and g all show certain deviations in keypoint prediction, with model f having the most accurate keypoint predictions and the highest tomato recognition confidence.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Examples of different experimental results.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g006.tif">
<alt-text content-type="machine-generated">Three panels of tomato plants under different detection models.   Panel (a) with label annotations highlights red tomatoes on the vine.   Panels (b), (c), and (d) show YOLOv5, YOLOv8, and YOLOv11 models, respectively, with blue arrows marking detected tomatoes.   Each model displays variable detection accuracy across different tomato clusters.</alt-text>
</graphic>
</fig>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>Examples of different experimental results.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g007.tif">
<alt-text content-type="machine-generated">Four rows of images compare tomato detection methods: YOLOv12-Pose, YOLOv8-Pose-cv, YOLOv8-Pose-ss, and a custom method labeled &#x201c;Ours.&#x201d; Each method is applied to tomatoes on plants, with variations in bounding boxes and annotations highlighting tomato positions.</alt-text>
</graphic>
</fig>
<p>
<xref ref-type="fig" rid="f8">
<bold>Figures&#xa0;8</bold>
</xref>, <xref ref-type="fig" rid="f9">
<bold>9</bold>
</xref> present a comparison of different models&#x2019; results on object recognition and keypoint detection in complex natural environments. Column (i) includes cases with hand background interference, clusters of tomatoes, and environmental occlusions; column (ii) shows scenarios with strong direct sunlight; column (iii) shows tomato clusters and fruit occlusions. Yellow arrows indicate false positives, while blue arrows denote other types of errors. In column (i), models c, d, e, and f exhibit three or more false positives; models b, g, and h have two false positives; models c, d, e, and g show missed detections of tomatoes; models b, f, and h show no missed detections; models b, d, e, and g have large deviations in keypoint predictions within multiple bounding boxes; models c, f, and h show large deviations within a single bounding box. In column (ii), models b, c, d, e, f, and g all have false positives under shadow conditions; models b, c, e, and g have false positives under direct sunlight; model g has two bounding boxes with large keypoint prediction deviations, while the other models have one such deviation each; model h has no false positives. In column (iii), models d, e, and f have false positives; models c, d, e, and f have missed detections; model g shows repeated detection of tomatoes; model h accurately recognizes every ripe tomato; models b, c, g, and h have large keypoint prediction deviations. Overall, model h outperforms the others. YOLOv8-LBP effectively addresses missegmentation caused by similar tomato fruit colors and occlusions in natural environments, providing more accurate keypoint localization, making it suitable for robotic harvesting applications.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>Comparative analysis of detection and prediction results of various models in complex natural scenes.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g008.tif">
<alt-text content-type="machine-generated">Three columns of images labeled (a) Label, (b) YOLOv5-Pose, (c) YOLOv8-Pose, and (d) YOLOv11-Pose showing tomatoes in various growth stages on the vine. Each column represents different pose estimation models with colored arrows and circles indicating detected points on the tomatoes.</alt-text>
</graphic>
</fig>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>Comparative analysis of detection and prediction results of various models in complex natural scenes.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1656381-g009.tif">
<alt-text content-type="machine-generated">Nine images display clusters of tomatoes on vines. The images are labeled (e) YOLOv12-Pose, (f) YOLOv8-Pose-cv, (g) YOLOv8-Pose-ss, and (h) Ours. Various colored arrows indicate points of interest on the tomatoes in each image, suggesting different detection methods or analyses.</alt-text>
</graphic>
</fig>
<p>To comprehensively evaluate the performance of different models, three commonly used computational metrics were selected for comparative analysis: the number of parameters (Params), Giga Floating-point Operations per Second (GFLOPs), and frames per second (FPS). The number of parameters (Params) represents the total count of all trainable parameters in the model, reflecting its complexity and storage requirements. GFLOPs measure the computational cost required for a single forward inference. FPS indicates how many image frames the model can process per second during actual operation and is an important metric for evaluating real-time performance. <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref> presents a comparison of these three computational metrics across different models. As shown in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>, YOLOv8-Pose-ss has the smallest number of parameters, the highest FPS, and the lowest computational cost (GFLOPs), indicating that our proposed model does not have an advantage in computational efficiency. Compared with the baseline model, our approach slightly increases the number of parameters (from 3.08M to 3.175M), raises GFLOPs by 0.1, and marginally improves FPS (from 96.15 to 99.01). Although computational cost increases slightly, this overhead is reasonable and acceptable, especially given the substantial improvements in accuracy: relative to the baseline model, object detection Precision improves by 4.5%, <italic>mAP</italic>
<sub>50</sub> by 1.1%, <italic>mAP</italic>
<sub>50&#x2212;95</sub> by 2.8%, and keypoint detection accuracy (<italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic>) by 3.3%. Compared with YOLOv8-Pose-ss, our model achieves 4.7% higher Precision (0.927 vs. 0.880), 2.3% higher <italic>mAP</italic>
<sub>50</sub> (0.917 vs. 0.894), 6.3% higher <italic>mAP</italic>
<sub>50&#x2212;95</sub> (0.710 vs. 0.647), and 12.2% higher <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic> (0.785 vs. 0.663). These results indicate that our approach achieves a favorable trade-off between accuracy and computational cost. Furthermore, with the continuous advancement of embedded device computing power and the widespread adoption of cloud computing, the impact of a slight increase in computational resources on practical deployment is gradually diminishing. Therefore, we believe the proposed method retains strong practical applicability and scalability while maintaining high accuracy.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Performance comparison of different models under typical evaluation metrics.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Model</th>
<th valign="middle" align="center">Parameters (M)</th>
<th valign="middle" align="center">FPS</th>
<th valign="middle" align="center">GFLOPs</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">YOLOv5-Pose</td>
<td valign="middle" align="center">2.58</td>
<td valign="middle" align="center">75.76</td>
<td valign="middle" align="center">7.3</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv8-Pose</td>
<td valign="middle" align="center">3.08</td>
<td valign="middle" align="center">96.15</td>
<td valign="middle" align="center">8.4</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv11-Pose</td>
<td valign="middle" align="center">2.65</td>
<td valign="middle" align="center">84.75</td>
<td valign="middle" align="center">6.6</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv12-Pose</td>
<td valign="middle" align="center">2.63</td>
<td valign="middle" align="center">73.5</td>
<td valign="middle" align="center">6.6</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv8-Pose-ss</td>
<td valign="middle" align="center">1.91</td>
<td valign="middle" align="center">113.64</td>
<td valign="middle" align="center">5.3</td>
</tr>
<tr>
<td valign="middle" align="center">YOLOv8-Pose-cv</td>
<td valign="middle" align="center">5.55</td>
<td valign="middle" align="center">109.9</td>
<td valign="middle" align="center">22.7</td>
</tr>
<tr>
<td valign="middle" align="center">Ours</td>
<td valign="middle" align="center">3.175</td>
<td valign="middle" align="center">99.01</td>
<td valign="middle" align="center">8.5</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Ablation experiments</title>
<p>To verify the effectiveness of each module, an ablation study was conducted based on the YOLOv8n-Pose model by gradually adding and replacing components. The results are shown in <xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>. Compared to the original YOLOv8n-Pose, adding the LSKA attention mechanism alone expanded the model&#x2019;s effective receptive field, enabling the extraction of richer feature information. This led to increases of 0.7% and 0.5% in <italic>mAP</italic>
<sub>50</sub> and <italic>mAP</italic>
<sub>50&#x2212;95</sub>, respectively. When adding the BiFPN module alone, the model&#x2019;s ability to process multi-scale feature information was enhanced, resulting in a 1.9% improvement in precision (<italic>P</italic>). When both LSKA and BiFPN were added simultaneously, the model not only obtained richer feature information but also leveraged BiFPN&#x2019;s ability to adaptively adjust the importance of these features, achieving more effective feature fusion. This allowed the model to better focus on the overall shape features of the target while reducing the interference of background textures and color similarity, leading to improvements of 4.5%, 1.1%, 2.8%, and 3.3% in <italic>P</italic>, <italic>mAP</italic>
<sub>50</sub>, <italic>mAP</italic>
<sub>50&#x2212;95</sub>, and <italic>mAP</italic>
<sub>50&#x2212;95</sub>&#x2212;<italic>kp</italic>, respectively. The ablation experiments demonstrate that the two improvements proposed in this study are effective.</p>
<table-wrap id="T6" position="float">
<label>Table&#xa0;6</label>
<caption>
<p>Ablation experiment results.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">LSKA</th>
<th valign="middle" align="center">BiFPN</th>
<th valign="middle" align="center">
<italic>P</italic>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50</sub>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50-95</sub>
</th>
<th valign="middle" align="center">
<italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic>
</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">&#xd7;</td>
<td valign="middle" align="center">&#xd7;</td>
<td valign="middle" align="center">0.882</td>
<td valign="middle" align="center">0.906</td>
<td valign="middle" align="center">0.682</td>
<td valign="middle" align="center">0.752</td>
</tr>
<tr>
<td valign="middle" align="center">&#x2713;</td>
<td valign="middle" align="center">&#xd7;</td>
<td valign="middle" align="center">0.881</td>
<td valign="middle" align="center">0.913</td>
<td valign="middle" align="center">0.687</td>
<td valign="middle" align="center">0.748</td>
</tr>
<tr>
<td valign="middle" align="center">&#xd7;</td>
<td valign="middle" align="center">&#x2713;</td>
<td valign="middle" align="center">0.901</td>
<td valign="middle" align="center">0.899</td>
<td valign="middle" align="center">0.679</td>
<td valign="middle" align="center">0.74</td>
</tr>
<tr>
<td valign="middle" align="center">&#x2713;</td>
<td valign="middle" align="center">&#x2713;</td>
<td valign="middle" align="center">0.927</td>
<td valign="middle" align="center">0.917</td>
<td valign="middle" align="center">0.71</td>
<td valign="middle" align="center">0.785</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s4" sec-type="conclusions">
<label>4</label>
<title>Conclusion</title>
<p>To achieve accurate detection of harvesting keypoints for ripe tomatoes in complex scenes, this paper draws on human pose estimation methods and proposes the YOLOv8-LBP model based on BiFPN and LSKA attention mechanisms to accomplish the recognition of ripe tomato fruits and the prediction of their harvesting keypoints. The weighted bidirectional feature pyramid network (BiFPN) is used to learn different input features, while the LSKA attention mechanism enlarges the model&#x2019;s effective receptive field, enabling more efficient multi-scale target information extraction and suppressing the influence of interfering features. Ablation experiments show that on the self-made ripe tomato dataset, the YOLOv8-LBP model improves <italic>P</italic>, <italic>mAP</italic>
<sub>50</sub>, <italic>mAP</italic>
<sub>50&#x2212;95</sub>, and <italic>mAP</italic>
<sub>50&#x2212;95</sub> &#x2212; <italic>kp</italic> by 4.5%, 1.1%, 2.8%, and 3.3% respectively, demonstrating significant enhancement; Only a small computational overhead was introduced, with the number of parameters increasing by 0.095 M, GFLOPs increasing by 0.1, and inference speed improving by 2.86 FPS. Experimental results confirm the effectiveness of the proposed improvements. The research outcomes provide effective support for enhancing ripe tomato harvesting capabilities in production and hold positive significance for the future development of smart agriculture.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>WY: Methodology, Writing &#x2013; review &amp; editing. XS: Investigation, Writing &#x2013; review &amp; editing, Writing &#x2013; original draft, Software. CK: Writing &#x2013; review &amp; editing, Data curation.</p>
</sec>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="ai-statement">
<title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors&#xa0;and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Cherry tomato detection for harvesting using multimodal perception and an improved yolov7-tiny neural network</article-title>. <source>Agronomy</source> <volume>14</volume>, <fpage>2320</fpage>&#x2013;<lpage>2336</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/agronomy14102320</pub-id>
</citation></ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Choi</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>Y. J.</given-names>
</name>
<name>
<surname>Bunn</surname> <given-names>J. M.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Tomato maturity evaluation using color image analysis</article-title>. <source>Trans. ASAE</source> <volume>38</volume>, <fpage>171</fpage>&#x2013;<lpage>176</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.13031/2013.27827</pub-id>
</citation></ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cubuk</surname> <given-names>E. D.</given-names>
</name>
<name>
<surname>Zoph</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Mane</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Vasudevan</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Le</surname> <given-names>Q. V.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Autoaugment: Learning augmentation strategies from data</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <fpage>113</fpage>&#x2013;<lpage>123</lpage>.</citation></ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Dslk-unet: Multiscale and large kernel convolution improved segmentation of skin lesions</article-title>. <source>Inf. Res.</source> <volume>50</volume>, <fpage>33</fpage>&#x2013;<lpage>41</lpage>.</citation></ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ding</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Scaling up your kernels to 31x31: Revisiting large kernel design in cnns</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <fpage>11963</fpage>&#x2013;<lpage>11975</lpage>.</citation></ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Feng</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Detection of grape cluster in stone hardening stage based on improved yolov5s</article-title>. <source>J. Chin. Agric. Mechanization</source> <volume>45</volume>, <fpage>240</fpage>&#x2013;<lpage>245</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.13733/j.jcam.issn.2095-5553.2024.08.035</pub-id>
</citation></ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Rapid palletizing recognition and grasp point detection based on improved yolov8-pose</article-title>. <source>J. Modular Mach. Tool Automatic Manufacturing Technol.</source> <volume>11</volume>, <fpage>125</fpage>&#x2013;<lpage>129</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.13462/j.cnki.mmtamt.2024.11.024.024</pub-id>
</citation></ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kang</surname> <given-names>S.-W.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>K.-C.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>Y.-J.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>D.-H.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Real-time pose estimation of oriental melon fruit-pedicel pairs using weakly localized fruit regions via class activation map</article-title>. <source>Comput. Electron. Agric.</source> <volume>237</volume>, <fpage>110734</fpage>&#x2013;<lpage>110745</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2025.110734</pub-id>
</citation></ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>D.-H.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>K.-C.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>Y.-J.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>2d pose estimation of multiple tomato fruit-bearing systems for robotic harvesting</article-title>. <source>Comput. Electron. Agric.</source> <volume>211</volume>, <fpage>108004</fpage>&#x2013;<lpage>108015</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2023.108004</pub-id>
</citation></ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname> <given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>25</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.1145/3065386</pub-id>
</citation></ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lau</surname> <given-names>K. W.</given-names>
</name>
<name>
<surname>Po</surname> <given-names>L.-M.</given-names>
</name>
<name>
<surname>Rehman</surname> <given-names>Y. A. U.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Large separable kernel attention: Rethinking the large kernel attention design in cnn</article-title>. <source>Expert Syst. Appl.</source> <volume>236</volume>, <fpage>121352</fpage>&#x2013;<lpage>121367</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.eswa.2023.121352</pub-id>
</citation></ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Ctda: an accurate and efficient cherry tomato detection algorithm in complex environments</article-title>. <source>Front. Plant Sci.</source> <volume>16</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2025.1492110</pub-id>, PMID: <pub-id pub-id-type="pmid">40182545</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>T.-Y.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Hariharan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Belongie</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Feature pyramid networks for object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. <fpage>2117</fpage>&#x2013;<lpage>2125</lpage>.</citation></ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Y-hrnet: Research on multi-category cherry tomato instance segmentation model based on improved yolov7 and hrnet fusion</article-title>. <source>Comput. Electron. Agric.</source> <volume>227</volume>, <fpage>109531</fpage>&#x2013;<lpage>109533</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2024.109531</pub-id>
</citation></ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Qi</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Qin</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Path aggregation network for instance segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. <fpage>8759</fpage>&#x2013;<lpage>8768</lpage>.</citation></ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A vision methodology for harvesting robot to detect cutting points on peduncles of double overlapping grape clusters in a vineyard</article-title>. <source>Comput. Industry</source> <volume>99</volume>, <fpage>130</fpage>&#x2013;<lpage>139</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compind.2018.03.017</pub-id>
</citation></ref>
<ref id="B17">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Maji</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Nagori</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Mathew</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Poddar</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Yolo-pose: Enhancing yolo for multi person pose estimation using object keypoint similarity loss</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <fpage>2637</fpage>&#x2013;<lpage>2646</lpage>.</citation></ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mochen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhenyuan</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Mingshi</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Qinglu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Jinxing</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Huawei</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Red-ripe strawberry recognition and peduncle detection based on improved yolov8-pose</article-title>. <source>Trans. Chin. Soc. Agric. Machinery</source> <volume>54</volume>, <fpage>244</fpage>&#x2013;<lpage>252</lpage>. doi:&#xa0;910.6041/j.issn.1000-1298.2023.S2.029
</citation></ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rong</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A peduncle detection method of tomato for autonomous harvesting</article-title>. <source>Complex Intelligent Syst.</source> <volume>8</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s40747-021-00522-7</pub-id>
</citation></ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sa</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Lehnert</surname> <given-names>C.</given-names>
</name>
<name>
<surname>English</surname> <given-names>A.</given-names>
</name>
<name>
<surname>McCool</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Dayoub</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Upcroft</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Peduncle detection of sweet pepper for autonomous crop harvesting&#x2014;combined color and 3-d information</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>2</volume>, <fpage>765</fpage>&#x2013;<lpage>772</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LRA.2017.2651952</pub-id>
</citation></ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Su</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Pw-yolo-pose: A novel algorithm for pose estimation of power workers</article-title>. <source>IEEE Access</source> <volume>12</volume>, <fpage>116841</fpage>&#x2013;<lpage>116860</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3437359</pub-id>
</citation></ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tan</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Pang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Le</surname> <given-names>Q. V.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Efficientdet: Scalable and efficient object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <fpage>10781</fpage>&#x2013;<lpage>10790</lpage>.</citation></ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tao</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Tomatodiverse:an open dataset for industrial tomato detection in complex natural environments</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.57760/sciencedb.13850</pub-id>
</citation></ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2025</year>). &#x201c;<article-title>Leff-yolo: A lightweight cherry tomato detection yolov8 network with enhanced feature fusion</article-title>,&#x201d; in <source>Advanced Intelligent Computing Technology and Applications</source>. Eds. <person-group person-group-type="editor">
<name>
<surname>Huang</surname> <given-names>D.-S.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>Y.</given-names>
</name>
</person-group> (<publisher-name>Springer Nature Singapore</publisher-name>, <publisher-loc>Singapore</publisher-loc>), <fpage>474</fpage>&#x2013;<lpage>488</lpage>.</citation></ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>A method for identifying grape stems using keypoints</article-title>. <source>Comput. Electron. Agric.</source> <volume>209</volume>, <fpage>107825</fpage>&#x2013;<lpage>107835</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2023.107825</pub-id>
</citation></ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Haiyan</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Challenges and countermeasures for the high-quality development of smart agriculture in China</article-title>. <source>J. smart Agric.</source> <volume>5</volume>, <fpage>6</fpage>&#x2013;<lpage>9</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.20028/j.zhnydk.2025.04.002</pub-id>
</citation></ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Ai</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Cha</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2024</year>a). <article-title>Target detection method for tomato harvesting robot based on shufflenetv2-yolov5</article-title>. <source>J. Heilongjiang Institute Eng.</source> <volume>38</volume>, <fpage>9</fpage>&#x2013;<lpage>15</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.19352/j.cnki.issn1671-4679.2024.05.002</pub-id>
</citation></ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>End-effector design and motion simulation of an apple-picking robot</article-title>. <source>Hydraulics Pneumatics Seals</source> <volume>43</volume>, <fpage>20</fpage>&#x2013;<lpage>25</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3969/j.issn.1008-0813.2023.06.005</pub-id>
</citation></ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bian</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2024</year>b). <article-title>Design and simulation of a vacuum-based end-effector for an apple-picking robot</article-title>. <source>Hydraulics Pneumatics Seals</source> <volume>44</volume>, <fpage>39</fpage>&#x2013;<lpage>46</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3969/j.issn.1008-0813.2024.01.007</pub-id>
</citation></ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Xue</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2024</year>). &#x201c;<article-title>Rfa-yolo-pose: A fusion algorithm for pose detection and object identification amidst complex crowds</article-title>,&#x201d; in <conf-name>2024 5th International Seminar on Artificial Intelligence</conf-name>. <fpage>966</fpage>&#x2013;<lpage>969</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/AINIT61980.2024.10581583</pub-id>
</citation></ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>He</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Bag of freebies for training object detection neural networks</article-title>. <source>arXiv preprint arXiv:1902.04103</source>.</citation></ref>
</ref-list>
</back>
</article>