<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2023.1243822</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A New Deep Learning-based Dynamic Paradigm Towards Open-World Plant Disease Detection</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Dong</surname>
<given-names>Jiuqing</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1992135"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Fuentes</surname>
<given-names>Alvaro</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/551374"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yoon</surname>
<given-names>Sook</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/595546"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Kim</surname>
<given-names>Hyongsuk</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/833601"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Jeong</surname>
<given-names>Yongchae</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2093017"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Park</surname>
<given-names>Dong Sun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/567101"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Electronic Engineering, Jeonbuk National University</institution>, <addr-line>Jeonju</addr-line>, <country>Republic of Korea</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Core Research Institute of Intelligent Robots, Jeonbuk National University</institution>, <addr-line>Jeonju</addr-line>, <country>Republic of Korea</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Department of Computer Engineering, Mokpo National University</institution>, <addr-line>Muan</addr-line>, <country>Republic of Korea</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Yongliang Qiao, University of Adelaide, Australia</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Preeta Sharan, The Oxford College of Engineering, India; Jun Liu, Shandong Provincial University Laboratory for Protected Horticulture, China; Catarina Silva, University of Coimbra, Portugal</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Hyongsuk Kim, <email xlink:href="mailto:hskim@jbnu.ac.kr">hskim@jbnu.ac.kr</email>; Sook Yoon, <email xlink:href="mailto:syoon@mokpo.ac.kr">syoon@mokpo.ac.kr</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1243822</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>09</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Dong, Fuentes, Yoon, Kim, Jeong and Park</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Dong, Fuentes, Yoon, Kim, Jeong and Park</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Plant disease detection has made significant strides thanks to the emergence of deep learning. However, existing methods have been limited to closed-set and static learning settings, where models are trained using a specific dataset. This confinement restricts the model&#x2019;s adaptability when encountering samples from unseen disease categories. Additionally, there is a challenge of knowledge degradation for these static learning settings, as the acquisition of new knowledge tends to overwrite the old when learning new categories. To overcome these limitations, this study introduces a novel paradigm for plant disease detection called open-world setting. Our approach can infer disease categories that have never been seen during the model training phase and gradually learn these unseen diseases through dynamic knowledge updates in the next training phase. Specifically, we utilize a well-trained unknown-aware region proposal network to generate pseudo-labels for unknown diseases during training and employ a class-agnostic classifier to enhance the recall rate for unknown diseases. Besides, we employ a sample replay strategy to maintain recognition ability for previously learned classes. Extensive experimental evaluation and ablation studies investigate the efficacy of our method in detecting old and unknown classes. Remarkably, our method demonstrates robust generalization ability even in cross-species disease detection experiments. Overall, this open-world and dynamically updated detection method shows promising potential to become the future paradigm for plant disease detection. We discuss open issues including classification and localization, and propose promising approaches to address them. We encourage further research in the community to tackle the crucial challenges in open-world plant disease detection. The code will be released at <ext-link ext-link-type="uri" xlink:href="https://github.com/JiuqingDong/OWPDD">https://github.com/JiuqingDong/OWPDD</ext-link>.</p>
</abstract>
<kwd-group>
<kwd>plant disease detection</kwd>
<kwd>incremental learning</kwd>
<kwd>open-world detection</kwd>
<kwd>out-of-distribution detection</kwd>
<kwd>dynamic paradigm</kwd>
</kwd-group>
<contract-num rid="cn001">2019R1A6A1A09031717, NRF-2021R1A2C1012174</contract-num>
<contract-num rid="cn002">1545027569</contract-num>
<contract-sponsor id="cn001">National Research Foundation of Korea<named-content content-type="fundref-id">10.13039/501100003725</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Korea Institute of Planning and Evaluation for Technology in Food, Agriculture and Forestry<named-content content-type="fundref-id">10.13039/501100014189</named-content>
</contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="7"/>
<equation-count count="3"/>
<ref-count count="56"/>
<page-count count="18"/>
<word-count count="10749"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Sustainable and Intelligent Phytoprotection</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Accurate and timely detection and diagnosis of plant diseases are crucial for preserving crop health and increasing agricultural productivity. However, traditional methods of plant disease detection primarily rely on skilled agricultural professionals who diagnose diseases based on visual symptoms and pathologic characteristics of pathogens. These methods suffer from limitations such as subjectivity, prolonged diagnosis time, and dependence on experienced experts (<xref ref-type="bibr" rid="B4">Dong et&#xa0;al., 2022</xref>). To address these limitations of traditional methods, plant disease detection based on image analysis and artificial intelligence has emerged as a hot research topic (<xref ref-type="bibr" rid="B45">Shoaib et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B54">Xu et&#xa0;al., 2023</xref>). This emerging approach utilizes images captured from various plant parts such as leaves and stems, followed by computer algorithms for image analysis and recognition, enabling automated detection and diagnosis of plant diseases. This method not only enhances the accuracy and efficiency of detection but also allows non-experts to participate in plant disease monitoring and diagnosis (<xref ref-type="bibr" rid="B35">Panchal et&#xa0;al., 2023</xref>).</p>
<p>A substantial body of published work attests to the success of deep learning in plant disease detection tasks (<xref ref-type="bibr" rid="B10">Fuentes et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B22">Li et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B34">Nazki et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B46">Singh et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B7">Fenu and Malloci, 2021</xref>; <xref ref-type="bibr" rid="B9">Fuentes et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B39">Qiao et&#xa0;al., 2022b</xref>). However, existing studies focus on fixed disease categories of specific species with all available annotations during the training phase. This training strategy is known as closed-set learning (<xref ref-type="bibr" rid="B53">Xiong et&#xa0;al., 2019</xref>). In this case, the model is more likely to classify suspicious regions as one of the categories it has already learned, rather than indicating the presence of an abnormal disease type (<xref ref-type="bibr" rid="B6">Du et&#xa0;al., 2022b</xref>). We show the potential risks associated with closed-set learning in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1A</bold>
</xref>. Note that &#x201c;known classes&#x201d; refer to the classes present in the training dataset, while &#x201c;unknown classes&#x201d; refer to the classes that exist in real-world scenarios but are either absent or unannotated in the training dataset.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Comparison of three different learning paradigms. <bold>(A)</bold> Closed-set-based models detect unknown diseases as known diseases; <bold>(B)</bold> Open-set based models can detect unknown diseases but do not learn them; <bold>(C)</bold> Open-world detector learns the known diseases and also autonomously detects unknown diseases. The identified unknown diseases are then provided as feedback to domain experts, who annotate these newly discovered labels. This valuable information is incorporated into the model during subsequent tasks, allowing it to adaptively update itself with new knowledge.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g001.tif"/>
</fig>
<p>In the concept of plant stress, unknown diseases may result in a large economic loss, and recognizing them is thus one of the fundamental demands (<xref ref-type="bibr" rid="B11">Geng et&#xa0;al., 2020</xref>). Therefore, unknown disease detection is more useful in most practical scenarios. The learning paradigm that can detect unknown classes is known as open-set learning (<xref ref-type="bibr" rid="B47">Vaze et&#xa0;al., 2021</xref>). <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1B</bold>
</xref> illustrates the open-set learning paradigm, which allows the model to detect instances that are currently unknown to the model. Developmental psychology (<xref ref-type="bibr" rid="B28">Livio, 2017</xref>) has revealed that the ability to recognize the unknown is crucial for stimulating curiosity, which in turn fuels the desire to learn new things. In the open-set learning paradigm, when the model detects unknown diseases and provides feedback to domain experts, it is important for the domain experts to pay attention to these disease samples and assign them appropriate category labels. This allows the model to further learn about these new diseases.</p>
<p>To learn these new diseases, one naive learning strategy is to combine the new and old data together and let the model learn again. However, as the number of tasks increases, the accumulated data volume becomes significantly large, resulting in high training costs. This approach may be feasible in the short term but is not sustainable as a long-term training strategy. Another learning method is to fine-tune the old model using new data. In this way, the model will quickly adapt to the new task but there is a risk of losing the ability to detect previously known classes. This prompts us to propose a new challenge: a new paradigm should be capable of recognizing instances of unknown diseases as unknown and gradually learning these unknown categories through incremental learning. <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1C</bold>
</xref> illustrates the workflow of this new paradigm.</p>
<p>Plant growth is a dynamic process, and plant disease dynamics are more complex than we imagined. During the plant growth cycle monitoring, unexpected diseases and pests are likely to emerge. Simultaneously, collecting all the existing plant diseases is difficult and even impossible for real-world applications (<xref ref-type="bibr" rid="B54">Xu et&#xa0;al., 2023</xref>). Given the dynamic nature of our world, the setup of open-world plant disease detection is more aligned with real-world applications compared to existing closed-set learning and open-set learning settings. Therefore, we need to introduce a new paradigm to continuously learn these unknown diseases instead of learning them all at once. In this paradigm, the model can detect unknown diseases and provide feedback to domain experts. Then, experts will label these unknown diseases. As and when more information about the identified unknown classes becomes available, the system should be able to incorporate them into its existing knowledge base. This iterative learning process will cycle throughout the model&#x2019;s lifecycle. In this paper, we propose an open-world detector for plant disease detection, aiming to achieve this goal.</p>
<p>The key contributions of our work as follows:</p>
<list list-type="simple">
<list-item>
<p>1. We introduce the concept of open-world problem formulation into plant disease detection for the first time, enabling a closer simulation of real-world application scenarios. Unlike all existing plant disease detectors, it dynamically expands the learned categories and actively responds to unknown diseases.</p>
</list-item>
<list-item>
<p>2. We introduce an unknown-aware region proposal network (UA-RPN) and conducted pre-training on various datasets. We find that the model pre-trained on LVIS (Large Vocabulary Instance Segmentation) (<xref ref-type="bibr" rid="B14">Gupta et&#xa0;al., 2019</xref>) dataset can exhibit superior performance across different experimental setups. Additionally, we propose a class-agnostic region of interest (ROI) head, which significantly improved the recall rate for unknown classes. Interestingly, the model trained on a dataset of tomato leaf diseases could even detect diseases in paprika fruit.</p>
</list-item>
<list-item>
<p>3. Our method also achieves class incremental detection of plant diseases. Additionally, we discuss the open issues associated with open-world plant disease detection and provide promising solutions. We believe that this open-world and dynamically updated detection method can become a new paradigm for future plant disease detection, and we encourage the research community to explore and address these open challenges.</p>
</list-item>
</list>
<p>Section 2 provides a detailed review of the deep learning techniques employed for plant anomaly detection and existing open-set and open-world deep learning approaches. Section 3 comprehensively describes the problem formulation, methodology, and evaluation framework of the novel paradigm we have introduced. In Section 4, experimental results are presented to demonstrate the effectiveness and expandability of our proposed approach. We have observed that the proposed method achieves cross-species disease detection. Furthermore, we discuss the open challenges concerning plant disease detection in the context of open-world detection. In the final section, we provide several conclusions to guide future researchers. In summary, this work establishes the foundation for open-world detection in intelligent agriculture and advocates for increased attention to incremental learning and unknown target detection within the community.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related works</title>
<p>In this section, we provide a brief overview of recent studies relevant to our proposed approach. Firstly, we delve into existing deep learning-based methods employed in plant disease detection. Furthermore, considering the limitations of the latest advancements in plant disease recognition, no previous work specifically addresses open-world detection. Consequently, we explore two closely related avenues: open-set detection and open-world detection.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Deep learning technics in plant disease detection</title>
<p>In recent years, various deep learning-based object detection algorithms have been applied in plant disease detection task (<xref ref-type="bibr" rid="B38">Qiao et&#xa0;al., 2022a</xref>; <xref ref-type="bibr" rid="B45">Shoaib et&#xa0;al., 2023</xref>). In the two-stage plant disease detection methods, <xref ref-type="bibr" rid="B8">Fuentes et&#xa0;al. (2017)</xref> first used Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>) to accurately locate tomato diseases and pests in a dataset consisting of 4800 images with 11 different classes. When using deep feature extractors like VGG-Net and ResNet, the mean average precision (mAP) was calculated as 88.66%. <xref ref-type="bibr" rid="B26">Liu and Wang (2021)</xref> suggested modifying the Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>) framework to automatically detect beet spot diseases by changing the parameters of the CNN model. <xref ref-type="bibr" rid="B37">Priyadharshini and Dolly (2023)</xref> provided a comparative investigation on tomato leaf disease detection and classification using RCNN (<xref ref-type="bibr" rid="B13">Girshick et&#xa0;al., 2014</xref>), Fast RCNN (<xref ref-type="bibr" rid="B12">Girshick, 2015</xref>) and Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>). <xref ref-type="bibr" rid="B33">Murugeswari et&#xa0;al. (2022)</xref> trained a model using 1500 images of healthy and diseased sugarcane leaves and deployed the model in an android application. <xref ref-type="bibr" rid="B43">Seetharaman and Mahendran (2022)</xref> proposed using a convolutional recurrent neural network for banana leaf disease detection. <xref ref-type="bibr" rid="B1">Alruwaili et&#xa0;al. (2022)</xref> proposed real-time faster region convolutional neural network (RTF-RCNN) for the real-time detection of tomato leaf diseases in video streams.</p>
<p>In the application of single-stage networks, <xref ref-type="bibr" rid="B55">Zhang et&#xa0;al. (2019)</xref> proposed a new method for detecting small agricultural pests by combining an improved version of the YOLOv3 algorithm with spatial pyramid pooling. This method addresses the low accuracy caused by the varying poses and scales of crop pests by applying deconvolution, oversampling, and convolution operations. <xref ref-type="bibr" rid="B31">Mathew and Mahesh (2022)</xref> used YOLOv5 to detect bell pepper leaf disease. <xref ref-type="bibr" rid="B48">Wang et&#xa0;al. (2022)</xref> optimized the lightweight YOLOv5 model for detecting peanut diseases. Additionally, <xref ref-type="bibr" rid="B4">Dong et&#xa0;al. (2022)</xref> evaluated the performance of different annotation strategies based on the YOLOv5 model.</p>
<p>During the training process, the aforementioned methods have access to all labels. However, they cannot locate and classify unknown diseases. In the task of plant disease classification, <xref ref-type="bibr" rid="B9">Fuentes et&#xa0;al. (2021)</xref> proposed an approach based on the concept of open-set domain adaptation to the task of plant disease recognition to allow existing systems to operate in new environments with unseen conditions and farms. To the best of our knowledge, there is currently no relevant work on detecting unknown diseases in plant disease detection tasks.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Out-of-distribution detection</title>
<p>The class in the training dataset refers to the &#x2018;known class&#x2019; while a class existing in the test dataset but not in the training dataset is termed an &#x2018;unknown class&#x2019;. Determining whether inputs are out-of-distribution (OOD) is an essential building block for safely deploying machine learning models in the open world. OOD detection is crucial for ensuring the reliability and usability of systems in the real world. <xref ref-type="bibr" rid="B18">Hendrycks and Gimpel (2016)</xref> proposed a baseline for OOD detection that relies on softmax confidence scores. However, such methods can be influenced by overconfidence in the posterior distribution of OOD data. <xref ref-type="bibr" rid="B27">Liu et&#xa0;al. (2020)</xref> demonstrated mathematically that the softmax confidence score is a biased scoring function that is not aligned with the density of the inputs and hence is not suitable for OOD detection.</p>
<p>The energy-based model maps each input to a single scalar that is lower for observed data and higher for unobserved ones (<xref ref-type="bibr" rid="B21">Lecun et&#xa0;al., 2006</xref>). <xref ref-type="bibr" rid="B27">Liu et&#xa0;al. (2020)</xref> first proposed a unified framework for OOD detection using energy scores. Unlike softmax confidence scores, energy scores are theoretically aligned with the probability density of the input and are less susceptible to issues of overconfidence. <xref ref-type="bibr" rid="B19">Joseph et&#xa0;al. (2021)</xref> were the first to apply energy-based OOD detection to object detection. In this paper, we follow the setup of (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>) and maintain a validation set to learn the energy distribution of both known and unknown classes.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Open-world object detection</title>
<p>Open-world object detection is an emerging topic in computer vision and has attracted extensive attention due to its practicability in the real world. Unlike OOD tasks that only focus on the identification of unknown classes, open-world tasks require models to learn new classes and recognize old classes. This learning process is also known as incremental learning. To our best knowledge, there have been only a few relevant works published in top-tier conferences and journals (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B15">Gupta et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B51">Wu et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B29">Ma et&#xa0;al., 2023a</xref>; <xref ref-type="bibr" rid="B30">Ma et&#xa0;al., 2023b</xref>; <xref ref-type="bibr" rid="B56">Zohar et&#xa0;al., 2023</xref>). Based on network architecture, these works can be categorized into methods based on Region Proposal Network (RPN) (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B51">Wu et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B30">Ma et&#xa0;al., 2023b</xref>) and methods based on Transformer (<xref ref-type="bibr" rid="B15">Gupta et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B29">Ma et&#xa0;al., 2023a</xref>; <xref ref-type="bibr" rid="B30">Ma et&#xa0;al., 2023b</xref>; <xref ref-type="bibr" rid="B56">Zohar et&#xa0;al., 2023</xref>).</p>
<p>To endow the model with the capacity of detecting unknown objects, <xref ref-type="bibr" rid="B19">Joseph et&#xa0;al. (2021)</xref> proposed the Open World Object Detection (ORE) method, in which an unknown auto-labeling RPN is designed to generate pseudo labels for unknown instances. <xref ref-type="bibr" rid="B15">Gupta et&#xa0;al. (2022)</xref> and <xref ref-type="bibr" rid="B56">Zohar et&#xa0;al. (2023)</xref> employed an attention mechanism to score candidate bounding boxes, enhancing the network&#x2019;s perception capability for unknown objects. <xref ref-type="bibr" rid="B29">Ma et&#xa0;al. (2023a)</xref> proposed a method that combines selective search and attention mechanisms to further enhance the retrieval capability for unknown objects. The underlying logic behind these methods is to enhance the proposal quality for unknown objects in order to obtain stronger weak supervision signals. However, methods based on attention mechanisms and selective search tend to be complex. Optimizing the perception capability for unknown objects through a simpler approach is indeed more desirable in practical engineering scenarios. Therefore, we improve the proposal quality of the network for unknown objects by using a pre-trained region proposal network (RPN), thereby enhancing the performance of open-world plant disease detection.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Challenges of real-world plant disease detection</title>
<p>Plant disease detection is a complex field that possesses distinct characteristics and challenges, particularly when considering the influence of diverse domains such as greenhouse conditions. Incremental learning serves as a crucial tool to address these challenges and enhance the accuracy and adaptability of disease detection systems.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Characteristics of plant disease detection</title>
<p>The process of plant disease detection is marked by several unique characteristics. Unlike some other domains, plant health is influenced by an intricate interplay of factors. Variations in features across plant species, genetic diversity, and environmental conditions lead to a diverse range of disease symptoms. These symptoms can be subtle, ranging from changes in leaf color and texture to wilting and necrosis. Additionally, the progression of diseases can vary widely, making it challenging to predict the trajectory and severity of an infection.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Challenges in diverse domains and greenhouse conditions</title>
<p>Diverse domains, such as greenhouse environments, introduce a set of challenges that impact plant disease detection. Greenhouses provide controlled conditions for plant growth, which can accelerate disease progression due to the close proximity of plants, regulated temperature, and humidity. The dynamic interactions between plants, pathogens, and the environment within greenhouses contribute to complex disease patterns that traditional, static models might struggle to capture. Moreover, the controlled environment can lead to rapid mutations in pathogens, adding further complexity to disease identification.</p>
<p>In a domain characterized by diverse symptoms, feature variations, environmental factors, and disease progression, previous models to detect plant disease can fall short. Our proposed approach, however, enables models to evolve alongside the evolving disease landscape. The adaptive nature of our approach allows models to incorporate new information, adapt to feature variations, and account for changing environmental conditions. As the disease patterns shift and pathogens mutate, incremental learning ensures that the detection system remains up-to-date and effective. This is particularly critical in greenhouse conditions, where rapid disease spread demands real-time monitoring and rapid response.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Problem formulation</title>
<p>In this section, we provide a formal definition of Open World Object Detection. In a closed-setting approach, a model is trained on a specific set of known classes and then tested on data collected from the same or similar environment such as <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msubsup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>M</mml:mi>
</mml:msubsup>
<mml:mo>&#x2282;</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msubsup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:msubsup>
<mml:mo>&#x2282;</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im3">
<mml:mi>X</mml:mi>
</mml:math>
</inline-formula> denotes the image samples in the dataset <inline-formula>
<mml:math display="inline" id="im4">
<mml:mi>D</mml:mi>
</mml:math>
</inline-formula>, and <inline-formula>
<mml:math display="inline" id="im5">
<mml:mi>C</mml:mi>
</mml:math>
</inline-formula> indicates the number of classes. However, real-world scenarios often involve new environments and the presence of unknown diseases that the model has not encountered before. Consequently, when tested on such data, the model may fail to perform accurately. In this context, the test dataset is <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:msubsup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:msubsup>
<mml:mo>&#x2282;</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>
<italic>(C</italic>&amp;<italic>U)</italic>, where <inline-formula>
<mml:math display="inline" id="im7">
<mml:mi>U</mml:mi>
</mml:math>
</inline-formula> denotes unknown classes in training phase. Therefore, in open-world disease detection, the primary target is to detect these unknown diseases.</p>
<p>After achieving the primary target, the model becomes capable of identifying diseases that were not part of the initial training set (unknown diseases). Our objective is for the model to learn these new classes in subsequent learning tasks while retaining its recognition ability for the classes learned int eh previous tasks. We define the initial training task as Task 1 and subsequent tasks as Task 2, Task 3, and so on. In Task 1, the training dataset, denoted as <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, consists of labeled samples for a number of <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> disease classes. However, during the inference process, the model may encounter instances of unknown diseases that were not seen during training. To address this, the model needs to accurately locate these unknown disease types and assign them the label &#x2018;unknown&#x2019;. These unknown disease instances will be presented to domain experts for annotation and will be used for training in Task 2. In Task 2, the number of new disease classes is denoted as <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. After completing Task 2, the set of known classes is updated to the previously known classes <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> along with the newly learned classes <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. However, during the inference process, the model still may encounter unknown diseases that do not belong to the known classes <inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Therefore, in addition to detecting the known classes, the model will continue to identify unknown diseases and assign them the label &#x2018;unknown&#x2019;. These unknown disease instances will be learned in Task 3.</p>
<p>This cycle of updating the model&#x2019;s knowledge continues throughout the entire lifecycle of the detector. In each task, the detector acquires new knowledge without forgetting the previously learned classes. This allows the model to continuously adapt and improve its detection capabilities by incorporating new information in a progressive manner.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Datasets and splits</title>
<p>After defining the open-world problem, it is necessary to search for suitable datasets to evaluate our method. In this study, we extended the tomato dataset used in previous works (<xref ref-type="bibr" rid="B10">Fuentes et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B9">Fuentes et&#xa0;al., 2021</xref>) to include 15 different classes, which were learned in Task 1, Task 2, and Task 3, respectively. To ensure a balanced distribution, we divided the classes equally, with 5 different classes assigned to each task. In Task 1, instances belonging to the classes of Task 2 and Task 3 were not available. Additionally, we aimed to investigate the performance of our model in cross-species training. For this purpose, we incorporated the paprika disease detection dataset (<xref ref-type="bibr" rid="B4">Dong et&#xa0;al., 2022</xref>) in Task 4. The tomato dataset originally consisted of 15 classes, while the paprika dataset contained 5 classes. To ensure the dataset&#x2019;s representation of real-world scenarios and to introduce complexity, we excluded images collected in a laboratory setting. This approach prevents potential overestimation of the model&#x2019;s performance and enhances the dataset&#x2019;s ability to simulate real-world conditions.</p>
<p>For each task, we employed a random selection process to designate 20% of the integrated dataset (combining tomato and paprika data) as the validation data. This allowed us to learn the distribution of known and unknown samples within this subset. Additionally, we randomly chose 20% of the data as the test set, which was used across all tasks. Here we aim to address the question: why do we test diseases from different species together? There are several reasons for this approach. Firstly, evaluating the performance of our model on different species&#x2019; diseases allows us to assess its generalization capability across species. In real-world scenarios, plant disease detection systems encounter various species and their associated diseases. By testing different species&#x2019; diseases together, we can effectively assess how well our model handles the challenges of detecting diseases across multiple species. This includes dealing with variations in symptoms, visual appearances, and disease patterns. Such evaluation helps us gain insights into the robustness and effectiveness of our model in practical applications where encounters with a diverse range of plant species are expected. Furthermore, successful detection of diseases from different species indicates that our model has acquired solid features of diseases as a concept. It demonstrates that the model&#x2019;s learning transcends species-specific information and can be effectively applied to diverse plant species. The dataset split and more specific details are presented in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. Unless otherwise specified, the training order of all experiments in this paper follows the sequence shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Task composition and data split in the proposed open-world plant disease detection protocol.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Task sequence</th>
<th valign="middle" align="center">Task 1</th>
<th valign="middle" align="center">Task 2</th>
<th valign="middle" align="center">Task 3</th>
<th valign="middle" align="center">Task 4</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Species</td>
<td valign="middle" align="center">Tomato</td>
<td valign="middle" align="center">Tomato</td>
<td valign="middle" align="center">Tomato</td>
<td valign="middle" align="center">Paprika</td>
</tr>
<tr>
<td valign="middle" align="center">Training Classes</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">5</td>
</tr>
<tr>
<td valign="middle" align="center">Categories</td>
<td valign="middle" align="center">magnesium deficiency,<break/>gray mold,<break/>leaf mold,<break/>yellow leaf curl virus,<break/>physical damage</td>
<td valign="middle" align="center">canker,<break/>plague,<break/>leaf miner,<break/>white fly,<break/>white fly egg</td>
<td valign="middle" align="center">wilt,<break/>chlorosis virus,<break/>stress,<break/>powdery mildew,<break/>old leaf</td>
<td valign="middle" align="center">blossom end rot,<break/>gray mold,<break/>powdery mildew,<break/>spider mite,<break/>spotting disease</td>
</tr>
<tr>
<td valign="middle" align="center">Training images</td>
<td valign="middle" align="center">3236</td>
<td valign="middle" align="center">1728</td>
<td valign="middle" align="center">1647</td>
<td valign="middle" align="center">2049</td>
</tr>
<tr>
<td valign="middle" align="center">Validation images</td>
<td valign="middle" align="center">2491</td>
<td valign="middle" align="center">2491</td>
<td valign="middle" align="center">2491</td>
<td valign="middle" align="center">2491</td>
</tr>
<tr>
<td valign="middle" align="center">Test images</td>
<td valign="middle" align="center">2493</td>
<td valign="middle" align="center">2493</td>
<td valign="middle" align="center">2493</td>
<td valign="middle" align="center">2493</td>
</tr>
<tr>
<td valign="middle" align="center">Known Classes</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">10</td>
<td valign="middle" align="center">15</td>
<td valign="middle" align="center">20</td>
</tr>
<tr>
<td valign="middle" align="center">Unknown Classes</td>
<td valign="middle" align="center">15</td>
<td valign="middle" align="center">10</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">0</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Architecture</title>
<p>In their study, <xref ref-type="bibr" rid="B3">Dhamija et&#xa0;al. (2020)</xref> found that two-stage networks outperform single-stage networks when it comes to detecting unknown objects. Motivated by this finding, we have chosen to implement our open-world detection model using the classic Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>), which is a two-stage network architecture. To enhance the representation of multi-scale features, we have incorporated the feature pyramid network (FPN) (<xref ref-type="bibr" rid="B24">Lin et&#xa0;al., 2017</xref>).</p>
<p>In <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>, we present an illustration of the Faster RCNN with the FPN network. Please note that our method, unlike the standard Faster RCNN, can detect unknown classes. This capability is achieved through a well-trained unknown perception Region Proposal Network (RPN) and a class-agnostic localization head. The unknown perception RPN is designed for automatic labeling of unknown objects, while the class-agnostic localization head is responsible for accurately localizing these unknown objects. Each of these components is explained in detail in the following subsections, providing a coherent understanding of their roles in our model.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Overview of our model where ResNet and FPN are constructed following the default approach in detectron2 (<xref ref-type="bibr" rid="B50">Wu et&#xa0;al., 2019</xref>). We illustrate the unknown-aware RPN and class-agnostic ROI head in the diagram. Unknown aware RPN modifies the labels of background candidate boxes with the highest object scores to &#x2018;unknown&#x2019;. The class-agnostic head focuses on regressing bounding boxes for disease regions without considering the disease category.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g002.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Well-trained unknowns-aware RPN</title>
<p>In the context of object detection tasks, the objective is to identify and localize objects of interest within an image. Traditional object detection models are typically trained on datasets that consist of known classes, assuming that all objects can be classified into predefined categories. However, real-world scenarios often present instances where the model encounters objects belonging to unknown or unseen classes.</p>
<p>To address the challenge of detecting unknown diseases, we introduce an additional &#x201c;unknown&#x201d; class during the training process. This class is assigned as a pseudo label &#x2018;unknown&#x2019; to proposals that have a high objectness score but do not overlap with any ground-truth objects. To generate high-quality proposal boxes, we directly train the detector on the object detection dataset to obtain well-initialized parameters. A well-trained RPN can generate highly accurate proposals or candidate object regions within an image. These proposals effectively filter out cluttered or background regions, enabling the model to focus solely on relevant object proposals. This capability helps in reducing false positives and improving overall detection accuracy. Additionally, a well-trained RPN can effectively handle objects of different sizes and shapes. It learns to generate proposals that encompass objects with varying aspect ratios, ensuring comprehensive coverage of the object space. This enables the model to effectively handle novel or unseen objects, thereby enhancing its performance and robustness in open-world scenarios. We further compare the performance of different pre-trained datasets in open-world plant disease detection.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Class-agnostic ROI head</title>
<p>Locating unknown diseases is an important issue in open-world detection tasks. Standard detectors are primarily designed for localizing objects of known classes, as they employ class-specific localization methods. For instance, detectors like Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>) and Mask RCNN (<xref ref-type="bibr" rid="B17">He et&#xa0;al., 2017</xref>) generate class-specific bounding boxes for each known class when the proposals enter their prediction heads.</p>
<p>To address the localization of novel objects, we introduce a class-agnostic Region of Interest (ROI) head in our object detection models. The class-agnostic ROI head treats region-based feature extraction and classification tasks independently of specific object classes. Unlike class-specific ROI heads that are designed to predict object classes for each region, the class-agnostic ROI head focuses solely on generating accurate bounding box regression outputs without considering the object categories. This makes it well-suited for open-world object detection scenarios where unknown or novel classes may appear.</p>
<p>Inspired by the learned objectness (<xref ref-type="bibr" rid="B20">Kuo et&#xa0;al., 2023</xref>), we utilize class-agnostic box regression heads instead. We have observed that class-agnostic ROI heads exhibit better generalization to unseen classes during inference. They are not biased towards specific object categories, allowing the model to adapt to new classes without the need for retraining or fine-tuning. Additionally, by removing the class-specific classification branch, the overall architecture becomes simpler and more streamlined. This modification not only reduces the computational complexity and memory requirements of the model but also enables more efficient handling of unknown classes.</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Alleviating forgetting</title>
<p>Catastrophic forgetting (<xref ref-type="bibr" rid="B16">Hayes et&#xa0;al., 2020</xref>) refers to the phenomenon observed in incremental learning, where a model trained on new data gradually loses or forgets the knowledge acquired from previously learned tasks or classes. This occurs when the new data heavily influences the model&#x2019;s parameters, leading to the overwriting or disrupting of previously learned information. To address catastrophic forgetting, several techniques have been proposed, such as parameter isolation (<xref ref-type="bibr" rid="B36">Prabhu et&#xa0;al., 2020</xref>), regularization (<xref ref-type="bibr" rid="B23">Li and Hoiem, 2017</xref>), and sample replay (<xref ref-type="bibr" rid="B41">Rebuffi et&#xa0;al., 2017</xref>). These techniques reinforce the model&#x2019;s memory of previous tasks or classes by incorporating previously observed samples during training. In this way, the model can maintain its performance on old tasks while learning new ones.</p>
<p>Sample replay is relatively straightforward compared to other techniques like parameter isolation or complex regularization strategies. It periodically included old samples in the training dataset, making integrating them into existing training pipelines easy. The simplest form of sample replay is randomly retaining training samples. This paper follows the sample replay strategy proposed by <xref ref-type="bibr" rid="B19">Joseph et&#xa0;al. (2021)</xref>, which is the simplest way of sample replay. After each incremental step, a balanced set of samples is stored randomly, and the model is fine-tuned. To ensure an adequate representation of each class, we guarantee a minimum of <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> instances for each class in the sample set. Generally, a larger <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> tends to result in better fine-tuning performance (an extreme case being the use of the entire dataset). However, this contradicts the original intention of dynamic learning in an open-world setting. To ensure a fair comparison among the models, we set <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mn>25</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for fine-tuning the model.</p>
</sec>
<sec id="s3_8">
<label>3.8</label>
<title>Evaluation metrics</title>
<p>We present a comprehensive evaluation protocol to assess the performance of an open-world detector in various aspects: identifying unknown classes, detecting known classes, and progressively learning new classes when labels are available for some unknown samples.</p>
<sec id="s3_8_1">
<label>3.8.1</label>
<title>Mean average precision score</title>
<p>mAP is the area under the precision-recall curve calculated for all classes. To evaluate the detection performance of known classes, we utilize the standard mean average precision (mAP) metric with an intersection over union (IoU) threshold of 0.5 [mAP@50, consistent with the existing literature (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B15">Gupta et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B51">Wu et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B29">Ma et&#xa0;al., 2023a</xref>; <xref ref-type="bibr" rid="B30">Ma et&#xa0;al., 2023b</xref>; <xref ref-type="bibr" rid="B56">Zohar et&#xa0;al., 2023</xref>)].</p>
<disp-formula>
<label>(1) </label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mtext mathvariant="bold-italic">AP</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn mathvariant="bold">1</mml:mn>
<mml:mrow>
<mml:mn mathvariant="bold">11</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn mathvariant="bold">0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn mathvariant="bold">0.1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mn mathvariant="bold">0.9</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mtext mathvariant="bold-italic">P</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>r</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mtext mathvariant="bold">max</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo>:</mml:mo>
<mml:mover accent="true">
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo>&#x2265;</mml:mo>
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
</mml:mrow>
</mml:munder>
<mml:mtext mathvariant="bold-italic">p</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mtext mathvariant="bold-italic">r</mml:mtext>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the maximum precision for any recall values greater than r, and <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mi>r</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the measured precision at recall <inline-formula>
<mml:math display="inline" id="im19">
<mml:mover accent="true">
<mml:mi>r</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>. Since the problem setting of open-world detectors is different from that of standard detectors, there are three forms of mAP, which are current classes mAP, previous classes mAP, and known classes mAP.</p>
</sec>
<sec id="s3_8_2">
<label>3.8.2</label>
<title>Unknown recall</title>
<p>We employ recall as the main metric for unknown object detection instead of the commonly used mAP. This is because all possible unknown object instances in the dataset are not annotated. Unknown recall is widely used in open-world object detection (<xref ref-type="bibr" rid="B15">Gupta et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B29">Ma et&#xa0;al., 2023a</xref>; <xref ref-type="bibr" rid="B30">Ma et&#xa0;al., 2023b</xref>; <xref ref-type="bibr" rid="B56">Zohar et&#xa0;al., 2023</xref>).</p>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mtext mathvariant="bold-italic">U</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext mathvariant="bold-italic">R</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="bold-italic">TPU</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="bold-italic">AU</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>U</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the true positive of unknown instances, and AU denotes all unknown instances for the current task.</p>
</sec>
<sec id="s3_8_3">
<label>3.8.3</label>
<title>Absolute open-set error</title>
<p>In addition, we employ the Absolute open-set error (A-OSE)  (<xref ref-type="bibr" rid="B32">Miller et&#xa0;al., 2018</xref>) metric to report the number of unknown objects that are misclassified as any of the known classes. This metric implicitly measures how effective the model is in handling unknown objects.</p>
<p>To facilitate readability, we use the abbreviations listed in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> to denote the evaluation metrics. The metrics include Unknown Recall and A-OSE, which assess the performance of the unknown classes, and Mean Average Precision (mAP), which evaluates the model&#x2019;s ability to detect the known classes. By employing these metrics, we can comprehensively evaluate and compare the model&#x2019;s performance across both known and unknown classes, providing a comprehensive assessment of its detection capabilities.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Abbreviation and meaning of the evaluation metrics.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="center">Abbreviation</th>
<th valign="top" align="center">Meaning</th>
<th valign="top" align="center">Others</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center">P</td>
<td valign="top" align="center">mean average precision score of previous task classes</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im21">
<mml:mo>&#x2191;</mml:mo>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="top" align="center">C</td>
<td valign="middle" align="center">mean average precision score of current task classes</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im22">
<mml:mo>&#x2191;</mml:mo>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="top" align="center">K</td>
<td valign="middle" align="center">mean average precision score of all known classes</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im23">
<mml:mo>&#x2191;</mml:mo>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="top" align="center">A</td>
<td valign="middle" align="center">Absolute open-set error</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im24">
<mml:mo>&#x2193;</mml:mo>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="top" align="center">U-R</td>
<td valign="middle" align="center">Unknown-Recall</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im25">
<mml:mo>&#x2191;</mml:mo>
</mml:math>
</inline-formula>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Arrows indicate expected trends. Up means that the larger the value, the better, and vice versa.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s4" sec-type="results">
<label>4</label>
<title>Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Implementation details</title>
<p>In the training task sequence, the model can only access the data from the current task. Known classes are defined as the classes in the current task as well as the previous tasks, while other classes are defined as unknown classes. For each image, the model generates only one unknown instance. We adopted the contrastive clustering loss proposed by ORE (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>) and used stochastic gradient descent to optimize the model, with a batch size set to 4. For each training task, we iterated 18,000 times, and for each fine-tuning task, we iterated 4,000 times. We used ResNeXt101 (<xref ref-type="bibr" rid="B52">Xie et&#xa0;al., 2017</xref>) as the final backbone. The entire training process for the project, conducted on 4 NVIDIA GeForce RTX 3090 GPUs, was completed in less than 12 hours. For more details, please refer to our code.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Overall results</title>
<p>
<xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref> compares our method with Faster RCNN (<xref ref-type="bibr" rid="B42">Ren et&#xa0;al., 2015</xref>) and ORE (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>) using the proposed open-world evaluation protocol. The 1-3 row in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref> showcases the result obtained by the standard Faster-RCNN. Note that we used the ResNet50 backbone on the ImageNet1K dataset as a pretraining backbone. We provide a brief overview of the training approach for Faster RCNN. Row 1: We trained Faster-RCNN using a static closed-set training strategy for a fair comparison. As anticipated, Faster-RCNN trained with the closed-set strategy demonstrated optimal results in closed-set evaluation metrics, because the model retrained with all known datasets for each task. However, the model&#x2019;s focus remains limited to known categories, incapable of identifying unknown targets, which contradicts the open-world setting. This experimental set allows researchers to grasp the upper-performance limits of the model in known-category recognition tasks. Hence, we employ &#x2018;Upper&#x2019; to denote the results of this experiment. Row 2: We trained the standard Faster-RCNN on Task 1, followed by Task 2, Task 3, and Task 4. After completing each task, the model&#x2019;s performance was evaluated through testing. In this scenario, the model was also unable to identify unknown diseases. We observed a significant decline in detection performance for previous classes during subsequent task learning with the standard Faster RCNN, which indicates that new knowledge quickly replaced old knowledge throughout the training process. In contrast, our method can successfully detect unknown classes and continuously learn new categories without the need to train from scratch. Row 3: We employed a sample replay strategy to train Faster-RCNN dynamically. This experimental set allows researchers to understand how much sample replay preserves the model&#x2019;s memory capabilities. We denote the results of this experiment as &#x2018;Faster-RCNN*&#x2019; in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Overall results of our method compared with the baseline approach.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Methods</th>
<th valign="middle" colspan="3" align="center">Task 1</th>
<th valign="middle" colspan="5" align="center">Task 2</th>
<th valign="middle" colspan="5" align="center">Task 3</th>
<th valign="middle" colspan="3" align="center">Task 4</th>
<th valign="middle" rowspan="2" align="center">Parameter</th>
</tr>
<tr>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Upper</td>
<td valign="middle" align="center">63.9</td>
<td valign="middle" align="center">2044</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">69.20</td>
<td valign="middle" align="center">60.54</td>
<td valign="middle" align="center">64.91</td>
<td valign="middle" align="center">1706</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">67.06</td>
<td valign="middle" align="center">48.66</td>
<td valign="middle" align="center">60.93</td>
<td valign="middle" align="center">705</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">63.88</td>
<td valign="middle" align="center">84.40</td>
<td valign="middle" align="center">69.02</td>
<td valign="middle" align="center">33M</td>
</tr>
<tr>
<td valign="middle" align="center">Faster RCNN</td>
<td valign="middle" align="center">63.9</td>
<td valign="middle" align="center">2044</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">6.6</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">24.8</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">3.6</td>
<td valign="middle" align="center">25.5</td>
<td valign="middle" align="center">10.9</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">1.8</td>
<td valign="middle" align="center">54.3</td>
<td valign="middle" align="center">14.9</td>
<td valign="middle" align="center">33M</td>
</tr>
<tr>
<td valign="middle" align="center">Faster RCNN*</td>
<td valign="middle" align="center">63.9</td>
<td valign="middle" align="center">2044</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">62.8</td>
<td valign="middle" align="center">40.3</td>
<td valign="middle" align="center">51.6</td>
<td valign="middle" align="center">2386</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">50.7</td>
<td valign="middle" align="center">32.2</td>
<td valign="middle" align="center">44.5</td>
<td valign="middle" align="center">1226</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">44.7</td>
<td valign="middle" align="center">62.2</td>
<td valign="middle" align="center">49.1</td>
<td valign="middle" align="center">33M</td>
</tr>
<tr>
<td valign="middle" align="center">ORE</td>
<td valign="middle" align="center">63.5</td>
<td valign="middle" align="center">2002</td>
<td valign="middle" align="center">14.2</td>
<td valign="middle" align="center">62.6</td>
<td valign="middle" align="center">39.1</td>
<td valign="middle" align="center">50.9</td>
<td valign="middle" align="center">2303</td>
<td valign="middle" align="center">4.7</td>
<td valign="middle" align="center">48.2</td>
<td valign="middle" align="center">31.0</td>
<td valign="middle" align="center">42.5</td>
<td valign="middle" align="center">1228</td>
<td valign="middle" align="center">5.1</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">62.7</td>
<td valign="middle" align="center">47.9</td>
<td valign="middle" align="center">33M</td>
</tr>
<tr>
<td valign="middle" align="center">Ours (a)</td>
<td valign="middle" align="center">65.4</td>
<td valign="middle" align="center">2124</td>
<td valign="middle" align="center">22.6</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">43.2</td>
<td valign="middle" align="center">53.2</td>
<td valign="middle" align="center">2190</td>
<td valign="middle" align="center">12.4</td>
<td valign="middle" align="center">50.6</td>
<td valign="middle" align="center">29.3</td>
<td valign="middle" align="center">43.5</td>
<td valign="middle" align="center">1572</td>
<td valign="middle" align="center">13.1</td>
<td valign="middle" align="center">44.5</td>
<td valign="middle" align="center">61.5</td>
<td valign="middle" align="center">48.8</td>
<td valign="middle" align="center">41M</td>
</tr>
<tr>
<td valign="middle" align="center">Ours (b)</td>
<td valign="middle" align="center">60.7</td>
<td valign="middle" align="center">1827</td>
<td valign="middle" align="center">24.0</td>
<td valign="middle" align="center">63.1</td>
<td valign="middle" align="center">46.3</td>
<td valign="middle" align="center">54.7</td>
<td valign="middle" align="center">3144</td>
<td valign="middle" align="center">8.9</td>
<td valign="middle" align="center">52.4</td>
<td valign="middle" align="center">28.2</td>
<td valign="middle" align="center">44.4</td>
<td valign="middle" align="center">1638</td>
<td valign="middle" align="center">
<bold>38.3</bold>
</td>
<td valign="middle" align="center">45.3</td>
<td valign="middle" align="center">67.0</td>
<td valign="middle" align="center">50.7</td>
<td valign="middle" align="center">41M</td>
</tr>
<tr>
<td valign="middle" align="center">Ours (c)</td>
<td valign="middle" align="center">62.4</td>
<td valign="middle" align="center">1932</td>
<td valign="middle" align="center">
<bold>25.3</bold>
</td>
<td valign="middle" align="center">62.7</td>
<td valign="middle" align="center">44.9</td>
<td valign="middle" align="center">53.8</td>
<td valign="middle" align="center">1961</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">52.7</td>
<td valign="middle" align="center">29.5</td>
<td valign="middle" align="center">45.0</td>
<td valign="middle" align="center">1211</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">47.1</td>
<td valign="middle" align="center">67.1</td>
<td valign="middle" align="center">52.1</td>
<td valign="middle" align="center">41M</td>
</tr>
<tr>
<td valign="middle" align="center">Ours (d)</td>
<td valign="middle" align="center">65.2</td>
<td valign="middle" align="center">1880</td>
<td valign="middle" align="center">23.8</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">47.0</td>
<td valign="middle" align="center">55.1</td>
<td valign="middle" align="center">1960</td>
<td valign="middle" align="center">
<bold>13.7</bold>
</td>
<td valign="middle" align="center">54.4</td>
<td valign="middle" align="center">31.3</td>
<td valign="middle" align="center">46.7</td>
<td valign="middle" align="center">1418</td>
<td valign="middle" align="center">15.7</td>
<td valign="middle" align="center">46.8</td>
<td valign="middle" align="center">66.9</td>
<td valign="middle" align="center">51.8</td>
<td valign="middle" align="center">41M</td>
</tr>
<tr>
<td valign="middle" align="center">Ours*</td>
<td valign="middle" align="center">
<bold>66.3</bold>
</td>
<td valign="middle" align="center">
<bold>1559</bold>
</td>
<td valign="middle" align="center">22.7</td>
<td valign="middle" align="center">
<bold>64.0</bold>
</td>
<td valign="middle" align="center">
<bold>48.9</bold>
</td>
<td valign="middle" align="center">
<bold>56.5</bold>
</td>
<td valign="middle" align="center">
<bold>1854</bold>
</td>
<td valign="middle" align="center">
<bold>13.4</bold>
</td>
<td valign="middle" align="center">
<bold>54.5</bold>
</td>
<td valign="middle" align="center">
<bold>35.8</bold>
</td>
<td valign="middle" align="center">
<bold>48.3</bold>
</td>
<td valign="middle" align="center">
<bold>1123</bold>
</td>
<td valign="middle" align="center">10.8</td>
<td valign="middle" align="center">
<bold>48.2</bold>
</td>
<td valign="middle" align="center">
<bold>70.0</bold>
</td>
<td valign="middle" align="center">
<bold>53.6</bold>
</td>
<td valign="middle" align="center">104M</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>For the notation of evaluation metrics, please refer to <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>. Bold font represents the optimal results. &#x201c;Upper&#x201d; represents training using the entire dataset, which theoretically serves as the performance upper bound. &#x201c;Faster RCNN*&#x201d; represents training using a dynamic paradigm.</p>
</fn>
<fn>
<p>Ours* is a larger model than ours (d), other settings are the same.</p>
</fn>
<fn>
<p>'-' indicates that the evaluation metric is not applicable to the current experimental setup.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Furthermore, our four variants, labeled as Ours (a), Ours (b), Ours (c), and Ours (d) utilized the ResNet-50-FPN backbone, but were pretrained on different datasets. Specifically, Ours (a) used Imagenet-1k (<xref ref-type="bibr" rid="B2">Deng et&#xa0;al., 2009</xref>), Ours (b) used COCO (<xref ref-type="bibr" rid="B25">Lin et&#xa0;al., 2014</xref>), Ours (c) used Object-365-v2 (<xref ref-type="bibr" rid="B44">Shao et&#xa0;al., 2019</xref>), and Ours (d) used the LVIS (<xref ref-type="bibr" rid="B14">Gupta et&#xa0;al., 2019</xref>) dataset. These experiments demonstrate that our method consistently outperforms the ORE (<xref ref-type="bibr" rid="B19">Joseph et&#xa0;al., 2021</xref>) baseline across all evaluation metrics. Additionally, we explored the ResNeXt101(<xref ref-type="bibr" rid="B52">Xie et&#xa0;al., 2017</xref>)architecture, an extension of ResNet, which introduced cardinality to enhance feature representation, making it potentially more powerful in capturing complex patterns and achieving better performance compared to ResNet101. To further improve the model&#x2019;s performance, we trained the ResNeXt-101-FPN on the LVIS dataset. The final row in the table shows the results of our method using the ResNeXt-101-FPN backbone pre-trained on the LVIS dataset, denoted as Ours* in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. Note that A-OSE scores and unknown recalls cannot be measured for Task 4 because of the absence of unknown ground truths. For a visual comparison with the baseline, we present the detection results for Task 1 in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>. Our model outperformed ORE in terms of known disease detection, demonstrating higher accuracy in <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3A, B</bold>
</xref>. Furthermore, when it comes to unknown diseases, our model excelled in reducing false positives as seen in <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3C, D</bold>
</xref>. Additionally, our model achieved precise localization for unknown diseases, as evident in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3E&#x2013;H</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Visualization results comparison between ORE and our model, both trained on Task 1. We present eight pairs of examples <bold>(A-H)</bold>. Best view in color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g003.tif"/>
</fig>
<p>Furthermore, in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>, we present additional qualitative results, showcasing a batch of images that were tested on our model across three tasks. Case A and Case E highlight the model&#x2019;s ability to remember previously learned classes, accurately classifying and locating diseases learned in Task 1. Cases B, C, and D demonstrate the model&#x2019;s capability to detect unknown diseases and progressively learn them. Although these instances were unknown in Task 1, the model gradually learned them in Task 2 and Task 3. Additionally, we include a set of failed cases where the model started to exhibit confusion in localizing old classes as new knowledge is introduced. These challenges will be addressed in future studies.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Qualitative results of our method on example images from our plant disease dataset. We present six groups of examples <bold>(A-F)</bold> from Task 1 to Task 3. Best view in color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g004.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Ablation experiment</title>
<p>To analyze the individual contributions of each component in our method, we conducted meticulous ablation experiments, and the results are presented in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Ablation results. PTD and CAH denote pre-trained dataset and class-agnostic head, respectively.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Architecture</th>
<th valign="middle" rowspan="2" align="center">PTD</th>
<th valign="middle" rowspan="2" align="center">CAH</th>
<th valign="middle" colspan="3" align="center">Task 1</th>
<th valign="middle" colspan="5" align="center">Task 2</th>
<th valign="middle" colspan="5" align="center">Task 3</th>
<th valign="middle" colspan="3" align="center">Task 4</th>
</tr>
<tr>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">63.5</td>
<td valign="middle" align="center">2002</td>
<td valign="middle" align="center">14.2</td>
<td valign="middle" align="center">62.6</td>
<td valign="middle" align="center">39.1</td>
<td valign="middle" align="center">50.9</td>
<td valign="middle" align="center">2303</td>
<td valign="middle" align="center">4.7</td>
<td valign="middle" align="center">48.2</td>
<td valign="middle" align="center">31.0</td>
<td valign="middle" align="center">42.5</td>
<td valign="middle" align="center">1228</td>
<td valign="middle" align="center">5.1</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">62.7</td>
<td valign="middle" align="center">47.9</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">67.6</td>
<td valign="middle" align="center">2173</td>
<td valign="middle" align="center">14.7</td>
<td valign="middle" align="center">63.6</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">53.3</td>
<td valign="middle" align="center">2176</td>
<td valign="middle" align="center">13.2</td>
<td valign="middle" align="center">51.5</td>
<td valign="middle" align="center">30.0</td>
<td valign="middle" align="center">44.3</td>
<td valign="middle" align="center">1688</td>
<td valign="middle" align="center">7.8</td>
<td valign="middle" align="center">44.8</td>
<td valign="middle" align="center">65.7</td>
<td valign="middle" align="center">50.0</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">65.4</td>
<td valign="middle" align="center">2124</td>
<td valign="middle" align="center">22.6</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">43.2</td>
<td valign="middle" align="center">53.2</td>
<td valign="middle" align="center">2190</td>
<td valign="middle" align="center">12.4</td>
<td valign="middle" align="center">50.6</td>
<td valign="middle" align="center">29.3</td>
<td valign="middle" align="center">43.5</td>
<td valign="middle" align="center">1572</td>
<td valign="middle" align="center">13.1</td>
<td valign="middle" align="center">44.5</td>
<td valign="middle" align="center">61.5</td>
<td valign="middle" align="center">48.8</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">59.0</td>
<td valign="middle" align="center">1906</td>
<td valign="middle" align="center">12.4</td>
<td valign="middle" align="center">60.3</td>
<td valign="middle" align="center">44.9</td>
<td valign="middle" align="center">52.6</td>
<td valign="middle" align="center">2075</td>
<td valign="middle" align="center">10.4</td>
<td valign="middle" align="center">52.9</td>
<td valign="middle" align="center">28.1</td>
<td valign="middle" align="center">44.6</td>
<td valign="middle" align="center">1468</td>
<td valign="middle" align="center">4.1</td>
<td valign="middle" align="center">45.6</td>
<td valign="middle" align="center">67.3</td>
<td valign="middle" align="center">51.0</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">60.7</td>
<td valign="middle" align="center">1827</td>
<td valign="middle" align="center">24.0</td>
<td valign="middle" align="center">63.1</td>
<td valign="middle" align="center">46.3</td>
<td valign="middle" align="center">54.7</td>
<td valign="middle" align="center">3144</td>
<td valign="middle" align="center">8.9</td>
<td valign="middle" align="center">52.4</td>
<td valign="middle" align="center">28.2</td>
<td valign="middle" align="center">44.4</td>
<td valign="middle" align="center">1638</td>
<td valign="middle" align="center">38.3</td>
<td valign="middle" align="center">45.3</td>
<td valign="middle" align="center">67.0</td>
<td valign="middle" align="center">50.7</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">Object365</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">62.0</td>
<td valign="middle" align="center">2066</td>
<td valign="middle" align="center">10.5</td>
<td valign="middle" align="center">62.1</td>
<td valign="middle" align="center">47.1</td>
<td valign="middle" align="center">54.6</td>
<td valign="middle" align="center">1971</td>
<td valign="middle" align="center">9.8</td>
<td valign="middle" align="center">53.2</td>
<td valign="middle" align="center">36.4</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">1300</td>
<td valign="middle" align="center">5.1</td>
<td valign="middle" align="center">47.8</td>
<td valign="middle" align="center">69.1</td>
<td valign="middle" align="center">53.1</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">Object365</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">62.4</td>
<td valign="middle" align="center">1932</td>
<td valign="middle" align="center">25.3</td>
<td valign="middle" align="center">62.7</td>
<td valign="middle" align="center">44.9</td>
<td valign="middle" align="center">53.8</td>
<td valign="middle" align="center">1961</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">52.7</td>
<td valign="middle" align="center">29.5</td>
<td valign="middle" align="center">45.0</td>
<td valign="middle" align="center">1211</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">47.1</td>
<td valign="middle" align="center">67.1</td>
<td valign="middle" align="center">52.1</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">64.6</td>
<td valign="middle" align="center">2036</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">61.9</td>
<td valign="middle" align="center">46.9</td>
<td valign="middle" align="center">54.4</td>
<td valign="middle" align="center">2156</td>
<td valign="middle" align="center">13.5</td>
<td valign="middle" align="center">55.0</td>
<td valign="middle" align="center">32.9</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">1566</td>
<td valign="middle" align="center">9.1</td>
<td valign="middle" align="center">47.1</td>
<td valign="middle" align="center">69.1</td>
<td valign="middle" align="center">52.6</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">65.2</td>
<td valign="middle" align="center">1880</td>
<td valign="middle" align="center">23.8</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">47.0</td>
<td valign="middle" align="center">55.1</td>
<td valign="middle" align="center">1960</td>
<td valign="middle" align="center">13.7</td>
<td valign="middle" align="center">54.4</td>
<td valign="middle" align="center">31.3</td>
<td valign="middle" align="center">46.7</td>
<td valign="middle" align="center">1418</td>
<td valign="middle" align="center">15.7</td>
<td valign="middle" align="center">46.8</td>
<td valign="middle" align="center">66.9</td>
<td valign="middle" align="center">51.8</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>IN1k denotes the Imagenet-1k dataset. For the notation of evaluation metrics, please refer to <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>.</p>
</fn>
<fn>
<p>'&#x221a;' and 'x' respectively indicate model with or without the class-agnostic head.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Backbone</title>
<p>We compared the FPN module with the C4 module on ResNet-50 (Row 1 and Row 2). The inclusion of FPN significantly enhanced the model&#x2019;s learning ability and memory capacity, as evidenced by improved performance in Task 1 (63.55% vs. 67.67%) and Task 3 (48.27% vs. 51.50%). Based on this observation, all subsequent experiments were performed using ResNet with FPN as the backbone network instead of the C4 structure.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Class-agnostic head</title>
<p>Our class-agnostic head played a crucial role in the model&#x2019;s performance. By not assigning specific class labels to detected objects, the class-agnostic head enabled the model to treat all objects as potential unknown classes. This means that if a detected object does not match any known class, it is more likely to be classified as an unknown object rather than misclassified into a known class. Consequently, the class-agnostic head improved the model&#x2019;s ability to recognize and recall unknown objects, thus enhancing overall performance in open-world scenarios. <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref> demonstrates that the class-agnostic head significantly improved the recall of unknown classes across different pretraining data. Moreover, <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref> indicates that the class-agnostic head remains effective even when used with larger networks.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Results of our method using larger model.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Architecture</th>
<th valign="middle" rowspan="2" align="center">PTD</th>
<th valign="middle" rowspan="2" align="center">CAH</th>
<th valign="middle" colspan="3" align="center">Task 1</th>
<th valign="middle" colspan="5" align="center">Task 2</th>
<th valign="middle" colspan="6" align="center">Task 3</th>
<th valign="middle" colspan="3" align="center">Task 4</th>
</tr>
<tr>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" colspan="2" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">63.5</td>
<td valign="middle" align="center">2002</td>
<td valign="middle" align="center">14.2</td>
<td valign="middle" align="center">62.6</td>
<td valign="middle" align="center">39.1</td>
<td valign="middle" align="center">50.9</td>
<td valign="middle" align="center">2303</td>
<td valign="middle" align="center">4.7</td>
<td valign="middle" align="center">48.2</td>
<td valign="middle" align="center">31.0</td>
<td valign="middle" align="center">42.5</td>
<td valign="middle" align="center">1228</td>
<td valign="middle" align="center">5.1</td>
<td valign="middle" colspan="2" align="center">42.9</td>
<td valign="middle" align="center">62.7</td>
<td valign="middle" align="center">47.9</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">67.6</td>
<td valign="middle" align="center">2173</td>
<td valign="middle" align="center">14.7</td>
<td valign="middle" align="center">63.6</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">53.3</td>
<td valign="middle" align="center">2176</td>
<td valign="middle" align="center">13.2</td>
<td valign="middle" align="center">51.5</td>
<td valign="middle" align="center">30.0</td>
<td valign="middle" align="center">44.3</td>
<td valign="middle" align="center">1688</td>
<td valign="middle" align="center">7.8</td>
<td valign="middle" colspan="2" align="center">44.8</td>
<td valign="middle" align="center">65.7</td>
<td valign="middle" align="center">50.0</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">65.2</td>
<td valign="middle" align="center">1880</td>
<td valign="middle" align="center">23.8</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">47.0</td>
<td valign="middle" align="center">55.1</td>
<td valign="middle" align="center">1960</td>
<td valign="middle" align="center">13.7</td>
<td valign="middle" align="center">54.4</td>
<td valign="middle" align="center">31.3</td>
<td valign="middle" align="center">46.7</td>
<td valign="middle" align="center">1418</td>
<td valign="middle" align="center">15.7</td>
<td valign="middle" colspan="2" align="center">46.8</td>
<td valign="middle" align="center">66.9</td>
<td valign="middle" align="center">51.8</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">64.8</td>
<td valign="middle" align="center">1815</td>
<td valign="middle" align="center">24.0</td>
<td valign="middle" align="center">65.4</td>
<td valign="middle" align="center">44.1</td>
<td valign="middle" align="center">54.7</td>
<td valign="middle" align="center">2448</td>
<td valign="middle" align="center">9.3</td>
<td valign="middle" align="center">54.2</td>
<td valign="middle" align="center">34.5</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">1418</td>
<td valign="middle" align="center">7.6</td>
<td valign="middle" colspan="2" align="center">48.3</td>
<td valign="middle" align="center">71.8</td>
<td valign="middle" align="center">54.1</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">&#x221a;</td>
<td valign="middle" align="center">66.3</td>
<td valign="middle" align="center">1559</td>
<td valign="middle" align="center">22.7</td>
<td valign="middle" align="center">64.0</td>
<td valign="middle" align="center">48.9</td>
<td valign="middle" align="center">56.5</td>
<td valign="middle" align="center">1854</td>
<td valign="middle" align="center">13.4</td>
<td valign="middle" align="center">54.5</td>
<td valign="middle" align="center">35.8</td>
<td valign="middle" align="center">48.3</td>
<td valign="middle" align="center">1123</td>
<td valign="middle" align="center">10.8</td>
<td valign="middle" colspan="2" align="center">48.2</td>
<td valign="middle" align="center">70.0</td>
<td valign="middle" align="center">53.6</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>For a fair comparison, we show the results of the baseline (R50-C4) and Resnet-50 with FPN (R50-FPN).</p>
</fn>
<fn>
<p>'&#x221a;' and 'x' respectively indicate model with or without the class-agnostic head.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Pretraining datasets</title>
<p>In order to investigate the influence of different pre-training datasets on our model, we conducted a series of experiments as outlined in <xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>. Our findings reveal that the model trained on the Imagenet-1k dataset exhibited better performance on the initial tasks. However, as the tasks progressed, this advantage gradually diminished. On the other hand, the model trained on the LVIS dataset showed an advantage in terms of unknown recall, with no significant drop in performance (mAP) for known class detection. Similarly, the model trained on the COCO dataset exhibited a similar trend, albeit with slightly lower performance.</p>
<table-wrap id="T6" position="float">
<label>Table&#xa0;6</label>
<caption>
<p>Results of different pre-training data in the open-world disease detection tasks.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" colspan="19" align="center">(A)</th>
</tr>
<tr>
<th valign="middle" rowspan="2" align="center">Architecture</th>
<th valign="middle" rowspan="2" align="center">PTD</th>
<th valign="middle" rowspan="2" align="center">CAH</th>
<th valign="middle" colspan="3" align="center">Task 1</th>
<th valign="middle" colspan="5" align="center">Task 2</th>
<th valign="middle" colspan="5" align="center">Task 3</th>
<th valign="middle" colspan="3" align="center">Task 4</th>
</tr>
<tr>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">A</th>
<th valign="middle" align="center">U-R</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">K</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">
<bold>67.6</bold>
</td>
<td valign="middle" align="center">2173</td>
<td valign="middle" align="center">
<bold>14.7</bold>
</td>
<td valign="middle" align="center">
<bold>63.6</bold>
</td>
<td valign="middle" align="center">42.9</td>
<td valign="middle" align="center">53.3</td>
<td valign="middle" align="center">2176</td>
<td valign="middle" align="center">13.2</td>
<td valign="middle" align="center">51.5</td>
<td valign="middle" align="center">30.0</td>
<td valign="middle" align="center">44.3</td>
<td valign="middle" align="center">1688</td>
<td valign="middle" align="center">7.8</td>
<td valign="middle" align="center">44.8</td>
<td valign="middle" align="center">65.7</td>
<td valign="middle" align="center">50.0</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">60.9</td>
<td valign="middle" align="center">
<bold>1950</bold>
</td>
<td valign="middle" align="center">11.8</td>
<td valign="middle" align="center">59.4</td>
<td valign="middle" align="center">44.0</td>
<td valign="middle" align="center">51.7</td>
<td valign="middle" align="center">2238</td>
<td valign="middle" align="center">9.4</td>
<td valign="middle" align="center">51.5</td>
<td valign="middle" align="center">27.6</td>
<td valign="middle" align="center">43.5</td>
<td valign="middle" align="center">1731</td>
<td valign="middle" align="center">1.3</td>
<td valign="middle" align="center">44.9</td>
<td valign="middle" align="center">66.3</td>
<td valign="middle" align="center">50.2</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">Object365</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">62.0</td>
<td valign="middle" align="center">2066</td>
<td valign="middle" align="center">10.5</td>
<td valign="middle" align="center">
<bold>62.1</bold>
</td>
<td valign="middle" align="center">
<bold>47.1</bold>
</td>
<td valign="middle" align="center">
<bold>54.6</bold>
</td>
<td valign="middle" align="center">
<bold>1971</bold>
</td>
<td valign="middle" align="center">9.8</td>
<td valign="middle" align="center">53.2</td>
<td valign="middle" align="center">
<bold>36.4</bold>
</td>
<td valign="middle" align="center">
<bold>47.6</bold>
</td>
<td valign="middle" align="center">
<bold>1300</bold>
</td>
<td valign="middle" align="center">5.1</td>
<td valign="middle" align="center">
<bold>47.8</bold>
</td>
<td valign="middle" align="center">
<bold>69.1</bold>
</td>
<td valign="middle" align="center">
<bold>53.1</bold>
</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">64.6</td>
<td valign="middle" align="center">2036</td>
<td valign="middle" align="center">13.0</td>
<td valign="middle" align="center">61.9</td>
<td valign="middle" align="center">46.9</td>
<td valign="middle" align="center">54.4</td>
<td valign="middle" align="center">2156</td>
<td valign="middle" align="center">13.5</td>
<td valign="middle" align="center">
<bold>55.0</bold>
</td>
<td valign="middle" align="center">32.9</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">1566</td>
<td valign="middle" align="center">
<bold>9.1</bold>
</td>
<td valign="middle" align="center">47.1</td>
<td valign="middle" align="center">69.1</td>
<td valign="middle" align="center">52.6</td>
</tr>
<tr>
<th valign="middle" colspan="19" align="center">(B)</th>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">
<bold>66.7</bold>
</td>
<td valign="middle" align="center">2042</td>
<td valign="middle" align="center">
<bold>14.9</bold>
</td>
<td valign="middle" align="center">63.2</td>
<td valign="middle" align="center">44.0</td>
<td valign="middle" align="center">53.6</td>
<td valign="middle" align="center">
<bold>2038</bold>
</td>
<td valign="middle" align="center">
<bold>13.2</bold>
</td>
<td valign="middle" align="center">52.1</td>
<td valign="middle" align="center">31.0</td>
<td valign="middle" align="center">45.0</td>
<td valign="middle" align="center">1538</td>
<td valign="middle" align="center">3.7</td>
<td valign="middle" align="center">44.9</td>
<td valign="middle" align="center">62.4</td>
<td valign="middle" align="center">49.2</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">61.5</td>
<td valign="middle" align="center">
<bold>1842</bold>
</td>
<td valign="middle" align="center">10.5</td>
<td valign="middle" align="center">61.9</td>
<td valign="middle" align="center">45.1</td>
<td valign="middle" align="center">53.5</td>
<td valign="middle" align="center">2367</td>
<td valign="middle" align="center">6.7</td>
<td valign="middle" align="center">52.1</td>
<td valign="middle" align="center">28.4</td>
<td valign="middle" align="center">44.2</td>
<td valign="middle" align="center">1538</td>
<td valign="middle" align="center">2.7</td>
<td valign="middle" align="center">45.5</td>
<td valign="middle" align="center">69.3</td>
<td valign="middle" align="center">51.4</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">Object365</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">62.6</td>
<td valign="middle" align="center">2077</td>
<td valign="middle" align="center">11.6</td>
<td valign="middle" align="center">62.5</td>
<td valign="middle" align="center">44.1</td>
<td valign="middle" align="center">53.3</td>
<td valign="middle" align="center">2035</td>
<td valign="middle" align="center">10.0</td>
<td valign="middle" align="center">53.7</td>
<td valign="middle" align="center">31.8</td>
<td valign="middle" align="center">46.4</td>
<td valign="middle" align="center">1546</td>
<td valign="middle" align="center">
<bold>8.2</bold>
</td>
<td valign="middle" align="center">47.7</td>
<td valign="middle" align="center">65.4</td>
<td valign="middle" align="center">52.1</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">64.1</td>
<td valign="middle" align="center">1910</td>
<td valign="middle" align="center">12.7</td>
<td valign="middle" align="center">
<bold>63.8</bold>
</td>
<td valign="middle" align="center">
<bold>45.9</bold>
</td>
<td valign="middle" align="center">
<bold>54.8</bold>
</td>
<td valign="middle" align="center">2338</td>
<td valign="middle" align="center">5.5</td>
<td valign="middle" align="center">
<bold>53.7</bold>
</td>
<td valign="middle" align="center">
<bold>32.2</bold>
</td>
<td valign="middle" align="center">
<bold>46.6</bold>
</td>
<td valign="middle" align="center">
<bold>1583</bold>
</td>
<td valign="middle" align="center">4.8</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">
<bold>72.1</bold>
</td>
<td valign="middle" align="center">
<bold>53.8</bold>
</td>
</tr>
<tr>
<th valign="middle" colspan="19" align="center">(C)</th>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">IN1k</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">
<bold>69.0</bold>
</td>
<td valign="middle" align="center">1832</td>
<td valign="middle" align="center">
<bold>18.1</bold>
</td>
<td valign="middle" align="center">
<bold>65.0</bold>
</td>
<td valign="middle" align="center">46.4</td>
<td valign="middle" align="center">55.7</td>
<td valign="middle" align="center">1894</td>
<td valign="middle" align="center">8.7</td>
<td valign="middle" align="center">52.8</td>
<td valign="middle" align="center">37.3</td>
<td valign="middle" align="center">47.6</td>
<td valign="middle" align="center">1231</td>
<td valign="middle" align="center">1.1</td>
<td valign="middle" align="center">45.9</td>
<td valign="middle" align="center">66.8</td>
<td valign="middle" align="center">51.1</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">58.0</td>
<td valign="middle" align="center">
<bold>1594</bold>
</td>
<td valign="middle" align="center">11.5</td>
<td valign="middle" align="center">60.6</td>
<td valign="middle" align="center">47.4</td>
<td valign="middle" align="center">54.0</td>
<td valign="middle" align="center">2029</td>
<td valign="middle" align="center">4.1</td>
<td valign="middle" align="center">52.5</td>
<td valign="middle" align="center">33.5</td>
<td valign="middle" align="center">46.2</td>
<td valign="middle" align="center">1542</td>
<td valign="middle" align="center">6.3</td>
<td valign="middle" align="center">45.7</td>
<td valign="middle" align="center">66.7</td>
<td valign="middle" align="center">51.0</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">Object365</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">61.0</td>
<td valign="middle" align="center">1774</td>
<td valign="middle" align="center">15.9</td>
<td valign="middle" align="center">63.6</td>
<td valign="middle" align="center">47.3</td>
<td valign="middle" align="center">55.5</td>
<td valign="middle" align="center">
<bold>1864</bold>
</td>
<td valign="middle" align="center">
<bold>9.6</bold>
</td>
<td valign="middle" align="center">51.8</td>
<td valign="middle" align="center">33.6</td>
<td valign="middle" align="center">45.7</td>
<td valign="middle" align="center">
<bold>1140</bold>
</td>
<td valign="middle" align="center">
<bold>7.8</bold>
</td>
<td valign="middle" align="center">46.5</td>
<td valign="middle" align="center">69.1</td>
<td valign="middle" align="center">52.2</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">x</td>
<td valign="middle" align="center">63.5</td>
<td valign="middle" align="center">1668</td>
<td valign="middle" align="center">12.7</td>
<td valign="middle" align="center">62.9</td>
<td valign="middle" align="center">
<bold>48.4</bold>
</td>
<td valign="middle" align="center">55.6</td>
<td valign="middle" align="center">1907</td>
<td valign="middle" align="center">5.5</td>
<td valign="middle" align="center">
<bold>54.4</bold>
</td>
<td valign="middle" align="center">
<bold>36.1</bold>
</td>
<td valign="middle" align="center">
<bold>48.3</bold>
</td>
<td valign="middle" align="center">1398</td>
<td valign="middle" align="center">4.5</td>
<td valign="middle" align="center">
<bold>48.8</bold>
</td>
<td valign="middle" align="center">
<bold>71.8</bold>
</td>
<td valign="middle" align="center">
<bold>54.6</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>We compared three different network models: ResNet50 with FPN, ResNet101 with FPN, and ResNetX101 with FPN. Bold font indicates the best results among each comparison group.</p>
</fn>
<fn>
<p>'x' denotes the model without the class-agnostic head.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>We attribute the benefits brought by the LVIS-based pre-trained models to two main factors. Firstly, the consistency of pre-training objectives played a significant role. The LVIS-based pre-trained models utilize training objectives that align closely with the target detection task, encompassing multi-label classification and bounding box regression. In contrast to ImageNet pre-trained models, these objectives are better suited for the target detection task, resulting in improved performance. Secondly, the richness of the data is a contributing factor. LVIS encompasses over 1,200 categories, whereas COCO only includes 80 categories. The significantly larger number of categories in LVIS provided a more diverse and comprehensive representation of objects across various domains. Consequently, this allowed the LVIS-based pretrained models to learn more comprehensive features and contextual information for different categories. Based on these observations, we argue that the pre-training model based on LVIS exhibited greater potential for subsequent tasks due to the alignment of training objectives and the broader representation of object categories.</p>
<p>Furthermore, we also trained and released these three models on the Object365 dataset using the Detectron2 framework. The Object365-v2 dataset (<xref ref-type="bibr" rid="B44">Shao et&#xa0;al., 2019</xref>) contains nearly 2 million images with over 10 million annotated bounding boxes. In terms of scale, Object365-v2 contains a greater number of instances compared to LVIS. However, we observed that pre-trained on the Object365-v2 dataset significantly boosts the performance of the COCO dataset in open-world evaluation settings, but its performance on plant disease datasets is slightly lower than the model pre-trained on LVIS dataset. As a result, we opted for the LVIS-based pre-trained model as the final choice for our work. Please note that fine-tuning the COCO dataset results using the Object365-v2 dataset is not the focus of this paper. We presented these results in our code repository.</p>
<p>Additionally, we performed experiments using larger models to enhance the performance of our model, as presented in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>. It was challenging to improve all performance metrics across all tasks simultaneously. However, in general, employing larger models, leveraging well-pretrained Region Proposal Networks (RPNs), and incorporating class-agnostic heads tended to yield better results.</p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Sensitivity analysis on training order</title>
<p>In the context of incremental learning tasks, the order in which tasks are presented to the model can significantly impact its performance and the overall learning process. The learning sequence plays a crucial role in addressing challenges such as knowledge forgetting, conflicting information, and fluctuations in performance. Recognizing the importance of the learning order, we conducted an investigation to understand the model&#x2019;s sensitivity to different training sequences.</p>
<p>By analyzing the results in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>, we observed that Faster RCNN achieved the highest performance on Task 1 and the lowest on Task 3 when detecting tomato diseases. This observation led us to infer that Task 1 might be relatively simpler, while Task 3 could pose more challenges in disease detection.</p>
<p>Following the principle of human learning from easy to difficult, we believe that the model should start learning from simple tasks. Therefore, in previous experimental settings, the default learning order was from Task 1 to Task 3. After learning the diseases of one species, the model continued to learn the cross-species detection task (Task 4).</p>
<p>However, to explore the sensitivity of our research model to the training order, we decided to deviate from the default sequence and adopt a different approach. We opted to initiate the learning process with the more difficult Task 3. By doing so, we aimed to observe how the model adapts and performs when confronted with the most challenging task from the start. Therefore, in this study, we rearranged the task sequence as follows: Task 3, Task 2, Task 1, and finally, Task 4.</p>
<p>This alternative task sequence enabled us to investigate the model&#x2019;s ability to learn and transfer knowledge in a non-conventional order, offering insights into its adaptability and potential for early tackling of more complex tasks.</p>
<p>The detection results of the model on the tomato disease dataset and the paprika disease dataset under different training orders are presented in <xref ref-type="table" rid="T7">
<bold>Table&#xa0;7</bold>
</xref>. Due to different task sequences, we can only compare the model&#x2019;s performance in detecting known classes after learning 15 tomato diseases. We also compared the model&#x2019;s memory ability to capture disease patterns and learning ability in cross-species detection tasks such as paprika. The memory ability is reflected in the mAP of previous classes (P), while the learning ability is reflected in the mAP of current classes (C). As expected, learning from more challenging task orders led to a slight performance degradation in the model for all aspects, even though the impact is not significant. This finding serves as a reminder to practitioners that the learning order of models should follow a progression from simpler to more difficult tasks in order to achieve optimal performance.</p>
<table-wrap id="T7" position="float">
<label>Table&#xa0;7</label>
<caption>
<p>Sensitivity analysis on the task training order.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Architecture</th>
<th valign="middle" rowspan="2" align="center">PTD</th>
<th valign="middle" rowspan="2" align="center">Order</th>
<th valign="middle" colspan="2" align="center">Task 1~3<break/>(Tomato)</th>
<th valign="middle" colspan="6" align="center">Task 4 (cross-species task)<break/>(Paprika)</th>
</tr>
<tr>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">&#x394;</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">&#x394;</th>
<th valign="middle" align="center">C</th>
<th valign="middle" align="center">&#x394;</th>
<th valign="middle" align="center">K</th>
<th valign="middle" align="center">&#x394;</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="top" align="center">IN1k</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">42.52</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">42.96</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">62.78</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">47.92</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="top" align="center">IN1k</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">41.37</td>
<td valign="middle" align="center">-1.15</td>
<td valign="middle" align="center">40.8</td>
<td valign="middle" align="center">-2.16</td>
<td valign="middle" align="center">63.03</td>
<td valign="middle" align="center">0.25</td>
<td valign="middle" align="center">46.4</td>
<td valign="middle" align="center">-1.52</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="top" align="center">IN1k</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">44.35</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">44.80</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">65.70</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">50.03</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="top" align="center">IN1k</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">42.38</td>
<td valign="middle" align="center">-1.97</td>
<td valign="middle" align="center">42.37</td>
<td valign="middle" align="center">-2.43</td>
<td valign="middle" align="center">58.64</td>
<td valign="middle" align="center">-7.06</td>
<td valign="middle" align="center">46.44</td>
<td valign="middle" align="center">-3.59</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">44.24</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">45.90</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">64.24</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">50.48</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-C4</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">41.46</td>
<td valign="middle" align="center">-2.78</td>
<td valign="middle" align="center">41.33</td>
<td valign="middle" align="center">-4.57</td>
<td valign="middle" align="center">62.16</td>
<td valign="middle" align="center">-2.08</td>
<td valign="middle" align="center">46.54</td>
<td valign="middle" align="center">-3.94</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">43.58</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">44.92</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">66.31</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">50.27</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">COCO</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">42.04</td>
<td valign="middle" align="center">-1.54</td>
<td valign="middle" align="center">41.84</td>
<td valign="middle" align="center">-3.08</td>
<td valign="middle" align="center">65.97</td>
<td valign="middle" align="center">-0.34</td>
<td valign="middle" align="center">47.87</td>
<td valign="middle" align="center">-2.40</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">47.69</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">47.19</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">69.12</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">52.67</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-50-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">43.17</td>
<td valign="middle" align="center">-4.52</td>
<td valign="middle" align="center">43.9</td>
<td valign="middle" align="center">-3.29</td>
<td valign="middle" align="center">68.23</td>
<td valign="middle" align="center">-0.89</td>
<td valign="middle" align="center">49.98</td>
<td valign="middle" align="center">-2.69</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">46.63</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">47.68</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">72.13</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">53.80</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">R-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">44.48</td>
<td valign="middle" align="center">-2.15</td>
<td valign="middle" align="center">44.37</td>
<td valign="middle" align="center">-3.31</td>
<td valign="middle" align="center">68.17</td>
<td valign="middle" align="center">-3.96</td>
<td valign="middle" align="center">50.32</td>
<td valign="middle" align="center">-3.48</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 1, Task 2, Task 3, Task 4</td>
<td valign="middle" align="center">48.31</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">48.87</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">71.86</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">54.61</td>
<td valign="middle" align="center">&#x2013;</td>
</tr>
<tr>
<td valign="middle" align="center">X-101-FPN</td>
<td valign="middle" align="center">LVIS</td>
<td valign="middle" align="center">Task 3, Task 2, Task 1, Task 4</td>
<td valign="middle" align="center">47.26</td>
<td valign="middle" align="center">-1.05</td>
<td valign="middle" align="center">46.56</td>
<td valign="middle" align="center">-2.31</td>
<td valign="middle" align="center">71.57</td>
<td valign="middle" align="center">-0.29</td>
<td valign="middle" align="center">52.81</td>
<td valign="middle" align="center">-1.80</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>We alternately show the performance of different training orders in the same model. &#x394; represents the differences caused by the order of learning. For the notation of evaluation metrics, please refer to <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>.</p>
</fn>
<fn>
<p>'-' indicates that the difference is not applicable to the current situation.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Cross-species detection</title>
<p>Our research has uncovered a fascinating discovery regarding the capabilities of our model, particularly in the context of cross-species disease detection. Despite being trained solely on a dataset of tomato diseases, our model exhibited the remarkable ability to identify and provide an initial assessment of affected regions in paprika fruit diseases. This intriguing finding is illustrated through two specific cases showcased in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5A</bold>
</xref>, namely Case 1 and Case 2.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Qualitative results on cross-species detection study. <bold>(A)</bold>. Training on tomato dataset and test on paprika dataset. <bold>(B)</bold>. Training on paprika dataset and test on tomato dataset. The sample number is indicated in the top left corner of each subplot. Best view in color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g005.tif"/>
</fig>
<p>Please note that our model has never been exposed to or trained on any tomato fruit diseases, only leaves, making its performance in detecting paprika fruit diseases all the more intriguing. The fact that the model can generalize its knowledge and effectively apply it to a different species highlights its versatility and potential for cross-species disease detection, which also demonstrates that our method learns the fundamental features of disease.</p>
<p>Furthermore, we conducted a similar experiment in which we trained a separate model using a paprika disease dataset and evaluated its performance on a test dataset consisting of tomato plants. The results were equally compelling. Our paprika-trained model successfully detected pests present on tomato leaves, as demonstrated by Case 8 in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5B</bold>
</xref>. This further reinforces the model&#x2019;s ability to transfer its learned knowledge across species boundaries and adapt it to different contexts.</p>
<p>To provide a comprehensive visualization of the model&#x2019;s cross-species detection capabilities, <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref> presents qualitative results of these experiments. These visual examples offer a glimpse into the model&#x2019;s ability to identify diseases and pests in species it has not been explicitly trained on, demonstrating its potential for broader applicability and practical use in real-world scenarios.</p>
</sec>
</sec>
<sec id="s5" sec-type="discussion">
<label>5</label>
<title>Discussion</title>
<p>The task of object detection is typically divided into two subtasks: classification and localization. In this section, we discuss the limitations of our method in these two subtasks, including open issues. Finally, we present several potential avenues for future research.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Localization</title>
<p>Addressing the localization problem of unknown objects is a key challenge in open-world object detection. The main difficulty lies in the lack of prior knowledge about the unknown classes in the model. As a result, it is challenging to directly learn their features and location information from the training data. Our method improved the model&#x2019;s ability to detect unknown diseases. However, qualitative experimental results showed that the unknown recall is still below 30%. A unified unknown detection evaluation protocol is even more difficult than finding unknown diseases. As shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, these unknown detection results are treated as false positive boxes under the current ground truth, even though our model has already localized these suspicious regions.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Qualitative results for unknown instances from our dataset. We present four pairs of examples <bold>(A&#x2013;D)</bold>. The first row displays the image and annotations, while the second row represents our detection results. Best view in color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g006.tif"/>
</fig>
<p>The controversy surrounding the evaluation criteria for unknown class localization stems from the lack of consistent standards and consensus. This controversy is formed when annotating datasets. Our previous work (<xref ref-type="bibr" rid="B4">Dong et&#xa0;al., 2022</xref>) discussed how to efficiently label plant diseases. We believe that different disease symptoms should adopt different labeling strategies, and we verified this scheme&#x2019;s effectiveness through several experiments. However, these annotation strategies and evaluation schemes were designed for known categories. To the best of our knowledge, no related work discusses the localization of unknown classes in plant disease detection tasks. Additionally, the definition and scope of unknown classes also introduce subjectivity and uncertainty, further contributing to the controversy of evaluation criteria. Therefore, further research and consensus-building are needed to establish consistent and fair evaluation criteria for assessing the localization performance of unknown classes.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Classification</title>
<p>We developed a dynamic open-world detector since plant growth is a dynamic process. However, plant disease dynamics are more complex than we imagined. Some diseases may exhibit different symptoms at different stages of growth, leading to a challenging feature expression. Additionally, different diseases may also exhibit similar symptoms at different stages, which can be due to different pathogens (such as bacteria, fungi, viruses, etc.) or environmental factors. We list some common examples of tomato diseases with similar symptoms:</p>
<p>Yellowing symptoms: Yellowing is a common symptom of many plant diseases, including viral infections, fungal diseases, and nutrient deficiencies. Different pathogens or causes may lead to yellowing of plant leaves or other tissues, but their pathological processes and treatment methods may differ completely.</p>
<p>Leaf spot diseases: Many pathogens can cause similar leaf spot diseases, such as fungal and bacterial leaf spots. They produce similar spots or patches on the leaves, but the pathogens and pathogenic mechanisms behind them are different.</p>
<p>Rotting symptoms: Rotting is a common symptom caused by various diseases or pathogens, including bacterial soft rot, fungal rot, and rotting caused by certain environmental factors. Although they manifest as the decay of plant tissues, the specific causes may be different.</p>
<p>Similar symptoms may also occur in the detection of cross-species diseases. In addition, the leaves of different plants are different in a healthy state. However, diseases may force the leaves of different species to deform to the same symptom at the final state. In this case, even experts also struggle to distinguish them. Therefore, deep learning models may still face the same challenges in accurately differentiating them. <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref> shows some cases with similar symptoms but different species. When testing for diseases on paprika leaves using a model trained on the tomato dataset, all suspicious regions should have been detected as unknown. However, some unknown regions are mistakenly detected as gray mold due to similar symptoms. Although the category is correct, these instances of gray mold are treated as false positives.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>Qualitative results for unknown instances from our dataset. These instances of gray mold should have been detected as unknown. The sample number is indicated in the top left corner of each subplot. Best view in color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1243822-g007.tif"/>
</fig>
<p>These pieces of evidence prove that deep learning models can offer advantages in distinguishing similar disease symptoms but are not infallible. Domain expertise and collaboration with experts remain critical in evaluating and validating the model&#x2019;s predictions. The model&#x2019;s success still depends on the availability of quality training data and the complexity of the differentiation task. Another limitation we encountered is the challenge of obtaining additional high-quality datasets for plant disease detection to validate generalizability further. Despite this constraint, we have conducted validation using the COCO dataset to showcase the method&#x2019;s performance. For more details, please refer to our code repository.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Future works</title>
<p>We present some promising approaches to tackle classification problems. Recently, <xref ref-type="bibr" rid="B5">Du et&#xa0;al. (2022a)</xref> introduced spatial-temporal unknown distillation (STUD), a model designed to detect unknown objects in videos by establishing spatial-temporal context. STUD (<xref ref-type="bibr" rid="B5">Du et&#xa0;al., 2022a</xref>) utilizes time series features to evaluate the relationship between the current frame and the reference frame, reducing the occurrence of classification errors. In real-world agricultural practices, the same species is commonly planted in one area. Therefore, considering the spatial-temporal context to determine the species category can effectively narrow down the range of disease classifications. Another intriguing direction is utilizing large visual language models (<xref ref-type="bibr" rid="B40">Radford et&#xa0;al., 2021</xref>), renowned for their impressive zero-shot detection capabilities, making them highly suitable for identifying unknown categories. A recent study (<xref ref-type="bibr" rid="B49">Wortsman et&#xa0;al., 2022</xref>) demonstrated that fine-tuning a large-scale visual language model through weight integration performs well not only on specific downstream tasks but also maintains its ability to recognize unknown targets. Consequently, embedding a large language-vision model into open-world detection tasks has the potential to enhance the model&#x2019;s robustness. We encourage the community to pay attention to these promising methods and apply them to plant disease detection tasks.</p>
</sec>
</sec>
<sec id="s6" sec-type="conclusions">
<label>6</label>
<title>Conclusions</title>
<p>In this study, we introduced a new paradigm called open world plant disease detector. This novel detection paradigm enables the detection of unknown diseases and allows for the dynamic updating of new knowledge. This paradigm breaks the closed-set, static open-set settings of conventional plant disease detectors. We observed that detectors trained on complex object detection datasets can enhance the detection performance for unknown classes, and the category-agnostic head further improved the recall rate for unknown diseases. Additionally, cross-species disease detection experiments have demonstrated that our model can comprehend the concept of diseases and successfully detect them across different species. Extensive ablation experiments validated the effectiveness of our proposed method. Furthermore, we thoroughly discussed the existing open challenges in plant disease detection and offered insightful perspectives. We strongly encourage researchers and practitioners to address the current challenges that remain.</p>
</sec>
<sec id="s7" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary files. Further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s8" sec-type="author-contributions">
<title>Author contributions</title>
<p>JD designed the method, performed the experiments, and wrote the manuscript. KH and SY advised in the design of the system and analysed the annotation strategies to find the best method for efficient plant disease detection. AF and DP provided support in the data collection and proofreading article. JY presents some conceptual suggestions and domain knowledge. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s9" sec-type="funding-information">
<title>Funding</title>
<p>This research was supported by Basic Science Research Program through the National Research Foundation of Korea(NRF) funded by the Ministry of Education (No. 2019R1A6A1A09031717); by the Korea Institute of Planning and Evaluation for Technology in Food, Agriculture and Forestry(IPET) and Korea Smart Farm R&amp;D Foundation (KosFarm) through Smart Farm Innovation Technology Development Program, funded by Ministry of Agriculture, Food and Rural Affairs(MAFRA) and Ministry of Science and ICT(MSIT), Rural Development Administration(RDA)(1545027569); and by the National Research Foundation of Korea(NRF) grant funded by the Korea government (MSIT). (NRF-2021R1A2C1012174).</p>
</sec>
<sec id="s10" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s11" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alruwaili</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Siddiqi</surname> <given-names>M. H.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Azad</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Alanazi</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>RTF-RCNN: An architecture for real-time tomato plant leaf diseases detection in video streaming using Faster-RCNN</article-title>. <source>Bioengineering</source> <volume>9</volume>, <elocation-id>565</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/bioengineering9100565</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Deng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Socher</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>L.-J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fei-Fei</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Imagenet: A large-scale hierarchical image database</article-title>,&#x201d; in <conf-name>2009 IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Miami, FL, USA</conf-loc>. <fpage>248</fpage>&#x2013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dhamija</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Gunther</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ventura</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Boult</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>The overlooked elephant of object detection: Open set</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</conf-name>, <conf-loc>Snowmass Village, CO, USA</conf-loc>. <fpage>1021</fpage>&#x2013;<lpage>1030</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Fuentes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>M. H.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Data-centric annotation analysis for plant disease detection: Strategy, consistency, and performance</article-title>. <source>Front. Plant Sci.</source> <volume>13</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2022.1037655</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Du</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Cai</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>b). <article-title>Vos: Learning what you don't know by virtual outlier synthesis</article-title>. In <source>International Conference on Learning Representations</source> (OpenReview, Online) (<publisher-name>arXiv preprint arXiv</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2202.01197</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Du</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Gozum</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>a). &#x201c;<article-title>Unknown-aware object detection: Learning what you don't know from videos in the wild</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>New Orleans, Louisiana, USA</conf-loc>. <fpage>13678</fpage>&#x2013;<lpage>13688</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fenu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Malloci</surname> <given-names>F. M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>DiaMOS plant: A dataset for diagnosis and monitoring plant disease</article-title>. <source>Agronomy</source> <volume>11</volume>, <fpage>2107</fpage>. doi: <pub-id pub-id-type="doi">10.3390/agronomy11112107</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fuentes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>D. S.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Anomaly detection of plant diseases and insects using convolutional neural networks</article-title>,&#x201d; in <conf-name>Proceedings of the International Society for Ecological Modelling Global Conference</conf-name>, <conf-loc>Jeju Island, South Korea</conf-loc>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fuentes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>D. S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Open set self and across domain adaptation for tomato disease recognition with deep learning techniques</article-title>. <source>Front. Plant Sci.</source> <volume>12</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2021.758027</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fuentes</surname> <given-names>A. F.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>D. S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>High-performance deep neural network-based tomato plant diseases and pests diagnosis system with refinement filter bank</article-title>. <source>Front. Plant Sci.</source> <volume>9</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2018.01162</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Geng</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>S.-J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Recent advances in open set recognition: A survey</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>43</volume>, <fpage>3614</fpage>&#x2013;<lpage>3631</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2020.2981604</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Fast r-cnn</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>, <conf-loc>Santiago, Chile</conf-loc>. <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Donahue</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Malik</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Columbus, OH, USA</conf-loc>. <fpage>580</fpage>&#x2013;<lpage>587</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Dollar</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Lvis: A dataset for large vocabulary instance segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>, <conf-loc>Long Beach, California, USA</conf-loc>. <fpage>5356</fpage>&#x2013;<lpage>5364</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Narayan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Joseph</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>F. S.</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Ow-detr: Open-world detection transformer</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>New Orleans, Louisiana, USA</conf-loc>. <fpage>9235</fpage>&#x2013;<lpage>9244</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hayes</surname> <given-names>T. L.</given-names>
</name>
<name>
<surname>Kafle</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Shrestha</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Acharya</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Kanan</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Remind your neural network to prevent catastrophic forgetting</article-title>,&#x201d; in <conf-name>Computer Vision&#x2013;ECCV 2020: 16th European Conference</conf-name>, <conf-loc>Glasgow, UK</conf-loc>, <conf-date>August 23&#x2013;28, 2020</conf-date>. <fpage>466</fpage>&#x2013;<lpage>483</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Gkioxari</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Mask r-cnn</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>, <conf-loc>Venice, Italy</conf-loc>. <fpage>2961</fpage>&#x2013;<lpage>2969</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hendrycks</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Gimpel</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A baseline for detecting misclassified and out-of-distribution examples in neural networks</article-title>. In <source>International Conference on Learning Representations</source> (<publisher-loc>San Juan, Puerto Rico</publisher-loc>: <publisher-name>arXiv preprint arXiv</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1610.02136</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Joseph</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>F. S.</given-names>
</name>
<name>
<surname>Balasubramanian</surname> <given-names>V. N.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Towards open world object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Nashville, United States</conf-loc>. <fpage>5830</fpage>&#x2013;<lpage>5840</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kuo</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Gu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Piergiovanni</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Angelova</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201c;<article-title>Open-vocabulary object detection upon frozen vision and language models</article-title>,&#x201d; in <conf-name>The Eleventh International Conference on Learning Representations</conf-name>, <conf-loc>Kigali, Rwanda</conf-loc>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lecun</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chopra</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hadsell</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Ranzato</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>A tutorial on energy-based learning</article-title>. <source>Predicting structured Data</source> (<publisher-loc>Massachusetts USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>) <volume>1</volume>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Hoiem</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Learning without forgetting</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>40</volume>, <fpage>2935</fpage>&#x2013;<lpage>2947</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2017.2773081</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Using deep learning for Image-Based different degrees of ginkgo leaf disease classification</article-title>. <source>Information</source> <volume>11</volume>, <fpage>95</fpage>. doi: <pub-id pub-id-type="doi">10.3390/info11020095</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>T.-Y.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Hariharan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Belongie</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Feature pyramid networks for object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>. <fpage>2117</fpage>&#x2013;<lpage>2125</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>T.-Y.</given-names>
</name>
<name>
<surname>Maire</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Belongie</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hays</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Perona</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Ramanan</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). &#x201c;<article-title>Microsoft coco: Common objects in context</article-title>,&#x201d; in <conf-name>European conference on computer vision</conf-name>, <conf-loc>Zurich, Switzerland</conf-loc>. <fpage>740</fpage>&#x2013;<lpage>755</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Plant diseases and pests detection based on deep learning: a review</article-title>. <source>Plant Methods</source> <volume>17</volume>, <fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s13007-021-00722-9</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Owens</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Energy-based out-of-distribution detection</article-title>. <source>Adv. Neural Inf. Process. Syst.</source>, <fpage>21464</fpage>&#x2013;<lpage>21475</lpage>. </citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Livio</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Why?: What makes us curious</source> (<publisher-loc>New York, United States</publisher-loc>: <publisher-name>Simon and Schuster</publisher-name>).</citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>b). &#x201c;<article-title>Annealing-based label-transfer learning for open world object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>. <fpage>11454</fpage>&#x2013;<lpage>11463</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T. H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>a). &#x201c;<article-title>CAT: LoCalization and identificAtion cascade detection transformer for open-world object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>. <fpage>19681</fpage>&#x2013;<lpage>19690</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mathew</surname> <given-names>M. P.</given-names>
</name>
<name>
<surname>Mahesh</surname> <given-names>T. Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Leaf-based disease detection in bell pepper plant using YOLO v5</article-title>. <source>Signal Image Video Process.</source> <volume>16</volume>, <fpage>841</fpage>&#x2013;<lpage>847</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11760-021-02024-y</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Miller</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Nicholson</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Dayoub</surname> <given-names>F.</given-names>
</name>
<name>
<surname>S&#xfc;nderhauf</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Dropout sampling for robust object detection in open-set conditions</article-title>,&#x201d; in <conf-name>2018 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>, <conf-loc>Brisbane, Australia</conf-loc>. <fpage>3243</fpage>&#x2013;<lpage>3249</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Murugeswari</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Anwar</surname> <given-names>Z. S.</given-names>
</name>
<name>
<surname>Dhananjeyan</surname> <given-names>V. R.</given-names>
</name>
<name>
<surname>Karthik</surname> <given-names>C. N.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Automated sugarcane disease detection using faster RCNN with an android application</article-title>,&#x201d; in <conf-name>2022 6th International Conference on Trends in Electronics and Informatics (ICOEI)</conf-name>, <conf-loc>Tirunelveli, India</conf-loc>. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nazki</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Fuentes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>D. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Unsupervised image translation using adversarial networks for improved plant disease recognition</article-title>. <source>Comput. Electron. Agric.</source> <volume>168</volume>, <elocation-id>105117</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2019.105117</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Panchal</surname> <given-names>A. V.</given-names>
</name>
<name>
<surname>Patel</surname> <given-names>S. C.</given-names>
</name>
<name>
<surname>Bagyalakshmi</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>I. R.</given-names>
</name>
<name>
<surname>Soni</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Image-based plant diseases detection using deep learning</article-title>. <source>Materials Today: Proc.</source> <volume>80</volume>, <fpage>3500</fpage>&#x2013;<lpage>3506</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.matpr.2021.07.281</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Prabhu</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Torr</surname> <given-names>P. H.</given-names>
</name>
<name>
<surname>Dokania</surname> <given-names>P. K.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Gdumb: A simple approach that questions our progress in continual learning</article-title>,&#x201d; in <conf-name>Computer Vision&#x2013;ECCV 2020: 16th European Conference</conf-name>, <conf-loc>Glasgow, UK</conf-loc>, <conf-date>August 23&#x2013;28, 2020</conf-date>. <fpage>524</fpage>&#x2013;<lpage>540</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Priyadharshini</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Dolly</surname> <given-names>D. R. J.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201c;<article-title>Comparative investigations on tomato leaf disease detection and classification using CNN, R-CNN, fast R-CNN and faster R-CNN</article-title>,&#x201d; in <conf-name>2023 9th International Conference on Advanced Computing and Communication Systems (ICACCS)</conf-name>, <conf-loc>TamilNadu, INDIA</conf-loc>. <fpage>1540</fpage>&#x2013;<lpage>1545</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qiao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Valente</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>He</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>AI, sensors and robotics in plant phenotyping and precision agriculture</article-title>. <source>Front. Plant Sci.</source> <volume>13</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2022.1064219</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Qiao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>He</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>b). &#x201c;<article-title>Deep learning based grape mildew disease severity classification</article-title>,&#x201d; in <conf-name>2022 ASABE Annual International Meeting</conf-name>, <conf-loc>Houston, Texas, USA</conf-loc>. <fpage>1</fpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Radford</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>J. W.</given-names>
</name>
<name>
<surname>Hallacy</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Ramesh</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Goh</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Agarwal</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Learning transferable visual models from natural language supervision</article-title>,&#x201d; in <conf-name>International conference on machine learning</conf-name>. <fpage>8748</fpage>&#x2013;<lpage>8763</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Rebuffi</surname> <given-names>S.-A.</given-names>
</name>
<name>
<surname>Kolesnikov</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sperl</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lampert</surname> <given-names>C. H.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>icarl: Incremental classifier and representation learning</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>. <fpage>2001</fpage>&#x2013;<lpage>2010</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ren</surname> <given-names>S.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Faster r-cnn: Towards real-time object detection with region proposal networks</article-title>,&#x201d; in <conf-name>Advances in neural information processing systems</conf-name>, <conf-loc>Montreal Convention Center, Montreal, Canada</conf-loc>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seetharaman</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Mahendran</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Leaf disease detection in banana plant using gabor extraction and region-based convolution neural network (RCNN)</article-title>. <source>J. Institution Engineers (India): Ser. A</source> <volume>103</volume>, <fpage>501</fpage>&#x2013;<lpage>507</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s40030-022-00628-2</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Shao</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). &#x201c;<article-title>Objects365: A large-scale, high-quality dataset for object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF international conference on computer vision</conf-name> (<publisher-loc>Seoul, Korea</publisher-loc>). <fpage>8430</fpage>&#x2013;<lpage>8439</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shoaib</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Ei-Sappagh</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ali</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Ullah</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Alenezi</surname> <given-names>F.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>An advanced deep learning models-based plant disease detection: A review of recent research</article-title>. <source>Front. Plant Sci.</source> <volume>14</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2023.1158933</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Singh</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Jain</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Jain</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Kayal</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Kumawat</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Batra</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>PlantDoc: a dataset for visual plant disease detection</article-title>,&#x201d; in <conf-name>Proceedings of the 7th ACM IKDD CoDS and 25th COMAD</conf-name> (<publisher-loc>Hyderabad, India</publisher-loc>).</citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vaze</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Vedaldi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Zisserman</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Open-set recognition: A good closed-set classifier is all you need</article-title>. In <source>International Conference on Learning Representations</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2110.06207</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Shang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>He</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Plant disease detection and classification method based on the optimized lightweight YOLOv5 model</article-title>. <source>Agriculture</source> <volume>12</volume>, <elocation-id>931</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/agriculture12070931</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wortsman</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ilharco</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>J. W.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Kornblith</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Roelofs</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). &#x201c;<article-title>Robust fine-tuning of zero-shot models</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>New Orleans, Louisiana, USA</conf-loc>. <fpage>7959</fpage>&#x2013;<lpage>7971</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Kirillov</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Massa</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Lo</surname> <given-names>W.-Y.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Detectron2</article-title>. <uri xlink:href="https://github.com/facebookresearch/detectron2">https://github.com/facebookresearch/detectron2</uri>.</citation>
</ref>
<ref id="B51">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Kang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>UC-OWOD: unknown-classified open world object detection</article-title>,&#x201d; in <conf-name>Computer Vision&#x2013;ECCV 2022: 17th European Conference</conf-name>, <conf-loc>Tel Aviv, Israel</conf-loc>, <conf-date>October 23&#x2013;27, 2022</conf-date>. <fpage>193</fpage>&#x2013;<lpage>210</lpage>.</citation>
</ref>
<ref id="B52">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Tu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Aggregated residual transformations for deep neural networks</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>. <fpage>1492</fpage>&#x2013;<lpage>1500</lpage>.</citation>
</ref>
<ref id="B53">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Xiong</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>From open set to closed set: Counting objects by spatial divide-and-conquer</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, <conf-loc>Seoul, Korea</conf-loc>. <fpage>8362</fpage>&#x2013;<lpage>8371</lpage>.</citation>
</ref>
<ref id="B54">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Fuentes</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yoon</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Embrace limited and imperfect training datasets: opportunities and challenges in plant disease recognition using deep learning</article-title>. <source>Front. Plant Science</source> <volume>14</volume>, <fpage>1225409</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2305.11533</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Three-channel convolutional neural networks for vegetable leaf disease recognition</article-title>. <source>Cogn. Syst. Res.</source> <volume>53</volume>, <fpage>31</fpage>&#x2013;<lpage>41</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cogsys.2018.04.006</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zohar</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>K.-C.</given-names>
</name>
<name>
<surname>Yeung</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201c;<article-title>PROB: Probabilistic objectness for open world object detection</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>. <fpage>11444</fpage>&#x2013;<lpage>11453</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>