<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2017.00252</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Automated Method to Determine Two Critical Growth Stages of Wheat: Heading and Flowering</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Sadeghi-Tehran</surname> <given-names>Pouria</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/380408/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sabermanesh</surname> <given-names>Kasra</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/381182/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Virlet</surname> <given-names>Nicolas</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/381183/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Hawkesford</surname> <given-names>Malcolm J.</given-names></name>
<xref ref-type="author-notes" rid="fn002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/26032/overview"/>
</contrib>
</contrib-group>
<aff><institution>Department of Plant Biology and Crop Sciences, Rothamsted Research</institution> <country>Harpenden, UK</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: John Doonan, Aberystwyth University, UK</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yuhui Chen, Samuel Roberts Noble Foundation, USA; Andrew French, University of Nottingham, UK</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Pouria Sadeghi-Tehran <email>pouria.sadeghi-tehran&#x00040;rothamsted.ac.uk</email></p></fn>
<fn fn-type="corresp" id="fn002"><p>Malcolm J. Hawkesford <email>malcolm.hawkesford&#x00040;rothamsted.ac.uk</email></p></fn>
<fn fn-type="other" id="fn003"><p>This article was submitted to Technical Advances in Plant Science, a section of the journal Frontiers in Plant Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>02</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>252</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>09</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>02</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Sadeghi-Tehran, Sabermanesh, Virlet and Hawkesford.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Sadeghi-Tehran, Sabermanesh, Virlet and Hawkesford</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Recording growth stage information is an important aspect of precision agriculture, crop breeding and phenotyping. In practice, crop growth stage is still primarily monitored by-eye, which is not only laborious and time-consuming, but also subjective and error-prone. The application of computer vision on digital images offers a high-throughput and non-invasive alternative to manual observations and its use in agriculture and high-throughput phenotyping is increasing. This paper presents an automated method to detect wheat heading and flowering stages, which uses the application of computer vision on digital images. The bag-of-visual-word technique is used to identify the growth stage during heading and flowering within digital images. Scale invariant feature transformation feature extraction technique is used for lower level feature extraction; subsequently, local linear constraint coding and spatial pyramid matching are developed in the mid-level representation stage. At the end, support vector machine classification is used to train and test the data samples. The method outperformed existing algorithms, having yielded 95.24, 97.79, 99.59% at early, medium and late stages of heading, respectively and 85.45% accuracy for flowering detection. The results also illustrate that the proposed method is robust enough to handle complex environmental changes (illumination, occlusion). Although the proposed method is applied only on identifying growth stage in wheat, there is potential for application to other crops and categorization concepts, such as disease classification.</p>
</abstract>
<kwd-group>
<kwd>image categorization</kwd>
<kwd>computer vision in agriculture</kwd>
<kwd>automated field phenotyping</kwd>
<kwd>automated growth stage observation</kwd>
<kwd>Field Scanalyzer</kwd>
<kwd>wheat heading stage</kwd>
<kwd>wheat flowering time</kwd>
</kwd-group>
<counts>
<fig-count count="11"/>
<table-count count="3"/>
<equation-count count="7"/>
<ref-count count="38"/>
<page-count count="14"/>
<word-count count="7295"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>An estimated doubling in required crop production is projected by 2,050 in order to meet the demand of the rapid growth human population (Tilman et al., <xref ref-type="bibr" rid="B28">2011</xref>). To achieve this, an approximate 38% increase over current increases in annual crop production rates is required, and on not much more arable land. Further concerns exist around not only achieving this target in a changing climate, but also achieving it sustainably, whereby reducing agricultural inputs to reduce the environmental degradation caused by our agricultural footprint (Tester and Langridge, <xref ref-type="bibr" rid="B26">2010</xref>). With wheat providing &#x02265;20% of the worlds calorie and protein intake (Braun et al., <xref ref-type="bibr" rid="B5">2010</xref>), the requirement to increase yield and production is widely recognized.</p>
<p>Breeding and precision agriculture, including information-based management of agricultural systems, are fundamental for achieving sustainable increases in wheat productivity and production. One component critical to both crop breeding and precision agriculture is the monitoring of developmental growth stages, as (i) it helps crop producers understand which phases of wheat development are most vulnerable to biotic and environmental stresses, and (ii) supports precision agriculture by helping making informed-decisions around which treatment should be applied, to what location and when to apply it. Two critical growth stages monitored in crops, including wheat, are heading date and flowering time, as cultivars with appropriate heading time to their target environment and life cycle duration will help maximize yield potential (Snape et al., <xref ref-type="bibr" rid="B24">2001</xref>; Zhang et al., <xref ref-type="bibr" rid="B37">2008</xref>).</p>
<p>The monitoring of heading and flowering stages are still primarily performed by human eye, which is labor-intensive and time-consuming, as these observations need to be performed on up to thousands of cultivars/varieties on a daily or bi-daily basis, given the importance in catching the starting date of these growth stages. Given that manual growth stage monitoring is also subjective, different observers may likely perceive the growth stage of the same plot differently, which introduces human-error into obtained data.</p>
<p>Computer vision offers an effective alternative for growth stage monitoring because of its low-cost (relative to man-hours invested in to manual observations) and the requirement for minimal human intervention. Computer vision has facilitated automation in high-throughput phenotyping, as well as areas of agriculture, such as disease detection (Pourreza et al., <xref ref-type="bibr" rid="B23">2015</xref>), weed identification (Guerrero et al., <xref ref-type="bibr" rid="B14">2012</xref>) and quality control (Alahi et al., <xref ref-type="bibr" rid="B1">2012</xref>; Valiente-Gonz&#x000E1;lez et al., <xref ref-type="bibr" rid="B29">2014</xref>). Despite the efforts of computer vision specialists over the past decades, developing reliable image-based model to identify and categorize images based on visual information is still difficult to achieve and remains an unsolved problem in the computer vision community. The visual recognition of object categories is a natural and trivial task for humans. Humans can recognize objects effortlessly even with changes in an object&#x00027;s appearance, such as viewing direction or a shadow being cast across the object. On the other hand, in computer vision it can be a challenging task to achieve such level of performance due to the difficulties inherent in the problem. Images are quite abstract and subjected to illumination, scale, deformation, background clutter, etc. Moreover, in computer vision, teaching a machine to distinguish and categorize objects is all about teaching it which differences in the image is matters and which don&#x00027;t, by scanning through diverse datasets, which is a computationally exhaustive process.</p>
<p>Computer vision has shown promise in detecting growth stages of crops. For seedling emergence, color segmentation approaches have been applied in maize (Yihang et al., <xref ref-type="bibr" rid="B34">2014</xref>) and oilseed rape (Yu et al., <xref ref-type="bibr" rid="B35">2013</xref>), using images acquired from a digital camera. Some approaches for observing later growth stages, such as heading date and flowering stage, have also been developed. Zhu et al. (<xref ref-type="bibr" rid="B38">2016</xref>), developed a method to detect wheat heading stage from RGB images using a two-step coarse-to-fine detection approach. For flowering stages, Guo et al. (<xref ref-type="bibr" rid="B15">2015</xref>) used object-recognition to detect flowering stages from rice panicles. Although the approaches by Zhu et al. (<xref ref-type="bibr" rid="B38">2016</xref>) and Guo et al. (<xref ref-type="bibr" rid="B15">2015</xref>) were effective on a single variety, within small patches of whole canopies, applications that are more versatile and that also can be applied on different varieties on larger scale canopies are required.</p>
<p>This study utilizes a novel visual-based approach to monitor heading and flowering stage of field grown wheat, through the automated learning of the visual consistency between classes of canopy images, in order to identify the critical growing stages of wheat (e.g., whether ears are emerging in canopies). This method searches through an image database to identify and retrieve images containing emerged wheat ears and ears at flowering stages. This visual-based approach is:
<list list-type="roman-lower">
<list-item><p>Not limited to specific wheat cultivars and is applicable to a variety of categorical wheat without implementing specific tuning for each category.</p></list-item>
<list-item><p>Robust to handle illumination changes and natural lightning conditions in the field.</p></list-item>
<list-item><p>Robust in distinguishing the early emerged ears, despite the color difference between ears and leaves being hardly distinguishable to the naked eye.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<title>2. Materials and methodology</title>
<p>The introduced technique is performed in four main steps (Figure <xref ref-type="fig" rid="F1">1</xref>):
<list list-type="bullet">
<list-item><p>Image acquisition: A RGB image is captured from 8 MP camera mounted inside the camera bay.</p></list-item>
<list-item><p>Pre-processing of the images to improve the contrast.</p></list-item>
<list-item><p>Extracting features that contain suitable information to discriminate images at the category level.</p></list-item>
<list-item><p>Classification: Images classified in different categories as specified.</p></list-item>
</list></p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Schematic representation of the method</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0001.tif"/>
</fig>
<p>Bag of Visual Words (BoVW) proved to be the leading strategy in computer vision applications such as image retrieval and image categorization (Csurka et al., <xref ref-type="bibr" rid="B9">2004</xref>); thus, it is being opted for the presented work. Categorizing digital images, embarks on extracting features and creating a visual vocabulary for the given dataset. It comprises of following states:
<list list-type="order">
<list-item><p>Extracting features.</p></list-item>
<list-item><p>Constructing visual vocabulary by clustering.</p></list-item>
<list-item><p>Using multi-class classifier for training using bags as feature vector.</p></list-item>
<list-item><p>For the testing image, obtain the nearest vocabulary based on the most optimum prediction of classifier.</p></list-item>
</list></p>
<p>However, in this study, several steps are integrated in the process to improve the overall performance compared to Csurka et al. (<xref ref-type="bibr" rid="B9">2004</xref>) described in Section 2.3. Our method treats canopy images acquired automatically in the field as a collection of unordered appearance descriptors extracted from local patches; then, quantizes them into discrete visual words. Each image is defined by a feature vector listing the number of regions which belongs to each cluster and are later used to train a classifier. In addition, the location information is taken into account which is one of the important factors in object recognition scenarios. In the final step, a linear Support Vector Machine (SVM) classifier is used to determine pre-defined classes (e.g., ear emergence, flowering). The experimental results show that the introduced method is capable of automatically identifying key wheat growing stage with high accuracy and efficiency (Section 3).</p>
<sec>
<title>2.1. Field experiment and image acquisition</title>
<p>Six wheat cultivars (<italic>Triticum aestivum</italic> L. cv. Avalon, Cadenza, Crusoe, Gatsby, Soissons and Maris Widgeon) were grown in the field at Rothamsted Research, Harpenden, UK, sown in Autumn 2015 and maturing in 2016. These cultivars were selected as they had different properties visible to the naked-eye (awns/no awns, differing wax properties, straight/floppy leaves, different ear morphology) (Figure <xref ref-type="fig" rid="F2">2</xref>). All cultivars were sown 20 October 2015, at a planting density of 350 plants/<italic>m</italic><sup>2</sup>. Nitrogen (N) treatments were applied as ammonium nitrate in the spring, at rates of 0 kg <italic>ha</italic><sup>&#x02212;1</sup> (residual soil N; N1), 100 kg <italic>ha</italic><sup>&#x02212;1</sup> (N2) and 200 kg <italic>ha</italic><sup>&#x02212;1</sup> (N3) (Figure <xref ref-type="fig" rid="F4">4</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Digital images of the six contrasting wheat cultivars (<italic><bold>Triticum aestivum</bold></italic> L. cv. Avalon, Cadenza, Crusoe, Gatsby, Soissons, and Maris Widgeon) used, at growth stage Z5.9</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0002.tif"/>
</fig>
<p>The Field Scanalyzer phenotyping platform (LemnaTec GmbH; Virlet et al., <xref ref-type="bibr" rid="B32">2017</xref>) was used to acquire all images (Figure <xref ref-type="fig" rid="F3">3</xref>). The Field Scanalyzer is a fully-automated, high-throughput, fixed-field phenotyping platform, carrying multiple sensors for non-invasive monitoring of plant growth, morphology, physiology and health. The on-board visible camera (color 12 bit Prosilica GT3300) was used to acquire RGB images at high-resolution (3,296 &#x000D7; 2,472 pixels). The camera is positioned perpendicular to the ground, and automatically adjusts to ensure a 2.5 m distance is maintained between the camera and canopy. The camera is set up in auto-exposure mode, to compensate for outdoor light changes. Wheat canopies were imaged daily during three stages of ear emergence: Stage 1 (Zadoks scale Z5.0; 3&#x02013;5 June 2016 Zadoks et al., <xref ref-type="bibr" rid="B36">1974</xref>); Stage 2 (Z5.3&#x02013;Z5.7; 7&#x02013;10 June 2016) and Stage 3 (&#x02265; Z5.9; 12&#x02013;14 June 2016), as well as flowering stage (14&#x02013;18 June 2016). In addition, illumination conditions were recorded during the image acquisition (Table <xref ref-type="table" rid="T1">1</xref>). Manual growth stages were recorded daily or on alternating days during heading and flowering. The growth stage of the plot was defined manually by the stage of &#x02265;50% of the plot. Videos and more information of the Field Scanalyzer platform can be accessed in our website: <ext-link ext-link-type="uri" xlink:href="http://www.rothamsted.ac.uk/field-scanalyzer">http://www.rothamsted.ac.uk/field-scanalyzer</ext-link>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>The Field Scanalyzer at Rothamsted research</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Digital images highlighting the impact of 0 kg <italic><bold>ha</bold></italic><sup><bold>&#x02212;1</bold></sup> (N1), 100 kg <italic><bold>ha</bold></italic><sup><bold>&#x02212;1</bold></sup> (N2) or 200 kg <italic><bold>ha</bold></italic><sup><bold>&#x02212;1</bold></sup> (N3) nitrogen fertilizer application on canopy complexity</bold>. Images were acquired 2.5 m above <italic>Triticum aestivum</italic> L. cv. Soissons canopies.</p></caption>
<graphic xlink:href="fpls-08-00252-g0004.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Date, start/end time, and PAR values of each images acquisition periods during ears emergence and flowering stages</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Date</bold></th>
<th valign="top" align="center"><bold>Start</bold></th>
<th valign="top" align="center"><bold>End</bold></th>
<th valign="top" align="center"><bold>PAR (<italic>&#x003BC;mol.m</italic><sup>&#x02212;2</sup>.<italic>s</italic><sup>&#x02212;1</sup>)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">03/06/2016</td>
<td valign="top" align="center">11:19:16</td>
<td valign="top" align="center">12:08:24</td>
<td valign="top" align="center">404 &#x000B1; 58</td>
</tr>
<tr>
<td valign="top" align="left">04/06/2016</td>
<td valign="top" align="center">11:53:11</td>
<td valign="top" align="center">12:42:41</td>
<td valign="top" align="center">512 &#x000B1; 67</td>
</tr>
<tr>
<td valign="top" align="left">05/06/2016</td>
<td valign="top" align="center">08:29:34</td>
<td valign="top" align="center">09:18:38</td>
<td valign="top" align="center">287 &#x000B1; 43</td>
</tr>
<tr>
<td valign="top" align="left">07/06/2016</td>
<td valign="top" align="center">13:33:40</td>
<td valign="top" align="center">14:25:07</td>
<td valign="top" align="center">1,037 &#x000B1; 95</td>
</tr>
<tr>
<td valign="top" align="left">08/06/2016</td>
<td valign="top" align="center">08:07:05</td>
<td valign="top" align="center">08:58:46</td>
<td valign="top" align="center">315 &#x000B1; 19</td>
</tr>
<tr>
<td valign="top" align="left">08/06/2016</td>
<td valign="top" align="center">18:05:53</td>
<td valign="top" align="center">18:38:33</td>
<td valign="top" align="center">528 &#x000B1; 180</td>
</tr>
<tr>
<td valign="top" align="left">09/06/2016</td>
<td valign="top" align="center">08:32:25</td>
<td valign="top" align="center">09:24:05</td>
<td valign="top" align="center">530 &#x000B1; 215</td>
</tr>
<tr>
<td valign="top" align="left">10/06/2016</td>
<td valign="top" align="center">07:43:46</td>
<td valign="top" align="center">08:37:33</td>
<td valign="top" align="center">461 &#x000B1; 33</td>
</tr>
<tr>
<td valign="top" align="left">12/06/2016</td>
<td valign="top" align="center">14:32:33</td>
<td valign="top" align="center">15:26:18</td>
<td valign="top" align="center">703 &#x000B1; 304</td>
</tr>
<tr>
<td valign="top" align="left">13/06/2016</td>
<td valign="top" align="center">09:49:04</td>
<td valign="top" align="center">10:41:12</td>
<td valign="top" align="center">555 &#x000B1; 110</td>
</tr>
<tr>
<td valign="top" align="left">14/06/2016</td>
<td valign="top" align="center">10:14:16</td>
<td valign="top" align="center">11:06:10</td>
<td valign="top" align="center">800 &#x000B1; 196</td>
</tr>
<tr>
<td valign="top" align="left">14/06/2016</td>
<td valign="top" align="center">15:01:19</td>
<td valign="top" align="center">15:33:55</td>
<td valign="top" align="center">569 &#x000B1; 121</td>
</tr>
<tr>
<td valign="top" align="left">16/06/2016</td>
<td valign="top" align="center">08:02:55</td>
<td valign="top" align="center">08:35:57</td>
<td valign="top" align="center">363 &#x000B1; 31</td>
</tr>
<tr>
<td valign="top" align="left">18/06/2016</td>
<td valign="top" align="center">10:51:31</td>
<td valign="top" align="center">11:44:17</td>
<td valign="top" align="center">919 &#x000B1; 238</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>PAR mean and standard deviation values are computed from the 54 scans collected during one acquisition periods</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>2.2. Image pre-processing and enhancement</title>
<p>The color of ears at early development stages are very similar to leaves and hardly discernable with the naked-eye (Figures <xref ref-type="fig" rid="F5">5A,C</xref>). In order to make the ears stand out in canopies and discriminate them from the background more easily, a pre-processing method is applied on plot images before extracting features, known as decorrelation stretching (DS). The decorrelation stretching technique enhances the color differences and increasing the image contrast in each plot image by removing the inter-channel correlation found in the pixels (Gillespie et al., <xref ref-type="bibr" rid="B12">1986</xref>). Therefore, it allows to see details such as ears that are otherwise too subtle for the naked-eye (Figures <xref ref-type="fig" rid="F5">5B,D</xref>). If the red, green, and blue values of pixels are treated coordinates in space, decorrelation stretch moves these points in space further apart, so they become much easier to see a difference between them.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>Digital image of wheat (<italic><bold>Triticum aestivum</bold></italic> L. cv. Soissons) canopy (A)</bold> before and <bold>(B)</bold> after enhancement of image contrast and application of decorrelation stretching. Scatterplot of every pixels normalized red, green blue (RGB) values from <bold>(C)</bold> the original image and <bold>(D)</bold> after applying the decorrelation stretch and contrast increase.</p></caption>
<graphic xlink:href="fpls-08-00252-g0005.tif"/>
</fig>
<p>The DS among the RGB channels is achieved through principle component analysis (PCA) to remove inter-channel correlation in an image. The application of PCA to the digital analysis of an image is based on first, calculating the covariance matrix between the three RGB bands. Then, obtaining eigenvectors and eigenvalues. Finally, rotating the original image vector to a new space by multiplying it by the eigenvectors (Equation 1) (Jolliffe, <xref ref-type="bibr" rid="B16">2002</xref>; Cerrillo-Cuenca and Sep&#x000FA;lveda, <xref ref-type="bibr" rid="B7">2015</xref>).</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>i</italic><sub><italic>n</italic></sub> is the image vector; <italic>n</italic> is the number of pixels; and <italic>R</italic> is the rotation matrix.</p>
<p>Campbell (<xref ref-type="bibr" rid="B6">1996</xref>) proposed a general framework consists of the following steps:</p>
<list list-type="roman-lower">
<list-item><p>Calculating <italic>p</italic><sub><italic>n</italic></sub> from Equation (1), eigenvalues and eigenvectors are obtained from the correlation matrix or alternatively from the covariance matrix.</p></list-item>
<list-item><p>Generating a stretch vector: diagonalize the covariance matrix composed by the inverse of the eigenvectors:
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where <italic>D</italic> is a diagonal matrix; <italic>v</italic> denotes each of the eigenvalues. Alternatively, <italic>D</italic> can be multiplied by an integer value that serves to achieve a higher contrast in the image (Alley, <xref ref-type="bibr" rid="B2">1996</xref>). Finally, the resultant matrix is applied to <italic>p</italic><sub><italic>n</italic></sub> (Equation 3). At this step, the matrix is re-centered and stretched its values to a maximum.
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
<list-item><p>The inverse transform is applied to map the colors back to the original space. The information is decorrelated into a new vector <italic>c</italic><sub><italic>n</italic></sub> composed of three matrices (RGB) (Equation 4)
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>D</mml:mi><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
<list-item><p>Finally, a standard deviation value is applied to visually increase the contrast (Alley, <xref ref-type="bibr" rid="B2">1996</xref>).</p></list-item>
</list>
</sec>
<sec>
<title>2.3. Bag of visual words construction</title>
<p>The first step of BoVW framework corresponds to feature extraction. Fixed length feature extraction techniques based on color (Swain and Ballard, <xref ref-type="bibr" rid="B25">1991</xref>; Chen et al., <xref ref-type="bibr" rid="B8">2010</xref>), texture (Duda et al., <xref ref-type="bibr" rid="B10">2000</xref>), shape (Mehrotra and Gary, <xref ref-type="bibr" rid="B21">1995</xref>), or a combination of two or more techniques, extract pixel values of an image only. These are excellent in comparing the overall image similarity (Angelov and Sadeghi-Tehran, <xref ref-type="bibr" rid="B3">2016</xref>); however, they are not scale or rotation invariant. Moreover, they are very sensitive to noise and illumination changes; thus, are unable to describe the object-based properties of the image content.</p>
<p>As opposed to global feature extraction methods mentioned earlier, local extraction algorithms are robust to partial visibility and clutter. It is an ideal candidate for object recognition, template matching and image mosaicing. There are several feature detector methods, which are scale and rotation invariant. They are also robust enough to handle illumination changes and resistant to geometry (Bay et al., <xref ref-type="bibr" rid="B4">2006</xref>; Leutenegger et al., <xref ref-type="bibr" rid="B18">2011</xref>; Alahi et al., <xref ref-type="bibr" rid="B1">2012</xref>). Among the proposed descriptors, Scale Invariant Feature Transform (SIFT) is selected due to its excellent performance attested in various applications (Mikolajczyk and Schmid, <xref ref-type="bibr" rid="B22">2005</xref>). It returns an <italic>N</italic> &#x000D7; 128 dimension image descriptor, where <italic>N</italic> is the number of features.</p>
<p>SIFT consists of Lowe (<xref ref-type="bibr" rid="B20">2004</xref>):
<list list-type="bullet">
<list-item><p><bold>Constructing a scale space</bold>: in this stage, location and scales of each keypoint are identified. Laplacian-of-Gaussian (LoG) is calculated for an image with various &#x003C3;. Due to change in &#x003C3;, LoG detects blobs of various sizes, then the local maxima can be found across the scale and space with a list of (<italic>x, y</italic>, &#x003C3;) values, which show there is a candidate keypoint at location <italic>x, y</italic> with scale of &#x003C3;. However, in order to reduce the computational complexity, SIFT uses Difference-of-Gaussian (DoG) which is a convolved image in scale space separated by a constant factor <italic>k</italic>:
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where <italic>I</italic>(<italic>x, y</italic>) is an input image; <italic>L</italic>(<italic>x, y, k&#x003C3;</italic>) is the scale space of an image; <italic>G</italic>(<italic>x, y, k&#x003C3;</italic>) is variable-scale Gaussian.</p>
<p><italic>D</italic> is computed by simple image subtraction and the Guassian image is sub-sampled by a factor of 2 and produces DoG for the sampled image. Once the DoG is computed, images are searched for local extrema over space and scale. For instance, one pixel is compared with its <italic>n</italic> &#x000D7; <italic>n</italic> neighborhood (<italic>n</italic> &#x0003D; 3 in our experiment) as well as 9 pixels in the next scale and 9 pixels in previous scales (Lowe, <xref ref-type="bibr" rid="B20">2004</xref>).</p></list-item>
<list-item><p><bold>Keypoint localization and filtering</bold>: Once the location of keypoints candidates are found, they are refined and some are eliminated to get a more accurate location of extrema. For instance, if the intensity at the extrema is less than a certain threshold (threshold &#x0003C;0.03) it is rejected. In addition, edges and low contrast regions are considered as bad keypoints and will be rejected.</p></list-item>
<list-item><p><bold>Orientation assignment</bold>: The orientation of each keypoint is obtained based on image gradient and local image gradient directions to achieve rotation invariance. Depending on the scale a neighborhood is taken around the keypoint location and the gradient magnitude and direction is calculated in that region.</p></list-item>
<list-item><p><bold>Keypoint descriptor</bold>: In order to generate a keypoint descriptor, the local image descriptor is computed for each keypoint based on image gradient magnitude and orientation at each image sample point in a region centered at keypoint. These samples build a 3D histogram of gradient location and orientation; with a 4 &#x000D7; 4 array location grid and 8 orientation bins in each sample, which creates 128 element dimensions of the keypoint descriptor, causing robustness against changes in scale and rotation (Figure <xref ref-type="fig" rid="F6">6</xref>).</p>
</list-item>
</list></p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>(A)</bold> A single keypoint candidate in the image; <bold>(B&#x02013;D)</bold> SIFT descriptor calculated at different scales of 4, 8, and 10; At each scale, the descriptor has 4 &#x000D7; 4 patches (color coded in yellow), which are rotated to the dominant orientation of the feature point. Each patch is represented in gradient magnitudes of eight directions, represented by yellow arrows inside each bin.</p></caption>
<graphic xlink:href="fpls-08-00252-g0006.tif"/>
</fig>
<p>The next step is to form clusters of similar features and assign them as visual words. The objective of constructing codebook is to relate features of testing images to the features previously extracted from the training image samples (Figure <xref ref-type="fig" rid="F7">7</xref>). Although in the field of unsupervised learning, clustering is a standard procedure, there is no single clustering algorithm that can be applied uniformly to all the application domains or address all related issues in a satisfactory manner. Here, a partition-based clustering approach known as <italic>K</italic>-means clustering is used to quantize each descriptor and generate a codebook. The process is iterative as follows (Lloyd, <xref ref-type="bibr" rid="B19">1982</xref>):</p>
<table-wrap position="float">
<label>Algorithm 1</label>
<caption><p><italic>K</italic>-means clustering procedure</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td align="left" valign="top">1: &#x000A0;Select <italic>K</italic> points as initial centers</td>
</tr>
<tr>
<td align="left" valign="top">2: &#x000A0;<bold>repeat</bold></td>
</tr>
<tr>
<td align="left" valign="top">3: &#x000A0;Assign each input data to its closest center</td>
</tr>
<tr>
<td align="left" valign="top">4: &#x000A0;Re-compute the center of each cluster by averaging all the members in the clusters</td>
</tr>
<tr>
<td align="left" valign="top">5: &#x000A0;<bold>until</bold></td>
</tr>
<tr>
<td align="left" valign="top">6: &#x000A0;convergence which means no pixel shifts from one cluster to another; centers do not change</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Schematic representation of the proposed method</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0007.tif"/>
</fig>
<p>In <italic>K</italic>-means the number of clusters is pre-defined beforehand and it should be large enough to identify relevant changes in each wheat cultivars. For an image having <italic>N</italic> features, the model will distribute the features with <italic>K</italic> clusters, which is the size of the visual vocabulary. We have been able to find the optimum numbers and get very good results with number of vocabulary (codebook) <italic>K</italic> &#x0003D; 2000 (Table <xref ref-type="table" rid="T2">2</xref>).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Comparison of different methods applied on the three heading stages</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Decorrelation pre-processing</bold></th>
<th valign="top" align="left"><bold>Feature extraction</bold></th>
<th valign="top" align="left"><bold>Coding method</bold></th>
<th valign="top" align="left"><bold>Spatial pyramid</bold></th>
<th valign="top" align="center"><bold>Vocabulary length</bold></th>
<th valign="top" align="center"><bold>Accuracy Z5.0 (%)</bold></th>
<th valign="top" align="center"><bold>Accuracy Z5.3&#x02013;Z5.7 (%)</bold></th>
<th valign="top" align="center"><bold>Accuracy &#x02265; Z5.9 (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="char" char=".">95.24</td>
<td valign="top" align="char" char=".">97.79</td>
<td valign="top" align="char" char=".">99.59</td>
</tr>
<tr>
<td valign="top" align="left">No</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="char" char=".">57.29</td>
<td valign="top" align="char" char=".">82.20</td>
<td valign="top" align="char" char=".">85.38</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">k-NN</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="char" char=".">90.54</td>
<td valign="top" align="char" char=".">94.48</td>
<td valign="top" align="char" char=".">96.97</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">No</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="char" char=".">92.90</td>
<td valign="top" align="char" char=".">96.54</td>
<td valign="top" align="char" char=".">96.90</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SURF</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="char" char=".">56.84</td>
<td valign="top" align="char" char=".">71.55</td>
<td valign="top" align="char" char=".">78.45</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">1,000</td>
<td valign="top" align="char" char=".">93.91</td>
<td valign="top" align="char" char=".">97.52</td>
<td valign="top" align="char" char=".">98.91</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">1,500</td>
<td valign="top" align="char" char=".">94.61</td>
<td valign="top" align="char" char=".">97.64</td>
<td valign="top" align="char" char=".">99.24</td>
</tr>
<tr>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">SIFT</td>
<td valign="top" align="left">LLC</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="center">2,500</td>
<td valign="top" align="char" char=".">94.90</td>
<td valign="top" align="char" char=".">97.59</td>
<td valign="top" align="char" char=".">99.49</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The codebook is used for quantizing features. A vector quantizer takes a feature vector and maps it to the index of the nearest code vector in a codebook. In our work, in order to project the descriptors onto the codebook elements, Local Linear Constraint (LLC) (Wang et al., <xref ref-type="bibr" rid="B33">2010</xref>) is used to generate a final vector which represents an image. LLC reduces the computational complexity to <italic>O</italic>(<italic>K</italic>&#x0002B;<italic>K</italic>) (where <italic>K</italic> is the length of the codebook; <italic>K</italic> &#x0003D; 2000 in this case) for each descriptor and can achieve acceptable image classification accuracy even with a linear SVM classifier (Wang et al., <xref ref-type="bibr" rid="B33">2010</xref>).</p>
<p>The main drawback of BoVW is that it is unable to capture spatial relationships between images. In order to preserve the spatial relations of the code vector Spatial Pyramid Matching (SPM) is implemented where the entire image is divided into levels. Each image is divided into spatial sub-regions and computes histograms of features from each sub-region. Each level divides the image into 2<sup><italic>l</italic></sup> &#x000D7; 2<sup><italic>l</italic>&#x02212;1</sup>; where <italic>l</italic> is level (Grauman and Darrell, <xref ref-type="bibr" rid="B13">2005</xref>; Lazebnik et al., <xref ref-type="bibr" rid="B17">2006</xref>). The features are computed locally for each grid and the spatial information is incorporated into histograms. A three level SPM is used with first, level 0 which comprises of a single histogram; level 1, comprising of 4 histograms, finally level 2, comprising of 16 histograms (Figure <xref ref-type="fig" rid="F7">7</xref>). The histogram from all the sub-region are concatenated together to generate the final representation of the image for classification. The result is a feature weighted histogram of 21 &#x000D7; <italic>K</italic> (number of words &#x0003D; 2000). Using such method will preserve the discriminative power of the descriptors; in addition, changes in the positioning of the objects and variations in the background will not affect the overall performance of the method.</p>
</sec>
<sec>
<title>2.4. Learning model</title>
<p>The construction of the model for our image annotation is based on the supervised machine learning principle. Supervised learning can be thought as learning by examples represented by a set of training-testing samples. In order to classifying unknown testing images, a certain number of training images are used for each class to train the classifier. A classifier approximates the mapping between the images and correctly labels the training set, called the <italic>training</italic> phase. After the model is trained, it is able to classify unknown image, into one of the learned class labels.</p>
<p>In our model, the complexity of visual categorization is reduced to two-class with positive and negative training patches. The SVM classifier is used as our classifier of choice as it is fast and can handle the long feature vectors generated by the SPM. During the training phase, labeled images (ears and background) are fed to the classifier and used to adapt a statistical decision procedure. Among many available classifiers, linear SVM with Hellinger kernel is used to predict the unlabeled test images and retrieve as much of the data as possible in a high ranked position. Feature vectors generated from each image are normalized to a unit Euclidean norm and used for a linear SVM classifier with the Hellinger kernel to compute the feature map (Vedaldi and Zisserman, <xref ref-type="bibr" rid="B31">2012</xref>).
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02032;</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02032;</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msqrt><mml:mo>;</mml:mo><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>;</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02032;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x02032;</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02032;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where <italic>n</italic> and <italic>n</italic>&#x02032; are normalized histograms; <italic>d</italic> &#x0003D; 42, 000</p>
<p>One-vs.-all strategy is chosen to train the SVM. Two classes are trained, each labels the sample inside one class as &#x0002B;1 and other samples (background) as -1. The SVM calculates the similarity of all trained classes and assigned the test image to the class with the highest similarity measure.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Experimental results and discussion</title>
<p>The experiment is divided into two sections of identifying ear emergence and flowering stages from the digital images acquired in the field. In the first section, ear emergence was tested at different time points, from early stages where only few spikelets are visible, to a more advanced stage where ears are fully emerged (Section 3.1). In the second part of the experiment, the method was tested to identify flowering growth stage during anthesis (Section 3.2). The training dataset for the ear emergence experiment includes images with ears at different emergence stages (positive class) and leaves, soil, etc. (negative class), which are manually cropped and stored in the dataset. On the other hand, the training dataset for the flowering experiment contains ears at different flowering time points (positive class) and ears before and after flowering (negative class). The collected dataset focuses on different challenges regardless of light conditions in the field and to demonstrate the robustness of the method to environmental changes. In addition, the versatility of the proposed technique were also tested by minimizing the number of cultivars as training patches, and evaluating the method on more varieties.</p>
<p>The research was conducted with the following specifications. System comprised of 24 GB RAM, Intel quad core Processor (3.40 GHz) with Windows 10 OS. The models have been developed in MATLAB (Mathworks Inc.); however, to improve the processing time, some of the algorithm, such as SIFT were written in C&#x0002B;&#x0002B; programming language. Utilities like VLFeat library (Vedaldi and Fulkerson, <xref ref-type="bibr" rid="B30">2010</xref>) to extract features as well as LibLinear library (Fan et al., <xref ref-type="bibr" rid="B11">2008</xref>) to train and test the SVM classifier. Using the above configured computer system, extracting features and generating code vectors from each training image approximately takes 0.45 s. However, the processing time increases to 5.4 s for each testing patch with resolution of 3,298 &#x000D7; 2,474 pixels.</p>
<p>Precision (Pr) and Recall (Re) are the most commonly used measurements to evaluate the performance of image retrieval systems. Thus, it is used in our experiment to quantitatively assess the precision of the proposed approach in detecting the two main growing stages of ear emergence and flowering. Precision is defined as the ratio of the number of retrieved relevant images <italic>N</italic><sub><italic>r</italic></sub> to the total number of retrieved images <italic>N</italic> (Equation 7); on the other hand, Recall is defined as the number of retrieved relevant images <italic>N</italic><sub><italic>r</italic></sub> over the total number of positive images <italic>N</italic><sub><italic>t</italic></sub> available in the database. In an ideal scenario, both Pr and Re should have high values (1). Therefore, instead of using Pr and Re individually, usually accuracy curve is used to characterize the performance of the retrieval system.</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo class="qopname">Pr</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>;</mml:mo><mml:mo class="qopname">Re</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<sec>
<title>3.1. Ear emergence</title>
<p>The learning process starts with 1,000 training image samples divided into 500 ears (positive class), which are manually cropped from full size canopy image and 500 background images (negative class). Figure <xref ref-type="fig" rid="F8">8</xref> shows image samples randomly selected from training patches which are not necessarily the same dimensions. Moreover, to observe the field challenges during data acquisition, ears are selected from different positions and illumination conditions (with or without occlusions and overlapping; sunny or cloudy days). Three different wheat cultivars are used as a training dataset including Avalon, Cadenza, and Soissons. Cadenza can present short awnlettes/scurs at the ear tip, although most of the times no awns are present in contrast to Soissons which is an awned variety. Although three wheat cultivars were used as a training dataset, six cultivars including Maris Widgeon, Avalon, and Gatsby are tested to highlight the versatility of the proposed technique.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Example of ear emergence training patches</bold>. Note that the training patches are not necessarily of equal size and resized for illustration purposes. <bold>(A)</bold> Examples of positive training patches of three different wheat cultivars (<italic>Triticum aestivum</italic> L. cv. Soissons, Avalon, and Cadenza). <bold>(B)</bold> Examples of negative training image patches.</p></caption>
<graphic xlink:href="fpls-08-00252-g0008.tif"/>
</fig>
<p>Ear identification was evaluated at three different time points of the emergence period, (i) at Z5.0, when the ears start to be visible (first spikelet of inflorescence visible), (ii) between Z5.3&#x02013;Z5.7, when 1/4 to 3/4 of the ears are emerged and (iii) at Z &#x02265; 5.9, when ears are fully emerged (Figure <xref ref-type="fig" rid="F8">8</xref>). Each time point was tested independently from datasets containing 80 images (40 with ears present and 40 without) of full size wheat canopies with the original resolution of 3,298 &#x000D7; 2,474 pixels.</p>
<p>The results for each ear development stage are shown in Table <xref ref-type="table" rid="T2">2</xref>. The accuracy of the method is evaluated using different techniques at different processing stages. (i) presence/absence of decorrelation processing, (ii) SIFT vs. SURF, (iii) LLC vs. KNN, (iv) presence/absence of spatial pyramid and (v) the vocabulary length. As shown in the Tables, the best performance was obtained using decorrelation pre-processing, SIFT, LLC coding, and a 2,000 entry codebook. The best performance at heading stage Z5.0 is 95.24%, and for heading stages Z5.3&#x02013;Z5.7 and &#x02265; Z5.9 are 97.79 and 99.59%, respectively (Figure <xref ref-type="fig" rid="F9">9</xref>). Out of the eight tested scenarios, we achieved accuracy of &#x0003E; 90% at Z5.0 and &#x0003E; 96% at Z5.9 in six scenarios. The impact of codebook size on the performance of the method was also investigated. It is clearly shown that the increasing number of codebook improves the accuracy; however, the accuracy plateaus at 2,000 visual words. Moreover, the low-level feature extraction and the decorrelation pre-processing technique has the biggest influence in the quality of results; especially in the early heading (Z5.0). The main conclusion is that mid-level feature coding and classification are highly impacted by the low level pre-processing and feature extraction techniques.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>Three ear development stages visually scored and used to evaluate the performance of the proposed method</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0009.tif"/>
</fig>
<p>Figure <xref ref-type="fig" rid="F10">10</xref> illustrates the performance evolution of the heading stage Z5.0 over the number of images in the training dataset. for both positive and negative data. Training patches of 50, 100, 300 were selected randomly apart from the full set when all 500 samples were used. The accuracy improves by increasing the number of training samples. The accuracy increased from 75.77 to 90.65% when the training dataset increased from 50 to 100. On the other hand, there was no substantial change in accuracy between 100 and 200 samples. However, the performance jumped by more than 5% from 90.80 to 95.24% when the dataset increased to 500.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>Accuracy of the proposed method over number of training image samples</bold>.</p></caption>
<graphic xlink:href="fpls-08-00252-g0010.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Flowering time</title>
<p>Similarly to ear emergence identification, two training classes were created, which comprised of three wheat cultivars (Soissons, Maris Widgeon, and Cadenza). The first class (positive class) contained 140 manually cropped images at flowering stage while the second class (negative class) contained the same number of images as the positive class, but with ears before and after flowering.</p>
<p>Figure <xref ref-type="fig" rid="F11">11</xref> shows randomly selected samples from the training patches. All training images were collected without considering the environmental changes and positioning or occlusion. As flowering development may be completed in only a few days, the beginning or intermediate stages can be easily missed. Therefore, all flowering images along the flowering duration were included. For the testing dataset, 108 full size canopy images were used with the original resolution, which includes 54 canopies with ears during flowering stage and 54 canopies with ears before or after flowering stage. The method selected to test the flowering stage was the one which produced the best result in the ear emergence experiment (decorrelation stretching, SIFT, LLC, and SPM algorithms with the vocabulary length of 2,000). The method was tested on each cultivar separately, as well as all three together. For all three cultivars, 38 images out of 54 images were retrieved correctly, which shows 82.54% accuracy. On the other hand, the accuracy when testing Soissons, Cadenza, and Widgeon individually was 76.72, 92.91, and 80.33%, respectively (Table <xref ref-type="table" rid="T3">3</xref>).</p>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p><bold>Example of flowering training patches</bold>. Examples of positive training patches of three different cultivars; (<italic>Triticum aestivum</italic> L. cv. Soissons, Maris Widgeon, and Cadenza), which contain flowering ears. Examples of background training patches which do not contain flowering ears.</p></caption>
<graphic xlink:href="fpls-08-00252-g0011.tif"/>
</fig>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p><bold>Comparison of flowering accuracy between three wheat cultivars</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Wheat cultivar</bold></th>
<th valign="top" align="center"><bold>No. training images</bold></th>
<th valign="top" align="center"><bold>No. testing images</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Cadenza</td>
<td valign="top" align="center">410</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">92.91</td>
</tr>
<tr>
<td valign="top" align="left">Soissons</td>
<td valign="top" align="center">410</td>
<td valign="top" align="center">23</td>
<td valign="top" align="center">76.72</td>
</tr>
<tr>
<td valign="top" align="left">Maris widgeon</td>
<td valign="top" align="center">410</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">80.33</td>
</tr>
<tr>
<td valign="top" align="left">All three</td>
<td valign="top" align="center">410</td>
<td valign="top" align="center">54</td>
<td valign="top" align="center">82.54</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.3. Discussion</title>
<p>To the best of our knowledge, few efforts have been made to automate the detection of crop growth stage (Thorp and Dierig, <xref ref-type="bibr" rid="B27">2011</xref>; Yu et al., <xref ref-type="bibr" rid="B35">2013</xref>; Guo et al., <xref ref-type="bibr" rid="B15">2015</xref>; Zhu et al., <xref ref-type="bibr" rid="B38">2016</xref>). Furthermore, the published methods have only been applied to small sections of the crops and generally tested only on a single cultivar. Unlike alternative methods, such as Yu et al. (<xref ref-type="bibr" rid="B35">2013</xref>), which used color properties to determine growth stages of maize, our approach uses rich feature collection techniques, such as SIFT, which carry suitable information to discriminate images at the category level on the canopy scale. The technique used by Guo et al. (<xref ref-type="bibr" rid="B15">2015</xref>) was only tested on two rice varieties individually at flowering stage and obtained just over 80% accuracy. However, our method integrated statistical variables, such as vector coding and spatial pyramid matching, which improved the accuracy and general versatility of the growth stage identification. On the other hand, their training system contained only flowering rice as the positive class and leaves as the negative class; failing to define rice before and after the flowering stage. This may have likely made their dataset more challenging because more variables would be added to the training dataset and distinguishing between non-flowering and flowering panicles would have added difficulty, potentially detecting false positives, ultimately reducing the accuracy of their method.</p>
<p>In our case, the accuracy of flowering detection is less than heading. This could be due to the size and color of anthers. The color of anthers can range from yellow to white depending on the cultivar, and the pale color of the anther has increased the sensitivity to over/under exposure as a result of changes in ambient illumination. Moreover, anthers are far smaller objects compared to wheat ears and are prone to noise, adding difficulty to detection them accurately. Nevertheless, the proposed method yielded greater accuracy than the existing method (Guo et al., <xref ref-type="bibr" rid="B15">2015</xref>).</p>
<p>Pre-processing is also an important factor in our method. Newly emerging ears are difficult to distinguish as they are nearly the same color as the canopy, making methods based on color features inadequate for this purpose. However, the use of color enhancement methods, such as decorrelation stretching, yields higher accuracy. In our case, the absence of decorrelation stretching, results a decrease in accuracy from 95.24% to 57.29% and from 99.59 to 85.38% at earliest and latest stage of heading, respectively. Moreover, applying decorrelation stretching as a color enhancement tool early in the process minimize various ambient light conditions. The other important factor is the low level feature extraction in the BoVW process. SIFT was replaced by SURF as an alternative technique; however, although SURF performs faster as a result of using integral images and Hessian Matrix (Bay et al., <xref ref-type="bibr" rid="B4">2006</xref>); SIFT still outperformed SURF (Table <xref ref-type="table" rid="T2">2</xref>) in our experiment. It has also been examined that SIFT showed more stability on blurry images and more robust to rotation and scale invariants (Mikolajczyk and Schmid, <xref ref-type="bibr" rid="B22">2005</xref>).</p>
<p>It should also be highlighted that the quality of the training dataset plays an important role in the overall performance. We aimed to define more scenarios for the system (e.g., ears at different positions, scales, and illumination conditions in the field, etc.). As shown in Figure <xref ref-type="fig" rid="F11">11</xref>, the accuracy of the ear emergence detection would increase by adding more training data. We would expect to improve the accuracy of the flowering experiment, by collecting data more frequently during the flowering period and increasing the size of the training dataset.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s4">
<title>4. Conclusion</title>
<p>We proposed an automated observing system using computer vision to determine two key growth stages in wheat: ear emergence and flowering time. The proposed method is capable of distinguishing the critical growth stages from the RGB images taken in the field. The approach demonstrated a high performance for identifying such development changes and was not affected by the environmental conditions or illumination invariants in the field.</p>
<p>In future work, we aim to test our proposed method on additional wheat genetic material and other species, and in addition, to investigate the effect of alternative computer vision techniques from features extraction to classification on the performance and overall accuracy. Finally, we aim to apply the proposed method on images acquired by Unmanned Aerial Vehicles (UAVs) to monitor large fields efficiently and believe it will dramatically accelerate the recording of such development stages.</p>
</sec>
<sec id="s5">
<title>Author contributions</title>
<p>PS proposed, developed, and tested the method; KS and NV planned and conducted the experiment; MH contributed to the revision of the manuscript and supervised the experiment; PS, KS, and NV contributed to writing the manuscript; all authors read and approved the final manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>Rothamsted Research receives support from the Biotechnology and Biological Sciences Research Council (BBSRC) of the UK as part of the 20:20 Wheat&#x000AE; project.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alahi</surname> <given-names>A.</given-names></name> <name><surname>Ortiz</surname> <given-names>R.</given-names></name> <name><surname>Vandergheynst</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <source>FREAK: Fast Retina Keypoint, 2012 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Providence, RI</publisher-loc>: <publisher-name>Providence Rhode Island Convention Center)</publisher-name>.</citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alley</surname> <given-names>R.</given-names></name></person-group> (<year>1996</year>). <source>Algorithm Theoretical Basis Document For: Decorrelation Stretch</source>. <publisher-loc>Pasadena, CA</publisher-loc>: <publisher-name>Jet Propulsion Laboratory</publisher-name>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Angelov</surname> <given-names>P.</given-names></name> <name><surname>Sadeghi-Tehran</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Look-a-Like: a fast content-based image retrieval approach using a hierarchically nested dynamically evolving image clouds and recursive local data density</article-title>. <source>Int. J. Intell. Syst.</source> <volume>32</volume>, <fpage>82</fpage>&#x02013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1002/int.21837</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bay</surname> <given-names>H.</given-names></name> <name><surname>Tuytelaars</surname> <given-names>T.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name></person-group> (<year>2006</year>). <article-title>SURF: Speeded Up Robust Features</article-title>, in <source>Computer Vision &#x02013; ECCV 2006: 9th European Conference on Computer Vision Graz, Austria, May 7&#x02013;13, 2006. Proceedings, Part I</source>, eds <person-group person-group-type="editor"><name><surname>Leonardis</surname> <given-names>A.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name> <name><surname>Pinz</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer Berlin Heidelberg</publisher-name>), <fpage>404</fpage>&#x02013;<lpage>417</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Braun</surname> <given-names>H. J.</given-names></name> <name><surname>Atlin</surname> <given-names>G.</given-names></name> <name><surname>Payne</surname> <given-names>T.</given-names></name></person-group> (<year>2010</year>). <article-title>Multi-location testing as a tool to identify plant response to global climate change</article-title>, in <source>Climate Change and Crop Production</source>, ed <person-group person-group-type="editor"><name><surname>Reynolds</surname> <given-names>M. P.</given-names></name></person-group>. <pub-id pub-id-type="doi">10.1079/9781845936334.0115</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>N.</given-names></name></person-group> (<year>1996</year>). <article-title>The decorrelation stretch transformation</article-title>. <source>Int. J. Remote Sens.</source> <volume>17</volume>, <fpage>1939</fpage>&#x02013;<lpage>1949</lpage>. <pub-id pub-id-type="doi">10.1080/01431169608948749</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cerrillo-Cuenca</surname> <given-names>E.</given-names></name> <name><surname>Sep&#x000FA;lveda</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>An assessment of methods for the digital enhancement of rock paintings: the rock art from the precordillera of Arica (Chile) as a case study</article-title>. <source>J. Archaeol. Sci.</source> <volume>55</volume>, <fpage>197</fpage>&#x02013;<lpage>208</lpage>. <pub-id pub-id-type="doi">10.1016/j.jas.2015.01.006</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>W.-T.</given-names></name> <name><surname>Liu</surname> <given-names>W.-C.</given-names></name> <name><surname>Chen</surname> <given-names>M.-S.</given-names></name></person-group> (<year>2010</year>). <article-title>Adaptive color feature extraction based on image color distributions</article-title>. <source>IEEE Trans. Image Process.</source> <volume>19</volume>, <fpage>2005</fpage>&#x02013;<lpage>2016</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2010.2051753</pub-id><pub-id pub-id-type="pmid">20519153</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Csurka</surname> <given-names>G.</given-names></name> <name><surname>Dance</surname> <given-names>C.</given-names></name> <name><surname>Fan</surname> <given-names>L.</given-names></name> <name><surname>Willamowski</surname> <given-names>J.</given-names></name> <name><surname>Bray</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <source>Visual Categorization with Bags of keypoints, Workshop on Statistical Learning in Computer Vision, ECCV.</source> <publisher-loc>Prague</publisher-loc>: <publisher-name>Czech Technical University</publisher-name>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Duda</surname> <given-names>R. O.</given-names></name> <name><surname>Hart</surname> <given-names>P. E.</given-names></name> <name><surname>Stork</surname> <given-names>D. G.</given-names></name></person-group> (<year>2000</year>). <source>Pattern Classification, 2nd Edn.</source>, <publisher-name>Wiley-Interscience</publisher-name>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>R.-E.</given-names></name> <name><surname>Chang</surname> <given-names>K.-W.</given-names></name> <name><surname>Hsieh</surname> <given-names>C.-J.</given-names></name> <name><surname>Wang</surname> <given-names>X.-R.</given-names></name> <name><surname>Lin</surname> <given-names>C.-J.</given-names></name></person-group> (<year>2008</year>). <article-title>LIBLINEAR: a library for large linear classification</article-title>. <source>J. Mach. Learn. Res.</source> <volume>9</volume>, <fpage>1871</fpage>&#x02013;<lpage>1874</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gillespie</surname> <given-names>A. R.</given-names></name> <name><surname>Kahle</surname> <given-names>A. B.</given-names></name> <name><surname>Walker</surname> <given-names>R. E.</given-names></name></person-group> (<year>1986</year>). <article-title>Color enhancement of highly correlated images. I. Decorrelation and HSI contrast stretches</article-title>. <source>Remote Sens. Environ.</source> <volume>20</volume>, <fpage>209</fpage>&#x02013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1016/0034-4257(86)90044-1</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Grauman</surname> <given-names>K.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <article-title>The pyramid match kernel: discriminative classification with sets of image features</article-title>, in <source>Tenth IEEE International Conference on Computer Vision (ICCV&#x00027;05) Volume 1.</source> (<publisher-loc>Beijing</publisher-loc>).</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guerrero</surname> <given-names>J. M.</given-names></name> <name><surname>Pajares</surname> <given-names>G.</given-names></name> <name><surname>Montalvo</surname> <given-names>M.</given-names></name> <name><surname>Romeo</surname> <given-names>J.</given-names></name> <name><surname>Guijarro</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Support vector machines for crop/weeds identification in maize fields</article-title>. <source>Exp. Syst. Appl.</source> <volume>39</volume>, <fpage>11149</fpage>&#x02013;<lpage>11155</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2012.03.040</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>W.</given-names></name> <name><surname>Fukatsu</surname> <given-names>T.</given-names></name> <name><surname>Ninomiya</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Automated characterization of flowering dynamics in rice using field-acquired time-series RGB images</article-title>. <source>Plant Methods</source> <volume>11</volume>:<fpage>7</fpage>. <pub-id pub-id-type="doi">10.1186/s13007-015-0047-9</pub-id><pub-id pub-id-type="pmid">25705245</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jolliffe</surname> <given-names>I. T.</given-names></name></person-group> (<year>2002</year>). <article-title>Principal component analysis and factor analysis</article-title>, in <source>Principal Component Analysis</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>150</fpage>&#x02013;<lpage>166</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lazebnik</surname> <given-names>S.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name> <name><surname>Ponce</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <source>Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR&#x00027;06)</source> (<publisher-loc>New York, NY</publisher-loc>).</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Leutenegger</surname> <given-names>S.</given-names></name> <name><surname>Chli</surname> <given-names>M.</given-names></name> <name><surname>Siegwart</surname> <given-names>R. Y.</given-names></name></person-group> (<year>2011</year>). <source>BRISK: Binary Robust invariant scalable keypoints, 2011 International Conference on Computer Vision</source> (<publisher-loc>Barcelona</publisher-loc>).</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lloyd</surname> <given-names>S.</given-names></name></person-group> (<year>1982</year>). <article-title>Least squares quantization in PCM</article-title>. <source>IEEE Trans. Inform. Theory</source> <volume>28</volume>, <fpage>129</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1109/TIT.1982.1056489</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lowe</surname> <given-names>D. G.</given-names></name></person-group> (<year>2004</year>). <article-title>Distinctive image features from scale-invariant keypoints</article-title>. <source>Int. J. Comput. Vis.</source> <volume>60</volume>, <fpage>91</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1023/B:VISI.0000029664.99615.94</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehrotra</surname> <given-names>R.</given-names></name> <name><surname>Gary</surname> <given-names>J. E.</given-names></name></person-group> (<year>1995</year>). <article-title>Similar-shape retrieval in shape data management</article-title>. <source>Computer</source> <volume>28</volume>, <fpage>57</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1109/2.410154</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolajczyk</surname> <given-names>K.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title>A performance evaluation of local descriptors</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>27</volume>, <fpage>1615</fpage>&#x02013;<lpage>1630</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2005.188</pub-id><pub-id pub-id-type="pmid">16237996</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pourreza</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>W. S.</given-names></name> <name><surname>Etxeberria</surname> <given-names>E.</given-names></name> <name><surname>Banerjee</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>An evaluation of a vision-based sensor performance in huanglongbing disease identification</article-title>. <source>Biosyst. Eng.</source> <volume>130</volume>, <fpage>13</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2014.11.013</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Snape</surname> <given-names>J.</given-names></name> <name><surname>Butterworth</surname> <given-names>K.</given-names></name> <name><surname>Whitechurch</surname> <given-names>E.</given-names></name> <name><surname>Worland</surname> <given-names>A. J.</given-names></name></person-group> (<year>2001</year>). <article-title>Waiting for fine times: genetics of flowering time in wheat</article-title>, in <source>Wheat in a Global Environment: Proceedings of the 6th International Wheat Conference, Budapest, Hungary</source> eds <person-group person-group-type="editor"><name><surname>Bed&#x000F6;</surname> <given-names>Z.</given-names></name> <name><surname>L&#x000E1;ng</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>67</fpage>&#x02013;<lpage>74</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swain</surname> <given-names>M. Z.</given-names></name> <name><surname>Ballard</surname> <given-names>D. H.</given-names></name></person-group> (<year>1991</year>). <article-title>Color indexing</article-title>. <source>Int. J. Comput. Vis.</source> <volume>7</volume>, <fpage>11</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1007/BF00130487</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tester</surname> <given-names>M.</given-names></name> <name><surname>Langridge</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Breeding technologies to increase crop production in a changing world</article-title>. <source>Science</source> <volume>327</volume>, <fpage>818</fpage>&#x02013;<lpage>822</lpage>. <pub-id pub-id-type="doi">10.1126/science.1183700</pub-id><pub-id pub-id-type="pmid">20150489</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thorp</surname> <given-names>K. R.</given-names></name> <name><surname>Dierig</surname> <given-names>D. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Color image segmentation approach to monitor flowering in lesquerella</article-title>. <source>Indust. Crops Products</source> <volume>34</volume>, <fpage>1150</fpage>&#x02013;<lpage>1159</lpage>. <pub-id pub-id-type="doi">10.1016/j.indcrop.2011.04.002</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tilman</surname> <given-names>D.</given-names></name> <name><surname>Balzer</surname> <given-names>C.</given-names></name> <name><surname>Hill</surname> <given-names>J.</given-names></name> <name><surname>Befort</surname> <given-names>B. L.</given-names></name></person-group> (<year>2011</year>). <article-title>Global food demand and the sustainable intensification of agriculture</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>108</volume>, <fpage>20260</fpage>&#x02013;<lpage>20264</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1116437108</pub-id><pub-id pub-id-type="pmid">22106295</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Valiente-Gonz&#x000E1;lez</surname> <given-names>J. M.</given-names></name> <name><surname>Andreu-Garc&#x000ED;a</surname> <given-names>G.</given-names></name> <name><surname>Potter</surname> <given-names>P.</given-names></name> <name><surname>Rodas-Jord&#x000E1;</surname> <given-names>&#x000C1;.</given-names></name></person-group> (<year>2014</year>). <article-title>Automatic corn (<italic>Zea mays</italic>) kernel inspection system using novelty detection based on principal component analysis</article-title>. <source>Biosyst. Eng.</source> <volume>117</volume>, <fpage>94</fpage>&#x02013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2013.09.003</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vedaldi</surname> <given-names>A.</given-names></name> <name><surname>Fulkerson</surname> <given-names>B.</given-names></name></person-group> (<year>2010</year>). <article-title>Vlfeat: an open and portable library of computer vision algorithms</article-title>, in <source>Proceedings of the 18th ACM International Conference on Multimedia</source> (<publisher-loc>Firenze</publisher-loc>), <fpage>1469</fpage>&#x02013;<lpage>1472</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vedaldi</surname> <given-names>A.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Efficient additive kernels via explicit feature maps</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>34</volume>, <fpage>480</fpage>&#x02013;<lpage>492</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2011.153</pub-id><pub-id pub-id-type="pmid">21808094</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Virlet</surname> <given-names>N.</given-names></name> <name><surname>Sabermanesh</surname> <given-names>K.</given-names></name> <name><surname>Sadeghi-Tehran</surname> <given-names>P.</given-names></name> <name><surname>Hawkesford</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Field Scanalyzer: an automated robotic field phenotyping platform for detailed crop monitoring</article-title>. <source>Funct. Plant Biol.</source> <volume>44</volume>, <fpage>143</fpage>&#x02013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1071/FP16163</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name> <name><surname>Lv</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>T.</given-names></name> <name><surname>Gong</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <source>Locality-constrained Linear Coding for image classification, 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>San Francisco, CA</publisher-loc>).</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yihang</surname> <given-names>F.</given-names></name> <name><surname>Tingting</surname> <given-names>C.</given-names></name> <name><surname>Ruifang</surname> <given-names>Z.</given-names></name> <name><surname>Xingyu</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <source>Automatic recognition of rape seeding emergence stage based on computer vision technology, 2014 The Third International Conference on Agro-Geoinformatics.</source> <publisher-loc>Beijing</publisher-loc>: <publisher-name>Friendship Hotel Beijing</publisher-name>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name> <name><surname>Qin</surname> <given-names>Y.</given-names></name> <name><surname>Zhuo</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Automatic image-based detection technology for two critical growth stages of maize: emergence and three-leaf stage</article-title>. <source>Agricult. For. Meteorol.</source> <fpage>174</fpage>&#x02013;<lpage>175</lpage>, 65&#x02013;84. <pub-id pub-id-type="doi">10.1016/j.agrformet.2013.02.011</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zadoks</surname> <given-names>J. C.</given-names></name> <name><surname>Chang</surname> <given-names>T. T.</given-names></name> <name><surname>Konzak</surname> <given-names>C. F.</given-names></name></person-group> (<year>1974</year>). <article-title>A decimal code for the growth stages of cereals</article-title>. <source>Weed Res.</source> <volume>14</volume>, <fpage>415</fpage>&#x02013;<lpage>421</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-3180.1974.tb01084.x</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Tian</surname> <given-names>J.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Chen</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Detection of quantitative trait loci for heading date based on the doubled haploid progeny of two elite chinese wheat cultivars</article-title>. <source>Genetica</source> <volume>135</volume>, <fpage>257</fpage>&#x02013;<lpage>265</lpage>. <pub-id pub-id-type="doi">10.1007/s10709-008-9274-6</pub-id><pub-id pub-id-type="pmid">18500653</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>In-field automatic observation of wheat heading stage using computer vision</article-title>. <source>Biosyst. Eng.</source> <volume>143</volume>, <fpage>28</fpage>&#x02013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2015.12.015</pub-id></citation>
</ref>
</ref-list>
<glossary>
<def-list>
<title>Abbreviations</title>
<def-item><term>BoVW</term>
<def><p>bag of visual words</p></def></def-item>
<def-item><term>SVM</term>
<def><p>support vector machine</p></def></def-item>
<def-item><term>SPM</term>
<def><p>spatial pyramid matching</p></def></def-item>
<def-item><term>RBF</term>
<def><p>radius basis kernel</p></def></def-item>
<def-item><term>LLC</term>
<def><p>local linear constraint</p></def></def-item>
<def-item><term>SIFT</term>
<def><p>scale invariant feature transform</p></def></def-item>
<def-item><term>DoG</term>
<def><p>difference-of-gaussian</p></def></def-item>
<def-item><term>LoG</term>
<def><p>laplacian-of-gaussian</p></def></def-item>
<def-item><term>PCA</term>
<def><p>principle component analysis</p></def></def-item>
<def-item><term>SURF</term>
<def><p>speeded up robust features</p></def></def-item>
<def-item><term>KNN</term>
<def><p>k-nearest-neighbors</p></def></def-item>
<def-item><term>UAV</term>
<def><p>unmanned aerial vehicle.</p></def></def-item>
</def-list>
</glossary>
</back>
</article>