<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2016.00594</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Motion-Based Feature for Event-Based Pattern Recognition</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Clady</surname> <given-names>Xavier</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/114467/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Maro</surname> <given-names>Jean-Matthieu</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/376783/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Barr&#x000E9;</surname> <given-names>S&#x000E9;bastien</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Benosman</surname> <given-names>Ryad B.</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/94237/overview"/>
</contrib>
</contrib-group>
<aff><institution>Centre National de la Recherche Scientifique, Institut National de la Sant&#x000E9; Et de la Recherche M&#x000E9;dicale, Institut de la Vision, Sorbonne Universit&#x000E9;s, UPMC University Paris 06</institution> <country>Paris, France</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Tobi Delbruck, ETH Zurich, Switzerland</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Dan Hammerstrom, Portland State University, USA; Rodrigo Alvarez-Icaza, IBM, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Xavier Clady <email>xavier.clady&#x00040;upmc.fr</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Neuromorphic Engineering, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>10</volume>
<elocation-id>594</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>09</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>12</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Clady, Maro, Barr&#x000E9; and Benosman.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Clady, Maro, Barr&#x000E9; and Benosman</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>This paper introduces an event-based luminance-free feature from the output of asynchronous event-based neuromorphic retinas. The feature consists in mapping the distribution of the optical flow along the contours of the moving objects in the visual scene into a matrix. Asynchronous event-based neuromorphic retinas are composed of autonomous pixels, each of them asynchronously generating &#x0201C;spiking&#x0201D; events that encode relative changes in pixels&#x00027; illumination at high temporal resolutions. The optical flow is computed at each event, and is integrated locally or globally in a speed and direction coordinate frame based grid, using speed-tuned temporal kernels. The latter ensures that the resulting feature equitably represents the distribution of the normal motion along the current moving edges, whatever their respective dynamics. The usefulness and the generality of the proposed feature are demonstrated in pattern recognition applications: local corner detection and global gesture recognition.</p>
</abstract>
<kwd-group>
<kwd>neuromorphic sensor</kwd>
<kwd>event-driven vision</kwd>
<kwd>pattern recognition</kwd>
<kwd>motion-based feature</kwd>
<kwd>speed-tuned integration time</kwd>
<kwd>histogram of oriented optical flow</kwd>
<kwd>corner detection</kwd>
<kwd>gesture recognition</kwd>
</kwd-group>
<contract-sponsor id="cn001">Horizon 2020 Framework Programme<named-content content-type="fundref-id">10.13039/100010661</named-content></contract-sponsor>
<counts>
<fig-count count="15"/>
<table-count count="2"/>
<equation-count count="24"/>
<ref-count count="90"/>
<page-count count="20"/>
<word-count count="12594"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>In computer vision, a feature is a more or less compact representation of visual information that is relevant to solve a task related to a given application (see Laptev, <xref ref-type="bibr" rid="B46">2005</xref>; Mikolajczyk and Schmid, <xref ref-type="bibr" rid="B52">2005</xref>; Mokhtarian and Mohanna, <xref ref-type="bibr" rid="B55">2006</xref>; Moreels and Perona, <xref ref-type="bibr" rid="B58">2007</xref>; Gil et al., <xref ref-type="bibr" rid="B29">2010</xref>; Dickscheid et al., <xref ref-type="bibr" rid="B22">2011</xref>; Gauglitz et al., <xref ref-type="bibr" rid="B27">2011</xref>). Building a feature consists in encoding information contained in the visual scene (global approach) or in a neighborhood of a point (local approach). It can represent static information (e.g., shape of an object, contour, etc.), dynamic information (e.g., speed and direction at the point, dynamic deformations, etc.) or both simultaneously.</p>
<p>In this article, we propose a motion-based feature computed on visual information provided by asynchronous image sensors known as neuromorphic retinas (see Delbr&#x000FC;ck et al., <xref ref-type="bibr" rid="B20">2010</xref>; Posch, <xref ref-type="bibr" rid="B74">2015</xref>). These cameras provide visual information as asynchronous event-based streams while conventional cameras output it as synchronous frame-based streams. The ATIS (&#x0201C;Asynchronous Time-based Image Sensor,&#x0201D; Posch et al., <xref ref-type="bibr" rid="B75">2010</xref>; Posch, <xref ref-type="bibr" rid="B74">2015</xref>), one of the neuromorphic visual sensors used in this work, is a time-domain encoding image sensor with QVGA resolution. It contains an array of fully autonomous pixels that combine an illuminance change detector circuit, associated to the PD1 photodiode, see Figure <xref ref-type="fig" rid="F1">1A</xref> and a conditional exposure measurement block, associated to the PD2 photodiode. The change detector individually and asynchronously initiates the measurement of an exposure/gray scale value only if a brightness change of a certain magnitude has been detected in the field-of-view of the respective pixel, as shown in the functional diagram of the ATIS pixel in Figures <xref ref-type="fig" rid="F1">1B</xref>, <xref ref-type="fig" rid="F2">2</xref>. The exposure measurement circuit encodes the absolute instantaneous pixel illuminance into the timing of asynchronous event pulses, more precisely into inter-event intervals. The DVS (&#x0201C;Dynamic Visual Sensor,&#x0201D; Lichtsteiner et al., <xref ref-type="bibr" rid="B50">2008</xref>; Serrano-Gotarredona and Linares-Barranco, <xref ref-type="bibr" rid="B83">2013</xref>), another neuromorphic camera used in this work, works in a similar manner but only the illuminance change detector is implemented and retina&#x00027;s spatial resolution is limited to 128 &#x000D7; 128<italic>pixels</italic>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>ATIS, Asynchronous Time-based Image Sensor: (A)</bold> The ATIS and its pixel array, made of 304 &#x000D7; 240 pixels (QVGA). PD1 is the change detector, PD2 is the grayscale measurement unit. <bold>(B)</bold> When a contrast change occurs in the visual scene, the ATIS outputs a change event (ON or OFF) and a grayscale event. <bold>(C)</bold> The spatio-temporal space of imaging events: static objects and scene background are acquired first. Then, dynamic objects trigger pixel-individual, asynchronous gray-level events after each change. Frames are absent from this acquisition process. Samples of generated images from the presented spatio-temporal space are shown in the upper part of the figure.</p></caption>
<graphic xlink:href="fnins-10-00594-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Illustration of the luminance change measured by a neuromorphic pixel, modeled as a cone-pixel (Debaecker et al., <xref ref-type="bibr" rid="B18">2010</xref>), viewing an moving edge</bold>. <bold>p</bold> are the coordinates of the center of the pixel, in the coordinate system related to the retina&#x00027;s plane. The emitted event in response to this luminance change is represented as a black dot in the coordinate system <italic>XYt</italic> related to the event space (in the lower-left part of the figure).</p></caption>
<graphic xlink:href="fnins-10-00594-g0002.tif"/>
</fig>
<p>Despite the recent introduction of neuromorphic cameras, numerous applications have already emerged in robotics (see Censi et al., <xref ref-type="bibr" rid="B11">2013</xref>; Delbr&#x000FC;ck and Lang, <xref ref-type="bibr" rid="B19">2013</xref>; Lagorce et al., <xref ref-type="bibr" rid="B43">2013</xref>; Clady et al., <xref ref-type="bibr" rid="B15">2014</xref>; Ni et al., <xref ref-type="bibr" rid="B63">2014</xref>; Milde et al., <xref ref-type="bibr" rid="B53">2015</xref>), shape tracking (see Drazen et al., <xref ref-type="bibr" rid="B23">2011</xref>; Ni et al., <xref ref-type="bibr" rid="B62">2015</xref>; Valeiras et al., <xref ref-type="bibr" rid="B86">2015</xref>), stereovision (cf. Rogister et al., <xref ref-type="bibr" rid="B78">2012</xref>; Carneiro et al., <xref ref-type="bibr" rid="B8">2013</xref>; Camu&#x000F1;as-Mesa et al., <xref ref-type="bibr" rid="B7">2014</xref>; Firouzi and Conradt, <xref ref-type="bibr" rid="B24">2015</xref>), corner detection (Clady et al., <xref ref-type="bibr" rid="B16">2015</xref>), or shape recognition (see P&#x000E9;rez-Carrasco et al., <xref ref-type="bibr" rid="B71">2013</xref>; Akolkar et al., <xref ref-type="bibr" rid="B2">2015</xref>; Orchard et al., <xref ref-type="bibr" rid="B67">2015a</xref>,<xref ref-type="bibr" rid="B68">b</xref>; Lee et al., <xref ref-type="bibr" rid="B48">2016</xref>). This strong interest in such a sensor is essentially due to its ability to provide visual information as a high temporal resolution, luminance-free, and non-redundant stream. This makes it a fitting for high-speed applications [e.g., gesture recognition as in Lee et al. (<xref ref-type="bibr" rid="B49">2014</xref>), high-speed object tracking as in Lagorce et al. (<xref ref-type="bibr" rid="B44">2014</xref>), Mueggler et al. (<xref ref-type="bibr" rid="B59">2015a</xref>)].</p>
<p>The proposed feature consists in mapping the distribution of the optical flow along the contours of the objects in the visual scene into a matrix (see Section 2). It can be computed locally or more globally according to the targeted applications. Indeed, in the experimental evaluations, we propose to demonstrate its usefulness and generality in various applications. It is used to locally detect corners (see Section 3) or to summarize global motion observed in a scene in order to recognize actions, here hand gestures for an application in human-machine interaction (see Section 4).</p>
</sec>
<sec id="s2">
<title>2. Motion-based feature</title>
<p>Visual event streams are generated asynchronously at a high temporal resolution, essentially by moving edges. They are thus especially suitable for visual motion flow or optical flow (OF) computation (Benosman et al., <xref ref-type="bibr" rid="B5">2014</xref>; Orchard and Etienne-Cummings, <xref ref-type="bibr" rid="B66">2014</xref>; Brosch et al., <xref ref-type="bibr" rid="B6">2015</xref>) along contours of objects. In the following sections, methods and mechanisms are proposed to estimate normal motion flows computed around events and to map them into a matrix in order to incrementally estimate scene motion distribution (locally or globally). This matrix will be considered as a feature. Its computation requires only the visual events provided by the change detectors of the retina (associated to photodiodes PD1 in Figure <xref ref-type="fig" rid="F1">1A</xref>), that can be defined as four components vectors:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold"><mml:mtext>e</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <bold>p</bold> &#x0003D; (<italic>x, y</italic>)<sup><italic>T</italic></sup> is the spatial coordinate of each event, <italic>t</italic>, its timestamp and <italic>pol</italic> &#x02208; {&#x02212;1, 1} is the polarity, which is equal to &#x02212;1/1 when the measured luminance decrease/increase is significant enough (see upper part of Figure <xref ref-type="fig" rid="F1">1B</xref>).</p>
<sec>
<title>2.1. Extracting normal visual motion</title>
<p>We use the event-based OF computation method proposed in Benosman et al. (<xref ref-type="bibr" rid="B5">2014</xref>) which is known for its robustness and its algorithmic efficiency (see Clady et al., <xref ref-type="bibr" rid="B15">2014</xref>, <xref ref-type="bibr" rid="B16">2015</xref>; Mueggler et al., <xref ref-type="bibr" rid="B60">2015b</xref>). More bio-inspired event-based OF computation methods such as Brosch et al. (<xref ref-type="bibr" rid="B6">2015</xref>) and Orchard and Etienne-Cummings (<xref ref-type="bibr" rid="B66">2014</xref>) can be used but they are computationally more expensive.</p>
<p>A function &#x003A3;<sub><italic>e</italic></sub> that maps to each <bold>p</bold> the time <italic>t</italic> is defined locally:</p>
<disp-formula id="E2"><mml:math id="M2"><mml:msub><mml:mrow><mml:mo>&#x003A3;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:mtable class="array"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle mathvariant='bold'><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>&#x021A6;</mml:mo><mml:mi>t</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Applying the inverse function theorem of calculus, the vector &#x02207;&#x003A3;<sub><italic>e</italic></sub> measures the rate and the direction of change of time with respect to space: it is the normal optical flow, noted <inline-formula><mml:math id="M3"><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, such as:</p>
<disp-formula id="E3"><mml:math id="M4"><mml:mo>&#x02207;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A3;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup></mml:math></disp-formula>
<p>This equation could be defined assuming that the surface described by the visual events (generated by a moving edge) in the space-time reference frame (<italic>XYt</italic>)<sup><italic>T</italic></sup> is continuous. This assumption is validated through a regularization process proposed in order to locally estimate this surface as a spatiotemporal plane (fitted directly on the local event stream). In this work the implementation proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) has been chosen because it proposes mechanisms to automatically adapt the temporal dimension of the local neighborhood to the edge&#x00027;s dynamics, and to reject estimations of optical flow probably wrong and due to noise. This algorithm allows us to consider a function that associates for each valid visual event <inline-formula><mml:math id="M5"><mml:mstyle mathvariant="bold"><mml:mtext>e</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">E</mml:mi></mml:mrow></mml:math></inline-formula>, a so-called visual motion event, noted <bold>v</bold><sub><bold>e</bold></sub>, such as:</p>
<disp-formula id="E4"><label>(2)</label><mml:math id="M6"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mtable class="array"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="-tex-caligraphic">E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold"><mml:mtext>e</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x021A6;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>e</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where (<italic>v</italic>, &#x003B8;)<sup><italic>T</italic></sup> corresponds to the intensity (i.e., speed) and the direction of the normal visual flow.</p>
<p>Remark 1. <italic>Note that the polarity of visual events is not conserved by the function (Equation 2). Indeed, in the applications proposed in this article, it is not useful to &#x0201C;memorize&#x0201D; if the visual flow has been computed on a positive or negative event stream. If required, the feature can be augmented in order to distinguish the distribution along &#x0201C;positive contours&#x0201D; from the one along &#x0201C;negative contours</italic>.&#x0201D;</p>
</sec>
<sec>
<title>2.2. Computing and updating the feature</title>
<p>As we said, the feature corresponds to the estimated distribution of the optical flow along the (local or global) contours in the visual scene. This distribution is evaluated on a grid-based sampling in the polar reference frame of the visual flow, such as it is subdivided into an interval set <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> where <inline-formula><mml:math id="M8"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> is an angle based interval and <inline-formula><mml:math id="M9"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> is an intensity based interval. Such a discretization of the velocity subspace is consistent with biologic observations about orientation (cf. Hubel and Wiesel, <xref ref-type="bibr" rid="B34">1962</xref>, <xref ref-type="bibr" rid="B35">1968</xref>) and speed (cf. Priebe et al., <xref ref-type="bibr" rid="B77">2006</xref>) selectivity in V1 cells and human psychophysical experiments about speed discrimination as in Orban et al. (<xref ref-type="bibr" rid="B65">1984</xref>) and Kime et al. (<xref ref-type="bibr" rid="B40">2014</xref>, <xref ref-type="bibr" rid="B39">2016</xref>). Here, we parametrize the grid sampling mostly according to these biologic observations and human psychophysical experiments. However, its ranges and precisions could be set in relation with targeted tasks, optimizing them according to given performance criteria. We define the centers <inline-formula><mml:math id="M10"><mml:msub><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the angle intervals such as: <inline-formula><mml:math id="M11"><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x003C0;</mml:mo><mml:mfrac><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x003C0;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, with <italic>i</italic> &#x02208; [0, <italic>N</italic><sub>&#x003B8;</sub> &#x02212; 1]; <inline-formula><mml:math id="M12"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x003C0;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> is the length of the interval and thus the angular precision of the grid. With <italic>N</italic><sub>&#x003B8;</sub> &#x0003D; 36, we barely reach the precision (&#x0007E;10&#x000B0;) observed for V1 simple cells (see Hubel and Wiesel, <xref ref-type="bibr" rid="B34">1962</xref>, <xref ref-type="bibr" rid="B35">1968</xref>). For the velocity intensity, we propose a non-regular speed-based sampling, where {<italic>v</italic><sup><italic>l</italic></sup>} are the centers of the speed based intervals on a logarithmic scale. The sampling is then operated such that <inline-formula><mml:math id="M13"><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, with <italic>i</italic> &#x02208; [0, <italic>N</italic><sub><italic>v</italic></sub> &#x02212; 1] and &#x003B3; &#x0003D; 1 &#x0002B; &#x003F5;<sub><italic>v</italic></sub> (&#x003F5;<sub><italic>v</italic></sub> &#x0003E; 0). This discretization strategy ensures an <italic>a priori</italic> constant relative precision in speed estimation: <inline-formula><mml:math id="M14"><mml:mfrac><mml:mrow><mml:mo>&#x00394;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003F5;</mml:mo></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Setting &#x003F5;<sub><italic>v</italic></sub> to 0.1 will barely correspond to the relative speed-discrimination threshold (10%) observed in human psychophysical experiments (see Orban et al., <xref ref-type="bibr" rid="B65">1984</xref>; Kime et al., <xref ref-type="bibr" rid="B40">2014</xref>, <xref ref-type="bibr" rid="B39">2016</xref>). <italic>v</italic><sub><italic>min</italic></sub> has been fixed to 1<italic>pixel</italic>.<italic>s</italic><sup>&#x02212;1</sup> and <italic>N</italic><sub><italic>v</italic></sub> to 73 in order that <inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is close to 1000<italic>pixels</italic>.<italic>s</italic><sup>&#x02212;1</sup>, i.e., inversely close to the temporal precision of the visual events, estimated over 1 ms (cf. Akolkar et al., <xref ref-type="bibr" rid="B2">2015</xref>). Motions with intensities less than <italic>v</italic><sub><italic>min</italic></sub> are then discarded: they are assumed as belonging to static or faraway objects in the background visual scene. Motions with intensities higher than <italic>v</italic><sub><italic>max</italic></sub> are also discarded because noise associated to their computation can <italic>a priori</italic> be considered as too high.</p>
<p>Finally, the feature, noted <inline-formula><mml:math id="M16"><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:math></inline-formula>, is defined as a matrix corresponding to this grid, and associated to a spatiotemporal point (<bold>p</bold>, <italic>t</italic>)<sup><italic>T</italic></sup> of the retina (or to the entire visual scene for a global approach), and computed as:</p>
<disp-formula id="E5"><label>(3)</label><mml:math id="M17"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mtable class="array"><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x000A0;&#x02003;&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">F</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>&#x021A6;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;&#x02003;&#x000A0;</mml:mtext><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where:
<list list-type="bullet">
<list-item><p><italic>w</italic><sub><italic>t</italic></sub> is a temporal exponentially decay function (or kernel), inspired by the synchrony measure of spike trains proposed in van Rossum (<xref ref-type="bibr" rid="B87">2001</xref>), such that:
<disp-formula id="E6"><label>(4)</label><mml:math id="M18"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where <italic>H</italic>(&#x000B7;) is the Heaviside step function and &#x003B1; parametrizes the global decreasing dynamic. In our experiments (see Sections 3 and 4), we fixed &#x003B1; to 0.8, i.e., close to 1 in order to mostly take into account the current edges while slightly smoothing them in order to make <bold>F</bold> less sensitive to both noise and missing data. This kernel gives indeed more weight (or a higher probability value) to events generated by current edges, i.e., the events with timings close to <italic>t</italic>, while also respecting an isoprobabilistic representation of the edges whatever their dynamics, as we will discuss below (see Section 2.3). Of course, other temporal kernels [Gaussian-based in Schreiber et al. (<xref ref-type="bibr" rid="B82">2003</xref>),&#x02026;] can be envisioned, but this one has the advantage of being causal and of leading to an incremental computation of the feature (see Equation 7).</p></list-item>
<list-item><p><italic>w</italic><sub><italic>s</italic></sub> is a spatial bivariate function, which can be defined as:</p>
<list list-type="order">
<list-item><p>in a global approach, <italic>w</italic><sub><italic>s</italic></sub>(<bold>p</bold> &#x02212; <bold>p</bold><sub><italic>j</italic></sub>) &#x0003D; 1, which gives an equitable representation to the edges whatever their spatial locations, or</p></list-item>
<list-item><p>in a local approach:
<disp-formula id="E7"><label>(5)</label><mml:math id="M19"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x003C0;</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mo>&#x003C3;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where &#x003C3;<sub><italic>s</italic></sub> implicitly parametrizes the spatial scale of a region of interest or neighborhood around the spatial location <bold>p</bold>; <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> then represents the local distribution of the normal velocities around the spatiotemporal location (<bold>p</bold>, <italic>t</italic>)<sup><italic>T</italic></sup>;</p></list-item>
</list>
</list-item>
<list-item><p><italic>w</italic><sub><italic>v</italic></sub> is the multiplication of two univariate Gaussian-like functions used to take into account potential imprecisions in the computation of the optical flow, defined as:
<disp-formula id="E8"><label>(6)</label><mml:math id="M20"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>with <inline-formula><mml:math id="M21"><mml:msup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in order to consider a relative speed imprecision, and &#x00398; set to 20&#x000B0;. So, even if an estimated motion belongs to a wrong interval because of noise, it will still contribute to the right element of the matrix, probably close.</p></list-item>
</list></p>
<p>As we said previously, the feature can be incrementally updated at each occurring visual motion event <bold>v</bold><sub><italic>i</italic></sub>, considering that <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> for all (<italic>v</italic><sup><italic>l</italic></sup>, &#x003B8;<sup><italic>l</italic></sup>)<sup><italic>T</italic></sup> (in order to consider, at time <italic>t</italic> &#x0003D; 0, an uniform distribution for the considered velocity-space), such as:</p>
<disp-formula id="E9"><label>(7)</label><mml:math id="M23"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>F</mml:mi></mml:mstyle><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mo>&#x003B8;</mml:mo><mml:mi>l</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>F</mml:mi></mml:mstyle><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mo>&#x003B8;</mml:mo><mml:mi>l</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;&#x02009;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msub><mml:mi>w</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>i</mml:mi></mml:mstyle></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mo>&#x003B8;</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mo>&#x003B8;</mml:mo><mml:mi>l</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Remark 2. <italic>The feature works like a voting matrix, i.e., each visual motion event votes for the speed and direction interval it belongs (and its neighboring intervals through the weighting kernel <italic>w</italic><sub><italic>v</italic></sub>, Equation 6). More visual events there are, more robust the feature will be. Conversely, the feature will be more sensitive to noise in low light or low contrast situations</italic>.</p>
<p><italic>In addition the feature</italic> <bold>F</bold> <italic>can be related to a probabilistic distribution while normalizing it to sum up to 1, i.e., to divide it with <inline-formula><mml:math id="M24"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munder><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></italic>.</p>
<p><italic>In the global approach</italic>, <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> <italic>is independent of</italic> <bold>p</bold>; <italic>it can then be noted</italic> <bold>F</bold><sub><italic>t</italic></sub><italic>. Note that the feature is noted</italic> <bold>F</bold> <italic>(without sub-index) in this article when the application context (local or global approach) is not relevant or obvious</italic>.</p>
</sec>
<sec>
<title>2.3. Speed-tuned vs. fixed decreasing strategies</title>
<p>Another important point to highlight is that the temporal decreasing function <italic>w</italic><sub><italic>t</italic></sub> (Equation 4) is related to the speed <italic>v</italic><sup><italic>l</italic></sup>. Indeed, <inline-formula><mml:math id="M25"><mml:msup><mml:mrow><mml:mo>&#x003C4;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></inline-formula> is the time during which an edge travels through a pixel or in other words, the estimated lifetime of its observation at a given location <bold>p</bold>, as already remarked in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) and Mueggler et al. (<xref ref-type="bibr" rid="B60">2015b</xref>). Including it as decay factor in the temporal kernel (Equations 4 and 7) provides a more isoprobabilistic representation of the moving edges in <bold>F</bold>, i.e., depending only of their contrasts whatever their respective dynamics.</p>
<p>In order to concretely illustrate this point, Figure <xref ref-type="fig" rid="F3">3</xref> represents two synchrony images <italic>I</italic> built integrating a visual event stream and with two different strategies for decay factor &#x003C4; (related to the speed or not), such as, for each occurring visual event <bold>e</bold><sub><italic>i</italic></sub>, <inline-formula><mml:math id="M29"><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003C4;</mml:mo></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mo>&#x003B4;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x0007C;</mml:mo><mml:mo>&#x0007C;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>p</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>i</mml:mi></mml:mstyle></mml:msub><mml:mo>&#x0007C;</mml:mo><mml:mo>&#x0007C;</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> where &#x003B4;(&#x000B7;) is the Dirac function. The left image (Figure <xref ref-type="fig" rid="F3">3A</xref>) results from this equation with a constant &#x003C4; &#x0003D; cst (whatever the dynamics of the edges), and the middle image (Figure <xref ref-type="fig" rid="F3">3B</xref>) with a speed-tuned <inline-formula><mml:math id="M30"><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula>. As shown in the right image (Figure <xref ref-type="fig" rid="F3">3C</xref>), which is the subtraction of both previous images without a speed-tuned factor the high-velocity edges (resulting from the moving and forward leg) are over-represented and the low-velocity edges (resulting from the backward leg) are under-represented in the corresponding synchrony image (Figure <xref ref-type="fig" rid="F3">3A</xref>). The moving edges are more equitably represented in the second synchrony image (Figure <xref ref-type="fig" rid="F3">3B</xref>) with a speed-tuned temporal kernel and, by extension, in feature <bold>F</bold>. Results in Section 3.2 show this equitable representation is very important to obtain accurate results.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Illustration of different strategies for the exponential decay function; comparison between synchrony images <italic><bold>I</bold></italic> built applying exponential temporal kernels with a constant decreasing factor (A)</bold> and with a speed-tuned decreasing factor <bold>(B)</bold> to an event stream (acquired from a visual scene containing a walking person). The comparison <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">cst</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msub></mml:math></inline-formula> <bold>(C)</bold> of both images shows that the second strategy provides a more isoprobabilistic representation of the edges (taking into account the observation lifetime of the moving edges as in Clady et al., <xref ref-type="bibr" rid="B16">2015</xref>; Mueggler et al., <xref ref-type="bibr" rid="B60">2015b</xref>) than the first one; the high-velocity edges (resulting from the moving and forward leg) are over-represented and the low-velocity edges (resulting from the backward leg) are under-represented in the left synchrony image. <bold>(A)</bold> <italic>I</italic><sub>&#x003C4; &#x0003D; cst</sub> with a fixed decreasing factor. <bold>(B)</bold> <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msub></mml:math></inline-formula> with a speed-tuned decreasing factor. <bold>(C)</bold> comparison <inline-formula><mml:math id="M28"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">cst</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msub></mml:math></inline-formula>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0003.tif"/>
</fig>
<p>The proposed strategy is also consistent with biological observations. Indeed Bair and Movshon (<xref ref-type="bibr" rid="B4">2004</xref>) showed that the effective integration time of the computations in direction-selective cells changes with stimulus speed; the integration time for slow motions is longer than that for fast motions. This is modeled in Equation (4) as a decay factor inversely proportional to the speed intensity.</p>
<p>The organization of the feature in a polar coordinate frame based grid, greatly facilitates its computation and its update. The representation of the visual motion information into speed and direction coordinates grants that each speed-tuned decay factor can be associated to an element of the grid, and not directly to the velocity associated to the occurring visual motion event. The latter indicates only which elements in the grid have to be incremented. A bio-inspired implementation can be envisioned where visual motion events are conveyed by selective lines (each line conveying only the motion events <bold>v</bold><sub><bold>e</bold></sub> included in its associated interval, <inline-formula><mml:math id="M34"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>) from a neuron layer computing the optical flow to a leaky integrate-and-fire (LIF) neural layer (cf. Gerstner and Kistler, <xref ref-type="bibr" rid="B28">2002</xref>), in which each neuron could be assimilated with an element of the feature; this selectivity of lines could result from the selectivity of neurons in the first neuron layer.</p>
<p>Indeed the following model (notations are inspired by Lee et al., <xref ref-type="bibr" rid="B48">2016</xref>) can be used to update the membrane potential of a LIF neuron for a given input event (or spike):</p>
<disp-formula id="E10"><label>(8)</label><mml:math id="M35"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003C4;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C4;<sub><italic>mp</italic></sub> is the membrane time constant, <italic>w</italic><sub><italic>k</italic></sub> is the synaptic weight of the k-th synapse (through which the input event or spike arrives) and <italic>w</italic><sub><italic>dyn</italic></sub> is a dynamic weight controlling a refractory period (see Gerstner and Kistler, <xref ref-type="bibr" rid="B28">2002</xref>; Lee et al., <xref ref-type="bibr" rid="B48">2016</xref> for more details). This model is very similar to the incremental updating equation of our feature, Equation (7). The only things missing are the dynamic weight <italic>w</italic><sub><italic>dyn</italic></sub> and a firing threshold <italic>V</italic><sub><italic>th</italic></sub> in order to output approximatively the value of the corresponding feature&#x00027;s element as an event stream (or spike train), and then approximatively following a rate-coding model. Here, the refractory period should be set close to 0 (probably as a small fraction of the integration time <inline-formula><mml:math id="M36"><mml:msup><mml:mrow><mml:mo>&#x003C4;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></inline-formula>), in order to allow (quasi-)simultaneous visual events in the neighborhood (i.e., the events generate by the same contour moving across several pixels in the neighborhood) to contribute equitably to the neuron&#x00027;s potential, i.e., the value of the corresponding element of the feature.</p>
<p>For the local approach, a leaky integrate-and-fire neural layer has to be implemented for each pixel; this neural layer collects the visual motion events from the receptive field, &#x003A9;<sub><bold>p</bold><sub><italic>i</italic></sub></sub> (defined as &#x02223;&#x02223; <bold>p</bold> &#x02212; <bold>p</bold><sub><italic>i</italic></sub> &#x02223;&#x02223;&#x0003C; 2&#x003C3;<sub><italic>s</italic></sub>) defined by the corresponding bi-variate spatial kernel (Equation 5). This local computation is detailed in Algorithm <xref ref-type="table" rid="T2">1</xref>. For the global approach, only one neural layer is required, collecting the visual motion events estimated over the entire retina.</p>
<table-wrap position="float" id="T2">
<caption><p><bold>Algorithm 1</bold> Computation of the local feature.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;1: <bold>for all</bold> pixel&#x00027;s location <bold>p</bold> &#x02208; Retina <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;2:&#x000A0;&#x000A0;&#x000A0;&#x000A0;Set <inline-formula><mml:math id="M31"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> for all (<italic>v</italic><sup><italic>l</italic></sup>, &#x003B8;<sup><italic>l</italic></sup>)<sup><italic>T</italic></sup></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;3: <bold>end for</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;4: <bold>for all</bold> event <bold>e</bold> &#x0003D; (<bold>p</bold>, <italic>t, pol</italic>)<sup><italic>T</italic></sup> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;5:&#x000A0;&#x000A0;&#x000A0;&#x000A0;Compute the current optical flow <inline-formula><mml:math id="M32"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>e</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (see Section 2.1).</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;6:&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>for all p</bold><sub><italic>i</italic></sub> &#x02208; &#x003A9;<sub><bold>p</bold></sub>, where &#x003A9;<sub><bold>p</bold><sub><italic>i</italic></sub></sub> is a spatial neighborhood such as &#x02223;&#x02223; <bold>p</bold> &#x02212; <bold>p</bold><sub><italic>i</italic></sub> &#x02223;&#x02223; &#x0003C; 2&#x003C3;<sub><italic>s</italic></sub>, <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;7:&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Update <bold>F</bold><sub><bold>p</bold><sub><italic>i</italic></sub>,<italic>t</italic><sub><italic>i</italic></sub></sub>: <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x003B8;</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>,</td>
</tr>
<tr>
<td>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;where <italic>t</italic><sub><italic>i</italic></sub> is the timing of the previous update of <bold>F</bold><sub><bold>p</bold><sub><italic>i</italic></sub>, <italic>t</italic></sub> (see Equation 7)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;8:&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>end for</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;9:&#x000A0;<bold>end for</bold></td></tr>
<tr><td align="left" valign="top">10:&#x000A0;Output <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub></td></tr>
</tbody>
</table>
</table-wrap>
<p>Finally, Figure <xref ref-type="fig" rid="F4">4</xref> shows that the distribution of optical flow representation in the global approach (Figure <xref ref-type="fig" rid="F4">4C</xref>) summarizes the principal motions observed in the visual scene. This property will allow us to propose a machine learning based approach to recognize gestures in Section 4. In the next Section, we will demonstrate that the local version can be also used to detect particular interest points, i.e., corners.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Illustration of the global motion-based feature for event-based vision: from the stream of events (A)</bold>, the optical flow <bold>(B)</bold> is extracted. The feature corresponds to the distribution of this optical flow <bold>(C)</bold> in a polar coordinate frame, and can be reduced into a more compact and scale-invariant representation, called Histogram of Oriented Optical Flow <bold>(D)</bold> (see Section 4.1). As we can see in <bold>(B,C)</bold>, the motions generated by the forward leg (magenta boxes), the backward leg (green boxes), and the rest of body (red boxes) corresponds to three distinct and representative modes in the proposed feature.</p></caption>
<graphic xlink:href="fnins-10-00594-g0004.tif"/>
</fig>
<p>Remark 3. <italic>If the photodiode of the retina&#x00027;s pixel is not square as for the ATIS&#x00027;s one (see Posch et al., <xref ref-type="bibr" rid="B75">2010</xref> and Figure <xref ref-type="fig" rid="F1">1A</xref>), the frequency of a set of events emitted by a pixel will be not the same when a contour moves horizontally or vertically in the pixel&#x00027;s field of view (contour&#x00027;s speed and contrast are considered equal in both cases), because the contour travels the same surface of the photodiode during different time periods. In this case, keeping a decay factor invariant whatever the direction of the motion will introduce a bias, favoring one direction over another, in</italic> <bold>F</bold>. <italic>To avoid this bias, a cone-pixel with an ellipse-based basis (and not a disk-based basis as illustrated in Figure <xref ref-type="fig" rid="F2">2</xref>) can be implicitly considered in a correcting function &#x003B1;<sub>&#x003B8;</sub>(&#x000B7;) introduced in Equations 4 and 7 (instead of the constant smoothing parameter &#x003B1;); it is depending on the direction &#x003B8;<sup><italic>l</italic></sup> of the visual motion and defined as</italic>:</p>
<disp-formula id="E11"><label>(9)</label><mml:math id="M37"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo class="qopname">cos</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>where</italic> &#x003B1; &#x02208; [0, 1] <italic>and</italic> <inline-formula><mml:math id="M38"><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt></mml:math></inline-formula> <italic>is the eccentricity of the ellipse, with <italic>a</italic> and <italic>b</italic> the width and the length of the photodiode, respectively. The second term of this equation increases the decay factor in the direction of the principal axis of the ellipse, rebalancing the representation of the moving edges in</italic> <bold>F</bold>.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Application to corner detection</title>
<p>In conventional frame-based vision, several techniques have been proposed that consist in determining points for which a measurement is locally optimal with respect to a criteria; in particular specific to corners. This measure can be computed by a cumulative process (Park et al., <xref ref-type="bibr" rid="B70">2004</xref>), using a self-similarity measure (Moravec, <xref ref-type="bibr" rid="B57">1980</xref>) derived from mathematical analysis [e.g., contour&#x00027;s local curvature (Mokhtarian and Suomela, <xref ref-type="bibr" rid="B56">1998</xref>), relying on an eigenvalue decomposition of a second-moment matrix (Harris and Stephens, <xref ref-type="bibr" rid="B32">1988</xref>)] or selected as the output from a machine learning process (Rosten and Drummond, <xref ref-type="bibr" rid="B79">2006</xref>).</p>
<p>In asynchronous event-based vision, Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) have proposed an algorithm based on the intersection of constraints principle (see Adelson and Movshon, <xref ref-type="bibr" rid="B1">1982</xref>); which considers corners as locations where the aperture problem can be solved locally. Since cameras have a finite aperture size, motion estimation is possible only for directions orthogonal to edges. Figure <xref ref-type="fig" rid="F5">5</xref> shows the ambiguity due to the finite aperture. This can be written as follows: if <bold>v</bold><sub><italic>n</italic></sub> is the normal component of the velocity vector to an edge at time <italic>t</italic> at a location <bold>p</bold>, then the real velocity vector is an element of the &#x0211D;<sup>2</sup> subspace spanned by the unit vector <bold>v</bold><sub><italic>t</italic></sub>, tangent to the edge at <bold>p</bold>. This subspace is defined as <inline-formula><mml:math id="M39"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> with &#x003B1; &#x02208; &#x0211D;. For a regular edge point, &#x003B1; can usually not be estimated. When two moving crossed gratings are superimposed to produce a coherent moving pattern, the velocity can be unambiguously estimated.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>The aperture problem allows estimating only the normal component <inline-formula><mml:math id="M40"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="bold">n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of the velocity of events generated by an edge</bold>. The tangential component <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is not recoverable. Any motion with the same component <inline-formula><mml:math id="M42"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> induces the same stimulus. These motions define the real plane subspace <inline-formula><mml:math id="M43"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. (extracted from Clady et al., <xref ref-type="bibr" rid="B16">2015</xref>).</p></caption>
<graphic xlink:href="fnins-10-00594-g0005.tif"/>
</fig>
<p>The geometry-based approach proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) consists in collecting planes, fitted directly on the event stream (as in Benosman et al., <xref ref-type="bibr" rid="B5">2014</xref> and Section 2.1) and considered as local observations of normal visual motions, around each visual event. This event is considered as a corner event (i.e., event generates at the spatiotemporal location of a corner) if most of the collected planes intersect as a straight line in (<italic>XYT</italic>)<sup><italic>T</italic></sup> reference frame, at a location temporally close to the event (see Figure <xref ref-type="fig" rid="F6">6</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>(A)</bold> An event <italic>e</italic> occurs at spatial location <bold>p</bold> at time <italic>t</italic> where two edges intersect. This configuration provides sufficient constraints to estimate the velocity <bold>v</bold> at <bold>p</bold> from the normal velocity vector <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M45"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> provided by the two edges. The velocity subspaces <inline-formula><mml:math id="M46"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M47"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> are derived from the normal vectors. <bold>(B)</bold> Vectors <inline-formula><mml:math id="M48"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M49"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> are computed by locally fitting two planes <bold>&#x003A0;</bold><sup>1</sup> and <bold>&#x003A0;</bold><sup>2</sup> on the events forming each edge over a space-time neighborhood. <inline-formula><mml:math id="M50"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M51"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> are extracted from the slope of (respectively) <bold>&#x003A0;</bold><sup>1</sup> and <bold>&#x003A0;</bold><sup>2</sup> at (<bold>p</bold>, <italic>t</italic>). (extracted from Clady et al., <xref ref-type="bibr" rid="B16">2015</xref>).</p></caption>
<graphic xlink:href="fnins-10-00594-g0006.tif"/>
</fig>
<sec>
<title>3.1. Feature-based approaches</title>
<p>In the local approach, normalized <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> is the distribution of the normal velocities along the contours around the spatiotemporal location (<bold>p</bold>, <italic>t</italic>)<sup><italic>T</italic></sup>. In an ideal case illustrated in Figure <xref ref-type="fig" rid="F7">7</xref>, if this location corresponds to a corner location, <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> is null execpt around two velocity coordinates, (<italic>v</italic><sup><italic>n</italic></sup>, &#x003B8;<sup><italic>n</italic></sup>)<sup><italic>T</italic></sup> and (<italic>v</italic><sup><italic>m</italic></sup>, &#x003B8;<sup><italic>m</italic></sup>)<sup><italic>T</italic></sup>, corresponding to both normal visual motions of the intersecting edges.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Illustration of the feature F<sub><bold>p</bold>, <italic><bold>t</bold></italic></sub> (right figure) computed at the spatiotemporal location (p, <italic><bold>t</bold></italic>)<sup><italic><bold>T</bold></italic></sup> of a corner (left figure) in an ideal case</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0007.tif"/>
</fig>
<sec>
<title>3.1.1. 2-maxima based decision</title>
<p>As we can see in this Figure, detecting corners (or junctions) will consist in determining if at least two local maxima in <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> are present. We first propose an algorithm in order to find the two first maxima i n <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> consisting in:
<list list-type="order">
<list-item><p>finding the maximum <italic>F</italic><sub><italic>max</italic></sub> and its velocity coordinates <inline-formula><mml:math id="M52"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub>,</p></list-item>
<list-item><p>inhibiting (set to zeros) all values in <bold>F</bold> for which the coordinates verify <inline-formula><mml:math id="M53"><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula>, with <italic>th</italic><sub>&#x003B8;</sub> &#x0003D; 20&#x000B0;, and</p></list-item>
<list-item><p>finding the maximum (second maximum) <inline-formula><mml:math id="M54"><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and its coordinates <inline-formula><mml:math id="M55"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> previously modified in step 2.</p></list-item>
</list></p>
<p>Finally, as an isoprobabilistic representation of the intersecting edges is assumed, both values of maxima, <italic>F</italic><sub><italic>max</italic></sub> and <inline-formula><mml:math id="M56"><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, should be close at the location of a corner (the difference would be essentially due to noise). Then we propose as selection criterion (noted <inline-formula><mml:math id="M57"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) to decide if a corner is present at (<bold>p</bold>, <italic>t</italic>)<sup><italic>T</italic></sup>:</p>
<disp-formula id="E12"><label>(10)</label><mml:math id="M58"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>with the threshold <inline-formula><mml:math id="M59"><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec>
<title>3.1.2. Velocity-constraint based decision</title>
<p>A second approach consists in considering each (<italic>v</italic><sup><italic>l</italic></sup>, &#x003B8;<sup><italic>l</italic></sup>)<sup><italic>T</italic></sup> (or noted <inline-formula><mml:math id="M60"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in a cartesian reference frame) as a velocity constraint <inline-formula><mml:math id="M61"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> weighted by the value <inline-formula><mml:math id="M62"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>; verifying (<bold>v</bold><sup><italic>l</italic></sup>)<sup><italic>T</italic></sup><bold>v</bold> &#x0003D; ||<bold>v</bold><sup><italic>l</italic></sup>||<sup>2</sup>, where <inline-formula><mml:math id="M63"><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the velocity of the corner.</p>
<p>A corner is present at location (<bold>p</bold>, <italic>t</italic>)<sup><italic>T</italic></sup> if <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> gives rise to a real solution to the equation:</p>
<disp-formula id="E13"><label>(11)</label><mml:math id="M64"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>W</mml:mi><mml:mi>A</mml:mi><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mi>W</mml:mi><mml:mstyle mathvariant="bold"><mml:mtext>B</mml:mtext></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where:
<list list-type="bullet">
<list-item><p><inline-formula><mml:math id="M65"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable class="array"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, with <italic>N</italic><sub><bold>v</bold></sub> &#x0003D; <italic>N</italic><sub><italic>v</italic></sub><italic>N</italic><sub>&#x003B8;</sub> the size of <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub>, i.e., the number of constraints,</p></list-item>
<list-item><p><inline-formula><mml:math id="M66"><mml:mstyle mathvariant="bold"><mml:mtext>B</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable class="array"><mml:mtr><mml:mtd><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M67"><mml:mi>W</mml:mi><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">diag</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
</list></p>
<p>Then the over-determined system can be solved if <italic>M</italic> &#x0003D; (<italic>WA</italic>)<sup><italic>T</italic></sup><italic>WA</italic> has a full rank, meaning that its two eigenvalues have to be significantly large. This significance is determined with the selection criterion established in Noble (<xref ref-type="bibr" rid="B64">1988</xref>):</p>
<disp-formula id="E14"><label>(12)</label><mml:math id="M68"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext class="textrm" mathvariant="normal">det</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">trace</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>with the threshold <inline-formula><mml:math id="M69"><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>.</p>
<p>Equation (11) is also solved with a least square minimization technique and solutions are considered as valid if <inline-formula><mml:math id="M70"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is greater than the threshold <inline-formula><mml:math id="M71"><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> usually set experimentally. Finally, a stream <inline-formula><mml:math id="M72"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">S</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of corner events (including features), noted <bold>c</bold> &#x0003D; (<bold>p</bold>, <bold>v</bold>, <italic>t</italic>, <bold>F</bold>)<sup><italic>T</italic></sup>, is outputted.</p>
<p>Remark 4. <italic>In order to be robust to noise, weak values in</italic> <bold>F</bold><sub><bold>p</bold>,<italic>t</italic></sub> <italic>are inhibited (associated equations are filtered out of the system): if</italic> <inline-formula><mml:math id="M73"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (with <italic>th</italic><sub><bold>F</bold></sub> &#x02208; [0, 1]), then <inline-formula><mml:math id="M74"><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>.</p>
<p>Remark 5. <italic>With the 2-maxima based decision approach, a corner event stream can also be obtained; the velocities of the detected corners can be estimated in a similar manner using only both maxima&#x00027;s coordinates, without weighting them. Furthermore, while the second approach is based on a (unnatural) mathematical analysis, the first decision method is closer to a time-based neural implementation; it could be implemented as a coincidence detector between two (or more) events, denoting the two-first (or more) maxima, outputted by the leaky integrate-and-fire neural layer assimilated to the feature</italic> <bold>F</bold> <italic>(see Discussion at the end of Section 2.3)</italic>.</p>
<p><italic>Note that neural networks have also been proposed in the literature (Cichocki and Unbehauen, <xref ref-type="bibr" rid="B14">1992</xref>) in order to solve similar systems of linear equations that are required in the velocity-constraint decision based method; VLSI implementations have even been proposed</italic>.</p>
<p>Remark 6. <italic>Note that the computation principle is quite similar to the one proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>); most mechanisms involved (kernels, filters, selection criteria) have been designed and set in a similar manner, in order to allow comparison in the fairest way possible (see next Section). The methods differ from each other essentially by the selection process of the velocity constraints. Through a time-based weighting process, Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) considers only constraints along edges intersecting the evaluated event. The methods proposed in this article consider all the edges in a spatial neighborhood even if they are not perfectly intersecting themselves at the evaluated location; however the spatial Gaussian-based weights <italic>w</italic><sub><italic>s</italic></sub>(&#x000B7;) implicitly perform a heuristic selection of the spatially closer edges, i.e., the most probable intersecting edges. So even if the location of their detected corner events should be consequently less precise, they should be close to a real corner; this is verified in the results presented in the next Section</italic>.</p>
</sec>
</sec>
<sec>
<title>3.2. Evaluations</title>
<p>In order to evaluate the detectors, we reproduced one of the experiments proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>), the one with the most quantitative evaluations. It consists into a swinging wired 3D cube shown to a neuromorphic camera (DVS, see Figure <xref ref-type="fig" rid="F8">8</xref>).</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Illustration of the experiment: a swinging 3D cube is shown to a neuromorphic camera</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0008.tif"/>
</fig>
<p>A complete accuracy evaluation, comparing the results obtained with the geometric-based method given in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) and the methods proposed in this article, is provided in Figure <xref ref-type="fig" rid="F9">9</xref>. The corner events parameters (spatial location and velocity) and the 11 corners ones (obtained with the ground-truth) are compared using different measures of errors. Each corner event is associated to the spatially closest ground-truth corner&#x00027;s trajectory.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>Precision evaluation of the corner detectors; the green plain curves correspond to the results obtained with the algorithm proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>); the blue dash-dotted blue curves to the velocity-constraint based decision proposed in Section 3.1.2 and the red dashed curves to the 2-maxima based decision proposed in the Section 3.1.1</bold>. The blue and red dotted curves correspond to the respective feature-based approaches but without speed-tuned temporal kernels. The left figure <bold>(A)</bold> represents the spatial location errors of the corner events compared to the manually-obtained ground-truth trajectories of the corners; the middle one <bold>(B)</bold> the relative error about the intensity of the estimated speed and the right one <bold>(C)</bold> its error in direction. Accuracies (X-axis fo the Figures) are given related to the considered percent (Y-axis) of the population of corner events detected with the different methods; e.g., with the method in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>), 80% of the corner events have a distance error in corner location &#x0003C;2 pixels compared to the ground truth, see plain green curve in <bold>(A)</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0009.tif"/>
</fig>
<p>In order to propose a fair evaluation, the thresholds used in the different methods have been set in order to detect the same number of corner events (1500) and other algorithms&#x00027; parameters have been set as the ones proposed in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) (see Remark 3.6). The distribution of the corner events per corner&#x00027;s trajectory is shown in Figure <xref ref-type="fig" rid="F11">11A</xref>. We can observe that the distributions using the geometric-based and the 2-maxima decision based methods are closely similar. However, the one obtained with the velocity-constraint decision based method is unbalanced, with a great number (close to the third of the corner events) of detections around a particular corner, corner number 5. This can be explained by the fact that the proposed method is less spatially precise than the geometric-based one (cf. the curves in Figure <xref ref-type="fig" rid="F9">9A</xref> and Remark 3.6) and, as we can see in Figure <xref ref-type="fig" rid="F10">10</xref>, the edges around this corner generated more events than the others because they are generated by &#x0201C;clean&#x0201D; intersecting edges, see Figure <xref ref-type="fig" rid="F8">8</xref>, and then verifying well the ideal conditions for the optical flow estimation, and because it is a X-junction. It is not the case for the corners number 1, 7, and 11, for example; the high speed of the cube (close to 500<italic>pix</italic>.<italic>s</italic><sup>&#x02212;1</sup>, i.e., inversely close to the precision of the event timings) and their badly shaped structures (they correspond to connections between the different wires constituting the cube) make their detection very hard due to the local bad quality of the event streams (in particular, there are numerous missing events as we can see in Figure <xref ref-type="fig" rid="F11">11</xref>).</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>Snapshots of the results obtained for the three compared detectors, projecting in a frame the visual events (black dots) and corner events (circles, associated to vectors representing the estimated speeds) over two short time periods (1 ms)</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0010.tif"/>
</fig>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p><bold>Distributions of the detected event corners related to the labeled corners. (A)</bold> Comparison between the three evaluated detectors. <bold>(B)</bold> Comparison with or without (black) speed-tuned temporal kernels.</p></caption>
<graphic xlink:href="fnins-10-00594-g0011.tif"/>
</fig>
<p>Remark 7. <italic>Note that accuracy results in Figure <xref ref-type="fig" rid="F9">9</xref> concern median evaluations over the 11 ground-truth corners. Each corner is associated to the spatially closest ground-truth corners trajectory. Each set of corner events (associated to a ground-truth corner) is sorted according to one of the evaluation criteria (type of errors). The Y%-most accurate corner events are then selected. Finally, the accuracy median value for this evaluation criterion is computed over all ground-truths corners. So these evaluations are <italic>a priori</italic> not (or weakly) biased by these differences in distributions</italic>.</p>
<p>We can observe that the detectors proposed in this article are influenced by the quantification of the grid; especially in the Figure <xref ref-type="fig" rid="F9">9C</xref> representing the angular precision of the estimated speed direction. Indeed a lot of corner events have a direction-related precision close to 5&#x000B0;, the half of the direction-related interval length. The velocity-constrainst based decision method is less clearly influenced because it takes into account more elements in the feature (not only the elements with the maximal values, but also their neighboring elements) to estimate the speed.</p>
<p>In addition, Figure <xref ref-type="fig" rid="F11">11B</xref> shows the detections distribution for both feature-based methods, with or without speed-tuned temporal kernels. In the approaches without speed-tuning, the temporal decreasing factor &#x003C4; has been fixed as <inline-formula><mml:math id="M75"><mml:mo>&#x003C4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>, where <italic>v</italic><sub><italic>mean</italic></sub> is the mean velocity computed over all corners and the stream duration (150 ms). Without speed-tuning, some corners are not or not often detected, in particular corners number 6 and 8. They correspond to X-junctions between two intersecting edges with quite different dynamics, because generated by front and back wires. Furthermore, the accuracy performances for the approaches without speed-tuned temporal kernels are significantly lower than the ones with speed-tuned kernels, as shown in Figure <xref ref-type="fig" rid="F9">9</xref>.</p>
<p>Finally, if we consider that a corner event detection is valid if the distance error is &#x0003C;3<italic>pixels</italic>, the geometric-based method generates only 2% of false alarms (with a median velocity error around 10% and a median direction error around 3&#x000B0; for the positive detections), while this rate rises to 8% and to 18% for the velocity-constraint decision and 2-maxima decision based methods, respectively (with a median velocity error around 10% and a median direction error around 8&#x000B0;, for both).</p>
<p>We have demonstrated that the proposed feature can be used (in its local approach) to detect corners in event streams. Even if the detectors are slightly less precise and more sensitive to the quality of the event streams than the other method proposed in the literature, our feature-based approaches are more efficient in terms of memory and computation loads.</p>
<p>Indeed the method in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) requires to memorize the stream of the visual motion events (see Equation 2) and spatiotemporal extrapolations of them (called &#x0201C;normal events&#x0201D;) and operates quite complex computations between them. In the approach presented in this article, the visual motion events are integrated directly in the neighboring features, and corner detection related computations are operated only using the feature at the spatiotemporal location of the current event. We have measured important differences in terms of computation time between their different implementations; e.g., for the event stream used for the above evaluations, the feature-based approaches are &#x0007E;10 times faster. Table <xref ref-type="table" rid="T1">1</xref> presents the distribution of mean computation times obtained with the different approaches and over 10 repetitions (for 1500 detections). But as the method in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) has been only implemented on Matlab (Matlab2015b), they should be taken with caution; it is indeed known that memory can be poorly managed on Matlab. Measuring the computation time without code lines dedicated to memory management (which is a crucial part of the method in Clady et al., <xref ref-type="bibr" rid="B16">2015</xref>), the gain is still around 40%. While the geometric-based method is only envisioned in Clady et al. (<xref ref-type="bibr" rid="B16">2015</xref>) for a real time implementation on massively parallel computers such as the SpiNNaker board (see Furber et al., <xref ref-type="bibr" rid="B26">2013</xref>; Orchard et al., <xref ref-type="bibr" rid="B67">2015a</xref>), the feature-based approaches run in real-time on a standard computer (in C&#x0002B;&#x0002B; on a Intel Core i7-4790K &#x00040; 4GHz, using only one core and without any optimization such as integer arithmetic instead of floating point based computations, e.g., Schraudolph, <xref ref-type="bibr" rid="B81">1999</xref>; Cawley, <xref ref-type="bibr" rid="B9">2000</xref>) for weakly complex visual scenes such as the one presented in this study.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Distribution of mean computation times (CT) with the different approaches (estimated on Matlab2015b)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>Total CT</bold></th>
<th valign="top" align="center"><bold><italic>%</italic> of CT OF estimation</bold></th>
<th valign="top" align="center"><bold><italic>%</italic> of CT feature computation</bold></th>
<th valign="top" align="center"><bold><italic>%</italic> of CT corner detection</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Velocity-constraint</td>
<td valign="top" align="center">76s.</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">2-maxima</td>
<td valign="top" align="center">75s.</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">Geometric</td>
<td valign="top" align="center">828s.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">99</td>
</tr>
<tr>
<td valign="top" align="left">Geometric</td>
<td valign="top" align="center">132s.</td>
<td valign="top" align="center">9</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">91</td>
</tr>
<tr>
<td valign="top" align="left">(w/o memory management)</td>
<td/>
<td/>
<td/>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<p>Beyond this operational asset, the greatest strength of the proposed feature-based approaches lies in fact that they lead to a solution of the corner detection issue on event streams based on classical event-based neural network models (leaky integrate-and-fire neural network, coincidence detectors, etc.) as it is highlighted in Section 2.3 and Remark 5.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Application to gesture recognition</title>
<p>Human movement analysis is an area of study that has been quickly expanding since the 1990&#x00027;s (see Moeslund et al., <xref ref-type="bibr" rid="B54">2006</xref>; Poppe, <xref ref-type="bibr" rid="B72">2007</xref>, <xref ref-type="bibr" rid="B73">2010</xref>). The evolution and miniaturization of both computers and motion capturing sensors have made motion analysis possible in a growing set of environments. They have enabled numerous applications in robotics, control, surveillance, medical purposes (Zhou and Hu, <xref ref-type="bibr" rid="B90">2008</xref>) or even in video-games with the Microsoft&#x00027;s Kinect (Han et al., <xref ref-type="bibr" rid="B31">2013</xref>). However, the available technologies and methods still present numerous limitations, discouraging their use in embedded systems. Conventional time-sampled acquisition is very problematic when implemented in mobile devices because the embedded cameras usually operate at a frame-rate of 30 to 60 Hz: normal speed gesture movements can not be properly captured. Increasing the frame rate would result in the overload of the recognition algorithm, only displacing the bottleneck from acquisition to post-processing. Furthermore, conventional cameras and infrared-based methods are perturbed by dynamic lighting and infra-red radiations emitted by the sun (cf. Pana&#x000EF;t&#x000E9; et al., <xref ref-type="bibr" rid="B69">2011</xref>). Because they both require light-controlled environments, those technologies are unsuitable for outdoor use.</p>
<p>Asynchronous event-based sensing technology is expected to overcome several limitations encountered by state-of-the-art gesture recognition systems, in particular for battery-powered, mobile devices. These vision sensors, due to their near continuous-time operation, allow capturing the complete and true dynamics of human motion during the whole gesture duration. Due to the pixel-individual style of acquisition and pre-processing of the visual information, and in contrast to practically all existing technologies, they will be also able to support device operation under uncontrolled lighting conditions, particularly in outdoor scenarios (cf. Simon-Chane et al., <xref ref-type="bibr" rid="B84">2016</xref>). Native redundancy suppression performed in event-based sensing and processing will ensure that computation can be performed in real time, while at the same time saving energy, decreasing system complexity.</p>
<p>Gesture recognition using neuromorphic camera has already been investigated by Lee et al. (<xref ref-type="bibr" rid="B49">2014</xref>). A stereo pair of DVS allows them to compute disparity in order to cluster the hand. Then, they use a tracking algorithm to extract the 2D trajectory of the movement. Finally the trajectory is sampled into directions, and the obtained sequence of directions is fed to a HMM classifier. This approach uses event-based information only during the first step (extraction of the location of the hand). In addition, with this type of multi-steps architecture, a failure in a step could result in the failure of the whole system.</p>
<p>Here we propose to demonstrate that our feature can be used to detect and recognize more directly gestures. Hoof-like features (see Section 4.1) are derived from the feature matrix and provided to a classification architecture that performs simultaneously detection and recognition. It is based on hybrid generative/discriminative classifiers (Lasserre et al., <xref ref-type="bibr" rid="B47">2006</xref>) in order to associate at each feature its probabilities to belong to the considered (hand) gestures or not, and these probabilities are integrated over time through a network of Bayes filters (Thrun et al., <xref ref-type="bibr" rid="B85">2008</xref>).</p>
<sec>
<title>4.1. A more compact and invariant representation</title>
<p>In order to reduce the dimensionality of the feature (it is often required in machine learning, in order to address the &#x0201C;curse of dimensionality&#x0201D; issue) and to provide (global speed- and) scale-invariance property to the gesture representation, <bold>F</bold> can be transformed into a more compact representation, noted <bold>h</bold> (<bold>h</bold><sub><bold>p</bold>,<italic>t</italic></sub> or <bold>h</bold><sub><italic>t</italic></sub>, in local or global approaches, respectively) and named hoof-like in reference to the Histogram of Oriented Optical Flow (HOOF) introduced by Chaudhry et al. (<xref ref-type="bibr" rid="B13">2009</xref>) in frame-based vision. This transformation consists in summing the intensities of the optical flow vectors with respect to their directions.</p>
<p>From the feature <bold>F</bold>, <inline-formula><mml:math id="M76"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> can be easily obtained:</p>
<disp-formula id="E15"><label>(13)</label><mml:math id="M77"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the global approach, normalization (to sum to 1) makes the hoof-like feature globally speed- and scale-invariant. Figure <xref ref-type="fig" rid="F4">4D</xref> represents the histogram of oriented optical flows computed globally on an event stream capturing a walking human (Figure <xref ref-type="fig" rid="F4">4A</xref>).</p>
</sec>
<sec>
<title>4.2. Classification architecture</title>
<p>We propose a classification architecture where the problem is framed as a Bayes filter, that is estimating the probabilities of gestures recursively over time using incoming measurements, given as the hoof-like features <inline-formula><mml:math id="M78"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> computed globally from every visual events [<bold>e</bold><sub>0</sub>, <bold>e</bold><sub><italic>k</italic></sub>].</p>
<p>Then we note the state <inline-formula><mml:math id="M79"><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow></mml:math></inline-formula>, the gesture (numerated <italic>i</italic>, <italic>i</italic> &#x02208; [1, <italic>K</italic>]) that the user is currently performing. A state <italic>g</italic><sup>0</sup> is added in <inline-formula><mml:math id="M80"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow></mml:math></inline-formula>, in order to consider the not-considered gestures or the instants while the user is not performing a hand gesture.</p>
<p>The camera observes the user&#x00027;s action and at each occurring feature estimates a distribution over the current state <inline-formula><mml:math id="M81"><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>:</p>
<disp-formula id="E16"><label>(14)</label><mml:math id="M82"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M83"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> is the observation of the gesture occurring at time <italic>t</italic><sub><italic>k</italic></sub>.</p>
<p>To estimate this probability, a time update and a measurement update are performed alternately. The time update updates the belief that the user is performing a specific gesture given previous information:</p>
<disp-formula id="E17"><label>(15)</label><mml:math id="M84"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The time update includes a transition probability from the previous state to the current state. As no-contextual information is available here, we assume that an user is likely to perform the same gesture, and at each timestamp has a large probability of transitioning to the same state:</p>
<disp-formula id="E18"><label>(16)</label><mml:math id="M85"><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mi>j</mml:mi></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x02223;</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo>&#x02223;</mml:mo></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mo>&#x02223;</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo>&#x02223;</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x02223;</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo>&#x02223;</mml:mo></mml:mrow></mml:mfrac><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mo>&#x003C4;</mml:mo><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>if&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x02223;</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo>&#x02223;</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x02223;</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo>&#x02223;</mml:mo></mml:mrow></mml:mfrac><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mo>&#x003C4;</mml:mo><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>with &#x003C4;<sub><italic>g</italic></sub> set to 150 ms, less than the half duration of shorter gestures. This assumption means that the gesture&#x00027;s certainty slowly decays over time, in the absence of corroborating information, converging to a uniform distribution (even if no event is observed).</p>
<p>The measurements update combines the previous belief with the newest observation to update each belief state, such as:</p>
<disp-formula id="E19"><label>(17)</label><mml:math id="M86"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x0221D;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In order to estimate <inline-formula><mml:math id="M87"><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, we propose a machine learning based approach to compute and select generative models for gesture. It is decomposed into two steps:
<list list-type="bullet">
<list-item><p>For the first step, we collect hoof-like features computed while the users (included in the training database, see Section 4.3.1) performed a gesture <italic>g</italic><sup><italic>i</italic></sup>, <italic>i</italic> &#x02208; [1, <italic>K</italic>]. Then a k-means algorithm is applied on them in order to compute <italic>N</italic> candidate models, noted <bold>m</bold><italic><sup>g<sup>i</sup></sup></italic>.</p></list-item>
<list-item><p>The second step consists in selecting from these candidate models, the ones that are the most discriminative against hoof-like features collected from the rest of the training event streams; these last features have been computed during other considered gestures (<italic>g</italic><sup><italic>j</italic></sup> with <italic>i</italic> &#x02260; <italic>j</italic>) or during other period times when users were not performing gestures. This selection is processed through a discrete Adaboost classifier.</p></list-item>
</list></p>
<p>Adaboost (Freund and Schapire, <xref ref-type="bibr" rid="B25">1996</xref>) is an iterative algorithm that finds, from a feature set, some weak but discriminative classification functions and combines them in a strong classification function:</p>
<disp-formula id="E20"><label>(18)</label><mml:math id="M88"><mml:mrow><mml:mi>B</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mtext>,&#x000A0;</mml:mtext><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mo>&#x003BB;</mml:mo><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msub><mml:mi>b</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mo>&#x003BB;</mml:mo><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mtext>,&#x000A0;otherwise</mml:mtext><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>B</italic> and <italic>b</italic> are the strong and weak classification functions, respectively, and &#x003BB; is a weight coefficient for each <italic>b</italic>. <italic>T</italic> is the threshold of the strong classifier <italic>B</italic>. The principle of the Adaboost algorithm is to select, at each iteration, a new weak classifier in favor of the instances (or features) misclassified by previous classifiers, through a weighting process attributing more influence to misclassified instances.</p>
<p>Note that a threshold value, noted <italic>th</italic><sub><italic>B</italic></sub>, can be defined (such as the condition in Equation 18 can be written: <inline-formula><mml:math id="M89"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) in order to optimize a particular classification performance. During the learning step, its default value is 1; this means a classification frontier at the middle of the margin (see Schapire et al., <xref ref-type="bibr" rid="B80">1998</xref>). Increasing or reducing its value correspond to moving the frontier closer or further to the positive class, respectively.</p>
<p>In literature, discriminative training of generative models, as we propose here, has been shown as efficient learning methods in numerous applications as object or human detection (Holub et al., <xref ref-type="bibr" rid="B33">2005</xref>; Negri et al., <xref ref-type="bibr" rid="B61">2008</xref>; Wang et al., <xref ref-type="bibr" rid="B89">2011</xref>), face or character recognition (Prevost et al., <xref ref-type="bibr" rid="B76">2005</xref>; Grabner et al., <xref ref-type="bibr" rid="B30">2007</xref>) or for medical purposes (Deselaers et al., <xref ref-type="bibr" rid="B21">2008</xref>; Wang et al., <xref ref-type="bibr" rid="B88">2015</xref>). The proposed classifier based on the training and the selection of generative models in a discriminative way, combines indeed the main characteristics of discriminative and generative approaches: discriminative power and generalization ability, respectively. The latter is in particular very important in our application, when a weak amount of labeled training data is available, see Section 4.3.1.</p>
<p>Following the framework described in Jing et al. (<xref ref-type="bibr" rid="B38">2008</xref>), we propose to design weak classifiers as generative ones, associated to each candidate models <inline-formula><mml:math id="M90"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>m</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> (<italic>s</italic> &#x02208; [1, <italic>N</italic>]):</p>
<disp-formula id="E21"><label>(19)</label><mml:math id="M91"><mml:mrow><mml:msubsup><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>h</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>m</mml:mi></mml:mstyle><mml:mi>s</mml:mi><mml:mrow><mml:msup><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>h</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>m</mml:mi></mml:mstyle><mml:mi>s</mml:mi><mml:mrow><mml:msup><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:msubsup><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>&#x003B8;</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:msup><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02265;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;otherwise</mml:mtext><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>d</italic>(&#x000B7;, &#x000B7;) is the Euclidean distance and <inline-formula><mml:math id="M92"><mml:msubsup><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> parametrizes the likelihood function <italic>f</italic> and is computed at each iteration of the algorithm through a maximum-likelihood estimation (taking into account the weights attributed to features).</p>
<p>During training, Adaboost based algorithm tends to select iteratively the most discriminative and complementary models for each gesture. We limit the number of selected models, such as the relative difference between F-measure (computed on training database, see Section 4.3.1) obtained at the corresponding iteration is superior or equal to 95% of its maximum (obtained with a greater number of iterations of Adaboost algorithm). Let us remind that F-measure is defined as <inline-formula><mml:math id="M93"><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">precision</mml:mtext></mml:mstyle><mml:mo>&#x000D7;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">recall</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">precision</mml:mtext></mml:mstyle><mml:mo>&#x0002B;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">recall</mml:mtext></mml:mstyle></mml:mrow></mml:mfrac></mml:math></inline-formula>. Optimizing it means also to determine a number of models for which an acceptable compromise between precision (the ratio of positive detections to instances belonging to performed gestures) and recall (the ratio of positive detections to all instances detected as belonging to gestures) is reached.</p>
<p>The probability <inline-formula><mml:math id="M94"><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is then estimated as proportional to a measure (&#x02208; [0, 1]) operated between the hoof-like feature and the set of selected models (applying a sigmoidal function to the output of the strong classifier):</p>
<disp-formula id="E22"><label>(20)</label><mml:math id="M95"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>with <italic>i</italic> &#x02208; [1, <italic>K</italic>] and <inline-formula><mml:math id="M96"><mml:msubsup><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the threshold obtained optimizing the F-measure. The probability associated to not-considered gesture (or no-gesture), noted <italic>g</italic><sup>0</sup>, is then defined as:</p>
<disp-formula id="E23"><label>(21)</label><mml:math id="M97"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">max</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>K</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Figure <xref ref-type="fig" rid="F12">12</xref> presents the obtained classification architecture. Finally a gesture&#x00027;s class <italic>G</italic><sub><italic>t</italic><sub><italic>k</italic></sub></sub> at each time is attributed from the distribution of probabilities, defined as:</p>
<disp-formula id="E24"><label>(22)</label><mml:math id="M99"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">argmax</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>K</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p><bold>Gesture Recognition Architecture: for each occurring hoof-like feature h<sub><italic>t</italic><sub><italic>k</italic></sub></sub>, the distribution of probabilities noted <italic>p</italic>(h<sub><italic>t</italic><sub><italic>k</italic></sub></sub> &#x02223; <italic>g</italic><sub><italic>t</italic><sub><italic>k</italic></sub></sub>), is estimated comparing the features to models computed and selected through an Adaboost-based learning process</bold>. Then the probabilities of gestures, noted <italic>p</italic>(<italic>g</italic><sub><italic>t</italic><sub><italic>k</italic></sub></sub> &#x02223; <bold>h</bold><sub><italic>t</italic><sub>0</sub>:<italic>t</italic><sub><italic>k</italic> &#x02212; 1</sub></sub>), are estimated recursively over time.</p></caption>
<graphic xlink:href="fnins-10-00594-g0012.tif"/>
</fig>
<p>Remark 8. <italic>Even if our implementation is based on a learning process not directly related to neural approaches (essentially due to the limited size of the database), we can observe that the resulting classification architecture could be fully implemented in an event-based framework. Through a rate-coding model, hoof-like features could be computed and transmitted from the leaky integrate-and-fire neural network, corresponding to the feature computation, as evoked in Section 2.3, to neural networks performing their comparison with gesture models (considering maybe another distance than the Euclidean one used here) and outputting positive events when they match; these positive events corresponding to the weak classifier responses (<inline-formula><mml:math id="M101"><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>). The coefficients <inline-formula><mml:math id="M102"><mml:msubsup><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> would be then assimilated to synaptic weights. The other operations, in particular involved in Bayes filters, would correspond to feedback lines and basic mathematical operations that can be modeled using precise timing and event-based paradigms as demonstrated in Lagorce and Benosman (<xref ref-type="bibr" rid="B42">2015</xref>)</italic>.</p>
</sec>
<sec>
<title>4.3. Results</title>
<sec>
<title>4.3.1. Experimental protocol</title>
<p>The protocol assumes that the users performed gestures in front of the camera. Event streams (using the ATIS camera) have been collected with nine users (young and middle-aged people working in the laboratory). All users are right-handed but the database could be extended to left-handed users by mirroring the sequences horizontally.</p>
<p>The hand is moving at a distance around 30 cm from the camera, approximatively. Note that this distance has been determined to ensure that the hand is fully viewed by the camera (see Figure <xref ref-type="fig" rid="F13">13A</xref>) considering the current optic lens (this distance should be reduced when a wider-angle lens will be implemented). Each gesture is repeated five times by each user, varying the hand speed.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p><bold>Illustration of the targeted human-machine interaction. (A)</bold> Example of a hand gesture performed in front of the camera. <bold>(B)</bold> ATIS camera embedded on a smartphone.</p></caption>
<graphic xlink:href="fnins-10-00594-g0013.tif"/>
</fig>
<p>Six gestures have been defined and correspond to a dictionary of coarse gestures; the gesture is defined by the global motion of the hand (hand moving to the left, to the right, upward, downward, opening, or closing). These gestures could match with the main controls we could intend to execute interacting with a smartphone or a tablet (navigating in a menu or a list, selecting/unselecting an object or an application), i.e., the targeted application (see Figure <xref ref-type="fig" rid="F13">13B</xref>). Furthermore, they constitute a dictionary for more complex gestures, successively combining these movements. In Figure <xref ref-type="fig" rid="F14">14</xref>, an iconic representation of these coarse gestures is presented in the second column.</p>
<fig id="F14" position="float">
<label>Figure 14</label>
<caption><p><bold>Iconic representations (second column) of the gestures (first column) and corresponding models selected by the Adaboost-based machine learning process</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0014.tif"/>
</fig>
<p>The training database is composed of the event streams collected with five users and the test database with the four other ones. During the evaluations (see next Section), a cross-validation is performed ten times (presented evaluations are the obtained mean values), putting randomly the users in the training or test databases. 30, 000 hoof-like features, computed on the training streams, are collected randomly and equitably in the time periods when gestures are performed (including the not-considered gestures or no gesture class) to train the Adaboost classifiers with a <italic>one</italic>-vs.<italic>-all</italic> strategy. An equal quantity is again randomly selected for the F-measure based optimization process and the selection of the number of models. Six hundred candidate models per gesture have been computed using k-means algorithm. The characteristics of the hoof-like features are the same as described in Section 2.2 (<italic>N</italic><sub>&#x003B8;</sub> &#x0003D; 36, etc).</p>
<p>A gesture is considered as detected when the duration of a time period with classified gestures (<italic>G</italic><sub><italic>t</italic><sub><italic>k</italic></sub></sub> &#x02260; 0 in Equation 22) is over 300 ms. This detection is counted as positive if this time period overlaps the manually labeled ground truth (with an overlap ratio superior to 0.5).</p>
</sec>
<sec>
<title>4.3.2. Evaluations</title>
<p>Figure <xref ref-type="fig" rid="F14">14</xref> represents the considered gestures and the models selected by Adaboost during a learning process (see Section 4.2). We can observe that the number of selected models is relatively weak (3 or 4). This means that the hoof-like features are able to represent well the gestures despite their (speed- and user-related) variability, mostly thanks to its speed- and scale-invariance property.</p>
<p>Another observation concerns the &#x0201C;shape&#x0201D; of the feature models. For most of them, they match well to the iconic representation of the corresponding motion; for example, for the motions to the left and to the right, most speed vectors are oriented to these respective directions, etc. However, some singularities have to be explained considering not only the global motion but also the directions of the principal contours of the human parts (hand, finger and arm) involved in the hand movement. For the opening hand motion, models obtained at iterations 1 and 2 highlight the motion of the thumb, for which the moving contours are prevalent in the feature. For the downward motion, the contours of the arm are too prevalent (see models obtained at iterations 2 and 3) because the camera viewed the user&#x00027;s bust (see Figure <xref ref-type="fig" rid="F13">13</xref>).</p>
<p>In terms of detection performance, we obtained a mean precision of 91% and a mean recall of 83% (F-measure &#x0003D; 0.85) which confirm the great discrimination power of the proposed feature. Note that the F-measures obtained during the optimization (to determine <italic>th</italic><sub><italic>B</italic></sub> and the number of models) are around 0.75. The greater value obtained at the final output highlights the filtering action of the Bayes filters.</p>
<p>Finally the confusion matrix given in Figure <xref ref-type="fig" rid="F15">15</xref> shows us the recognized gestures among the positive detections. The downward and closing hand gestures are obviously a little confused because the similarity of the hand&#x00027;s and the fingers&#x00027; motions, respectively. The confusion of other gestures with the opening hand is probably due to the fact that the gesture is hard to detect, probably because the larger proportion of the movement involved the other fingers than the thumb and their moving contours generated few visual events (because in folded positions; the finger-skin vs. palm-skin contrast changes are weakly captured, see Remark 2.2). Indeed, in order to optimize the F-measure, the proposed process tends to select a low threshold compared to others (3 or 4 times lower); this means that classification frontier defined for this gesture tends to include other gestures. Hence, these gestures are sometimes misclassified as opening hand.</p>
<fig id="F15" position="float">
<label>Figure 15</label>
<caption><p><bold>Confusion matrix (expressed in percent) showing the recognized gestures (columns) related to the performed gestures (lines), among the positive detections</bold>.</p></caption>
<graphic xlink:href="fnins-10-00594-g0015.tif"/>
</fig>
<p>In further developments, we expect to improve these performances combining this global feature with locally computed ones, taking into account their relative spatio-temporal relationships. This should help us to better distinct the global motion of the hand and the local motions of the fingers, and hence better detect and categorize gestures.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5. Conclusion and discussion</title>
<p>In this article, we have proposed a motion-based feature for event-based vision. It consists in encoding the local or global visual information provided a neuromorphic camera in a grid-sampled map of optical flow. Collecting optical flow (or visual motion events) computed around each visual event in a neighborhood or in the entire retina, this map represents their current probabilistic distribution in a speed- and direction-coordinates frame.</p>
<p>Two event-based pattern recognition frameworks have been developed in order to demonstrate its usefulness for such tasks. The first one is dedicated to detection of specific interest points, corners. Two feature-based approaches have been developed and evaluated. Formulated as an intersection of constraints issue, this fundamental task in computer vision can be resolved operating with the information encoded in the proposed local feature. The second one consists in a hand gesture recognition system for human-machine interaction, in particular with mobile devices. More compact and scale-invariant representations (called hoof-like features) of the motion observed in the visual scene, are extracted directly from the global version of the proposed feature, and feed a classification architecture, based on a discriminative learning schema of gestures&#x00027; generative models and framed as a Bayes filter. Evaluations show that this feature has sufficient descriptive power to solve such pattern recognition problems. Other extensions or derivations of the proposed feature can be also envisioned in further developments, in order to address other pattern recognition issues. For example, summing the elements of the feature, with respect to their directions and without weighting them by corresponding speed, will result into another compact form, similar to the hog (histogram of oriented gradients) feature proposed by Dalal and Triggs (<xref ref-type="bibr" rid="B17">2005</xref>). This feature and its derivations have been demonstrated as very efficient for many pattern recognition tasks in frame-based vision. To evaluate it in event-based vision would required to design event-based and dedicated classification architecture(s).</p>
<p>It is interesting to notice that our motion-based feature allows us to detect features defined by &#x0201C;static&#x0201D; properties, i.e., corners, and recognize dynamic actions, i.e., gestures, in visual scenes. All required information for both tasks are provided by a local computation of optical flow; this information is precisely encoded in the primary area (V1) of the visual cortex via the selectivity of V1 neurons. We underline also that the proposed frameworks are fully incremental and could be implemented as event-based neural networks, in particular thanks to speed and direction coordinates frame based representation of the visual motion information.</p>
<p>Such polar coordinate frame based representations have been already investigated for computer vision; e.g., based on bank of Gabor filters, using whether synchronous frame-based (Lades et al., <xref ref-type="bibr" rid="B41">1993</xref>; Jain et al., <xref ref-type="bibr" rid="B37">1997</xref>; Lyons et al., <xref ref-type="bibr" rid="B51">1998</xref>, etc.) or asynchronous event-based (Akolkar et al., <xref ref-type="bibr" rid="B2">2015</xref>) visual information. Works about natural image statistics (Hyvarinen et al., <xref ref-type="bibr" rid="B36">2009</xref>) showed that similar decompositions of visual information emerge naturally from independent component analysis applied on patches collected on natural images. Recently, a work in Chandrapala and Shi (<xref ref-type="bibr" rid="B12">2016</xref>) encoding more directly local event streams as local spatiotemporal surfaces (Lagorce et al., <xref ref-type="bibr" rid="B45">2016</xref>), showed that an unsupervised learning process applied on a relatively large database acquired with a neuromorphic camera, leads to a similar result: basic and local feature extractors coding contours&#x00027; speed and direction. Moreover, other works (Cedras and Shah, <xref ref-type="bibr" rid="B10">1995</xref>; Chaudhry et al., <xref ref-type="bibr" rid="B13">2009</xref>; Ahad et al., <xref ref-type="bibr" rid="B3">2012</xref>, etc.) in frame-based vision have shown that optical flow is a valuable information to encode in features for pattern recognition tasks.</p>
<p>In addition, the work presented in this article supports the proposition that optical flow&#x00027;s speed and direction based grid is not only a powerful manner for encoding visual information in pattern recognition tasks, but it plays also a key role at a computational level when dealing with asynchronous event-based streams. Indeed we have shown that, to compute the distribution of optical flow along current edges, we need to take into account their respective dynamics, in order to ensure that the moving edges are equitably represented in the feature (whatever their own dynamics). The discretization of the visual motion information into the proposed speed- and direction-based grid allows us to incorporate directly the required speed-tuned temporal kernels in the structure of the computational architecture computing the feature. We have in addition proposed that this architecture can be implemented as a leaky integrate-and-fire neural layer, wherein neurons have then speed-tuned integration times; so it could be further integrated as the first layer in a spiking neural network using back-propagation based deep learning technique, as the one recently proposed by Lee et al. (<xref ref-type="bibr" rid="B48">2016</xref>) wherein LIF neurons are also used.</p>
<p>Finally, in the asynchronous event-based multilayer architectures proposed recently in Chandrapala and Shi (<xref ref-type="bibr" rid="B12">2016</xref>) and Lagorce et al. (<xref ref-type="bibr" rid="B45">2016</xref>), the integration times are tuned as increasing at higher layers. In addition, in our gesture recognition architecture, we have set the integration time in Bayes filters regarding the gesture durations, not the dynamics of the visual information. Further, investigations could address the following issue: when (or at what level in hierarchical models) the integration times should be tuned not regarding the dynamics of the perceived information, but other temporal considerations or dynamics, maybe related to a targeted task or action, or maybe related to other perceptive, learning, or memory functions.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>XC developed the theory for feature and performed experiments and analysis for corner detection. XC, JM, and SB designed the experiments, performed analysis and interpreted data for gesture recognition. XC wrote the article and JM, SB, and RB helped to edit the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>The authors are grateful to Jacques Chartier-Kastler for his help in collecting the database, the members of our research team who participate at it as &#x0201C;users,&#x0201D; Chronocam&#x00027;s team (<ext-link ext-link-type="uri" xlink:href="http://www.chronocam.com/">http://www.chronocam.com/</ext-link>) for designing and providing the camera, Germain Haessig for designing and 3D-printing the camera&#x00027;s supports for mobile devices, Xavier Lagorce for the photographies of the embedded camera (see <ext-link ext-link-type="uri" xlink:href="http://www.ecomode-project.eu">http://www.ecomode-project.eu</ext-link>) and Camille Simon-Chane who has checked the manuscript for spelling, grammar, punctuation, etc. This work received funding from the European Unions Horizon 2020 research and innovation programme under grant agreement N&#x000B0;644096.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adelson</surname> <given-names>E.</given-names></name> <name><surname>Movshon</surname> <given-names>J.</given-names></name></person-group> (<year>1982</year>). <article-title>Phenomenal coherence of moving visual patterns</article-title>. <source>Nature</source> <volume>200</volume>, <fpage>523</fpage>&#x02013;<lpage>525</lpage>. <pub-id pub-id-type="doi">10.1038/300523a0</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahad</surname> <given-names>M. A. R.</given-names></name> <name><surname>Tan</surname> <given-names>J. K.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Ishikawa</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Motion history image: its variants and applications</article-title>. <source>Mach. Vis. Appl.</source> <volume>23</volume>, <fpage>255</fpage>&#x02013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-010-0298-4</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Akolkar</surname> <given-names>H.</given-names></name> <name><surname>Meyer</surname> <given-names>C.</given-names></name> <name><surname>Clady</surname> <given-names>Z.</given-names></name> <name><surname>Marre</surname> <given-names>O.</given-names></name> <name><surname>Bartolozzi</surname> <given-names>C.</given-names></name> <name><surname>Panzeri</surname> <given-names>S.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>What can neuromorphic event-driven precise timing add to spike-based pattern recognition?</article-title> <source>Neural Comput.</source> <volume>27</volume>, <fpage>561</fpage>&#x02013;<lpage>593</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00703</pub-id><pub-id pub-id-type="pmid">25602775</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bair</surname> <given-names>W.</given-names></name> <name><surname>Movshon</surname> <given-names>J. A.</given-names></name></person-group> (<year>2004</year>). <article-title>Adaptive temporal integration of motion in direction-selective neurons in macaque visual cortex</article-title>. <source>J. Neurosci.</source> <volume>24</volume>, <fpage>7305</fpage>&#x02013;<lpage>7323</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0554-04.2004</pub-id><pub-id pub-id-type="pmid">15317857</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Clercq</surname> <given-names>C.</given-names></name> <name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Bartolozzi</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Event-based visual flow</article-title>. <source>IEEE Trans. Neural Netw. Learn. Systems</source> <volume>25</volume>, <fpage>407</fpage>&#x02013;<lpage>417</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2013.2273537</pub-id><pub-id pub-id-type="pmid">24807038</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brosch</surname> <given-names>T.</given-names></name> <name><surname>Tschechne</surname> <given-names>S.</given-names></name> <name><surname>Neumann</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>On event-based optical flow detection</article-title>. <source>Front. Neurosci.</source> <volume>9</volume>:<fpage>137</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2015.00137</pub-id><pub-id pub-id-type="pmid">25941470</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Camu&#x000F1;as-Mesa</surname> <given-names>L. A.</given-names></name> <name><surname>Serrano-Gotarredona</surname> <given-names>T.</given-names></name> <name><surname>Ieng</surname> <given-names>S. H.</given-names></name> <name><surname>Benosman</surname> <given-names>R. B.</given-names></name> <name><surname>Linares-Barranco</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>On the use of orientation filters for 3d reconstruction in event-driven stereo vision</article-title>. <source>Front. Neurosci.</source> <volume>8</volume>:<fpage>48</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00048</pub-id><pub-id pub-id-type="pmid">24744694</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carneiro</surname> <given-names>J.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>Event-based 3d reconstruction from neuromorphic retinas</article-title>. <source>Neural Netw.</source> <volume>45</volume>, <fpage>27</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2013.03.006</pub-id><pub-id pub-id-type="pmid">23545156</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cawley</surname> <given-names>G. C.</given-names></name></person-group> (<year>2000</year>). <article-title>On a fast, compact approximation of the exponential function</article-title>. <source>Neural Comput.</source> <volume>12</volume>, <fpage>2009</fpage>&#x02013;<lpage>2012</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300015033</pub-id><pub-id pub-id-type="pmid">10976136</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cedras</surname> <given-names>C.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>1995</year>). <article-title>Motion-based recognition: a survey</article-title>. <source>Image Vis. Comput.</source> <volume>13</volume>, <fpage>129</fpage>&#x02013;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1016/0262-8856(95)93154-K</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Censi</surname> <given-names>A.</given-names></name> <name><surname>Strubel</surname> <given-names>J.</given-names></name> <name><surname>Brandli</surname> <given-names>C.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Scaramuzza</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>Low-latency localization by active led markers tracking using a dynamic vision sensor</article-title>, in <source>IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source> (<publisher-loc>Tokyo</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>891</fpage>&#x02013;<lpage>898</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chandrapala</surname> <given-names>T. N.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name></person-group> (<year>2016</year>). <article-title>Invariant feature extraction from event based stimuli</article-title>. arXiv:1604.04327.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chaudhry</surname> <given-names>R.</given-names></name> <name><surname>Ravichandran</surname> <given-names>A.</given-names></name> <name><surname>Hager</surname> <given-names>G.</given-names></name> <name><surname>Vidal</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>Histograms of oriented optical flow and binet-cauchy kernels on nonlinear dynamical systems for the recognition of human actions</article-title>, in <source>IEEE Conference on Computer Vision and Pattern Recognition, 2009. CVPR 2009.</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1932</fpage>&#x02013;<lpage>1939</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cichocki</surname> <given-names>A.</given-names></name> <name><surname>Unbehauen</surname> <given-names>R.</given-names></name></person-group> (<year>1992</year>). <article-title>Neural networks for solving systems of linear equations and related problems</article-title>. <source>IEEE Trans. Circ. Syst. I Fundam. Theor. Appl.</source> <volume>39</volume>, <fpage>124</fpage>&#x02013;<lpage>138</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Clercq</surname> <given-names>C.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Houseini</surname> <given-names>F.</given-names></name> <name><surname>Randazzo</surname> <given-names>M.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Asynchronous visual event-based time-to-contact</article-title>. <source>Front. Neurosci.</source> <volume>8</volume>:<fpage>9</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00009</pub-id><pub-id pub-id-type="pmid">24570652</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Asynchronous event-based corner detection and matching</article-title>. <source>Neural Netw.</source> <volume>66</volume>, <fpage>91</fpage>&#x02013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2015.02.013</pub-id><pub-id pub-id-type="pmid">25828960</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dalal</surname> <given-names>N.</given-names></name> <name><surname>Triggs</surname> <given-names>B.</given-names></name></person-group> (<year>2005</year>). <article-title>Histograms of oriented gradients for human detection</article-title>, in <source>IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR&#x00027;05)</source>, <volume>Vol. 1</volume> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>886</fpage>&#x02013;<lpage>893</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Debaecker</surname> <given-names>T.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Ieng</surname> <given-names>S. H.</given-names></name></person-group> (<year>2010</year>).&#x0201C;Image sensor model using geometric algebra: from calibration to motion estimation,&#x0201D; in <source>Geometric Algebra Computing</source>, eds <person-group person-group-type="editor"><name><surname>Bayro-Corrochano</surname> <given-names>E.</given-names></name> <name><surname>Scheuermann</surname> <given-names>G.</given-names></name></person-group> (<publisher-loc>London</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>), <fpage>277</fpage>&#x02013;<lpage>297</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Lang</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Robotic goalie with 3 ms reaction time at 4% cpu load using event-based dynamic vision sensor</article-title>. <source>Front. Neurosci.</source> <volume>7</volume>:<fpage>223</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2013.00223</pub-id><pub-id pub-id-type="pmid">24311999</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Linares-Barranco</surname> <given-names>B.</given-names></name> <name><surname>Culurciello</surname> <given-names>E.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Activity-driven, event-based vision sensors</article-title>, in <source>Proceedings of 2010 IEEE International Symposium on Circuits and Systems (ISCAS)</source> (<publisher-loc>Paris</publisher-loc>), <fpage>2426</fpage>&#x02013;<lpage>2429</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deselaers</surname> <given-names>T.</given-names></name> <name><surname>Heigold</surname> <given-names>G.</given-names></name> <name><surname>Ney</surname> <given-names>H.</given-names></name></person-group> (<year>2008</year>). <article-title>Svms, gaussian mixtures, and their generative/discriminative fusion</article-title>, in <source>19th International Conference on Pattern Recognition. ICPR 2008</source> (<publisher-loc>Tampa, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dickscheid</surname> <given-names>T.</given-names></name> <name><surname>Schindler</surname> <given-names>F.</given-names></name> <name><surname>F&#x000F6;rstner</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>Coding images with local features</article-title>. <source>Int. J. Comput. Vis.</source> <volume>94</volume>, <fpage>154</fpage>&#x02013;<lpage>174</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-010-0340-z</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Drazen</surname> <given-names>D.</given-names></name> <name><surname>Lichtsteiner</surname> <given-names>P.</given-names></name> <name><surname>H&#x000E4;fliger</surname> <given-names>P.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Jensen</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Toward real-time particle tracking using an event-based dynamic vision sensor</article-title>. <source>Exp. Fluids</source> <volume>51</volume>, <fpage>1465</fpage>&#x02013;<lpage>1469</lpage>. <pub-id pub-id-type="doi">10.1007/s00348-011-1207-y</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Firouzi</surname> <given-names>M.</given-names></name> <name><surname>Conradt</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Asynchronous event-based cooperative stereo matching using neuromorphic silicon retinas</article-title>. <source>Neural Process. Lett.</source> <volume>43</volume>, <fpage>311</fpage>&#x02013;<lpage>326</lpage>. <pub-id pub-id-type="doi">10.1007/s11063-015-9434-5</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Freund</surname> <given-names>Y.</given-names></name> <name><surname>Schapire</surname> <given-names>R. E.</given-names></name></person-group> (<year>1996</year>). <article-title>Experiments with a new boosting algorithm</article-title>, in <source>Icml</source>, ed <person-group person-group-type="editor"><name><surname>Kaufmann</surname> <given-names>M.</given-names></name></person-group> (<publisher-loc>Bari</publisher-loc>), Vol. <volume>96</volume>, <fpage>148</fpage>&#x02013;<lpage>156</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Furber</surname> <given-names>S.</given-names></name> <name><surname>Lester</surname> <given-names>D.</given-names></name> <name><surname>Plana</surname> <given-names>L.</given-names></name> <name><surname>Garside</surname> <given-names>J.</given-names></name> <name><surname>Painkras</surname> <given-names>E.</given-names></name> <name><surname>Temple</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Overview of the spinnaker system architecture</article-title>. <source>IEEE Trans. Comput.</source> <volume>62</volume>, <fpage>2454</fpage>&#x02013;<lpage>2467</lpage>. <pub-id pub-id-type="doi">10.1109/TC.2012.142</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gauglitz</surname> <given-names>S.</given-names></name> <name><surname>H&#x000F6;llerer</surname> <given-names>T.</given-names></name> <name><surname>Turk</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Evaluation of interest point detectors and feature descriptors for visual tracking</article-title>. <source>Int. J. Comput. Vis.</source> <volume>94</volume>, <fpage>335</fpage>&#x02013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-011-0431-5</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Kistler</surname> <given-names>W. M.</given-names></name></person-group> (<year>2002</year>). <source>Spiking Neuron Models: Single Neurons, Populations, Plasticity</source>. <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gil</surname> <given-names>A.</given-names></name> <name><surname>Mozos</surname> <given-names>O. M.</given-names></name> <name><surname>Ballesta</surname> <given-names>M.</given-names></name> <name><surname>Reinoso</surname> <given-names>O.</given-names></name></person-group> (<year>2010</year>). <article-title>A comparative evaluation of interest point detectors and local descriptors for visual slam</article-title>. <source>Mach. Vis. Appl.</source> <volume>21</volume>, <fpage>905</fpage>&#x02013;<lpage>920</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-009-0195-x</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Grabner</surname> <given-names>H.</given-names></name> <name><surname>Roth</surname> <given-names>P. M.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name></person-group> (<year>2007</year>). <article-title>Eigenboosting: combining discriminative and generative information</article-title>, in <source>2007 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Minneapolis, MN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Shao</surname> <given-names>L.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Shotton</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Enhanced computer vision with microsoft kinect sensor: a review</article-title>. <source>IEEE Trans. Cybernet.</source> <volume>43</volume>, <fpage>1318</fpage>&#x02013;<lpage>1334</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2013.2265378</pub-id><pub-id pub-id-type="pmid">23807480</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Harris</surname> <given-names>C.</given-names></name> <name><surname>Stephens</surname> <given-names>M.</given-names></name></person-group> (<year>1988</year>). <article-title>A combined corner and edge detector</article-title>, in <source>Proceedings of the 4th Alvey Vision Conference</source> (<publisher-loc>Manchester</publisher-loc>), <fpage>147</fpage>&#x02013;<lpage>151</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Holub</surname> <given-names>A. D.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2005</year>). <article-title>Combining generative models and fisher kernels for object recognition</article-title>, in <source>Tenth IEEE International Conference on Computer Vision (ICCV&#x00027;05)</source>, <volume>Vol. 1</volume> (<publisher-loc>Beijing</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>136</fpage>&#x02013;<lpage>143</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hubel</surname> <given-names>D. H.</given-names></name> <name><surname>Wiesel</surname> <given-names>T. N.</given-names></name></person-group> (<year>1962</year>). <article-title>Receptive fields, binocular interaction and functional architecture in the cat&#x00027;s visual cortex</article-title>. <source>J. Physiol.</source> <volume>160</volume>, <fpage>106</fpage>&#x02013;<lpage>154</lpage>. <pub-id pub-id-type="pmid">14449617</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hubel</surname> <given-names>D. H.</given-names></name> <name><surname>Wiesel</surname> <given-names>T. N.</given-names></name></person-group> (<year>1968</year>). <article-title>Receptive fields and functional architecture of monkey striate cortex</article-title>. <source>J. Physiol.</source> <volume>195</volume>, <fpage>215</fpage>&#x02013;<lpage>243</lpage>. <pub-id pub-id-type="pmid">4966457</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hyvarinen</surname> <given-names>A.</given-names></name> <name><surname>Hurri</surname> <given-names>J.</given-names></name> <name><surname>Hoyer</surname> <given-names>P. O.</given-names></name></person-group> (<year>2009</year>). <source>Natural Image Statistics: A Probabilistic Approach to Early Computational Vision</source>. <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>A. K.</given-names></name> <name><surname>Ratha</surname> <given-names>N. K.</given-names></name> <name><surname>Lakshmanan</surname> <given-names>S.</given-names></name></person-group> (<year>1997</year>). <article-title>Object detection using gabor filters</article-title>. <source>Pattern Recognit.</source> <volume>30</volume>, <fpage>295</fpage>&#x02013;<lpage>309</lpage>. <pub-id pub-id-type="doi">10.1016/S0031-3203(96)00068-4</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jing</surname> <given-names>Y.</given-names></name> <name><surname>Pavlovi&#x00107;</surname> <given-names>V.</given-names></name> <name><surname>Rehg</surname> <given-names>J. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Boosted bayesian network classifiers</article-title>. <source>Mach. Learn.</source> <volume>73</volume>, <fpage>155</fpage>&#x02013;<lpage>184</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-008-5065-7</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kime</surname> <given-names>S.</given-names></name> <name><surname>Galluppi</surname> <given-names>F.</given-names></name> <name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Lorenceau</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Psychophysical assessment of perceptual performance with varying display frame rates</article-title>. <source>J. Disp. Technol.</source> <volume>12</volume>, <fpage>1372</fpage>&#x02013;<lpage>1382</lpage>. <pub-id pub-id-type="doi">10.1109/JDT.2016.2603222</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kime</surname> <given-names>S.</given-names></name> <name><surname>Galluppi</surname> <given-names>F.</given-names></name> <name><surname>Lorenceau</surname> <given-names>J.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Exploring speed discrimination of visual stimuli at a high frame rate</article-title>, in <source>Annual Meeting of the Society For Neuroscience(SFN)</source> (<publisher-loc>Washington, DC</publisher-loc>).</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lades</surname> <given-names>M.</given-names></name> <name><surname>Vorbruggen</surname> <given-names>J. C.</given-names></name> <name><surname>Buhmann</surname> <given-names>J.</given-names></name> <name><surname>Lange</surname> <given-names>J.</given-names></name> <name><surname>von der Malsburg</surname> <given-names>C.</given-names></name> <name><surname>Wurtz</surname> <given-names>R. P.</given-names></name> <etal/></person-group>. (<year>1993</year>). <article-title>Distortion invariant object recognition in the dynamic link architecture</article-title>. <source>IEEE Trans. Comp.</source> <volume>42</volume>, <fpage>300</fpage>&#x02013;<lpage>311</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Stick: spike time interval computational kernel, a framework for general purpose computation using neurons, precise timing, delays, and synchrony</article-title>. <source>Neural Comput.</source> <volume>27</volume>, <fpage>2261</fpage>&#x02013;<lpage>2317</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00783</pub-id><pub-id pub-id-type="pmid">26378879</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>Event-based features for robotic vision</article-title>, in <source>IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source> (<publisher-loc>Tokyo</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4214</fpage>&#x02013;<lpage>4219</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Meyer</surname> <given-names>C.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Filliat</surname> <given-names>D.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Asynchronous event-based multikernel algorithm for high-speed visual features tracking</article-title>. <source>Trans. Neural Netw. Learn. Syst.</source> <volume>26</volume>, <fpage>1710</fpage>&#x02013;<lpage>1720</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2014.2352401</pub-id><pub-id pub-id-type="pmid">25248193</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Orchard</surname> <given-names>G.</given-names></name> <name><surname>Galluppi</surname> <given-names>F.</given-names></name> <name><surname>Shi</surname> <given-names>B.</given-names></name> <name><surname>Benosman</surname> <given-names>R</given-names></name></person-group> (<year>2016</year>). <article-title>Hots: a hierarchy of event-based time-surfaces for pattern recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://ieeexplore.ieee.org/abstract/document/7508476/">http://ieeexplore.ieee.org/abstract/document/7508476/</ext-link> <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2574707</pub-id><pub-id pub-id-type="pmid">27411216</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laptev</surname> <given-names>I.</given-names></name></person-group> (<year>2005</year>). <article-title>On space-time interest points</article-title>. <source>Int. J. Comput. Vis.</source> <volume>64</volume>, <fpage>107</fpage>&#x02013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2003.1238378</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lasserre</surname> <given-names>J. A.</given-names></name> <name><surname>Bishop</surname> <given-names>C. M.</given-names></name> <name><surname>Minka</surname> <given-names>T. P.</given-names></name></person-group> (<year>2006</year>). <article-title>Principled hybrids of generative and discriminative models</article-title>, in <source>2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR&#x00027;06)</source>, <volume>Vol. 1</volume> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>87</fpage>&#x02013;<lpage>94</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>J. H.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Training deep spiking neural networks using backpropagation</article-title>. <source>Front. Neurosci.</source> <volume>10</volume>:<fpage>508</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2016.00508</pub-id><pub-id pub-id-type="pmid">27877107</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>J. H.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name> <name><surname>Park</surname> <given-names>P. K.</given-names></name> <name><surname>Shin</surname> <given-names>C.-W.</given-names></name> <name><surname>Ryu</surname> <given-names>H. E.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Real-time gesture interface based on event-driven processing from stereo silicon retinas</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>25</volume>, <fpage>2250</fpage>&#x02013;<lpage>2263</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2014.2308551</pub-id><pub-id pub-id-type="pmid">25420246</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lichtsteiner</surname> <given-names>P.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name></person-group> (<year>2008</year>). <article-title>A 128<sup>&#x0002A;</sup>128 120dB 15us latency asynchronous temporal contrast vision sensor</article-title>. <source>IEEE J. Solid State Circ.</source> <volume>43</volume>, <fpage>566</fpage>&#x02013;<lpage>576</lpage>. <pub-id pub-id-type="doi">10.1109/JSSC.2007.914337</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lyons</surname> <given-names>M.</given-names></name> <name><surname>Akamatsu</surname> <given-names>S.</given-names></name> <name><surname>Kamachi</surname> <given-names>M.</given-names></name> <name><surname>Gyoba</surname> <given-names>J.</given-names></name></person-group> (<year>1998</year>). <article-title>Coding facial expressions with gabor wavelets</article-title>, in <source>Proceedings of Third IEEE International Conference on Automatic Face and Gesture Recognition</source> (<publisher-loc>Nara</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>200</fpage>&#x02013;<lpage>205</lpage>.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolajczyk</surname> <given-names>K.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title>A performance evaluation of local descriptors</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>27</volume>, <fpage>1615</fpage>&#x02013;<lpage>1630</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2005.188</pub-id><pub-id pub-id-type="pmid">16237996</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Milde</surname> <given-names>M.</given-names></name> <name><surname>Bertrand</surname> <given-names>O. J. N.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Egelhaaf</surname> <given-names>M.</given-names></name> <name><surname>Chicca</surname> <given-names>E.</given-names></name></person-group> (<year>2015</year>). <article-title>Bioinspired event-driven collision avoidance algorithm based on optic flow</article-title>, in <source>Event-Based Control, Communication, and Signal Processing (EBCCSP)</source> (<publisher-loc>Krakow</publisher-loc>).</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moeslund</surname> <given-names>T. B.</given-names></name> <name><surname>Hilton</surname> <given-names>A.</given-names></name> <name><surname>Kruger</surname> <given-names>V.</given-names></name></person-group> (<year>2006</year>). <article-title>A survey of advances in vision-based human motion capture and analysis</article-title>. <source>Comput. Vis. Image Underst.</source> <volume>104</volume>, <fpage>90</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2006.08.002</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mokhtarian</surname> <given-names>F.</given-names></name> <name><surname>Mohanna</surname> <given-names>F.</given-names></name></person-group> (<year>2006</year>). <article-title>Performance evaluation of corner detectors using consistency and accuracy measureness</article-title>. <source>Comput. Vis. Image Understand.</source> <volume>102</volume>, <fpage>81</fpage>&#x02013;<lpage>94</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2005.11.001</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mokhtarian</surname> <given-names>F.</given-names></name> <name><surname>Suomela</surname> <given-names>R.</given-names></name></person-group> (<year>1998</year>). <article-title>Robust image corner detection through curvature scale space</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>20</volume>, <fpage>1376</fpage>&#x02013;<lpage>1381</lpage>.</citation>
</ref>
<ref id="B57">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Moravec</surname> <given-names>H.</given-names></name></person-group> (<year>1980</year>). <source>Obstacle Avoidance and Navigation in the Real World by a Seeing Robot Rover.</source> Technical report, CMU-RI-TR-80-03, Robotics Institute, Carnegie Mellon University and doctoral dissertation, Stanford University.</citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreels</surname> <given-names>P.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2007</year>). <article-title>Evaluation of features detectors and descriptors based on 3d objects</article-title>. <source>Int. J. Comput. Vis.</source> <volume>73</volume>, <fpage>263</fpage>&#x02013;<lpage>284</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-006-9967-1</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mueggler</surname> <given-names>E.</given-names></name> <name><surname>Baumli</surname> <given-names>N.</given-names></name> <name><surname>Fontana</surname> <given-names>F.</given-names></name> <name><surname>Scaramuzza</surname> <given-names>D.</given-names></name></person-group> (<year>2015a</year>). <article-title>Towards evasive maneuvers with quadrotors using dynamic vision sensors</article-title>, in <source>European Conference on Mobile Robots (ECMR)</source> (<publisher-loc>Paris</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mueggler</surname> <given-names>E.</given-names></name> <name><surname>Forster</surname> <given-names>C.</given-names></name> <name><surname>Baumli</surname> <given-names>N.</given-names></name> <name><surname>Gallego</surname> <given-names>G.</given-names></name> <name><surname>Scaramuzza</surname> <given-names>D.</given-names></name></person-group> (<year>2015b</year>). <article-title>Lifetime estimation of events from dynamic vision sensors</article-title>, in <source>2015 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4874</fpage>&#x02013;<lpage>4881</lpage>.</citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Negri</surname> <given-names>P.</given-names></name> <name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Hanif</surname> <given-names>S. M.</given-names></name> <name><surname>Prevost</surname> <given-names>L.</given-names></name></person-group> (<year>2008</year>). <article-title>A cascade of boosted generative and discriminative classifiers for vehicle detection</article-title>. <source>EURASIP J. Adv. Signal Process.</source> <volume>2008</volume>:<fpage>136</fpage>. <pub-id pub-id-type="doi">10.1155/2008/782432</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ni</surname> <given-names>Z.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>R&#x000E9;gnier</surname> <given-names>S.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Visual tracking using neuromorphic asynchronous event-based cameras</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00720</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ni</surname> <given-names>Z.</given-names></name> <name><surname>Pacoret</surname> <given-names>C.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>R&#x000E9;gnier</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <source>Haptic Feedback Teleoperation of Optical Tweezers</source>. <publisher-name>John Wiley and Sons</publisher-name>.</citation>
</ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noble</surname> <given-names>J.</given-names></name></person-group> (<year>1988</year>). <article-title>Finding corners</article-title>. <source>Image Vis. Comput.</source> <volume>6</volume>, <fpage>121</fpage>&#x02013;<lpage>128</lpage>.</citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orban</surname> <given-names>G. A.</given-names></name> <name><surname>Wolf</surname> <given-names>J. d.</given-names></name> <name><surname>Maes</surname> <given-names>H.</given-names></name></person-group> (<year>1984</year>). <article-title>Factors influencing velocity coding in the human visual system</article-title>. <source>Vis. Res.</source> <volume>24</volume>, <fpage>33</fpage>&#x02013;<lpage>39</lpage>. <pub-id pub-id-type="pmid">6695505</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orchard</surname> <given-names>G.</given-names></name> <name><surname>Etienne-Cummings</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Bioinspired visual motion estimation</article-title>. <source>Proc. IEEE</source> <volume>102</volume>, <fpage>1520</fpage>&#x02013;<lpage>1536</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2014.2346763</pub-id></citation>
</ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Orchard</surname> <given-names>G.</given-names></name> <name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Furber</surname> <given-names>S. B.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Galluppi</surname> <given-names>F.</given-names></name></person-group> (<year>2015a</year>). <article-title>Real-time event-driven spiking neural network object recognition on the spinnaker platform</article-title>, in <source>IEEE International Symposium on Circuits and Systems (ISCAS)</source> (<publisher-loc>Lisbon</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2413</fpage>&#x02013;<lpage>2416</lpage>.</citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orchard</surname> <given-names>G.</given-names></name> <name><surname>Meyer</surname> <given-names>C.</given-names></name> <name><surname>Etienne-Cummings</surname> <given-names>R.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Thakor</surname> <given-names>N.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015b</year>). <article-title>Hfirst: a temporal approach to object recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>37</volume>, <fpage>2028</fpage>&#x02013;<lpage>2040</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2392947</pub-id><pub-id pub-id-type="pmid">26353184</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pana&#x000EF;t&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Usciati</surname> <given-names>T.</given-names></name> <name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Haliyo</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>An experimental study of the kinect&#x00027;s depth sensor</article-title>, in <source>IEEE International Symposium on Robotic and Sensors Environment</source> (<publisher-loc>Montreal</publisher-loc>).</citation>
</ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>S.</given-names></name> <name><surname>Ahmad</surname> <given-names>M.</given-names></name> <name><surname>Seung-Hak</surname> <given-names>R.</given-names></name> <name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name></person-group> (<year>2004</year>). <article-title>Image corner detection using radon transform</article-title>, in <source>Computational Science and Its Applications, Vol. 3046, Lecture Notes in Computer Science</source>, eds <person-group person-group-type="editor"><name><surname>Lagano</surname> <given-names>A.</given-names></name> <name><surname>Gavrilova</surname> <given-names>M.</given-names></name> <name><surname>Kumar</surname> <given-names>V.</given-names></name> <name><surname>Mun</surname> <given-names>Y.</given-names></name> <name><surname>Tan</surname> <given-names>C.</given-names></name> <name><surname>Gervasi</surname> <given-names>O.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>948</fpage>&#x02013;<lpage>955</lpage>.</citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>P&#x000E9;rez-Carrasco</surname> <given-names>J. A.</given-names></name> <name><surname>Zhao</surname> <given-names>B.</given-names></name> <name><surname>Serrano</surname> <given-names>C.</given-names></name> <name><surname>Acha</surname> <given-names>B.</given-names></name> <name><surname>Serrano-Gotarredona</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Mapping from frame-driven to frame-free event-driven vision systems by low-rate rate coding and coincidence processing - application to feedforward convnets</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>35</volume>, <fpage>2706</fpage>&#x02013;<lpage>2719</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2013.71</pub-id><pub-id pub-id-type="pmid">24051730</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poppe</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <article-title>Vision-based Human motion analysis: an overview</article-title>. <source>Comput. Vis. Image Underst.</source> <volume>108</volume>, <fpage>4</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2006.10.016</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poppe</surname> <given-names>R.</given-names></name></person-group> (<year>2010</year>). <article-title>A survey on vision-based human action recognition</article-title>. <source>Image Vis. Comput.</source> <volume>28</volume>, <fpage>976</fpage>&#x02013;<lpage>990</lpage>. <pub-id pub-id-type="doi">10.1016/j.imavis.2009.11.014</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Posch</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Bioinspired vision sensing</article-title>, <source>Biologically Inspired Computer Vision: Fundamentals and Applications</source>, eds <person-group person-group-type="editor"><name><surname>Crist&#x000F3;bal</surname> <given-names>G.</given-names></name> <name><surname>Perrinet</surname> <given-names>L.</given-names></name> <name><surname>Keil</surname> <given-names>M. S.</given-names></name></person-group> (<publisher-loc>Weinheim</publisher-loc>: <publisher-name>Wiley-VCH Verlag GmbH &#x00026; Co.</publisher-name>). <pub-id pub-id-type="doi">10.1002/9783527680863.ch2</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Matolin</surname> <given-names>D.</given-names></name> <name><surname>Wohlgenannt</surname> <given-names>R.</given-names></name></person-group> (<year>2010</year>). <article-title>High-DR frame-free PWM imaging with asynchronous AER intensity encoding and focal-plane temporal redundancy suppression</article-title>, in <source>Proceedings of 2010 IEEE International Symposium on Circuits and Systems (ISCAS)</source> (<publisher-loc>Paris</publisher-loc>), <fpage>2430</fpage>&#x02013;<lpage>2433</lpage>.</citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prevost</surname> <given-names>L.</given-names></name> <name><surname>Oudot</surname> <given-names>L.</given-names></name> <name><surname>Moises</surname> <given-names>A.</given-names></name> <name><surname>Michel-Sendis</surname> <given-names>C.</given-names></name> <name><surname>Milgram</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>Hybrid generative/discriminative classifier for unconstrained character recognition</article-title>. <source>Pattern Recognit. Lett.</source> <volume>26</volume>, <fpage>1840</fpage>&#x02013;<lpage>1848</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2005.03.005</pub-id></citation>
</ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Priebe</surname> <given-names>N. J.</given-names></name> <name><surname>Lisberger</surname> <given-names>S. G.</given-names></name> <name><surname>Movshon</surname> <given-names>J. A.</given-names></name></person-group> (<year>2006</year>). <article-title>Tuning for spatiotemporal frequency and speed in directionally selective neurons of macaque striate cortex</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>2941</fpage>&#x02013;<lpage>2950</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3936-05.2006</pub-id><pub-id pub-id-type="pmid">16540571</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogister</surname> <given-names>P.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Lichtsteiner</surname> <given-names>P.</given-names></name> <name><surname>Delbr&#x000FC;ck</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>Asynchronous event-based binocular stereo matching</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>23</volume>, <fpage>347</fpage>&#x02013;<lpage>353</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2011.2180025</pub-id><pub-id pub-id-type="pmid">24808513</pub-id></citation>
</ref>
<ref id="B79">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rosten</surname> <given-names>E.</given-names></name> <name><surname>Drummond</surname> <given-names>T.</given-names></name></person-group> (<year>2006</year>). <article-title>Machine learning for high-speed corner detection</article-title>, in <source>European Conference on Computer Vision</source> (<publisher-loc>Graz</publisher-loc>), Vol. <volume>1</volume>, <fpage>430</fpage>&#x02013;<lpage>443</lpage>.</citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schapire</surname> <given-names>R. E.</given-names></name> <name><surname>Freund</surname> <given-names>Y.</given-names></name> <name><surname>Bartlett</surname> <given-names>P.</given-names></name> <name><surname>Lee</surname> <given-names>W. S.</given-names></name></person-group> (<year>1998</year>). <article-title>Boosting the margin: A new explanation for the effectiveness of voting methods</article-title>. <source>Ann. Stat.</source> <volume>26</volume>, <fpage>1651</fpage>&#x02013;<lpage>1686</lpage>.</citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schraudolph</surname> <given-names>N. N.</given-names></name></person-group> (<year>1999</year>). <article-title>A fast, compact approximation of the exponential function</article-title>. <source>Neural Comput.</source> <volume>11</volume>, <fpage>853</fpage>&#x02013;<lpage>862</lpage>. <pub-id pub-id-type="pmid">10226185</pub-id></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schreiber</surname> <given-names>S.</given-names></name> <name><surname>Fellous</surname> <given-names>J. M.</given-names></name> <name><surname>Whitmer</surname> <given-names>D.</given-names></name> <name><surname>Tiesinga</surname> <given-names>P.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2003</year>). <article-title>A new correlation based measure of spike timing reliability</article-title>. <source>Neurocomputing</source> <volume>52</volume>, <fpage>925</fpage>&#x02013;<lpage>931</lpage>. <pub-id pub-id-type="doi">10.1016/S0925-2312(02)00838-X</pub-id><pub-id pub-id-type="pmid">20740049</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Serrano-Gotarredona</surname> <given-names>T.</given-names></name> <name><surname>Linares-Barranco</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>A 128x128 1.5% 20 contrast sensitivity 0.9% 20 fpn 3 &#x003BC;s latency 4 mw asynchronous frame-free dynamic vision sensor using transimpedance preamplifiers</article-title>. <source>IEEE J. Solid State Circ.</source> <volume>48</volume>, <fpage>827</fpage>&#x02013;<lpage>838</lpage>.</citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simon-Chane</surname> <given-names>C.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Posch</surname> <given-names>C.</given-names></name> <name><surname>Benosman</surname> <given-names>R. B.</given-names></name></person-group> (<year>2016</year>). <article-title>Event-based tone mapping for asynchronous time-based image sensor</article-title>. <source>Front. Neurosci.</source> <volume>10</volume>:<fpage>391</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2016.00391</pub-id><pub-id pub-id-type="pmid">27642275</pub-id></citation>
</ref>
<ref id="B85">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thrun</surname> <given-names>S.</given-names></name> <name><surname>Burgard</surname> <given-names>W.</given-names></name> <name><surname>Fox</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <source>Probabilistic Robotics</source>. <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B86">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Valeiras</surname> <given-names>D. R.</given-names></name> <name><surname>Lagorce</surname> <given-names>X.</given-names></name> <name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Bartolozzi</surname> <given-names>C.</given-names></name> <name><surname>Ieng</surname> <given-names>S.-H.</given-names></name> <name><surname>Benosman</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>An asynchronous neuromorphic event-driven visual part-based shape tracking</article-title>. <source>Trans. Neural Netw. Learn. Syst.</source> <volume>26</volume>, <fpage>3045</fpage>&#x02013;<lpage>3059</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2015.2401834</pub-id><pub-id pub-id-type="pmid">25794399</pub-id></citation>
</ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Rossum</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>A novel spike distance</article-title>. <source>Neural Comput.</source> <volume>13</volume>, <fpage>751</fpage>&#x02013;<lpage>763</lpage>. <pub-id pub-id-type="doi">10.1162/089976601300014321</pub-id><pub-id pub-id-type="pmid">11255567</pub-id></citation>
</ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Luo</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Discriminative and generative vocabulary tree: with application to vein image authentication and recognition</article-title>. <source>Image Vis. Comput.</source> <volume>34</volume>, <fpage>51</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1016/j.imavis.2014.10.014</pub-id></citation>
</ref>
<ref id="B89">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Clady</surname> <given-names>X.</given-names></name> <name><surname>Granata</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>A human detection system for proxemics interaction</article-title>, in <source>Proceedings of the 6th International Conference on Human-Robot Interaction</source> (<publisher-loc>Lausanne</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>285</fpage>&#x02013;<lpage>286</lpage>.</citation>
</ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>H.</given-names></name> <name><surname>Hu</surname> <given-names>H.</given-names></name></person-group> (<year>2008</year>). <article-title>Human motion tracking for rehabilitation&#x00027;s survey</article-title>. <source>Biomed. Signal Process. Control</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.bspc.2007.09.001</pub-id></citation>
</ref>
</ref-list>
</back>
</article>