<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mar. Sci.</journal-id>
<journal-title>Frontiers in Marine Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mar. Sci.</abbrev-journal-title>
<issn pub-type="epub">2296-7745</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmars.2021.785357</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Marine Science</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Revealing Sea Turtle Behavior in Relation to Fishing Gear Using Color-Coded Spatiotemporal Motion Patterns With Deep Neural Networks</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Reavis</surname> <given-names>Janie L.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1388676/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Demir</surname> <given-names>H. Seckin</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1522450/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Witherington</surname> <given-names>Blair E.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Bresette</surname> <given-names>Michael J.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Blain Christen</surname> <given-names>Jennifer</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/338994/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Senko</surname> <given-names>Jesse F.</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/394567/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ozev</surname> <given-names>Sule</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Life Sciences, Arizona State University</institution>, <addr-line>Tempe, AZ</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Electrical, Computer and Energy Engineering, Arizona State University</institution>, <addr-line>Tempe, AZ</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Inwater Research Group</institution>, <addr-line>Jensen Beach, FL</addr-line>, <country>United States</country></aff>
<aff id="aff4"><sup>4</sup><institution>School for the Future of Innovation in Society, Arizona State University</institution>, <addr-line>Tempe, AZ</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: David M. P. Jacoby, Lancaster University, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Duane Edgington, Monterey Bay Aquarium Research Institute (MBARI), United States; Gail Schofield, Queen Mary University of London, United Kingdom</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Janie L. Reavis <email>jreavis3&#x00040;asu.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Marine Megafauna, a section of the journal Frontiers in Marine Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>25</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>785357</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Reavis, Demir, Witherington, Bresette, Blain Christen, Senko and Ozev.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Reavis, Demir, Witherington, Bresette, Blain Christen, Senko and Ozev</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license></permissions>
<abstract><p>Incidental capture, or bycatch, of marine species is a global conservation concern. Interactions with fishing gear can cause mortality in air-breathing marine megafauna, including sea turtles. Despite this, interactions between sea turtles and fishing gear&#x02014;from a behavior standpoint&#x02014;are not sufficiently documented or described in the literature. Understanding sea turtle behavior in relation to fishing gear is key to discovering how they become entangled or entrapped in gear. This information can also be used to reduce fisheries interactions. However, recording and analyzing these behaviors is difficult and time intensive. In this study, we present a machine learning-based sea turtle behavior recognition scheme. The proposed method utilizes visual object tracking and orientation estimation tasks to extract important features that are used for recognizing behaviors of interest with green turtles (<italic>Chelonia mydas</italic>) as the study subject. Then, these features are combined in a color-coded feature image that represents the turtle behaviors occurring in a limited time frame. These spatiotemporal feature images are used along a deep convolutional neural network model to recognize the desired behaviors, specifically evasive behaviors which we have labeled &#x0201C;reversal&#x0201D; and &#x0201C;U-turn.&#x0201D; Experimental results show that the proposed method achieves an average F1 score of 85% in recognizing the target behavior patterns. This method is intended to be a tool for discovering why sea turtles become entangled in gillnet fishing gear.</p></abstract>
<kwd-group>
<kwd>green turtle</kwd>
<kwd><italic>Chelonia mydas</italic></kwd>
<kwd>behavior recognition</kwd>
<kwd>color-coding</kwd>
<kwd>spatiotemporal features</kwd>
<kwd>neural network</kwd>
<kwd>machine learning</kwd>
<kwd>motion</kwd>
</kwd-group>
<contract-num rid="cn001">1837473</contract-num>
<contract-sponsor id="cn001">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="3"/>
<equation-count count="1"/>
<ref-count count="38"/>
<page-count count="8"/>
<word-count count="5734"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Incidental capture of non-target animal species, termed bycatch, in fisheries is a global ecological threat to marine wildlife (Estes et al., <xref ref-type="bibr" rid="B14">2011</xref>). Fisheries bycatch poses a threat to air-breathing animals such as sea turtles. One such gear, gillnets, can create an ecological barrier that does not naturally occur, so there is likely no evolutionary mechanism that causes avoidance (Casale, <xref ref-type="bibr" rid="B7">2011</xref>). Various approaches have been proposed to reduce bycatch rates of sea turtles and other marine megafauna (Wang et al., <xref ref-type="bibr" rid="B35">2010</xref>; Lucchetti et al., <xref ref-type="bibr" rid="B25">2019</xref>; Demir et al., <xref ref-type="bibr" rid="B11">2020</xref>). Attempted solutions include: marine policy that sets bycatch limits for fisheries (Moore et al., <xref ref-type="bibr" rid="B26">2009</xref>); acoustic deterrents similar to pingers used to prevent dolphin bycatch; buoyless nets and illuminated nets, which have shown promising results for reducing bycatch in coastal net fisheries (Wang et al., <xref ref-type="bibr" rid="B35">2010</xref>; Peckham et al., <xref ref-type="bibr" rid="B28">2016</xref>). These bycatch reduction approaches can involve changing the technical design of gear or introducing novel visual or acoustic stimuli, which also changes gear configuration. However, as a part of the design process, effectiveness of different types of stimuli must be analyzed by observing the associated behavioral response of sea turtles, which has not been clearly documented in previous studies. Analyzing sea turtle interactions with fishing gear and bycatch reduction technologies (BRTs) is not an easy task, since it requires the researcher to monitor the experiment underwater for long periods while identifying and recording sea turtle behaviors and ensuring the study subject&#x00027;s safety. Even when experiments are recorded with GoPros, short battery life requires constant monitoring of each camera view, and the subsequent manual behavioral analysis is time-intensive for researchers. Fortunately, with the developments in computer vision-based approaches, recognition of certain behaviors can be performed automatically after training this convolutional neural network with behavioral data.</p>
<p>Various approaches have been proposed to complete the behavior recognition task for different applications involving humans or animals (Bodor et al., <xref ref-type="bibr" rid="B5">2003</xref>; Porto et al., <xref ref-type="bibr" rid="B29">2013</xref>; Ijjina and Chalavadi, <xref ref-type="bibr" rid="B21">2017</xref>; Nweke et al., <xref ref-type="bibr" rid="B27">2018</xref>; Yang et al., <xref ref-type="bibr" rid="B38">2018</xref>; Chakravarty et al., <xref ref-type="bibr" rid="B8">2019</xref>). Some of the recognition algorithms analyze the data captured using wearable sensors (Nweke et al., <xref ref-type="bibr" rid="B27">2018</xref>; Chakravarty et al., <xref ref-type="bibr" rid="B8">2019</xref>). While the sensors used in these type of experiments provide valuable information about the activities of interest, they are not applicable in our context as they need to be located on the subject&#x00027;s body in a controlled environment. Various methods use vision based approaches for the behavior recognition task (Bodor et al., <xref ref-type="bibr" rid="B5">2003</xref>; Porto et al., <xref ref-type="bibr" rid="B29">2013</xref>; Ijjina and Chalavadi, <xref ref-type="bibr" rid="B21">2017</xref>; Yang et al., <xref ref-type="bibr" rid="B38">2018</xref>). Earlier examples of the vision based methods employ hand-crafted features for analyzing the activities (Bodor et al., <xref ref-type="bibr" rid="B5">2003</xref>; Porto et al., <xref ref-type="bibr" rid="B29">2013</xref>). While these approaches can perform well for differentiating basic behaviors, they are not very efficient in recognizing complex activities. With the advancements in the machine learning field, recent studies employ deep neural networks (DNN) successfully for the activity recognition task (Ijjina and Chalavadi, <xref ref-type="bibr" rid="B21">2017</xref>; Yang et al., <xref ref-type="bibr" rid="B38">2018</xref>). Although DNNs provide powerful representations to analyze complex data sets, end-to-end training approaches usually require large amounts of data samples and a large number of network coefficients. In this study, we propose a hybrid approach for the sea turtle behavior recognition task. We use domain knowledge for determining base features to recognize certain behaviors and convert them into color-coded spatiotemporal 3-D images to train deep convolutional neural networks (CNN). In our application, we are specifically interested in recognizing &#x0201C;U-Turn&#x0201D; and &#x0201C;Reversal&#x0201D; behaviors of turtles, since they are important indicators of effectiveness of the given stimuli. In order to recognize these behaviors, we combine turtle location, velocity, and orientation information in spatiotemporal images and use these images as inputs to a CNN architecture.</p>
<p>In the U-turn behavior, the turtle makes a u-shaped maneuver in a short amount of time possibly due to an external visual stimulus. In Reversal behavior, the turtle moves backwards while facing forward rather than changing its orientation. These are avoidance behaviors exhibited by sea turtles when faced with a barrier or other deterrent. To differentiate these behaviors from each other and from other motion patterns, we use turtle location, speed, and orientation information. In order to extract those features and combine them as an input to a deep neural network based architecture, we propose the recognition system shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Overview of the proposed approach.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-08-785357-g0001.tif"/>
</fig>
<p>Here, we explain how we conduct the physical experiments and provide an overview of the proposed behavior recognition framework and explain the functional blocks. We then present the results of our comparative study for the object tracking task on the turtle dataset that we collected. We then report the performance of the proposed orientation estimation network and behavior recognition network followed by an explanation of the anticipated results and utility for conservation purposes.</p></sec>
<sec id="s2">
<title>2. Materials and Experimental Set Up</title>
<sec>
<title>2.1. Animal Acquisition and Facility Maintenance</title>
<p>All sea turtles used in this study were captured by Inwater Research Group (IRG) via dip net, entangling net, or hand capture after entrainment in the intake canal at the St. Lucie Nuclear Power Plant in Jensen Beach, FL. Capture of these turtles is necessary for returning them to the open ocean. For our choice trials, we included healthy juvenile and subadult green (<italic>C. mydas</italic>) turtles with a standard carapace length of less than 78 cm. After the IRG team removed turtles from the canal and collected biometric data, all turtles were kept in separate 6 ft diameter holding tanks with circulating seawater from the canal. Turtles were not held for more than 72 h.</p></sec>
<sec>
<title>2.2. Computer Setup</title>
<p>Captured data have been processed using a computer with Intel(R) Core(TM) i7-9750H processor and NVIDIA RTX 2070 GPU unit and 16GB RAM. For deep neural network architectures, we have used Keras Libraries<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>.</p></sec></sec>
<sec sec-type="methods" id="s3">
<title>3. Methods</title>
<sec>
<title>3.1. Animal Behavior Experiments and Analysis</title>
<p>We conducted all tank experiments in a 13.9 x 2.3 x 1.5 m concrete tank beside the intake canal at the St. Lucie Nuclear Power Plant in Jensen Beach, Florida (<xref ref-type="fig" rid="F2">Figure 2</xref>). The two treatments we used in the development of this method consisted of a gillnet vs. no gillnet set up during the day and at night, meaning a turtle was given the choice between a pathway with a gillnet fully blocking it or a pathway with nothing in it (see <xref ref-type="fig" rid="F2">Figure 2</xref>). The variable being changed is time of day with darkness being the most important factor in nighttime experiments. Each turtle was used in three consecutive 15- min trials with the same treatment. All trials were recorded using GoPro Hero8 cameras from 4 different viewpoints, although this study focused on behaviors recorded from the primary overhead view, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Turtle behavior was analyzed from the recordings rather than in real-time due to the need to monitor turtle safety. Here, we specifically focus on the novel turtle avoidance behavior identified in relation to the gillnet deployed in the treatment sector: Reversal and U-turn. A Reversal occurred when a turtle made contact with the gillnet and then escaped by moving backward with its rear flippers and maintaining a forward-facing orientation. A U-turn involved a 180 degree turn within a 3-s period. Here, we only classify U-turns that occur near the barrier of interest (i.e., the gillnet or treatment area containing the gillnet).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Experimental tank at the St. Nuclear Power Plant in Jensen Beach, FL. A juvenile green turtle (<italic>C. mydas</italic>) is participating in a daytime net vs. no net trial. Image captured from the primary overhead camera used to record all experiments.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-08-785357-g0002.tif"/>
</fig></sec>
<sec>
<title>3.2. Related Work</title>
<p>Our behavior recognition approach requires the turtle location information in every frame. Thus, we included an object tracking method as part of the design. The visual object tracking problem has long been studied in the computer vision field. Early methods have commonly used correlation based approaches and hand-crafted features for the tracking task. In Ross et al. (<xref ref-type="bibr" rid="B32">2007</xref>), the authors proposed a method (IVT) that employs an incremental principal component analysis algorithm to achieve low dimensional subspace representations of the target object for tracking purposes. In Babenko et al. (<xref ref-type="bibr" rid="B1">2009</xref>), a multiple instance learning (MILTrack) framework was used for object tracking where Haar-like features were used for discriminating the positive and negative image sets. In Bolme et al. (<xref ref-type="bibr" rid="B6">2010</xref>), an adaptive correlation based algorithm (MOSSE) that calculates the optimal filter for the desired Gauss-shaped correlation output was proposed. In another approach (Bao et al., <xref ref-type="bibr" rid="B2">2012</xref>), Bao et al. modeled the target by using a sparse approximation over a template set (L1APG). In this method, an &#x02113;-1 norm related minimization problem was solved iteratively to achieve the sparse representation. In Gundogdu et al. (<xref ref-type="bibr" rid="B17">2015</xref>), an adaptive ensemble of simple correlation filters (TBOOST) was used to generate tracking decisions by switching among the individual correlators in a computationally efficient manner. Henriques et al. (<xref ref-type="bibr" rid="B20">2014</xref>) presents a method to use Kernelized Correlation Filters (KCF) operating on histogram of oriented gradients, where the key idea is to include all the cyclic shift versions of the target patch in the sample set, and train the network in Fourier Domain efficiently. In Danelljan et al. (<xref ref-type="bibr" rid="B9">2015</xref>), authors propose a discriminative correlation filter based approach (SRDCF) where they use a spatial regularization function that penalizes filter coefficients residing outside the target region. In Demir and Cetin (<xref ref-type="bibr" rid="B12">2016</xref>), authors propose a &#x0201C;co-difference&#x0201D; feature-based tracking algorithm (CODIFF) to efficiently represent and match image parts. This idea is further extended in Demir and Adil (<xref ref-type="bibr" rid="B10">2018</xref>) by including a part based approach (P-CODIFF) to achieve robustness against rotations and shape deformations. In Bertinetto et al. (<xref ref-type="bibr" rid="B3">2016</xref>), the authors propose a method (STAPLE) to combine both correlation based and color based representations to construct a model that is robust to intensity changes and deformations. More recent methods use CNNs for the tracking task. Siamese network based methods have achieved remarkable results for the object tracking benchmarks (Kristan et al., <xref ref-type="bibr" rid="B23">2019</xref>, <xref ref-type="bibr" rid="B22">2020</xref>; Li et al., <xref ref-type="bibr" rid="B24">2019</xref>). In our experiments, we compared the performance of various state-of-the-art tracking algorithms on our dataset and used the best performing method for our application. Detailed results are given in section 4.1.</p>
<p>As a part of our design, we also estimated turtle orientation to differentiate some of the behavior patterns. Various methods have been proposed to estimate the orientation of animals (Wagner et al., <xref ref-type="bibr" rid="B34">2013</xref>), humans (Raza et al., <xref ref-type="bibr" rid="B30">2018</xref>), and other objects (Hara et al., <xref ref-type="bibr" rid="B18">2017</xref>). Similar to the tracking and behavior recognition problems, deep CNNs have successfully been used for the orientation estimation problem as well. In our method, a lightweight CNN architecture is employed to estimate the turtle orientation.</p></sec>
<sec>
<title>3.3. Proposed Method</title>
<p>In this study, we intended to successfully recognize <italic>U-turn</italic> and <italic>Reversal</italic> behaviors of sea turtles. To differentiate these behaviors from each other and from other motion patterns, we use turtle location, speed, and orientation information. In order to extract those features and combine them as an input to a deep neural network based architecture, we propose the recognition system shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<p>The turtle location and speed were calculated by the visual object tracker and the turtle orientation calculated by the angle estimation network are combined to generate color-coded spatiotemporal images. The images are used by another network as the input for the behavior recognition task. Details of these building blocks are given in the subsections below.</p>
<sec>
<title>3.3.1. Visual Object Tracker</title>
<p>The purpose of the visual object tracking block is to find the object location and size in every frame based on a given initial bounding box. Object location found by the visual object tracker is used to calculate the motion velocity vector (<bold>v</bold>). Bounding box output is also used to crop the object region from the image for the angle estimation network. <bold>v<sub>n</sub></bold> is calculated from the current object location <bold>p<sub>n</sub></bold> and the previous object location <bold>p<sub>n&#x02212;1</sub></bold> as shown in Equation (1).</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>v</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>n</mml:mi></mml:mstyle></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>In order to employ a successful object tracking algorithm in the proposed framework, we performed a comparison between the state-of-the-art visual object trackers IVT (Ross et al., <xref ref-type="bibr" rid="B32">2007</xref>), MILTrack (Babenko et al., <xref ref-type="bibr" rid="B1">2009</xref>), MOSSE (Bolme et al., <xref ref-type="bibr" rid="B6">2010</xref>), L1APG (Bao et al., <xref ref-type="bibr" rid="B2">2012</xref>), TBOOST (Gundogdu et al., <xref ref-type="bibr" rid="B17">2015</xref>), KCF (Henriques et al., <xref ref-type="bibr" rid="B20">2014</xref>), SRDCF (Danelljan et al., <xref ref-type="bibr" rid="B9">2015</xref>), P-CODIFF (Demir and Adil, <xref ref-type="bibr" rid="B10">2018</xref>), Staple (Bertinetto et al., <xref ref-type="bibr" rid="B3">2016</xref>), and SiamMargin (Kristan et al., <xref ref-type="bibr" rid="B23">2019</xref>) on our turtle dataset. Based on the quantitative metrics, we used the best performing tracker. Details of the experimental results are given in section 4.</p></sec>
<sec>
<title>3.3.2. Orientation Estimation Network</title>
<p>We built a relatively small CNN architecture for detecting the orientation of the turtle. The network topology is summarized in <xref ref-type="table" rid="T1">Table 1</xref>. Note that we use two outputs for representing the angle values on the unit circle so that we can use the MSE loss function without any modifications. We could use a single output for the angle value. However, we would need to redefine the loss function to prevent penalizing the jumps between 0&#x000B0; and 360&#x000B0;. For training the network coefficients, we annotated nearly 25,000 turtle images with bounding box and orientation labels. We extended this number by rotating the turtle images by 30 to 330 degrees with 30 degree steps and included associated orientation labels based on the rotation angle.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Topology of orientation estimation network.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Name</bold></th>
<th valign="top" align="left"><bold>Explanation</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">imageinput</td>
<td valign="top" align="left">64x64x1 input images</td>
</tr>
<tr>
<td valign="top" align="left">conv_1</td>
<td valign="top" align="left">8 3x3x1 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_1</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">avgpool2d_1</td>
<td valign="top" align="left">2x2 average pooling with stride 2</td>
</tr>
<tr>
<td valign="top" align="left">conv_2</td>
<td valign="top" align="left">16 3x3x8 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_2</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">avgpool2d_2</td>
<td valign="top" align="left">2x2 average pooling with stride 2</td>
</tr>
<tr>
<td valign="top" align="left">conv_3</td>
<td valign="top" align="left">32 3x3x16 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_3</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">avgpool2d_3</td>
<td valign="top" align="left">2x2 average pooling with stride 2</td>
</tr>
<tr>
<td valign="top" align="left">conv_4</td>
<td valign="top" align="left">64 3x3x32 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_4</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">avgpool2d_4</td>
<td valign="top" align="left">2x2 average pooling with stride 2</td>
</tr>
<tr>
<td valign="top" align="left">conv_5</td>
<td valign="top" align="left">128 3x3x64 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_5</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">avgpool2d_5</td>
<td valign="top" align="left">2x2 average pooling with stride 2</td>
</tr>
<tr>
<td valign="top" align="left">conv_6</td>
<td valign="top" align="left">128 3x3x128 convolutions, stride 1</td>
</tr>
<tr>
<td valign="top" align="left">relu_6</td>
<td valign="top" align="left">ReLU layer</td>
</tr>
<tr>
<td valign="top" align="left">dropout</td>
<td valign="top" align="left">20% dropout</td>
</tr>
<tr>
<td valign="top" align="left">fc</td>
<td valign="top" align="left">2 Fully connected layers</td>
</tr>
<tr>
<td valign="top" align="left">regressionoutput</td>
<td valign="top" align="left">Regression with MSE loss function</td>
</tr>
</tbody>
</table>
</table-wrap></sec>
<sec>
<title>3.3.3. Color Coding</title>
<p>This block generates spatiotemporal feature images based on the visual object tracking output and estimates the turtle orientation. We basically aim to represent the turtle behavior occurring over a time period as an RGB image. In order to generate this image, we draw the path of the turtle using the visual object tracking result. However, we also include the orientation and speed information using hue and value channels of the hue-saturation-value (HSV) color space. The angular difference between the velocity vector (<italic>v</italic><sub><italic>n</italic></sub>) direction and the turtle orientation (&#x003B8;<sub><italic>n</italic></sub>) is used for determining the hue channel, while the magnitude of the velocity vector is used for value channel. An example output of the color coding block is given in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Color coded spatiotemporal feature images generated using turtle velocity vector and orientation information.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-08-785357-g0003.tif"/>
</fig></sec>
<sec>
<title>3.3.4. Behavior Recognition Network</title>
<p>This block aims to recognize the target turtle behaviors using the color-coded spatiotemporal feature images. Since we formulate the behavior recognition task as a vision based classification problem, we adopt a widely used network architecture, ResNet50 (He et al., <xref ref-type="bibr" rid="B19">2016</xref>), for this task. In order to train and test the network, we used a dataset consisting of 172 sequences with U-turn, Reversal, and random motions. This dataset is further extended with rotated, shifted, and symmetric versions of the sequences. Since we have a relatively small dataset, we employed the transfer learning approach where we use the coefficients pre-trained on the ImageNet (Deng et al., <xref ref-type="bibr" rid="B13">2009</xref>) dataset. We modified the last two fully-connected layers for our behavior recognition task so that the network gives a decision between three behavior classes. The coefficients in the last two layers are trained using our training set.</p></sec></sec></sec>
<sec id="s4">
<title>4. Experimental Results</title>
<sec>
<title>4.1. Object Tracking and Behavior Recognition</title>
<p>For our visual object tracking experiments, we compared several state-of-the-art algorithms on a dataset consisting of 59 sequences with nearly 25,000 frames. We use Center Location Error (CLE) and Overlap Ratio (OR) as two base metrics which are widely used in object tracking problems (Wu et al., <xref ref-type="bibr" rid="B37">2013</xref>). CLE is the Euclidean distance between the ground truth location and the predicted location, while OR denotes the overlap ratio of predicted bounding box and ground truth bounding box. Based on these metrics, we generated the success and precision plots. The precision plot shows the ratio of frames where CLE is smaller than a certain threshold. The success plot shows the ratio of frames where OR is higher than a given threshold. <xref ref-type="fig" rid="F4">Figure 4</xref> shows the performance results of the compared algorithms.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Success and precision plots of various methods. <bold>(A)</bold> Success vs. overlap threshold plots of various methods. <bold>(B)</bold> Precision vs. localization error plots of various methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-08-785357-g0004.tif"/>
</fig>
<p>Based on this comparative analysis, we determined that the SiamMargin (Kristan et al., <xref ref-type="bibr" rid="B23">2019</xref>) algorithm achieved the highest success and precision graphs among the compared algorithms on the turtle dataset. Therefore, we used this algorithm in our visual object tracking block.</p>
<p>For the orientation estimation experiments, we used 70 percent of the images as training samples, and the rest for the validation and test samples. We used the batch size as 64, initial learning rate as 1e-3, and the number of epochs as 50. In every 20 epochs, we dropped the learning rate by using the drop factor value of 0.1. With these parameters, the model achieved a mean error value of 12.4 degrees on the test set.</p>
<p>In our final set of experiments, we used color-coded spatiotemporal feature images to recognize turtle behaviors. For these experiments, we similarly used 70 percent of the behavior sequences in our dataset to create spatiotemporal motion patterns and trained the last two fully connected layers of the ResNet50 architecture. Then, we used the test sequences to create similar spatiotemporal motion patterns using the outputs of SiamMargin tracker and orientation estimation network that we trained in the previous step. Based on the behavior recognition network outputs, we achieved the prediction results given in <xref ref-type="table" rid="T2">Table 2</xref>. Corresponding Precision, Recall, and F1 Scores for each behavior are presented in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Normalized confusion matrix showing the percentage of actual and predicted classes for 3 different behaviors.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th/>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Predicted</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center"><bold>U-Turn</bold></th>
<th valign="top" align="center"><bold>Reversal</bold></th>
<th valign="top" align="center"><bold>Random</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Actual</td>
<td valign="top" align="left">U-Turn</td>
<td valign="top" align="center">83.3</td>
<td valign="top" align="center">3.7</td>
<td valign="top" align="center">13.0</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Reversal</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">82.8</td>
<td valign="top" align="center">17.2</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Random</td>
<td valign="top" align="center">6.7</td>
<td valign="top" align="center">5.0</td>
<td valign="top" align="center">88.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Precision, Recall, and F1 Scores of defined behaviors.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1 Score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">U-Turn</td>
<td valign="top" align="center">0.833</td>
<td valign="top" align="center">0.925</td>
<td valign="top" align="center">0.877</td>
</tr>
<tr>
<td valign="top" align="left">Reversal</td>
<td valign="top" align="center">0.828</td>
<td valign="top" align="center">0.905</td>
<td valign="top" align="center">0.865</td>
</tr>
<tr>
<td valign="top" align="left">Random</td>
<td valign="top" align="center">0.883</td>
<td valign="top" align="center">0.745</td>
<td valign="top" align="center">0.808</td>
</tr>
<tr>
<td valign="top" align="left">Macro-Avg</td>
<td valign="top" align="center">0.848</td>
<td valign="top" align="center">0.858</td>
<td valign="top" align="center">0.85</td>
</tr>
</tbody>
</table>
</table-wrap></sec>
<sec>
<title>4.2. Anticipated Behavioral Results Using This Method</title>
<p>Studying the effectiveness of bycatch reduction technologies (BRTs) is a difficult task when conditions are less than ideal for recording sea turtle interactions with fishing gear and BRTs in the field and behavioral data requires intensive analysis by researchers even when it can be obtained. Therefore, using behavioral data from controlled experiments to train this convolutional neural network improves the process. We intend to use this initial study to discover if sea turtles do, in fact, recognize fishing nets as a barrier, in which case they would likely avoid the net with a U-turn when they can see them (presumably during the day). We expect to identify more Reversal behaviors during night trials when sea turtles most likely cannot see the net before them. These behaviors can last as little as 3 to 5 s, so in one 15-min trial a sea turtle can perform dozens to hundreds of behaviors that require recording by a researcher. With most treatments involving at least 15 sea turtles at 3 trials each, it becomes a time-intensive project with natural human error that comes along with watching hours of behavior videos. This algorithm can identify these behaviors and enable a comparison between U-turn and Reversal behaviors in daytime and nighttime trials.</p></sec></sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<sec>
<title>5.1. Future Uses and Related Behaviors</title>
<p>While this method has been created and tested exclusively on behavioral data in a controlled setting, we intend to use this method on field trials in the future. Given that most gillnet fisheries operate at night (Wang et al., <xref ref-type="bibr" rid="B35">2010</xref>), obtaining high resolution footage of sea turtle interactions is challenging. In particular, we plan on assessing video footage of <italic>in situ</italic> sea turtle interactions with gillnet fisheries as a future step of this research project.</p>
<p>We also recognize that the reversal and U-turn behaviors observed here are likely not exclusive to gillnet avoidance. While we were unable to find literature outlining these specific behaviors, we suspect that reversals and U-turns are evident in other common sea turtle interactions, such as mating (e.g., avoidance behavior by females during courtship) (Frick et al., <xref ref-type="bibr" rid="B15">2000</xref>), predator avoidance (Wirsing et al., <xref ref-type="bibr" rid="B36">2008</xref>), and competition over food or habitat resources (Gaos et al., <xref ref-type="bibr" rid="B16">2021</xref>). Additionally, because this method was created for overhead video, drone footage of sea turtle interactions would be an ideal way to collect behavioral data in the field and subsequently detect the behaviors of interest in other contexts, which has become a common technique for capturing sea turtle behavior (Schofield et al., <xref ref-type="bibr" rid="B33">2019</xref>). For example, studies have captured overhead drone footage of sea turtle courtship behavior (Bevan et al., <xref ref-type="bibr" rid="B4">2016</xref>; Rees et al., <xref ref-type="bibr" rid="B31">2018</xref>). In the future, our machine learning method could be used to detect these behaviors in relation to intraspecific aggression, predator avoidance, and other important interactions captured by drone footage.</p></sec>
<sec>
<title>5.2. Conclusion</title>
<p>In this study, we developed a behavior recognition framework for sea turtles using color-coded spatiotemporal motion patterns. Our approach uses visual object tracking and CNN based orientation estimation blocks to generate spatiotemporal feature images and processes them to recognize certain behaviors. Our experiments demonstrate that the proposed method achieves an average F1 score of 85% on recognizing the behaviors of interest.</p></sec></sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p></sec>
<sec id="s7">
<title>Ethics Statement</title>
<p>The animal study was reviewed and approved by Arizona State University Institutional Animal Care and Use Committee.</p></sec>
<sec id="s8">
<title>Author Contributions</title>
<p>JR and HD primarily wrote the manuscript. JR collected the data (i.e., turtle videos) to be used for training the neural network and analyzed behaviors. BW and MB provided the facility and subjects for data collection, also assisting in data collection. HD created the neural network with assistance from SO and JB. JS, SO, and JB edited the manuscript. All authors contributed to the article and approved the submitted version.</p></sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This material was based upon work supported by the National Science Foundation under Grant No. 1837473. The research was also supported by Inwater Research Group and Florida Power and Light Company. The work on protected species was conducted under Florida FWC Marine Turtle Permit 20-125. This project was funded in part by a grant awarded from the Sea Turtle Grants Program. The Sea Turtle Grants Program is funded from proceeds from the sale of the Florida Sea Turtle License Plate. Learn more at <ext-link ext-link-type="uri" xlink:href="http://www.helpingseaturtles.org">www.helpingseaturtles.org</ext-link>. This work was also partially funded by the National Fish and Wildlife Foundation.</p>
</sec>
<sec id="s10"> <title>Author Disclaimer</title>
<p>The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the opinions or policies of the U.S. Government or the National Fish and Wildlife Foundation and its funding sources. Mention of trade names or commercial products does not constitute their endorsement by the U.S. Government, or the National Fish and Wildlife Foundation or its funding sources.</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec>
</body>
<back>
<ack><p>We would like to thank Don Chuy and Grupo Tortuguero for building the nets used in this study. We would also like to thank Dale DeNardo for academic support and advising on the publication process.</p>
</ack><sec sec-type="supplementary-material" id="s12">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmars.2021.785357/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmars.2021.785357/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Video_1.AVI" id="SM1" mimetype="video/avi" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Video_2.AVI" id="SM2" mimetype="video/avi" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Babenko</surname> <given-names>B.</given-names></name> <name><surname>Yang</surname> <given-names>M.-H.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Visual tracking with online multiple instance learning,</article-title> in <source>Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on</source> (<publisher-loc>San Diego, CA</publisher-loc>), <fpage>983</fpage>&#x02013;<lpage>990</lpage>.<pub-id pub-id-type="pmid">26843855</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Ling</surname> <given-names>H. L.</given-names></name> <name><surname>Ji</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>Real time robust l1 tracker using accelerated proximal gradient approach,</article-title> in <source>Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on</source> (<publisher-loc>Providence, RI</publisher-loc>), <fpage>1830</fpage>&#x02013;<lpage>1837</lpage>.<pub-id pub-id-type="pmid">27992474</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bertinetto</surname> <given-names>L.</given-names></name> <name><surname>Valmadre</surname> <given-names>J.</given-names></name> <name><surname>Golodetz</surname> <given-names>S.</given-names></name> <name><surname>Miksik</surname> <given-names>O.</given-names></name> <name><surname>Torr</surname> <given-names>P. H. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Staple: complementary learners for real-time tracking,</article-title> in <source>The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Oxford, UK</publisher-loc>).</citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bevan</surname> <given-names>E.</given-names></name> <name><surname>Wibbels</surname> <given-names>T.</given-names></name> <name><surname>Navarro</surname> <given-names>E.</given-names></name> <name><surname>Rosas</surname> <given-names>M.</given-names></name> <name><surname>Najera</surname> <given-names>B.</given-names></name> <name><surname>Illescas</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Using unmanned aerial vehicle (uav) technology for locating, identifying, and monitoring courtship and mating behaviour in the green turtle (<italic>Chelonia mydas</italic>)</article-title>. <source>Herpetol. Rev.</source> <volume>47</volume>, <fpage>27</fpage>&#x02013;<lpage>32</lpage>.</citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bodor</surname> <given-names>R.</given-names></name> <name><surname>Jackson</surname> <given-names>B.</given-names></name> <name><surname>Papanikolopoulos</surname> <given-names>N.</given-names></name></person-group> (<year>2003</year>). <article-title>Vision-based human tracking and activity recognition,</article-title> in <source>Proc. of the 11th Mediterranean Conf. on Control and Automation</source> (<publisher-loc>Minneapolis, MN</publisher-loc>), Vol. 1.</citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bolme</surname> <given-names>D.</given-names></name> <name><surname>Beveridge</surname> <given-names>J.</given-names></name> <name><surname>Draper</surname> <given-names>B.</given-names></name> <name><surname>Lui</surname> <given-names>Y. M.</given-names></name></person-group> (<year>2010</year>). <article-title>Visual object tracking using adaptive correlation filters,</article-title> in <source>Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on</source> (<publisher-loc>San Francisco, CA</publisher-loc>), <fpage>2544</fpage>&#x02013;<lpage>2550</lpage>.<pub-id pub-id-type="pmid">31170074</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casale</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>Sea turtle by-catch in the mediterranean</article-title>. <source>Fish Fish.</source> <volume>12</volume>, <fpage>299</fpage>&#x02013;<lpage>316</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-2979.2010.00394.x</pub-id><pub-id pub-id-type="pmid">24764769</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chakravarty</surname> <given-names>P.</given-names></name> <name><surname>Cozzi</surname> <given-names>G.</given-names></name> <name><surname>Ozgul</surname> <given-names>A.</given-names></name> <name><surname>Aminian</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>A novel biomechanical approach for animal behaviour recognition using accelerometers</article-title>. <source>Methods Ecol. Evol.</source> <volume>10</volume>, <fpage>802</fpage>&#x02013;<lpage>814</lpage>. <pub-id pub-id-type="doi">10.1111/2041-210X.13172</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Danelljan</surname> <given-names>M.</given-names></name> <name><surname>H&#x000E4;ger</surname> <given-names>G.</given-names></name> <name><surname>Khan</surname> <given-names>F. S.</given-names></name> <name><surname>Felsberg</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning spatially regularized correlation filters for visual tracking,</article-title> in <source>2015 IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Santiago</publisher-loc>), <fpage>4310</fpage>&#x02013;<lpage>4318</lpage>.</citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Demir</surname> <given-names>H. S.</given-names></name> <name><surname>Adil</surname> <given-names>O. F.</given-names></name></person-group> (<year>2018</year>). <article-title>Part-based co-difference object tracking algorithm for infrared videos,</article-title> in <source>2018 25th IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Athens</publisher-loc>), <fpage>3723</fpage>&#x02013;<lpage>3727</lpage>.</citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Demir</surname> <given-names>H. S.</given-names></name> <name><surname>Blain Christen</surname> <given-names>J.</given-names></name> <name><surname>Ozev</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Energy-efficient image recognition system for marine life</article-title>. <source>IEEE Trans. Comput. Aided Design Integr. Circuits Syst.</source> <volume>39</volume>, <fpage>3458</fpage>&#x02013;<lpage>3466</lpage>. <pub-id pub-id-type="doi">10.1109/TCAD.2020.3012745</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Demir</surname> <given-names>H. S.</given-names></name> <name><surname>Cetin</surname> <given-names>A. E.</given-names></name></person-group> (<year>2016</year>). <article-title>Co-difference based object tracking algorithm for infrared videos,</article-title> in <source>2016 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Phoenix, AZ</publisher-loc>), <fpage>434</fpage>&#x02013;<lpage>438</lpage>.</citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>ImageNet: a large-scale hierarchical image database,</article-title> in <source>CVPR09</source> (<publisher-loc>Miami, FL</publisher-loc>).</citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Estes</surname> <given-names>J. A.</given-names></name> <name><surname>Terborgh</surname> <given-names>J.</given-names></name> <name><surname>Brashares</surname> <given-names>J. S.</given-names></name> <name><surname>Power</surname> <given-names>M. E.</given-names></name> <name><surname>Berger</surname> <given-names>J.</given-names></name> <name><surname>Bond</surname> <given-names>W. J.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Trophic downgrading of planet earth</article-title>. <source>Science</source> <volume>333</volume>, <fpage>301</fpage>&#x02013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1126/science.1205106</pub-id><pub-id pub-id-type="pmid">21764740</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frick</surname> <given-names>M.</given-names></name> <name><surname>Slay</surname> <given-names>C.</given-names></name> <name><surname>Quinn</surname> <given-names>C.</given-names></name> <name><surname>Windham-Reid</surname> <given-names>A.</given-names></name> <name><surname>Duley</surname> <given-names>P.</given-names></name> <name><surname>Ryder</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2000</year>). <article-title>Aerial observations of courtship behavior in loggerhead sea turtles (caretta caretta) from Southeastern Georgia and Northeastern Florida</article-title>. <source>J. Herpetol.</source> <volume>34</volume>, <fpage>153</fpage>&#x02013;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.2307/1565255</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gaos</surname> <given-names>A. R.</given-names></name> <name><surname>Johnson</surname> <given-names>C. E.</given-names></name> <name><surname>McLeish</surname> <given-names>D. B.</given-names></name> <name><surname>King</surname> <given-names>C. S.</given-names></name> <name><surname>Senko</surname> <given-names>J. F.</given-names></name></person-group> (<year>2021</year>). <article-title>Interactions Among Hawaiian Hawksbills Suggest Prevalence of Social Behaviors in Marine Turtles</article-title>. <source>Chelonian Conserv. Biol.</source> <pub-id pub-id-type="doi">10.2744/CCB-1481.1</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gundogdu</surname> <given-names>E.</given-names></name> <name><surname>Ozkan</surname> <given-names>H.</given-names></name> <name><surname>Demir</surname> <given-names>H. S.</given-names></name> <name><surname>Ergezer</surname> <given-names>H.</given-names></name> <name><surname>Akagunduz</surname> <given-names>E.</given-names></name> <name><surname>Pakin</surname> <given-names>S. K.</given-names></name></person-group> (<year>2015</year>). <article-title>Comparison of infrared and visible imagery for object tracking: Toward trackers with superior ir performance,</article-title> in <source>Computer Vision and Pattern Recognition Workshops, 2015 IEEE Conference on</source> (<publisher-loc>Boston, MA</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hara</surname> <given-names>K.</given-names></name> <name><surname>Vemulapalli</surname> <given-names>R.</given-names></name> <name><surname>Chellappa</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Designing deep convolutional neural networks for continuous object orientation estimation</article-title>. <source>arXiv [preprint]</source> arXiv:1702.01499.</citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>.<pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henriques</surname> <given-names>J.F.</given-names></name> <name><surname>Caseiro</surname> <given-names>R.</given-names></name> <name><surname>Martins</surname> <given-names>P.</given-names></name> <name><surname>Batista</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>High-speed tracking with kernelized correlation filters,</article-title> in <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, <volume>37</volume>, <fpage>583</fpage>&#x02013;<lpage>596</lpage>.<pub-id pub-id-type="pmid">26353263</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ijjina</surname> <given-names>E. P.</given-names></name> <name><surname>Chalavadi</surname> <given-names>K. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Human action recognition in rgb-d videos using motion sequence information and deep learning</article-title>. <source>Pattern Recogn.</source> <volume>72</volume>, <fpage>504</fpage>&#x02013;<lpage>516</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2017.07.013</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kristan</surname> <given-names>M.</given-names></name> <name><surname>Leonardis</surname> <given-names>A.</given-names></name> <name><surname>Matas</surname> <given-names>J.</given-names></name> <name><surname>Felsberg</surname> <given-names>M.</given-names></name> <name><surname>Pflugfelder</surname> <given-names>R.</given-names></name> <name><surname>K&#x000E4;m&#x000E4;r&#x000E4;inen</surname> <given-names>J. K.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>The eighth visual object tracking VOT2020 challenge results,</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>547</fpage>&#x02013;<lpage>601</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-68238-5_39</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kristan</surname> <given-names>M.</given-names></name> <name><surname>Matas</surname> <given-names>J.</given-names></name> <name><surname>Leonardis</surname> <given-names>A.</given-names></name> <name><surname>Felsberg</surname> <given-names>M.</given-names></name> <name><surname>Pflugfelder</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>The seventh visual object tracking vot2019 challenge results,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision Workshops</source> (<publisher-loc>Seoul</publisher-loc>).</citation></ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Zhang</surname> <given-names>F.</given-names></name> <name><surname>Xing</surname> <given-names>J.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Siamrpn&#x0002B;&#x0002B;: Evolution of siamese visual tracking with very deep networks,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>4282</fpage>&#x02013;<lpage>4291</lpage>.</citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lucchetti</surname> <given-names>A.</given-names></name> <name><surname>Bargione</surname> <given-names>G.</given-names></name> <name><surname>Petetta</surname> <given-names>A.</given-names></name> <name><surname>Vasapollo</surname> <given-names>C.</given-names></name> <name><surname>Virgili</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Reducing sea turtle bycatch in the mediterranean mixed demersal fisheries</article-title>. <source>Front. Mar. Sci.</source> <volume>6</volume>:<fpage>387</fpage>. <pub-id pub-id-type="doi">10.3389/fmars.2019.00387</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moore</surname> <given-names>J.</given-names></name> <name><surname>Wallace</surname> <given-names>B.</given-names></name> <name><surname>Lewison</surname> <given-names>R.</given-names></name> <name><surname>Zydelis</surname> <given-names>R.</given-names></name> <name><surname>Cox</surname> <given-names>T.</given-names></name> <name><surname>Crowder</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>A review of marine mammal, sea turtle and seabird bycatch in usa fisheries and the role of policy in shaping management</article-title>. <source>Mar. Policy</source> <volume>33</volume>, <fpage>435</fpage>&#x02013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.1016/j.marpol.2008.09.003</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nweke</surname> <given-names>H. F.</given-names></name> <name><surname>Teh</surname> <given-names>Y. W.</given-names></name> <name><surname>Al-Garadi</surname> <given-names>M. A.</given-names></name> <name><surname>Alo</surname> <given-names>U. R.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep learning algorithms for human activity recognition using mobile and wearable sensor networks: State of the art and research challenges</article-title>. <source>Expert Syst. Appl.</source> <volume>105</volume>, <fpage>233</fpage>&#x02013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2018.03.056</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peckham</surname> <given-names>S.</given-names></name> <name><surname>Lucero-Romero</surname> <given-names>J.</given-names></name> <name><surname>Maldonado-D&#x000ED;az</surname> <given-names>D.</given-names></name> <name><surname>Rodr&#x000ED;guez-S&#x000E1;nchez</surname> <given-names>A.</given-names></name> <name><surname>Senko</surname> <given-names>J.</given-names></name> <name><surname>Wojakowski</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Buoyless nets reduce sea turtle bycatch in coastal net fisheries</article-title>. <source>Conserv. Lett.</source> <volume>9</volume>, <fpage>114</fpage>&#x02013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.1111/conl.12176</pub-id><pub-id pub-id-type="pmid">23883577</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Porto</surname> <given-names>S. M.</given-names></name> <name><surname>Arcidiacono</surname> <given-names>C.</given-names></name> <name><surname>Anguzza</surname> <given-names>U.</given-names></name> <name><surname>Cascone</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>A computer vision-based system for the automatic detection of lying behaviour of dairy cows in free-stall barns</article-title>. <source>Biosyst. Eng.</source> <volume>115</volume>, <fpage>184</fpage>&#x02013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2013.03.002</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Raza</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Rehman</surname> <given-names>S.-U.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Bao</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Appearance based pedestrians&#x00027; head pose and body orientation estimation using deep learning</article-title>. <source>Neurocomputing</source> <volume>272</volume>, <fpage>647</fpage>&#x02013;<lpage>659</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2017.07.029</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rees</surname> <given-names>A.</given-names></name> <name><surname>Avens</surname> <given-names>L.</given-names></name> <name><surname>Ballorain</surname> <given-names>K.</given-names></name> <name><surname>Bevan</surname> <given-names>E.</given-names></name> <name><surname>Broderick</surname> <given-names>A.</given-names></name> <name><surname>Carthy</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>The potential of unmanned aerial systems for sea turtle research and conservation: a review and future directions</article-title>. <source>Endanger. Species Res.</source> <volume>35</volume>, <fpage>81</fpage>&#x02013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.3354/esr00877</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ross</surname> <given-names>D. A.</given-names></name> <name><surname>Lim</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>R.-S.</given-names></name> <name><surname>Yang</surname> <given-names>M.-H.</given-names></name></person-group> (<year>2007</year>). <article-title>Incremental learning for robust visual tracking</article-title>. <source>Int. J. Comput. Vision</source> <volume>77</volume>, <fpage>125</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-007-0075-7</pub-id><pub-id pub-id-type="pmid">22868649</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schofield</surname> <given-names>G.</given-names></name> <name><surname>Esteban</surname> <given-names>N.</given-names></name> <name><surname>Katselidis</surname> <given-names>K. A.</given-names></name> <name><surname>Hays</surname> <given-names>G. C.</given-names></name></person-group> (<year>2019</year>). <article-title>Drones for research on sea turtles and other marine vertebrates&#x02014;a review</article-title>. <source>Biol. Conserv.</source> <volume>238</volume>:<fpage>108214</fpage>. <pub-id pub-id-type="doi">10.1016/j.biocon.2019.108214</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wagner</surname> <given-names>R.</given-names></name> <name><surname>Thom</surname> <given-names>M.</given-names></name> <name><surname>Gabb</surname> <given-names>M.</given-names></name> <name><surname>Limmer</surname> <given-names>M.</given-names></name> <name><surname>Schweiger</surname> <given-names>R.</given-names></name> <name><surname>Rothermel</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Convolutional neural networks for night-time animal orientation estimation,</article-title> in <source>2013 IEEE Intelligent Vehicles Symposium (IV)</source> (<publisher-loc>Gold Coast, QLD</publisher-loc>), <fpage>316</fpage>&#x02013;<lpage>321</lpage>.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Fisler</surname> <given-names>S.</given-names></name> <name><surname>Swimmer</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <article-title>Developing visual deterrents to reduce sea turtle bycatch in gill net fisheries</article-title>. <source>Mar. Ecol. Prog. Ser.</source> <volume>408</volume>, <fpage>241</fpage>&#x02013;<lpage>250</lpage>. <pub-id pub-id-type="doi">10.3354/meps08577</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wirsing</surname> <given-names>A.</given-names></name> <name><surname>Abernethy</surname> <given-names>R.</given-names></name> <name><surname>Heithaus</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Speed and maneuverability of adult loggerhead turtles (caretta caretta) under simulated predatory attack: do the sexes differ?</article-title> <source>J. Herpetol.</source> <volume>42</volume>, <fpage>411</fpage>&#x02013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1670/07-1661.1</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Lim</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>M.-H.</given-names></name></person-group> (<year>2013</year>). <article-title>Online object tracking: a benchmark,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Portland, OR</publisher-loc>), <fpage>2411</fpage>&#x02013;<lpage>2418</lpage>.</citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Q.</given-names></name> <name><surname>Xiao</surname> <given-names>D.</given-names></name> <name><surname>Lin</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Feeding behavior recognition for group-housed pigs with the faster r-cnn</article-title>. <source>Comput. Electron. Agric.</source> <volume>155</volume>, <fpage>453</fpage>&#x02013;<lpage>460</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2018.11.002</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://keras.io">https://keras.io</ext-link>.</p></fn>
</fn-group>
</back>
</article>