<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mar. Sci.</journal-id>
<journal-title>Frontiers in Marine Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mar. Sci.</abbrev-journal-title>
<issn pub-type="epub">2296-7745</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmars.2023.1129852</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Marine Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Estimating catch rates in real time: Development of a deep learning based <italic>Nephrops</italic> (<italic>Nephrops norvegicus</italic>) counter for demersal trawl fisheries</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Avsar</surname>
<given-names>Ercan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2060222"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Feekings</surname>
<given-names>Jordan P.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1555086"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Krag</surname>
<given-names>Ludvig Ahm</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Technical University of Denmark, Institute of Aquatic Resources, Section for Fisheries Technology</institution>, <addr-line>Hirtshals</addr-line>, <country>Denmark</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Computer Engineering Department, Dokuz Eylul University</institution>, <addr-line>Izmir</addr-line>, <country>T&#xfc;rkiye</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Hongsheng Bi, University of Maryland, College Park, United States</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Nikos Petrellis, University of Peloponnese, Greece; Amaya Alvarez, Mediterranean Institute for Advanced Studies (CSIC), Spain</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Ercan Avsar, <email xlink:href="mailto:erca@aqua.dtu.dk">erca@aqua.dtu.dk</email>
</p>
</fn>
<fn fn-type="other" id="fn002">
<p>This article was submitted to Ocean Observation, a section of the journal Frontiers in Marine Science</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1129852</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>02</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Avsar, Feekings and Krag</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Avsar, Feekings and Krag</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Demersal trawling is largely a blind process where information on catch rates and compositions is only available once the catch is taken onboard the vessel. Obtaining quantitative information on catch rates of target species while fishing can improve a fisheries economic and environmental performance as fishers would be able to use this information to make informed decisions during fishing. Despite there are real-time underwater monitoring systems developed for this purpose, the video data produced by these systems is not analyzed in near real-time. In other words, the user is expected to watch the video feed continuously to evaluate catch rates and composition. This is obviously a demanding process in which quantification of the fish counts will be of a qualitative nature. In this study, underwater footages collected using an in-trawl video recording system were processed to detect, track, and count the number of individuals of the target species, <italic>Nephrops norvegicus</italic>, entering the trawl in real-time. The detection was accomplished using a You Only Look Once v4 (YOLOv4) algorithm. Two other variants of the YOLOv4 algorithm (tiny and scaled) were included in the study to compare their effects on the accuracy of the subsequent steps and overall speed of the processing. SORT algorithm was used as the tracker and any <italic>Nephrops</italic> that cross the horizontal level at 4/5 of the frame height were counted as catch. The detection performance of the YOLOv4 model provided a mean average precision (mAP@50) value of 97.82%, which is higher than the other two variants. However, the average processing speed of the tiny model is the highest with 253.51 frames per second. A correct count rate of 80.73% was achieved by YOLOv4 when the total number of <italic>Nephrops</italic> are considered in all the test videos. In conclusion, this approach was successful in processing underwater images in real time to determine the catch rates of the target species. The approach has great potential to process multiple species simultaneously in order to provide quantitative information not only on the target species but also bycatch and unwanted species to provide a comprehensive picture of the catch composition.</p>
</abstract>
<kwd-group>
<kwd>demersal trawling</kwd>
<kwd>
<italic>Nephrops</italic> counting</kwd>
<kwd>object detection</kwd>
<kwd>object tracking</kwd>
<kwd>sort</kwd>
<kwd>underwater video processing</kwd>
<kwd>YOLO</kwd>
</kwd-group>
<contract-num rid="cn001">33112-P-18-051</contract-num>
<contract-num rid="cn002">33113-I-22-187</contract-num>
<contract-num rid="cn003">7553521</contract-num>
<contract-sponsor id="cn001">European Maritime and Fisheries Fund<named-content content-type="fundref-id">10.13039/100014510</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Ministeriet for F&#xf8; devarer, Landbrug og Fiskeri<named-content content-type="fundref-id">10.13039/100008396</named-content>
</contract-sponsor>
<contract-sponsor id="cn003">Horizon 2020<named-content content-type="fundref-id">10.13039/501100007601</named-content>
</contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="6"/>
<equation-count count="4"/>
<ref-count count="59"/>
<page-count count="15"/>
<word-count count="7569"/>
</counts>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<title>Introduction</title>
<p>Demersal trawling is an effective way of catching various species. However, usage of demersal trawls is challenged by several factors such as high bycatch rates and negative effects on the biomass and biodiversity (<xref ref-type="bibr" rid="B10">Eigaard et&#xa0;al., 2017</xref>). In addition, disturbance of the seabed by bottom trawls results in aqueous CO<sub>2</sub> emissions which may inhibit marine carbon cycling after years of continuous trawling (<xref ref-type="bibr" rid="B37">Sala et&#xa0;al., 2021</xref>). Despite the presence of such concerns, demersal trawling is critical for catching economically valuable commercial species like shrimp, whitefish, and <italic>Nephrops</italic>.</p>
<p>
<italic>Nephrops</italic> excavate burrows in mud or mud/sand substrates and emerge at specific times to feed, mate and maintain their burrows, among others (<xref ref-type="bibr" rid="B44">Tully and Hillis, 1995</xref>; <xref ref-type="bibr" rid="B1">Aguzzi and Sard&#xe0;, 2008</xref>; <xref ref-type="bibr" rid="B12">Feekings et&#xa0;al., 2015</xref>). Their behavior is influential on catch rates when trawling as they need to be outside of the burrows to be caught (<xref ref-type="bibr" rid="B28">Main and Sangster, 1985</xref>). Besides, <italic>Nephrops</italic>-directed bottom trawling is known to have high discard rate which eventually causes not only economic loss but also loss of undersized individuals (<xref ref-type="bibr" rid="B4">Bergmann et&#xa0;al., 2002</xref>). In addition to these issues, is demersal trawling a blind process, meaning that the catch and size composition is unknown until the trawl is taken onboard after hours of trawling.</p>
<p>Advancements in underwater camera technologies may provide solutions to some limitations in demersal trawling. In particular, such cameras allow for recognition, counting and measurement of the individuals making it possible to understand the catch rates of <italic>Nephrops</italic> and unwanted species. Even though there are different tasks such as species identification and length measurement (<xref ref-type="bibr" rid="B45">Underwood et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B46">Underwood et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B2">Allken et&#xa0;al., 2021</xref>), and segmentation of the fish from the background (<xref ref-type="bibr" rid="B33">Prados et&#xa0;al., 2017</xref>) accomplished using in-trawl camera systems, they do not concern determining the catch composition in real time. The real-time processing of video footage collected by underwater in-trawl cameras is important to quantify catch rates of the target species. This information is valuable for the fishermen as it provides insight about the ongoing fishing process and further enable active search for better catch rates during the fishing operation. Deep learning-based methods enable automated extraction of such information. In fisheries research, deep learning is mostly used for processing visual data collected either onboard or by using underwater cameras. However, the main issue related with deep learning methods is the substantiality of the associated computation amount which brings about drawbacks like latency in processing and requirement of hardware with sufficient computational capacity. To address this issue, various deep learning models with different sizes have been developed, and they can be applied to different problems. A review of related literature is provided in Section 2. There are deep learning-based methods available that are applicable to underwater videos collected by in-trawl cameras for real-time detection and counting of <italic>Nephrops</italic>. A fast and accurate video processing system in <italic>Nephrops</italic> fisheries is useful for generating the spatial distribution of catch items as well as determining the number of <italic>Nephrops</italic> caught.</p>
<p>In this study, a real-time processing pipeline for underwater videos to determine the number of <italic>Nephrops</italic> caught during demersal trawling is proposed as such information will provide a strong decision tool for fishers to optimize their catching operation. The processed video footages were collected by an in-trawl camera developed earlier (<xref ref-type="bibr" rid="B41">Sokolova et&#xa0;al., 2021b</xref>). The algorithm for <italic>Nephrops</italic> counting has three major steps that are <italic>i</italic>) <italic>Nephrops</italic> detection, <italic>ii)</italic> tracking of the detected <italic>Nephrops</italic>, and <italic>iii</italic>) determining the true tracks accounted for <italic>Nephrops</italic> catches. The accurate detection of <italic>Nephrops</italic> in the video frames is important as the subsequent steps rely on the detected <italic>Nephrops</italic>. The detection has been accomplished using You Only Look Once v4 (YOLOv4) model which is known to be a fast deep learning model for object detection operating at high frames-per-second (FPS) values. In addition, two variants of YOLOv4, namely, YOLOv4-Tiny and YOLOv4-Scaled are used separately for <italic>Nephrops</italic> detection, and their effects on the tracking, counting, and the overall processing speed are observed and compared. The second step, tracking detections, is necessary for making association between the detections in the consecutive video frames. Simple Online Realtime Tracking (SORT) algorithm is used as the object tracker. For benchmarking purposes, the tracking performance of SORT is compared with two other object tracking algorithms, those being Minimum Output Sum of Squared Error (MOSSE) and DeepSORT. Finally, tracked objects satisfying some predefined conditions are considered as a <italic>Nephrops</italic> catch. These steps are illustrated in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. In this study we address the following research questions:</p>
<list list-type="bullet">
<list-item>
<p>How do the different YOLO-based object detection methods affect the overall speed and accuracy of the counting process?</p>
</list-item>
<list-item>
<p>What is the range of the processing speed of the proposed algorithm, and can it be considered as real-time under different circumstances?</p>
</list-item>
<list-item>
<p>Is it possible to provide simple decision parameters for the fishers during trawling operation?</p>
</list-item>
<list-item>
<p>What is the relation between the precision of the object detection and rate of correct <italic>Nephrops</italic> counts?</p>
</list-item>
</list>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Overview of the algorithm steps.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-10-1129852-g001.tif"/>
</fig>
</sec>
<sec id="s2">
<title>Related work</title>
<p>Utilization of deep learning methods in computer vision applications has become widespread in recent years due to their major advantage of automated feature extraction. However, the deep learning models typically possess many computational layers with high numbers of parameters. Performing all the calculations throughout all layers of the network takes time and hence the latency becomes an issue when the input data needs to be processed in real time.</p>
<p>Depending on the type of the problem (e.g. image classification, object detection, instance segmentation), there are various techniques to reduce the computational cost of the deep learning models while keeping the model performance as high as possible. For instance, MobileNets are efficient models developed to be used in hardware with limited computational resources (<xref ref-type="bibr" rid="B17">Howard et&#xa0;al., 2017</xref>) and can be used as a standalone classifier for animal classification in underwater images (<xref ref-type="bibr" rid="B24">Liu et&#xa0;al., 2019</xref>). Together with two other improved versions (<xref ref-type="bibr" rid="B38">Sandler et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B16">Howard et&#xa0;al., 2019</xref>) and single shot object detectors (SSD), they have more diverse applications such as detection of sea cucumbers (<xref ref-type="bibr" rid="B55">Yao et&#xa0;al., 2019</xref>), underwater objects with different scales (<xref ref-type="bibr" rid="B56">Zhang et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B50">Wang et&#xa0;al., 2022b</xref>), and <italic>Nephrops</italic> burrows (<xref ref-type="bibr" rid="B31">Naseer et&#xa0;al., 2020</xref>).</p>
<p>Another object detection method with many versions is YOLO, which is known for being very fast and accurate at the same time (<xref ref-type="bibr" rid="B36">Redmon et&#xa0;al., 2015</xref>). It can predict the bounding box coordinates and the corresponding confidence scores with one single network. There are numerous YOLO versions dedicated to operating on underwater images for detection of various objects such as starfish, shrimp, crab, scallop, and waterweed (<xref ref-type="bibr" rid="B25">Liu et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B57">Zhao et&#xa0;al., 2022</xref>). Among these models, the recently proposed model, YOLO-fish was designed for fish detection and is reported to be performing close to YOLOv4 model on two different public datasets (<xref ref-type="bibr" rid="B30">Muksit et&#xa0;al., 2022</xref>). Even though it is claimed to be a lightweight model the associated number of parameters and the detection time are between those of YOLOv3 and YOLOv4 (<xref ref-type="bibr" rid="B30">Muksit et&#xa0;al., 2022</xref>). In another study, an underwater imaging system to develop and test a lightweight YOLO model for automated fish behavior analysis was introduced (<xref ref-type="bibr" rid="B18">Hu et&#xa0;al., 2021</xref>). In that study, a modified version of YOLOv3-Lite model was proposed, and its detection performance as well as the prediction speed were compared with other state of the art models. It was shown that the proposed model works at 240 FPS processing speed while detecting the fish with higher precision and recall values.</p>
<p>Changing the detection scale, increasing the number of anchor boxes, or defining a new loss function are some of the modifications that can be done in the YOLO network structure (<xref ref-type="bibr" rid="B34">Raza and Hong, 2020</xref>). Moreover, combining the output of the YOLO model with other information sources such as optical flow and Gaussian mixture models is another strategy to obtain an improved detection in underwater images (<xref ref-type="bibr" rid="B19">Jalal et&#xa0;al., 2020</xref>).</p>
<p>In addition to underwater image and video processing methods, there are different applications to identify fish types on the vessel. Such studies involve usage of image classifiers based on convolutional neural networks (CNN) (<xref ref-type="bibr" rid="B58">Zheng et&#xa0;al., 2018</xref>) or instance segmentation networks such as Mask R-CNN (<xref ref-type="bibr" rid="B13">French et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B43">Tseng et&#xa0;al., 2020</xref>). Such segmentation operations are also useful in making morphological measurements on underwater fish images (<xref ref-type="bibr" rid="B32">Petrellis, 2021</xref>). This approach may be practical when the aim is to get an estimate of the individual fish sizes and weights in the catch.</p>
<p>The existing studies focus on either improving the detection performance, the computational load in individual images or application of the deep learning models to a new problem domain. In particular, object detection and tracking are widely studied today in various problem domains such as face recognition (<xref ref-type="bibr" rid="B47">Vijaya Kumar and Mahammad Shafi, 2022</xref>), processing of aerial images (<xref ref-type="bibr" rid="B11">ElTantawy and Shehata, 2020</xref>; <xref ref-type="bibr" rid="B54">Wu et&#xa0;al., 2022</xref>), and maritime surveillance (<xref ref-type="bibr" rid="B20">Jin et&#xa0;al., 2020</xref>). Despite the presence of many studies with different purposes and strategies, the number of studies concerning the real-time processing while tracking and counting the detected fish is very limited. In a study that is aimed to serve as a precursor to fish counting tasks, deep learning was used to classify the environmental conditions (<xref ref-type="bibr" rid="B42">Soom et&#xa0;al., 2022</xref>). According to the detected conditions, some traditional image processing methods were applied to the image to detect the presence/absence of fish. Even though no object detection and tracking were involved, the processing speed and power consumption of the proposed algorithm was evaluated on different hardware with various specifications.</p>
<p>On the other hand, there exists tracking algorithms developed for underwater objects like fish schools (<xref ref-type="bibr" rid="B23">Liu et&#xa0;al., 2022</xref>). In that work, a ResNet50 model was used as the feature extractor and an amendment detection module was proposed to improve the object detection and hence the performance of the tracking. The proposed model was compared with four different tracking algorithms, and it was shown that it outperforms the others in three out of four metrics. In two other studies, an experimental setup was prepared for collecting video footage using a web cam placed above a small fish tank. The fish in the tank were detected by YOLOv3-Tiny model that is trained on the specific dataset. Next, the tracking of the detections was accomplished by optical flow (<xref ref-type="bibr" rid="B29">Mohamed et&#xa0;al., 2020</xref>) or Euclidean distance (<xref ref-type="bibr" rid="B48">Wageeh et&#xa0;al., 2021</xref>). In these studies, tracking performances are provided poorly with no clear definition of a fish count and a correct track. In another study about fish tracking, an end-to-end model was proposed to detect and track the fish in a tank and determine the abnormal behaviors (<xref ref-type="bibr" rid="B51">Wang et&#xa0;al., 2022a</xref>). For the detection task, a modified version of YOLOv5 was used and the tracking was accomplished by SiamRPN++. The proposed model was shown to be operating at 84 FPS with higher detection performance than the other object detectors.</p>
<p>As can be understood from the existing studies, there are many efforts for object detection and tracking in underwater videos. However, the number of applications aimed at counting specific individuals by tracking them is very limited. One example can be the method based on Mask R-CNN to detect and count the catch items during trawling (<xref ref-type="bibr" rid="B39">Sokolova et&#xa0;al., 2021a</xref>). In that study, the detections and catch counts were collected under four classes, namely, <italic>Nephrops</italic>, round fish, flat fish, and other. The study involves detailed experiments about different data augmentation methods together with tracking and counting of the catch belonging to the specified classes. Though, it focuses on improvement of the object detection performance, overlooking the detection speed of the algorithm.</p>
<p>Current study differs from previous studies in <italic>i)</italic> counting of <italic>Nephrops</italic> in real-time by detecting and tracking them in underwater videos, <italic>ii)</italic> comparing the effects of three different YOLO models to the performances at every stage of the algorithm as well as the overall processing speed, and <italic>iii)</italic> showing the possibility of real-time monitoring and automated description of the catch items during trawling.</p>
</sec>
<sec id="s3" sec-type="materials|methods">
<title>Materials and methods</title>
<sec id="s3_1">
<title>The video dataset</title>
<p>The dataset used in this study consists of five videos collected using an underwater image acquisition system mounted at the codend entrance of a demersal trawl that allows in-trawl observation during fishing (<xref ref-type="bibr" rid="B41">Sokolova et&#xa0;al., 2021b</xref>). The videos were recorded on June 27, 2020, in Skagerrak on commercial <italic>Nephrops</italic> grounds where the catch in each haul were length measured to provide size and count for all caught species The footages have different durations and <italic>Nephrops</italic> ground truth counts. The object densities in the videos are different and such a diversity allows for better performance estimation for real-world applications. The details about the videos are provided in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. The stereo camera of the image acquisition system was set to record videos with a resolution of 1280 &#xd7; 720 pixels at 60 frames per second (FPS). Only the videos from the right camera were used for processing the frames as the entire data output from the stereo camera is useful for generating depth maps which is not within the scope of this study.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Details of the video footages.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">Duration (min)</th>
<th valign="top" align="center">Total <italic>Nephrops</italic> (no.)</th>
<th valign="top" align="center">FPS</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">
<bold>Video 1</bold>
</td>
<td valign="top" align="center">00:55</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>Video 2</bold>
</td>
<td valign="top" align="center">01:31</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>Video 3</bold>
</td>
<td valign="top" align="center">07:30</td>
<td valign="top" align="center">36</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>Video 4</bold>
</td>
<td valign="top" align="center">08:10</td>
<td valign="top" align="center">40</td>
<td valign="top" align="center">60</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>Video 5</bold>
</td>
<td valign="top" align="center">06:29</td>
<td valign="top" align="center">23</td>
<td valign="top" align="center">60</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<title>
<italic>Nephrops</italic> detection models</title>
<p>Among various versions of YOLO, the fourth version (YOLOv4) is efficient and stable with various applications in different domains (<xref ref-type="bibr" rid="B7">Bochkovskiy et&#xa0;al., 2020</xref>). The object detection task is considered as a regression problem by YOLOv4, and it eliminates the necessity of using large mini-batches during training. It optimizes the trade-off between the detection speed and accuracy, which means that it is possible to obtain accurate detections at high FPS values. Therefore, YOLOv4 has been selected as the primary model for <italic>Nephrops</italic> detection in this study. In addition, two variants of this model, YOLOv4-Tiny and YOLOv4-Scaled, are used to compare their performances.</p>
<p>YOLOv4 uses a CSPDarknet53 model as the feature extractor backbone. It contains 29 convolutional layers and has advantages like high receptive field and a large number of parameters that are required for an accurate object detection (<xref ref-type="bibr" rid="B7">Bochkovskiy et&#xa0;al., 2020</xref>). The output feature maps of the CSPDarknet53 are passed through a multi-scale max-pooling operation. This operation is implemented by a spatial pyramid pooling (SPP) layer where outputs of four max-pooling operations with kernel sizes 1x1, 5x5, 9x9, and 13x13 are concatenated. Processing with the SPP layer is important for increasing the receptive field and separate the contextual features. YOLOv4 also uses features at different levels of the feature extractor backbone. To accomplish this, feature maps from three layers of the CSPDarknet53 model are input to the path aggregation network (PANet) in which the features are fused both in top-down and bottom-up directions. Such an aggregation allows for simultaneous utilization of localization information present in the lower level features and semantic information in the higher level features. The extracted features with this structure are then passed through a YOLOv3 head to predict bounding box locations and the corresponding confidence scores. To improve generalization and reduce the risk of overfitting, two new methods are introduced in the algorithm: Mosaic and Self-Adversarial Training (SAT). In addition, a continuously differentiable and smooth function Mish is used as the activation between the layers of the network.</p>
<p>YOLOv4-Tiny is a lightweight version of the original YOLOv4 architecture. The major differences are in the numbers of anchor boxes and the convolutional layers in the backbone. Specifically, the tiny model has six anchor boxes while the original version has nine. Also, the number of YOLO prediction layers was reduced from three to two, which allows higher prediction speed while performing poor on the small objects. The scaled version of YOLOv4 (YOLOv4-Scaled) introduces modifications in the backbone and neck structures of the YOLOv4 architecture (<xref ref-type="bibr" rid="B49">Wang et&#xa0;al., 2020</xref>). In particular, the first CSP layer in the CSPDarknet53 backbone was replaced by a Darknet residual layer. In addition, up and down feature scaling operations in the PANet and pooling operations in the SPP module are enhanced by CSP blocks that ultimately may decrease the computation cost by 40%.</p>
</sec>
<sec id="s3_3">
<title>Tracking and counting of the detected <italic>nephrops</italic>
</title>
<p>Since the main goal of the study is to automatically count the number <italic>Nephrops</italic> entering the trawl, the detected <italic>Nephrops</italic> should be tracked as they appear in the frames. To accomplish this, an algorithm to make association between the detections in the consecutive frames should be implemented. This is done by object tracking algorithms that are particularly useful when the object of interest is occluded or not detected for a certain number of frames.</p>
<p>Simple Online and Real-time Tracking (SORT) is the object tracking method used in this study (<xref ref-type="bibr" rid="B5">Bewley et&#xa0;al., 2016</xref>). SORT uses 2D motion information for modeling the state (i.e. bounding box location, area, and aspect ratio) of each track in the video. Kalman filter with a linear velocity model predicts the state of the tracks for the next frame (<xref ref-type="bibr" rid="B21">Kalman, 1960</xref>). The association between the detections and the predicted tracks is accomplished by applying the Hungarian algorithm (<xref ref-type="bibr" rid="B22">Kuhn, 1955</xref>) on the cost matrix whose entries are the IoU values between the detections and predictions. In order to highlight the suitability of the SORT algorithm for real time <italic>Nephrops</italic> tracking, the performance of two other tracking methods, MOSSE and DeepSORT, are tested as well. Details of this comparison are given in Section 4.4.</p>
<p>Due to occlusions or inaccuracy of the object detector model, the target objects may not be detected in all frames when they are in the field of view of the camera. These discontinuities in the detection constitute a challenge for the tracking process. SORT algorithm is capable of predicting the bounding box coordinates in case of such discontinuities. However, if a track is not associated with a detection for 30 consecutive frames, then this track is considered finished. This means that the finished track will not be considered for association with the new detections anymore.</p>
<p>In order to determine the count for the <italic>Nephrops</italic> catches, the tracks output by the SORT tracker are checked. This is done with the help of a horizontal level defined at the top 4/5 of the frame height. When the <italic>Nephrops</italic> are leaving the frame from the bottom, they are partly visible, and this may cause the object tracker to assign different identities to the same <italic>Nephrops</italic> as they are about to disappear. Such an identity switch may generate false positive counts if the horizontal threshold is set to be the bottom of the frame. This is the reason for selecting a level different than the bottom of the frame.</p>
<p>In particular, any track satisfying at least one of the following conditions increases the counter by one:</p>
<list list-type="roman-lower">
<list-item>
<p>
<italic>The track with the lower level of the associated bounding box crosses the horizontal level</italic>. When the <italic>Nephrops</italic> is tracked successfully with no occlusions or distortions, this condition is easily satisfied. This is the most common condition.</p>
</list-item>
<list-item>
<p>
<italic>The track with the center of the associated bounding box crosses the horizontal level</italic>. Due to occlusions, tracking of some <italic>Nephrops</italic> are initialized after the lower level of their bounding box is below the horizontal level. This condition is useful for counting such <italic>Nephrops</italic>.</p>
</list-item>
<list-item>
<p>
<italic>The track with the height of the associated bounding box is greater than 2/3 of the frame height</italic>. Some <italic>Nephrops</italic> pass very close to the camera causing them to appear very large and in small number of frames. In such cases, the first two conditions cannot be satisfied. So this condition allows for detecting these <italic>Nephrops</italic>.</p>
</list-item>
</list>
<p>One sample counting instance for each condition are given in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Illustration of the counting conditions. Two consecutive frames in the columns. <bold>(A&#x2013;C)</bold> correspond to the conditions <italic>i, ii, iii</italic>, respectively.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-10-1129852-g002.tif"/>
</fig>
</sec>
<sec id="s3_4">
<title>Model training</title>
<p>The models mentioned in Section 3.2 are trained using an image dataset generated by the frames extracted from the videos included in this study. The majority of the frames in the videos do not contain any objects and are consequently not useful for the training process. Therefore, a manual selection of the frames with some objects is required. A total number of 4044 images were selected according to the presence of <italic>Nephrops</italic>, fish, or others. After the selection of frames, the bounding boxes for the objects in all the frames were manually labeled using the VIA annotation tool (<xref ref-type="bibr" rid="B9">Dutta and Zisserman, 2019</xref>). Since the aim is to count the number of <italic>Nephrops</italic> entering the gear, any object other than <italic>Nephrops</italic> was labeled as <italic>other.</italic> Therefore, the object detection step is considered as a binary detection problem.</p>
<p>The dataset was randomly divided into training and test sets with proportions of 87.5% and 12.5%, respectively. Next, 1000 images were generated using the Copy-Paste (CP) augmentation method and added to the training set (<xref ref-type="bibr" rid="B14">Ghiasi et&#xa0;al., 2021</xref>). When performing the CP augmentation, pixel values corresponding to the masks of the objects in the source images were pasted onto the destination images. To improve the diversity in the augmented images, some geometric transformations were applied to the images as explained in (<xref ref-type="bibr" rid="B39">Sokolova et&#xa0;al., 2021a</xref>). The details, like number of images and the object instances in the image dataset after the augmentation are given in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>, and three sample images are provided in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Numbers of images and instances from both classes in the training and test sets used in the object detection step.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">Images</th>
<th valign="top" align="center">
<italic>Nephrops</italic> Instances</th>
<th valign="top" align="center">
<italic>Other</italic> Instances</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">
<bold>Training Set</bold>
</td>
<td valign="top" align="center">4538</td>
<td valign="top" align="center">3766</td>
<td valign="top" align="center">8014</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>Test Set</bold>
</td>
<td valign="top" align="center">506</td>
<td valign="top" align="center">204</td>
<td valign="top" align="center">775</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Samples from the image dataset. <bold>(A)</bold> An image with a <italic>Nephrops</italic> instance. <bold>(B)</bold> An image with some <italic>other</italic> instances. <bold>(C)</bold> An image with copy-paste augmentation.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-10-1129852-g003.tif"/>
</fig>
<p>The darknet framework was used for the training of the models (<xref ref-type="bibr" rid="B35">Redmon, 2016</xref>). The training and testing were performed on a Tesla A100 GPU with 40 GB RAM, CUDA 11.1, and cudnn v8.0.4.30. All the coding was done with Python v3.9.12 following the instructions and model configuration files made available at (<xref ref-type="bibr" rid="B6">Bochkovskiy, 2022</xref>). Some of the hyperparameters regarding the models and their training are listed in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. Note that all the models were trained for 6000 iterations and the weights yielding the best detection performance were used in the subsequent steps.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Summary of the model settings.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">Network Size</th>
<th valign="top" align="center">Initial Learning Rate</th>
<th valign="top" align="center">Momentum</th>
<th valign="top" align="center">Decay</th>
<th valign="top" align="center">Training Epochs</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">
<bold>YOLOv4</bold>
</td>
<td valign="top" align="center">416</td>
<td valign="top" align="center">0.00100</td>
<td valign="top" align="center">0.949</td>
<td valign="top" align="center">0.0005</td>
<td valign="top" align="center">6000</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>YOLOv4-Tiny</bold>
</td>
<td valign="top" align="center">416</td>
<td valign="top" align="center">0.00261</td>
<td valign="top" align="center">0.900</td>
<td valign="top" align="center">0.0005</td>
<td valign="top" align="center">6000</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>YOLOv4-Scaled</bold>
</td>
<td valign="top" align="center">640</td>
<td valign="top" align="center">0.00100</td>
<td valign="top" align="center">0.949</td>
<td valign="top" align="center">0.0005</td>
<td valign="top" align="center">6000</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5">
<title>Performance evaluation metrics</title>
<p>The performances of each step in the study are evaluated and reported separately in Section 4. To evaluate the object detection performance, different <italic>mAP</italic> values are calculated for each of the models using the test set. <italic>mAP</italic> is a quantification of the detection performance by comparing the amount of overlap between the ground truth and predicted bounding boxes. It is a widely used metric and has good representation of the detection performance as it considers both the prediction confidence score and the intersection over union (<italic>IoU</italic>) values. First, the confidence scores for the bounding boxes are converted into class labels for different threshold values. This allows to obtain a confusion matrix for each threshold and hence calculate the precision and recall values using the True Positive (<italic>TP</italic>), False Positive (<italic>FP</italic>), and False Negative (<italic>FN</italic>) in each matrix given by the following equations.</p>
<disp-formula>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here the subscript <italic>n</italic> represents different confidence score thresholds. The multiple (recall, precision) points correspond to a curve in 2D space (precision-recall curve), and the average precision (<italic>AP</italic>) value is the weighted mean of the precisions with the weights being the changes in the recall values.</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:msub>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<p>This <italic>AP</italic> calculation procedure is repeated for all classes separately in the dataset. The average of all the <italic>AP</italic> values is defined as the <italic>mAP</italic> which can be obtained by</p>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>c</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>c</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>c</italic> represents the number of classes in the dataset and <italic>AP<sub>i</sub>
</italic> is the <italic>AP</italic> value for the <italic>i<sup>th</sup>
</italic> class.</p>
<p>The <italic>mAP</italic> value can be computed for different IoU thresholds that affects the shape of the precision-recall curves. As a convention, the <italic>mAP</italic> value is calculated for <italic>IoU</italic> = 0.50 (<italic>mAP</italic>@.50). However, for benchmarking purposes, <italic>mAP</italic> values at different IoU thresholds are calculated and averaged as well. In this study, three <italic>mAP</italic> values are provided as the detection performance of the models: <italic>mAP</italic>@.50, <italic>mAP</italic>@.75, and <italic>mAP</italic>@.50:.05:.95 (<italic>mAP</italic> values averaged for the thresholds from 0.50 to 0.95 with steps of 0.05). In addition, since the purpose is to track and count the <italic>Nephrops</italic> only, the <italic>AP</italic> values belonging to <italic>Nephrops</italic> class (<italic>AP<sub>nep</sub>
</italic>) are also given for the same IoU thresholds.</p>
<p>Having obtained the tracks as the algorithm output as explained in Section 3.3, the tracking performance metrics were calculated. Among the calculated metrics, multi-object tracking accuracy (MOTA) is a combination of three error types namely, number of misses, false positives, and mismatches. It is obtained by normalizing the total of these three errors by the number of ground truth tracks. In calculation of MOTA, only the track locations are used. In other words, no bounding box information is considered in MOTA. To overcome this situation, another metric called multi-object tracking precision (MOTP) is defined. MOTP is the average overlap between the bounding boxes of predictions and ground truths. Mostly tracked (MT) and mostly lost (ML) are two quality measures that consider the ratio of successfully tracked frames for an object. A track is MT if it is tracked for at least 80% of its life span. If the tracking ratio is less than 20%, then is called ML. Within the context of object tracking, it is also desirable to obtain tracks preserving their identities with small numbers of untracked frames. Therefore, it is possible to mention two more metrics here. Identity switch (ID-Sw) is the total number of tracks changing their identity for the same ground truth object. Fragmentation is the number of interruptions in the track where no tracking is made. Finally, higher order tracking accuracy (HOTA) combines errors originating from both association and detection (<xref ref-type="bibr" rid="B27">Luiten et&#xa0;al., 2021</xref>). Specifically, it is the geometric mean of association accuracy and detection accuracy.</p>
</sec>
</sec>
<sec id="s4" sec-type="results">
<title>Results</title>
<sec id="s4_1">
<title>Detection performance of the models</title>
<p>The <italic>mAP</italic> and <italic>AP<sub>nep</sub>
</italic> values for different IoU thresholds for all three models are given in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>. These values are obtained by passing the test set samples in the image dataset introduced in Section 3.1 through the trained models. Note that the best weights determined during the training phase are used for prediction on the test set which can be considered as a regularization step to avoid overfitting. In other words, the weights calculated in the subsequent iterations are not considered for <italic>Nephrops</italic> detection. The best weights are obtained at iterations 4962, 5245, and 4113 for YOLOv4, YOLOv4-Tiny, and YOLOv4-Scaled, respectively.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Performance comparison of the detector models.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" rowspan="2" align="left"/>
<th valign="top" colspan="3" align="center">
<italic>mAP</italic> (%)</th>
<th valign="top" colspan="3" align="center">
<italic>AP<sub>nep</sub>
</italic> (%)</th>
</tr>
<tr>
<th valign="top" align="center">@.50</th>
<th valign="top" align="center">@.75</th>
<th valign="top" align="center">@.50:.05:.95</th>
<th valign="top" align="center">@.50</th>
<th valign="top" align="center">@.75</th>
<th valign="top" align="center">@.50:.05:.95</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">
<bold>YOLOv4</bold>
</td>
<td valign="top" align="center">
<bold>97.82</bold>
</td>
<td valign="top" align="center">85.58</td>
<td valign="top" align="center">71.89</td>
<td valign="top" align="center">97.84</td>
<td valign="top" align="center">91.37</td>
<td valign="top" align="center">74.76</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>YOLOv4-Tiny</bold>
</td>
<td valign="top" align="center">95.10</td>
<td valign="top" align="center">73.06</td>
<td valign="top" align="center">62.71</td>
<td valign="top" align="center">94.57</td>
<td valign="top" align="center">76.95</td>
<td valign="top" align="center">64.28</td>
</tr>
<tr>
<td valign="top" align="left">
<bold>YOLOv4-Scaled</bold>
</td>
<td valign="top" align="center">97.55</td>
<td valign="top" align="center">
<bold>88.10</bold>
</td>
<td valign="top" align="center">
<bold>72.28</bold>
</td>
<td valign="top" align="center">
<bold>98.47</bold>
</td>
<td valign="top" align="center">
<bold>94.05</bold>
</td>
<td valign="top" align="center">
<bold>75.97</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Best values are provided in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>In most of the performance metrics, YOLOv4-Scaled outperforms the other two models. Nevertheless, the differences between YOLOv4 and YOLOv4-Scaled are minor which precludes suggesting the best model for all cases. For the threshold <italic>IoU</italic> = 0.5, the scaled version is slightly better at detection of the <italic>Nephrops</italic>, but when the <italic>AP</italic> values for both classes are considered, YOLOv4 has a higher <italic>mAP</italic> value. This means that YOLOv4-Scaled is not as precise as YOLOv4 when detecting the objects from the other class. On the other hand, the difference between the performances of YOLOv4-Tiny and the other two models is smaller when <italic>IoU</italic> = 0.5. This indicates that the tiny version is capable of detecting the bounding boxes but not with as high IoU values as those obtained by the other models.</p>
</sec>
<sec id="s4_2">
<title>Tracking and counting performance of the models</title>
<p>Note that only the tracks satisfying the count conditions were involved in the tracking performance calculation because these are the tracks used in counting performance calculation as well. In addition, the tracking metrics were obtained for all five videos separately, but their average values are provided here as one single clustered column chart (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>). The MOTA, MOTP, and HOTA values are given as percentages (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4A</bold>
</xref>) and the rest are number of tracks (<xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4B</bold>
</xref>).</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Tracking performances associated with the detectors. <bold>(A)</bold> Percentage values for MOTA, MOTP, and HOTA, <bold>(B) </bold>MT, ML, ID-Sw, and Fragmentation numbers averaged over the test videos.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-10-1129852-g004.tif"/>
</fig>
<p>The <italic>Nephrops</italic> counts output by the algorithm associated with the tracks are given in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>. The numbers of true positive counts are reported together with the numbers of false positive and false negative counts together with the correct count rates for each individual video. The lowest total number of false positives is achieved by YOLOv4-Scaled which has the highest false negative tracks as well. Therefore, it is possible to explain the low false positive rate by its inefficiency in generating tracks satisfying the count conditions. The lowest amount of false tracks are achieved by YOLOv4 which also has the highest true positives. Specifically, the related F-scores calculated on the total counts for YOLOv4, Tiny, and Scaled versions are 85.44%, 80.21%, and 74.44%, respectively.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Detailed numbers of counts obtained by the detection models.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left"/>
<th valign="top" align="center"/>
<th valign="top" align="center">Video-1</th>
<th valign="top" align="center">Video-2</th>
<th valign="top" align="center">Video-3</th>
<th valign="top" align="center">Video-4</th>
<th valign="top" align="center">Video-5</th>
<th valign="top" align="center">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left"/>
<td valign="middle" align="left">
<bold>Ground Truth</bold>
</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">6</td>
<td valign="middle" align="center">36</td>
<td valign="middle" align="center">40</td>
<td valign="middle" align="center">23</td>
<td valign="middle" align="center">109</td>
</tr>
<tr>
<td valign="middle" rowspan="5" align="left">
<bold>YOLOv4</bold>
</td>
<td valign="middle" align="left">
<bold>Output</bold>
</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">39</td>
<td valign="middle" align="center">31</td>
<td valign="middle" align="center">19</td>
<td valign="middle" align="center">97</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>True Positives</bold>
</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">34</td>
<td valign="middle" align="center">27</td>
<td valign="middle" align="center">19</td>
<td valign="middle" align="center">88</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Positives</bold>
</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">9</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Negatives</bold>
</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">13</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">21</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>Correct Count Rate (%)</bold>
</td>
<td valign="middle" align="center">100.00</td>
<td valign="middle" align="center">66.67</td>
<td valign="middle" align="center">94.44</td>
<td valign="middle" align="center">67.50</td>
<td valign="middle" align="center">82.61</td>
<td valign="middle" align="center">80.73</td>
</tr>
<tr>
<td valign="middle" rowspan="5" align="left">
<bold>YOLOv4-Tiny</bold>
</td>
<td valign="middle" align="left">
<bold>Output</bold>
</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">33</td>
<td valign="middle" align="center">24</td>
<td valign="middle" align="center">18</td>
<td valign="middle" align="center">83</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>True Positives</bold>
</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">31</td>
<td valign="middle" align="center">21</td>
<td valign="middle" align="center">18</td>
<td valign="middle" align="center">77</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Positives</bold>
</td>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">6</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Negatives</bold>
</td>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">19</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">32</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>Correct Count Rate (%)</bold>
</td>
<td valign="middle" align="center">75.00</td>
<td valign="middle" align="center">66.67</td>
<td valign="middle" align="center">86.11</td>
<td valign="middle" align="center">52.50</td>
<td valign="middle" align="center">78.26</td>
<td valign="middle" align="center">70.64</td>
</tr>
<tr>
<td valign="middle" rowspan="5" align="left">
<bold>YOLOv4-Scaled</bold>
</td>
<td valign="middle" align="left">
<bold>Output</bold>
</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">27</td>
<td valign="middle" align="center">19</td>
<td valign="middle" align="center">18</td>
<td valign="middle" align="center">71</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>True Positives</bold>
</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">25</td>
<td valign="middle" align="center">17</td>
<td valign="middle" align="center">18</td>
<td valign="middle" align="center">67</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Positives</bold>
</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">0</td>
<td valign="middle" align="center">4</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>False Negatives</bold>
</td>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">11</td>
<td valign="middle" align="center">23</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">42</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>Correct Count Rate (%)</bold>
</td>
<td valign="middle" align="center">75.00</td>
<td valign="middle" align="center">66.67</td>
<td valign="middle" align="center">69.44</td>
<td valign="middle" align="center">42.50</td>
<td valign="middle" align="center">78.26</td>
<td valign="middle" align="center">61.46</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<title>Processing speed comparison of the models</title>
<p>The required amount of calculations in the model and the hardware specifications are the two major factors affecting the processing speed. The calculation amounts are determined at the design stage of the models, and this can be adjusted to some degree by changing the input image sizes which is also named as network size (see <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>). Typically, a larger network size in the model yields better object detection, sacrificing the processing speed and vice versa. The input image size for the YOLOv4-Scaled model was adjusted to be higher than the other two models to improve its detection accuracy. Such an adjustment allowed for obtaining a similar accuracy with YOLOv4 model and hence benchmarking their tracking, counting and speed performances.</p>
<p>The FPS values for each model and video are summarized in <xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>. As expected, the YOLOv4-Tiny model is the fastest in all the videos because it has a reduced number of computational layers to enhance its speed. The slowest model is YOLOv4-Scaled. The reason for its lower FPS values is related with its larger network size. However, a smaller network size for this model would cause lower detection and tracking performances eventually yielding a lower number of true positive counts.</p>
<table-wrap id="T6" position="float">
<label>Table&#xa0;6</label>
<caption>
<p>Comparison of image processing speed between models in frames per second (mean [min-max]).</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="center">Video-1</th>
<th valign="top" align="center">Video-2</th>
<th valign="top" align="center">Video-3</th>
<th valign="top" align="center">Video-4</th>
<th valign="top" align="center">Video-5</th>
<th valign="top" align="center">Average</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">
<bold>YOLOv4</bold>
</td>
<td valign="middle" align="center">116.49<break/>[65-123]</td>
<td valign="middle" align="center">115.64<break/>[76-123]</td>
<td valign="middle" align="center">116.67<break/>[75-123]</td>
<td valign="middle" align="center">114.77<break/>[69-123]</td>
<td valign="middle" align="center">115.76<break/>[62-122]</td>
<td valign="middle" align="center">115.87<break/>[69.4-122.8]</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>YOLOv4-Tiny</bold>
</td>
<td valign="middle" align="center">267.51<break/>[84-323]</td>
<td valign="middle" align="center">248.58<break/>[96-267]</td>
<td valign="middle" align="center">251.22<break/>[76-318]</td>
<td valign="middle" align="center">251.50<break/>[75-316]</td>
<td valign="middle" align="center">248.72<break/>[91-311]</td>
<td valign="middle" align="center">253.51<break/>[84.4-307.0]</td>
</tr>
<tr>
<td valign="middle" align="left">
<bold>YOLOv4-Scaled</bold>
</td>
<td valign="middle" align="center">78.93<break/>[39-80]</td>
<td valign="middle" align="center">79.51<break/>[51-81]</td>
<td valign="middle" align="center">80.31<break/>[40-82]</td>
<td valign="middle" align="center">79.93<break/>[44-82]</td>
<td valign="middle" align="center">80.73<break/>[48-82]</td>
<td valign="middle" align="center">79.88<break/>[44.4-81.4]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<title>Benchmarking with other trackers</title>
<p>To evaluate the suitability of SORT, two other object tracking algorithms were tested on the same dataset. One of these methods is based on a correlation filter, namely, Minimum Output Sum of Squared Error (MOSSE) filter (<xref ref-type="bibr" rid="B8">Bolme et&#xa0;al., 2010</xref>). The reason for selecting this object tracker is that its processing speed is claimed to reach 669 FPS (<xref ref-type="bibr" rid="B8">Bolme et&#xa0;al., 2010</xref>). In addition, usage of MOSSE was shown to be one of the effective trackers tested in underwater videos (<xref ref-type="bibr" rid="B26">Lopez-Marcano et&#xa0;al., 2021</xref>). The MOSSE algorithm initializes a correlation filter based on a detected object in a frame. Next, in the subsequent frames, the algorithm looks for a location having the highest correlation with the initially detected object. Due to the changes in appearance of the same <italic>Nephrops</italic> instances throughout the video, the <italic>Nephrops</italic> detection used for generating the correlation filter is updated every fifth frame. This approach was implemented earlier for tracking of yellowfin bream in underwater videos (<xref ref-type="bibr" rid="B26">Lopez-Marcano et&#xa0;al., 2021</xref>).</p>
<p>The other tracker evaluated is DeepSORT, an improved version of the SORT algorithm (<xref ref-type="bibr" rid="B53">Wojke et&#xa0;al., 2017</xref>). DeepSORT uses the appearance information of the detected objects together with their motion information in 2D. The motion information is quantified by the Mahalanobis distance between the detected bounding box centroids and the Kalman filter predictions under a constant velocity model. On the other hand, the appearance features for each detection are obtained by passing the bounding box region through a pre-trained CNN containing two convolutional and six residual layers. The minimum cosine distance between the appearance features of the detections and the last 100 features of each track is determined as the second metric used by DeepSORT. For the benchmarking experiments, the resources and the instructions made available in the official repository of DeepSORT are utilized (<xref ref-type="bibr" rid="B52">Wojke, 2019</xref>).</p>
<p>Instead of reporting the full detailed results for benchmarking trackers, only MOTA, HOTA, correct count rate, average FPS values, and F-scores for YOLOv4 model are provided (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>). Evaluation of these metrics is sufficient for comparing the trackers by understanding their overall performance.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Some performance metrics obtained by three different object tracking algorithms.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-10-1129852-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s5" sec-type="discussion">
<title>Discussion</title>
<p>A major challenge in demersal trawling is the lack of information about the catch entering the gear during fishing. This study demonstrates a full pipeline to acquire, process and display catch information for <italic>Nephrops</italic>, in close to real-time, to act as a decision tool for the fisher during the fishing operation. The applicability of such tools in commercial trawling and their potential improvements is discussed below.</p>
<p>One advantage of the proposed algorithm is the powerful image acquisition system that provides mostly sediment-free clear videos for being processed in the subsequent steps (<xref ref-type="bibr" rid="B41">Sokolova et&#xa0;al., 2021b</xref>; <xref ref-type="bibr" rid="B40">Sokolova et&#xa0;al., 2022</xref>). In the existing literature for underwater image processing, there are some papers where the effects of preprocessing on underwater images are analyzed for improving the detection performance (<xref ref-type="bibr" rid="B15">Han et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B59">Zhou et&#xa0;al., 2022</xref>). But the preprocessing requires some time, degrading the overall processing speed. In addition, there are different types of degradations such as low contrast and color distortion present in the underwater images (<xref ref-type="bibr" rid="B3">An et&#xa0;al., 2021</xref>). Our method does not require any preprocessing to enhance the detection accuracy because the image acquisition system is robust and capable of capturing clear videos with adjustable illumination (<xref ref-type="bibr" rid="B41">Sokolova et&#xa0;al., 2021b</xref>).</p>
<sec id="s5_1">
<title>Evaluation of the algorithm steps</title>
<p>Since the followed strategy is tracking-by-detection, successful <italic>Nephrops</italic> detection is expected to imply more accurate tracking which eventually may result in better <italic>Nephrops</italic> counts. Hence, achieving high <italic>mAP</italic> is critical at the object detection step. The performances of object detector models may be considered as sufficiently successful for an accurate tracking and counting task because all three models have <italic>mAP</italic> @.50 values above 95% (<xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>). In addition, the <italic>Nephrops</italic> detection performance, <italic>AP<sub>nep</sub>
</italic> value, associated with YOLOv4-Scaled model is the highest indicating a better detection capability of <italic>Nephrops</italic>. However, this situation is in connection with the increased size of the YOLOv4-Scaled model which slows down its respective detection speed (<xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>).</p>
<p>In the literature, there are numerous metrics defined for evaluating the performance of an object tracking algorithm. For simplicity, only those metrics commonly mentioned in the object tracking literature are provided in this paper. Among the three models, YOLOv4 model has the best values for MOTA, MT, ML, and HOTA. For a detection model, having higher MT and lower ML track count means that their associated successive detections are good enough to attain a valid track. This idea is also supported by the high accuracy values in MOTA and HOTA. On the other hand, an identity switch can be the source of a false positive count provided that the switching happens somewhere close to the horizontal level defined for counting conditions. As for the MOTP, it is very close for three of the models. This means that they have nearly the same level of success in bounding box localization throughout the tracks and cannot be used as a distinguishing factor for commenting on the counting performance.</p>
<p>Finally, it is possible to mention the performance for total <italic>Nephrops</italic> counts and the processing speeds of the method. Checking only the total counts at the end of the video may be misleading since some <italic>Nephrops</italic> are not counted while there may be multiple counts for some others. Therefore, checking the false positive and false negative counts together with the true positives gives better insight about the counting performance. The quantification of these three types of tracks is done by calculating the F-scores for each detector model. In addition, the rates for correct counts in each video are provided. At this point, it is notable that the correct count rates for Video-4 are relatively low when compared to the other four videos. The reason for such a remarkable difference is that Video-4 has some sediments degrading the visibility of the objects in the video. This situation highlights the importance of sediment-free video acquisition. Furthermore, when <xref ref-type="table" rid="T4">
<bold>Tables&#xa0;4</bold>
</xref>, <xref ref-type="table" rid="T5">
<bold>5</bold>
</xref> are considered together, it is possible to conclude that high performance at the object detection step does not always imply better correct count rates. This is apparent for the YOLOv4-Scaled model which has a very high detection rate but fails to achieve good count performance.</p>
<p>As for the processing speed, it is measured in terms of FPS. It is the type of the detector model that has a major impact on the overall duration of processing a frame. In addition, updating the object tracks by the SORT algorithm takes some time. During the experiments on the videos, it was observed that, on average, 1.6% of the total processing duration of the frames are used by SORT tracking algorithm when YOLOv4 is used as the object detector. However, tracking is effective only when there is a tracked object in the frame. Nevertheless, the maximum processing speed related with three of the models is higher than the FPS value of the input video (<xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>). This means that the detectors are capable of running at real-time processing speed, but this speed may be reduced when there is a tracked object in the video. On average, the processing speeds of YOLOv4-Scaled is slightly below the real time threshold while the other two models are fast enough to be considered real-time.</p>
<p>The benchmarking results of SORT with MOSSE and DeepSORT trackers revealed that SORT is a better tracker for this application in terms of tracking accuracy, <italic>Nephrops</italic> counting, and processing speed. The major problem with the MOSSE tracker is the requirement for updating the correlation filters frequently. This process slows down the procedure considerably. On the other hand, tracking without any correlation filter update step, MOSSE is quite inefficient for this problem because the <italic>Nephrops</italic> individuals float and rotate under the influence of water flow causing their appearance to be changed as they are in the field of view of the camera. As for DeepSORT, it is more accurate than MOSSE in terms of counting performance. However, the CNN-based feature extraction step slows down the overall tracking speed and eventually causes the slowest processing.</p>
</sec>
<sec id="s5_2">
<title>Implications for the <italic>nephrops</italic> fishing</title>
<p>Demersal trawling is a blind process today, which means that fishers do not know if they are catching the target species during trawling operation. This study constitutes a basis for addressing this problem by outputting the target catch count with a real-time speed. In other words, it demonstrates the possibility of providing the <italic>Nephrops</italic> catch amount throughout the trawling operation. Such information is useful for not only improving the catch rates of the target species but also reducing the bycatch amounts, oil and energy consumption, and ultimately improve the economic, environmental, and social sustainability of the fishery.</p>
</sec>
<sec id="s5_3">
<title>Further development</title>
<p>The first step for further improvement of the proposed method is to run it on an edge device with limited computational power. Note that the reported results in this study were obtained using a powerful processing unit (Section 3.4). In real world applications, it may not be practical to access such a computer. Therefore, experimentation with an edge device, which is more accessible onboard commercial fishing vessels, is one of the improvement plans with high priority. The change of the processing platform may not affect the correct count rates, but will have an influence on the overall processing speed. Nevertheless, the achieved speed with YOLOv4-Tiny model is promising and it may still perform sufficiently fast on an edge device.</p>
<p>When there is a tracked object in the video, the tracking speed drops considerably. In other words, tracking step is a bottleneck in the procedure. However, SORT is known to be one of the fast tracking algorithms in the literature, which is also supported by the benchmarking results. In case of requiring higher speed, skipping some intermediate frames may be helpful at cost of degradation in the count accuracy. This may contribute to the compensation of the speed loss due to the edge device. Besides, even if there is a small delay, the achieved processing speed may be considered as a significant improvement when compared to hours of delay associated with the current situation, where information on catch rates and compositions is only available once the catch is taken onboard the vessel.</p>
<p>In the longer term, the method may be extended to detect and count more species and contribute to a larger scale in fisheries. However, this requires generation of a larger video dataset containing more diverse species. In addition, the edge processing unit may be connected to the stereo camera directly by integrating them inside the underwater camera box. This may be coupled with a wireless transceiver device that transmits the count information, e.g. acoustically to a screen onboard. This key information is sufficient for the fisher to decide whether to continue fishing in the same area.</p>
</sec>
</sec>
<sec id="s6" sec-type="conclusion">
<title>Conclusion</title>
<p>This study demonstrates the possibility of using state-of-the-art deep learning methods to develop real-time decision tools for the trawl fisheries demonstrated here as a <italic>Nephrops</italic> counter. In particular, the experiments are carried out with three different object detector models on underwater videos collected by an in-trawl camera. The detection, tracking, and counting performances as well as the processing speeds associated with these models are calculated. According to the obtained results, it is possible to conclude that such a system is promising for improving the sustainability of trawl fisheries.</p>
</sec>
<sec id="s7" sec-type="data-availability">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.11583/DTU.21769442">https://doi.org/10.11583/DTU.21769442</ext-link>.</p>
</sec>
<sec id="s8" sec-type="ethics-statement">
<title>Ethics statement</title>
<p>Ethical review and approval was not required for the animal study because all data collection was conducted during trawl fishing at sea which do not require an ethical permit or animal welfare approval.</p>
</sec>
<sec id="s9" sec-type="author-contributions">
<title>Author contributions</title>
<p>EA: methodology, coding, manuscript writing. JF: conceptualization, supervision, manuscript writing and editing. LK: funding acquisition, conceptualization, supervision, manuscript writing and editing. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s10" sec-type="funding-information">
<title>Funding</title>
<p>This work has received funding from the European Maritime and Fisheries Fund (EMFF), the Ministry of Food, Agriculture and Fisheries of Denmark, and the European Union&#x2019;s Horizon 2020 research and innovation program as part of the projects: Development of a real-time catch monitoring system with automatic detection of the catch composition to minimize catch of unwanted species and sizes [AutoCatch (33112-P-18-051)], Udvikling af SELEKTive redskaber og teknologier til kommercielle fiskerier [SELEKT (33113-I-22-187)], and Smart fisheries technologies for an efficient, compliant and environmentally friendly fishing sector [SMARTFISH (agreement no: 7553521)].</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>The authors thank the skipper and the crew on DTU&#x2019;s research vessel RV Havfisken for assistance in data collection at sea.</p>
</ack>
<sec id="s11" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s12" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aguzzi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sard&#xe0;</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>A history of recent advancements on nephrops norvegicus behavioral and physiological rhythms</article-title>. <source>Rev. Fish Biol. Fish</source> <volume>18</volume>, <fpage>235</fpage>&#x2013;<lpage>248</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S11160-007-9071-9/FIGURES/6</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Allken</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Rosen</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Handegard</surname> <given-names>N. O.</given-names>
</name>
<name>
<surname>Malde</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A deep learning-based method to identify and count pelagic and mesopelagic fishes from trawl camera images</article-title>. <source>ICES J. Mar. Sci.</source> <volume>78</volume>, <fpage>3780</fpage>&#x2013;<lpage>3792</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ICESJMS/FSAB227</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>An</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Hao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Application of computer vision in fish intelligent feeding system&#x2013;a review</article-title>. <source>Aquac Res.</source> <volume>52</volume>, <fpage>423</fpage>&#x2013;<lpage>437</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/ARE.14907</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bergmann</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Wieczorek</surname> <given-names>S. K.</given-names>
</name>
<name>
<surname>Moore</surname> <given-names>P. G.</given-names>
</name>
<name>
<surname>Atkinson</surname> <given-names>R. J. A.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Discard composition of the nephrops fishery in the Clyde Sea area, Scotland</article-title>. <source>Fish Res.</source> <volume>57</volume>, <fpage>169</fpage>&#x2013;<lpage>183</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/S0165-7836(01)00345-9</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bewley</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Ge</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ott</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ramos</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Upcroft</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Simple online and realtime tracking</article-title>,&#x201d; in <conf-name>Proceedings - International Conference on Image Processing, ICIP</conf-name>, , <conf-date>2016-August</conf-date>. <fpage>3464</fpage>&#x2013;<lpage>3468</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICIP.2016.7533003</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Bochkovskiy</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>) <source>GitHub - AlexeyAB/darknet: YOLOv4 / scaled-YOLOv4 / YOLO - neural networks for object detection (Windows and Linux version of darknet )</source>. Available at: <uri xlink:href="https://github.com/AlexeyAB/darknet">https://github.com/AlexeyAB/darknet</uri> (Accessed <access-date>November 14, 2022</access-date>).</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bochkovskiy</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.-Y.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>H.-Y. M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>YOLOv4: Optimal speed and accuracy of object detection</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.2004.10934</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bolme</surname> <given-names>D. S.</given-names>
</name>
<name>
<surname>Beveridge</surname> <given-names>J. R.</given-names>
</name>
<name>
<surname>Draper</surname> <given-names>B. A.</given-names>
</name>
<name>
<surname>Lui</surname> <given-names>Y. M.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Visual object tracking using adaptive correlation filters</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</conf-name>. <fpage>2544</fpage>&#x2013;<lpage>2550</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2010.5539960</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dutta</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Zisserman</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>The VIA annotation software for images, audio and video</article-title>,&#x201d; in <conf-name>Proceedings of the 27th ACM International Conference on Multimedia MM &#x2019;19</conf-name>, <conf-loc>New York, NY, USA</conf-loc>. <fpage>2276</fpage>&#x2013;<lpage>2279</lpage> (Association for Computing Machinery). doi:&#xa0;<pub-id pub-id-type="doi">10.1145/3343031.3350535</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eigaard</surname> <given-names>O. R.</given-names>
</name>
<name>
<surname>Bastardie</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Hintzen</surname> <given-names>N. T.</given-names>
</name>
<name>
<surname>Buhl-Mortensen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Buhl-Mortensen</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Catarino</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>The footprint of bottom trawling in European waters: distribution, intensity, and seabed integrity</article-title>. <source>ICES J. Mar. Sci.</source> <volume>74</volume>, <fpage>847</fpage>&#x2013;<lpage>865</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ICESJMS/FSW194</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>ElTantawy</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shehata</surname> <given-names>M. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Local null space pursuit for real-time moving object detection in aerial surveillance</article-title>. <source>Signal Image Video Process</source> <volume>14</volume>, <fpage>87</fpage>&#x2013;<lpage>95</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S11760-019-01528-Y/FIGURES/3</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Feekings</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Christensen</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Jonsson</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Frandsen</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Ulmestrand</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Munch-Petersen</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>The use of at-sea-sampling data to dissociate environmental variability in Norway lobster (Nephrops norvegicus) catches to improve resource exploitation efficiency within the Skagerrak/Kattegat trawl fishery</article-title>. <source>Fish Oceanogr</source> <volume>24</volume>, <fpage>383</fpage>&#x2013;<lpage>392</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/FOG.12116</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>French</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Mackiewicz</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Fisher</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Holah</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Kilburn</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Campbell</surname> <given-names>N.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Deep neural networks for analysis of fisheries surveillance video and automated monitoring of fish discards</article-title>. <source>ICES J. Mar. Sci.</source> <volume>77</volume>, <fpage>1340</fpage>&#x2013;<lpage>1353</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ICESJMS/FSZ149</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ghiasi</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Srinivas</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>T.-Y.</given-names>
</name>
<name>
<surname>Cubuk</surname> <given-names>E. D.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Simple copy-paste is a strong data augmentation method for instance segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <fpage>2918</fpage>&#x2013;<lpage>2928</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2012.07177</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Underwater image processing and object detection based on deep CNN method</article-title>. <source>J. Sens</source> <volume>2020</volume>, <page-range>1&#x2013;20</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2020/6707328</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Howard</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sandler</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.-C.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Searching for MobileNetV3</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.1905.02244</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Howard</surname> <given-names>A. G.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Kalenichenko</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Weyand</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>MobileNets: Efficient convolutional neural networks for mobile vision applications</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.1704.04861</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Real-time nondestructive fish behavior detecting in mixed polyculture system using deep-learning and low-cost devices</article-title>. <source>Expert Syst. Appl.</source> <volume>178</volume>, <elocation-id>115051</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.ESWA.2021.115051</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jalal</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Salman</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Mian</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shortis</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Shafait</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Fish detection and species classification in underwater environments using deep learning with temporal information</article-title>. <source>Ecol. Inform</source> <volume>57</volume>, <elocation-id>101088</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.ECOINF.2020.101088</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Vision-based target tracking for unmanned surface vehicle considering its motion features</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>132655</fpage>&#x2013;<lpage>132664</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2020.3010327</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kalman</surname> <given-names>R. E.</given-names>
</name>
</person-group> (<year>1960</year>). <article-title>A new approach to linear filtering and prediction problems</article-title>. <source>J. Basic Eng.</source> <volume>82</volume>, <fpage>35</fpage>&#x2013;<lpage>45</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1115/1.3662552</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuhn</surname> <given-names>H. W.</given-names>
</name>
</person-group> (<year>1955</year>). <article-title>The Hungarian method for the assignment problem</article-title>. <source>Naval Res. Logistics Q.</source> <volume>2</volume>, <fpage>83</fpage>&#x2013;<lpage>97</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/NAV.3800020109</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>He</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Gu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A robust underwater multiclass fish-school tracking algorithm</article-title>. <source>Remote Sens.</source> <volume>14</volume>, <elocation-id>4106</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/RS14164106</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Hou</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Real-time marine animal images classification by embedded system based on mobilenet and transfer learning</article-title>,&#x201d; in <conf-name>OCEANS 2019 - Marseille</conf-name>. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/OCEANSE.2019.8867190</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Towards domain generalization in underwater object detection</article-title>,&#x201d; in <conf-name>2020 IEEE International Conference on Image Processing (ICIP)</conf-name>. <fpage>1971</fpage>&#x2013;<lpage>1975</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICIP40778.2020.9191364</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lopez-Marcano</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Jinks</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Buelow</surname> <given-names>C. A.</given-names>
</name>
<name>
<surname>Brown</surname> <given-names>C. J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Kusy</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Automatic detection of fish and tracking of movement for ecology</article-title>. <source>Ecol. Evol.</source> <volume>11</volume>, <fpage>8254</fpage>&#x2013;<lpage>8263</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/ECE3.7656</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luiten</surname> <given-names>J.</given-names>
</name>
<name>
<surname>O&#x161;ep</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Dendorfer</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Torr</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Geiger</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Leal-Taix&#xe9;</surname> <given-names>L.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>HOTA: A higher order metric for evaluating multi-object tracking</article-title>. <source>Int. J. Comput. Vis.</source> <volume>129</volume>, <fpage>548</fpage>&#x2013;<lpage>578</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S11263-020-01375-2/FIGURES/18</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Main</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sangster</surname> <given-names>G. I.</given-names>
</name>
</person-group> (<year>1985</year>). &#x201c;<article-title>The behaviour of the Norway lobster, nephrops norvegicus (L.), during trawling</article-title>,&#x201d; in <source>Scottish Fisheries research report</source>, vol. <volume>43</volume>. (<publisher-loc>Aberdeen</publisher-loc>: <publisher-name>Department of Agriculture and Fisheries for Scotland</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>23</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mohamed</surname> <given-names>H. E. D.</given-names>
</name>
<name>
<surname>Fadl</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Anas</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Wageeh</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Elmasry</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Nabil</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>MSR-YOLO: Method to enhance fish detection and tracking in fish farms</article-title>. <source>Proc. Comput. Sci.</source> <volume>170</volume>, <fpage>539</fpage>&#x2013;<lpage>546</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.PROCS.2020.03.123</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muksit</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Hasan</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Hasan Bhuiyan Emon</surname> <given-names>M. F.</given-names>
</name>
<name>
<surname>Haque</surname> <given-names>M. R.</given-names>
</name>
<name>
<surname>Anwary</surname> <given-names>A. R.</given-names>
</name>
<name>
<surname>Shatabda</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>YOLO-fish: A robust fish detection model to detect fish in realistic underwater environment</article-title>. <source>Ecol. Inform</source> <volume>72</volume>, <elocation-id>101847</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.ECOINF.2022.101847</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Naseer</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Baro</surname> <given-names>E. N.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S. D.</given-names>
</name>
<name>
<surname>Gordillo</surname> <given-names>Y. V.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Automatic detection of nephrops norvegicus burrows in underwater images using deep learning</article-title>,&#x201d; in <conf-name>2020 Global Conference on Wireless and Optical Technologies (GCWOT)</conf-name>. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/GCWOT49901.2020.9391590</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Petrellis</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Measurement of fish morphological features through image processing and deep learning techniques</article-title>. <source>Appl. Sci.</source> <volume>11</volume>, <elocation-id>4416</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/APP11104416</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Prados</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Garcia</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Gracias</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Neumann</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Vagstol</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Real-time fish detection in trawl nets</article-title>,&#x201d; in <conf-name>OCEANS 2017</conf-name>, <conf-loc>Aberdeen</conf-loc>, <conf-date>2017-October</conf-date>. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/OCEANSE.2017.8084760</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Raza</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Fast and accurate fish detection design with improved YOLO-v3 model and transfer learning</article-title>. <source>Int. J. Advanced Comput. Sci. Appl.</source> <volume>11</volume>, <fpage>7</fpage>&#x2013;<lpage>16</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.14569/IJACSA.2020.0110202</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Redmon</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>) <source>Darknet: Open source neural networks in c</source>. Available at: <uri xlink:href="https://pjreddie.com/darknet/">https://pjreddie.com/darknet/</uri> (Accessed <access-date>November 14, 2022</access-date>).</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Redmon</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Divvala</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Farhadi</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>You only look once: Unified, real-time object detection</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.1506.02640</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sala</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Mayorga</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Bradley</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Cabral</surname> <given-names>R. B.</given-names>
</name>
<name>
<surname>Atwood</surname> <given-names>T. B.</given-names>
</name>
<name>
<surname>Auber</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Protecting the global ocean for biodiversity, food and climate</article-title>. <source>Nature</source> <volume>592</volume>, <fpage>397</fpage>&#x2013;<lpage>402</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41586-021-03371-z</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sandler</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Howard</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhmoginov</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.-C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>MobileNetV2: Inverted residuals and linear bottlenecks</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.1801.04381</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sokolova</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Momp&#xf3; Alepuz</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Mariani</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Galeazzi</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Krag</surname> <given-names>L. A.</given-names>
</name>
</person-group> (<year>2021</year>a). <article-title>A deep learning approach to assist sustainability of demersal trawling operations</article-title>. <source>Sustainability</source> <volume>13</volume>, <elocation-id>12362</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/SU132212362</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sokolova</surname> <given-names>M.</given-names>
</name>
<name>
<surname>O&#x2019;Neill</surname> <given-names>F. G.</given-names>
</name>
<name>
<surname>Savina</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Krag</surname> <given-names>L. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Test and development of a sediment suppressing system for catch monitoring in demersal trawls</article-title>. <source>Fish Res.</source> <volume>251</volume>, <elocation-id>106323</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.FISHRES.2022.106323</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sokolova</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Mariani</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Krag</surname> <given-names>L. A.</given-names>
</name>
</person-group> (<year>2021</year>b). <article-title>Towards sustainable demersal fisheries: NepCon image acquisition system for automatic nephrops norvegicus detection</article-title>. <source>PloS One</source> <volume>16</volume>, <elocation-id>e0252824</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/JOURNAL.PONE.0252824</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Soom</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Pattanaik</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Leier</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Tuhtan</surname> <given-names>J. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Environmentally adaptive fish or no-fish classification for river video fish counters using high-performance desktop and embedded hardware</article-title>. <source>Ecol. Inform</source> <volume>72</volume>, <elocation-id>101817</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.ECOINF.2022.101817</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tseng</surname> <given-names>C.-H.</given-names>
</name>
<name>
<surname>Kuo</surname> <given-names>Y.-F.</given-names>
</name>
<name>
<surname>Tseng</surname> <given-names>C.-H.</given-names>
</name>
<name>
<surname>Kuo</surname> <given-names>Y.-F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Detecting and counting harvested fish and identifying fish types in electronic monitoring system videos using deep convolutional neural networks</article-title>. <source>ICES J. Mar. Sci.</source> <volume>77</volume>, <fpage>1367</fpage>&#x2013;<lpage>1378</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ICESJMS/FSAA076</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tully</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Hillis</surname> <given-names>J. P.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Causes and spatial scales of variability in population structure of nephrops norvegicus (L.) in the Irish Sea</article-title>. <source>Fish Res.</source> <volume>21</volume>, <fpage>329</fpage>&#x2013;<lpage>347</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/0165-7836(94)00303-E</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Underwood</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Rosen</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Engas</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Eriksen</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Deep vision: An in-trawl stereo camera makes a step forward in monitoring the pelagic community</article-title>. <source>PloS One</source> <volume>9</volume>, <elocation-id>e112304</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/JOURNAL.PONE.0112304</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Underwood</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Rosen</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Eng&#xe5;s</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Jorgensen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Fern&#xf6;</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Species-specific residence times in the aft part of a pelagic survey trawl: implications for inference of pre-capture spatial distribution using the deep vision system</article-title>. <source>ICES J. Mar. Sci.</source> <volume>75</volume>, <fpage>1393</fpage>&#x2013;<lpage>1404</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/ICESJMS/FSX233</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vijaya Kumar</surname> <given-names>D. T. T.</given-names>
</name>
<name>
<surname>Mahammad Shafi</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A fast feature selection technique for real-time face detection using hybrid optimized region based convolutional neural network</article-title>. <source>Multimed Tools Appl.</source>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S11042-022-13728-9</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wageeh</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Mohamed</surname> <given-names>H. E. D.</given-names>
</name>
<name>
<surname>Fadl</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Anas</surname> <given-names>O.</given-names>
</name>
<name>
<surname>ElMasry</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Nabil</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>YOLO fish detection with euclidean tracking in fish farms</article-title>. <source>J. Ambient Intell. Humaniz Comput.</source> <volume>12</volume>, <fpage>5</fpage>&#x2013;<lpage>12</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S12652-020-02847-6/FIGURES/6</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>C. Y.</given-names>
</name>
<name>
<surname>Bochkovskiy</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>H. Y. M.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Scaled-YOLOv4: Scaling cross stage partial network</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition</conf-name>. <fpage>13024</fpage>&#x2013;<lpage>13033</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.2011.08036</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>He</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Shao</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>b). <article-title>A novel attention-based lightweight network for multiscale object detection in underwater images</article-title>. <source>J. Sens</source> <volume>2022</volume>, <elocation-id>2582687</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2022/2582687</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>Real-time detection and tracking of fish abnormal behavior based on improved YOLOV5 and SiamRPN++</article-title>. <source>Comput. Electron Agric.</source> <volume>192</volume>, <elocation-id>106512</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.COMPAG.2021.106512</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wojke</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2019</year>) <source>GitHub - nwojke/deep_sort: Simple online realtime tracking with a deep association metric</source>. Available at: <uri xlink:href="https://github.com/nwojke/deep_sort">https://github.com/nwojke/deep_sort</uri> (Accessed <access-date>November 14, 2022</access-date>).</citation>
</ref>
<ref id="B53">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wojke</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Bewley</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Paulus</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Simple online and realtime tracking with a deep association metric</article-title>,&#x201d; in <conf-name>Proceedings - International Conference on Image Processing, ICIP</conf-name>, , <conf-date>2017-September</conf-date>. <fpage>3645</fpage>&#x2013;<lpage>3649</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arxiv.1703.07402</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Deep learning for unmanned aerial vehicle-based object detection and tracking: A survey</article-title>. <source>IEEE Geosci Remote Sens Mag</source> <volume>10</volume>, <fpage>91</fpage>&#x2013;<lpage>124</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/MGRS.2021.3115137</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Qiu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Application of improved MobileNet-SSD on underwater sea cucumber detection robot</article-title>,&#x201d; in <conf-name>2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC)</conf-name>. <fpage>402</fpage>&#x2013;<lpage>407</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IAEAC47372.2019.8997970</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>W.</given-names>
</name>
<name>
<surname>He</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Lightweight underwater object detection based on YOLO v4 and multi-scale attentional feature fusion</article-title>. <source>Remote Sens (Basel)</source> <volume>13</volume>, <page-range>1&#x2013;22</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13224706</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>An improved YOLO algorithm for fast and accurate underwater object detection</article-title>. <source>Symmetry (Basel)</source> <volume>14</volume>, <page-range>1&#x2013;16</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/sym14081669</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). &#x201c;<article-title>Fish recognition from a vessel camera using deep convolutional neural network and data augmentation</article-title>,&#x201d; in <conf-name>2018 OCEANS - MTS/IEEE Kobe Techno-Oceans, OCEANS - Kobe 2018</conf-name>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/OCEANSKOBE.2018.8559314</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>An underwater target recognition method based on improved YOLOv4 in complex marine environment</article-title>. <source>Syst. Sci. Control. Eng.</source> <volume>10</volume>, <fpage>590</fpage>&#x2013;<lpage>602</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1080/21642583.2022.2082579</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>