<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Sig. Proc.</journal-id>
<journal-title>Frontiers in Signal Processing</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Sig. Proc.</abbrev-journal-title>
<issn pub-type="epub">2673-8198</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">765006</article-id>
<article-id pub-id-type="doi">10.3389/frsip.2021.765006</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Signal Processing</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning-Based Object Tracking <italic>via</italic> Compressed Domain Residual Frames</article-title>
<alt-title alt-title-type="left-running-head">El Khoury et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Object Tracking Using Residual Frames</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>El Khoury</surname>
<given-names>Karim</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1469303/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Samelson</surname>
<given-names>Jonathan</given-names>
</name>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1448585/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Macq</surname>
<given-names>Beno&#xee;t</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1506803/overview"/>
</contrib>
</contrib-group>
<aff>Institute of Information and Communication Technologies, Electronics and Applied Mathematics, Universit&#xe9; catholique de Louvain, <addr-line>Louvain-la-Neuve</addr-line>, <country>Belgium</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1230969/overview">Lukas Esterle</ext-link>, Aarhus University, Denmark</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1178210/overview">Stefania Colonnese</ext-link>, Sapienza University of Rome, Italy</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1169249/overview">Ionut Schiopu</ext-link>, Vrije University Brussel, Belgium</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Karim El Khoury, <email>karim.elkhoury@uclouvain.be</email>
</corresp>
<fn fn-type="equal" id="fn1">
<label>
<sup>&#x2020;</sup>
</label>
<p>These authors have contributed equally to this work and share first authorship</p>
</fn>
<fn fn-type="other">
<p>This article was submitted to Image Processing, a section of the journal Frontiers in Signal Processing</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>1</volume>
<elocation-id>765006</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 El Khoury, Samelson and Macq.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>El Khoury, Samelson and Macq</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>The extensive rise of high-definition CCTV camera footage has stimulated both the data compression and the data analysis research fields. The increased awareness of citizens to the vulnerability of their private information, creates a third challenge for the video surveillance community that also has to encompass privacy protection. In this paper, we aim to tackle those needs by proposing a deep learning-based object tracking solution via compressed domain residual frames. The goal is to be able to provide a public and privacy-friendly image representation for data analysis. In this work, we explore a scenario where the tracking is achieved directly on a restricted part of the information extracted from the compressed domain. We utilize exclusively the residual frames already generated by the video compression codec to train and test our network. This very compact representation also acts as an information filter, which limits the amount of private information leakage in a video stream. We manage to show that using residual frames for deep learning-based object tracking can be just as effective as using classical decoded frames. More precisely, the use of residual frames is particularly beneficial in simple video surveillance scenarios with non-overlapping and continuous traffic.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>video compression</kwd>
<kwd>residual frames</kwd>
<kwd>video surveillance</kwd>
<kwd>object detection</kwd>
<kwd>object tracking</kwd>
<kwd>privacy protection</kwd>
<kwd>HOTA</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>According to Cisco&#x2019;s Visual Networking Index report in 2017, global data consumption has been increasing exponentially for the past decade, with video data accounting for 80% of the worldwide traffic<xref ref-type="fn" rid="fn2">
<sup>1</sup>
</xref>. One of the largest growing types of video data consumption is video surveillance traffic, which is set to achieve a seven-fold increase by 2022 to account for a total of 3% of the worldwide Internet traffic. This substantial surge in video surveillance data had created three major needs. The first major need is to be able to transfer and store the data which calls for the use of innovative video compression codecs. The second major need is to be able to analyze the large flow of data, which calls for the use of machine learning and more specifically deep learning algorithms. Lastly, the third major need, which is especially relevant when working with video surveillance footage, is to able to preserve the privacy of the individuals involved in the captured scenes. The main motivation of this paper is to address these three needs by using an inexpensive, low-storage and privacy-friendly image representation that can therefore be made publicly available for traffic analysis.</p>
<p>In this work, we utilize an inexpensive compressed domain image representation already generated by the video compression codec: residual frames. Also known as the prediction error, residual frames are the difference between the prediction of a frame at time <italic>t&#x2b;1</italic> using the frame at time <italic>t</italic> and the original frame at time <italic>t&#x2b;1</italic>. Residual frames not only have a low-storage cost but they also act as an information filter by only keeping the movement regions of interest (ROI) between two consecutive frames. In this research, we choose to work exclusively on the residual frames to train and test a deep learning-based object detector and tracker. This will allow us to not only store our data in the compressed format but also provide a privacy-friendly data source for deep learning-based object tracking. Deep learning-based object tracking trained and tested solely on residual frames is a new approach explored by this paper. This research&#x2019;s main contribution is to show that using residual frames as an image representation for a deep learning-based object tracking can be just as effective as using decoded frames while limiting the amount of private information leakage in a video stream.</p>
<p>This paper is organized as follows. In <xref ref-type="sec" rid="s2">Section 2</xref>, we detail the state of the art of combining data compression, data analysis and data protection. In <xref ref-type="sec" rid="s3">Section 3</xref> we put forward the utilized materials and methods by presenting the compression algorithm and both object detectors and trackers used. Then we introduce the HOTA evaluation metric, the different datasets that we have worked with, and the two experiments that have been conducted. In <xref ref-type="sec" rid="s4">Section 4</xref> we present the detailed individual results of the two experiments. Thereafter, in <xref ref-type="sec" rid="s5">Section 5</xref>, we expose the benefits and drawbacks resulting from the two experiments, examine the privacy-friendly capabilities of our solution, and comment on key limitations that have impacted this research. Finally, in <xref ref-type="sec" rid="s6">Section 6</xref>, we conclude the paper by summarizing the results and outcomes of our research and proposing several potential further&#x20;work.</p>
</sec>
<sec id="s2">
<title>2 Related Works</title>
<p>The challenges of combining the two needs of video compression and video analysis is a topic that has already been addressed in the literature. The Moving Picture Experts Group (MPEG) recently created an ad hoc group dedicated to the standardization of Video Coding for Machines (VCM) (<xref ref-type="bibr" rid="B9">Duan et&#x20;al., 2020</xref>). The VCM group&#x2019;s inception came after the realization that traditional video compression codecs were not optimal for deep learning feature extraction. The aim of the VCM group is to create a video compression codec tailored to machine vision rather than human perception. The proposed codec managed to achieve, at lower bit-rate costs, much better detection accuracy in most cases and more visually pleasing decoded videos than the High Efficiency Video Codec (HEVC). Another proposition within the same scope proposed a hybrid framework that combined convolutional neural networks (CNN) with classical background subtraction techniques (<xref ref-type="bibr" rid="B14">Kim et&#x20;al., 2018</xref>). The proposed framework was made up of a two-step process. The first step was to identify the ROI using a background subtraction algorithm on all frames. The second step was to apply a CNN classifier to the ROI. They managed to achieve a classification accuracy of up to&#x20;85%.</p>
<p>In addition, other works have also looked at taking advantage of already generated compressed domain motion vectors to improve the efficiency of deep learning networks. Researchers proposed to work on a CNN-based detector combined with compressed domain motion vectors to lower the power consumption of classical deep learning-based detectors (<xref ref-type="bibr" rid="B32">Ujiie et&#x20;al., 2018</xref>). They utilized the inexpensive motion vectors already generated by the video compression codec to speed up the detection process in the predicted frames. Using the MOT16 benchmark, they obtained a MOTA score of 88% while also cutting the detection frequency by twelve times. Other researchers have also explored the use of the compressed domain motion vectors, but concentrated their efforts on improving the efficiency of their CNN-based object tracker (<xref ref-type="bibr" rid="B16">Liu et&#x20;al., 2019</xref>). They manage to achieve a tracker that is six times faster than the state-of-the-art online multi-object tracking (MOT) methods.</p>
<p>Another challenge that has been addressed in the literature is to combine the two needs of data analysis and data privacy. Researchers have tried to simplify the problem by proposing to tackle specific features and excluding them from the frames as a binary decision. An image scrambling method for privacy-friendly video surveillance showed that it was possible to scramble the frame&#x2019;s ROI to hide critical information in the observed scene (<xref ref-type="bibr" rid="B10">Dufaux and Ebrahimi, 2006</xref>). Other related work take up the challenge of combining data compression and data privacy. Researchers developed a custom license plate recognition and facial recognition software to encrypt the specific ROI before encoding and sending out the video sequence (<xref ref-type="bibr" rid="B6">Carrillo et&#x20;al., 2008</xref>).</p>
<p>Although the presented works tackle at least one of the three major needs, none of them attempt to tackle all three major needs in one unified solution.</p>
</sec>
<sec id="s3">
<title>3 Materials and Methods</title>
<p>In this section we present the compression algorithm used, the object detectors and trackers that we have worked with, our evaluation metric and datasets, and finally our experimental setup. All publicly available source codes used in this work are made available at <ext-link ext-link-type="uri" xlink:href="https://github.com/JonathanSamelson/ResidualsTracking">https://github.com/JonathanSamelson/ResidualsTracking</ext-link>.</p>
<sec id="s3-1">
<title>3.1 Compression Algorithm</title>
<p>Inter-frame video codecs use the temporal redundancies of a video sequence to compress it. This is achieved by segmenting the video sequences into reference frames (also called I frames) and predicted frames (also called P or B frames). The reference frames consist of sending the full intra-frame image whereas the predicted frames are generated by a process called block matching. Block matching divides the frame into several non-overlapping blocks of predetermined size and assigns a motion vector (also known as a displacement vector) to each block by identifying the location of that block in the previous frame. The motion vectors paired with the latest original stored frame enable us to make a prediction on the upcoming frame and subsequently generate the frame prediction error (also called residual frame) by subtracting the latest original frame from the predicted&#x20;frame.</p>
<p>The most widely used inter-frame video compression formats such as HEVC (<xref ref-type="bibr" rid="B28">Sullivan et&#x20;al., 2012</xref>), VVC (<xref ref-type="bibr" rid="B13">Huang et&#x20;al., 2020</xref>), and VP9 (<xref ref-type="bibr" rid="B22">Mukherjee et&#x20;al., 2013</xref>) all rely on a quadtree structure for their block matching process called adaptive block matching. The goal is to have variable block sizes that depend on the scene depicted in the frame. Ideally, we would like to have large blocks that represent inanimate sections of the frame for the background, and small blocks that represent movement areas, for the foreground. This would lower the encoding cost per frame, as it would reduce the number of encoded blocks and motion vectors per frame. The adaptive block matching process can also be seen as an image content filter given that it highlights ROI in the frame (movement areas) over the inanimate sections of the frame. <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> shows a sample image of the residual frame (A) as well as the respective adaptive block matching quadtree (B) generated for the two successive decoded frames shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Sample image of the residual frame <bold>(A)</bold> and the respective adaptive block matching quadtree <bold>(B)</bold> generated from two successive frames.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Sample image of two successive decoded frames <bold>(A)</bold> and <bold>(B)</bold>.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g002.tif"/>
</fig>
<p>For this work we developed our own open-source generic adaptive block matching algorithm inspired by the works of (<xref ref-type="bibr" rid="B33">Vermaut et&#x20;al., 2001</xref>; <xref ref-type="bibr" rid="B1">Barjatya, 2004</xref>). Our algorithm works similarly to the widely used standardized inter-frame compression formats such as presented in (<xref ref-type="bibr" rid="B7">Chien et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B39">Zhang et&#x20;al., 2019</xref>). This allows us to generate the motion vectors and corresponding residual frames needed for our study without accessing, editing, and testing all the available inter-frame compression video codecs. The algorithm needs to know only three preset values: the size of the largest possible block, the size of the smallest possible block, and the sensitivity threshold. The algorithm starts by calculating the motion vectors for the largest blocks. Once it has done so, it goes through the motion vectors individually from left to right and from top to bottom and looks at each block&#x2019;s individual neighbors. If the absolute value of the difference of the motion vector and the averages of its neighbors are greater than the preset threshold, the concerned block is split. Otherwise, this motion vector is confirmed and becomes final. The act of splitting means that the block containing the motion vector will be divided into four equal blocks and the process continues recursively until we reach the preset smallest block size. The detailed algorithm is shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Block diagram of the generic adaptive block matching algorithm (ABMA).</p>
</caption>
<graphic xlink:href="frsip-01-765006-g003.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>3.2 Object Detectors</title>
<p>In this work we chose to compare the performance on the two image representations (residual and decoded frames) with the help of two trained object detectors: YOLOv4 and tiny YOLOv4. Over the last few years, You Only Look Once (YOLO) has been one of the state-of-the-art single-stage detectors on datasets such as MS-COCO (<xref ref-type="bibr" rid="B15">Lin et&#x20;al., 2014</xref>), thanks to its different updates and versions. As the name suggests, the whole frame is scanned in a single evaluation, making the inference faster and allowing the detector to achieve real-time performance. To do so, YOLO divides the frame into pre-defined grid cells, each responsible for detecting objects thanks to YOLO&#x2019;s anchor boxes (prior boxes) of different shapes. Thus, each cell can produce multiple predictions containing the bounding box dimensions as well as the object class and certainty.</p>
<p>In its fourth revision, the authors present many improvements that make YOLO faster and more robust (<xref ref-type="bibr" rid="B5">Bochkovskiy et&#x20;al., 2020</xref>). Among the most important ones that improve training, are CutMix and Mosaic data augmentation techniques (<xref ref-type="bibr" rid="B38">Yun et&#x20;al., 2019</xref>), DropBlock regularization method, Cross mini-Batch Normalization (CmBN), and Self-Adversarial Training (SAT). To improve inference time, they notably introduced Mish activation function (<xref ref-type="bibr" rid="B21">Misra, 2019</xref>), SPP-block (<xref ref-type="bibr" rid="B12">He et&#x20;al., 2015</xref>), Cross-stage partial connections (CSP) (<xref ref-type="bibr" rid="B34">Wang et&#x20;al., 2020b</xref>), PAN path-aggregation block (<xref ref-type="bibr" rid="B17">Liu et&#x20;al., 2018</xref>), and Multi-input weighted residual connections (MiWRC).</p>
<p>The tiny version of YOLO follows the same principles, with a drastically reduced network size. Basically, the number of convolutional layers in the backbone are scaled down, as is the number of anchor boxes that make the predictions. Consequently, it infers the result much quicker but often leads to models with lower accuracy.</p>
<p>One could consider applying a thresholding method to this light representation to obtain a binary image and find the contours of the objects using the algorithms presented in (<xref ref-type="bibr" rid="B29">Suzuki and Abe, 1985</xref>; <xref ref-type="bibr" rid="B26">Ren et&#x20;al., 2002</xref>). Such techniques are commonly used on top of frame differencing methods. Yet, they only work in simple scenarios, since close objects are often seen as one. Deep learning-based methods are more convenient to cover more complex scenarios such as dense traffic or crowded scenes, as they are able to recognize the object shapes.</p>
</sec>
<sec id="s3-3">
<title>3.3 Object Trackers</title>
<p>The role of an object tracker is to associate the object detections with the same identities over successive frames. We chose to run two object trackers on top of our two detector models, resulting in a combination of four algorithms on both decoded and residual frames.</p>
<p>IOU tracker (<xref ref-type="bibr" rid="B4">Bochinski et&#x20;al., 2017</xref>) is a very simple algorithm that relies on the assumption that detections of an object highly overlap on successive frames, resulting in a high Intersection Over Union (IOU) score. Although this is true for high refresh-rate video sequences, this is less the case for videos from traffic surveillance cameras, which run at lower refresh-rates and where vehicles travel a larger distance between two frames. To address this constraint, we chose the Kalman-IOU tracker (KIOU) instead, where a Kalman filter is added to better estimate object location and speed. This filter also lets you retain a history of the objects and re-identify them in case of missing detections. It was slightly adapted in order to work in an online tracking context, i.e.,&#x20;to work simultaneously with the detector.</p>
<p>The second tracker we chose is Simple Online and Realtime Tracking (SORT) (<xref ref-type="bibr" rid="B3">Bewley et&#x20;al., 2016</xref>). This algorithm also includes a Kalman filter to predict existing targets&#x2019; locations. It computes an assignment cost matrix between those predictions and the provided detections on the current frame using the IOU distance and then solves it optimally using the Hungarian algorithm.</p>
<p>Both trackers are localization-based trackers as they use only a fast statistical approach based on localization of bounding boxes. There also exists more complex trackers called feature-based trackers, such as DeepSORT (<xref ref-type="bibr" rid="B37">Wojke et&#x20;al., 2017</xref>), which also base their predictions on objects&#x2019; appearance information. However, detailed features such as vehicle models, brands, and colors cannot be distinguished in residual frames. Therefore, we limited our exploration to the former kind of trackers.</p>
</sec>
<sec id="s3-4">
<title>3.4 HOTA Evaluation Metric</title>
<p>In this section, we introduce the Higher Order Tracking Accuracy (HOTA) evaluation metric developed in detail by (<xref ref-type="bibr" rid="B19">Luiten et&#x20;al., 2020</xref>). This metric was used to assess the performance of our detector/tracker combinations on the multi-object tracking task. In their work, they allow to measure the performance of the two stages (the detection and the association) evenly in a single metric. They also show that the HOTA metric should be used instead of the MOTA metric (<xref ref-type="bibr" rid="B2">Bernardin and Stiefelhagen, 2008</xref>) because the latter is biased towards detection. Ground-truth annotations and predictions are matched thanks to the Hungarian algorithm, provided that their similarity score is above a threshold <italic>&#x3b1;</italic>.</p>
<p>The HOTA score is computed by means of sub-metrics that can also be used for deeper analysis. The matching is done at the detection level in each frame based on the similarity score. The matched pairs of detections are called the true positives (TP). Predictions that are not matched with a ground-truth detection are called false positives (FP). Likewise, ground-truth detections that are not matched with a prediction are called false negatives (FN). Then, the detection precision (DetPr), recall (DetRe), and accuracy (DetA) are obtained using <xref ref-type="disp-formula" rid="e1">Eqs 1</xref>&#x2013;<xref ref-type="disp-formula" rid="e3">3</xref> respectively. More specifically, the detection recall measures the performance in finding all the ground-truth detections while the detection precision evaluates how well the predictor does not produce extra detections.<disp-formula id="e1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">D</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">D</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">N</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(2)</label>
</disp-formula>
<disp-formula id="e3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">D</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">N</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>The tracks&#x2019; association can be computed for each matched detection. This is done by evaluating the alignment between the predicted detection&#x2019;s track and ground-truth detection&#x2019;s track. Then, the matching detections between the two tracks are called true positives (TPA). The remaining detections from the predicted track are the false positives (FPA) and the ones from the ground-truth track are the false negatives (FNA). When the best matching tracks are found, association precision (AssPr), recall (AssRe), and association (AssA) are computed using in <xref ref-type="disp-formula" rid="e4">Eqs 4</xref>&#x2013;<xref ref-type="disp-formula" rid="e6">6</xref> respectively. This time, the recall tells how well the predictor does not split the tracks of the objects whereas the precision assesses how it avoids merging the tracks of different objects.<disp-formula id="e4">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="[" close="}">
<mml:mrow>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mtext>&#x2009;AssRe&#x2009;</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">N</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m6">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>&#x2009;AssPr&#x2009;</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>In short, HOTA<sub>
<italic>&#x3b1;</italic>
</sub> combines the detection and the association accuracies. Each of them measures the overall quality of their own stage and can be broken down to obtain the recalls and the precisions. This is illustrated in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>. The final HOTA score is the average of the nineteen HOTA scores computed at each similarity score threshold (ranging from 0.05 to 0.95). In the case of bounding boxes, the similarity score (<italic>S</italic>) is the Intersection Over Union (IOU). Additionally, the localization accuracy (LocA) measures the overall spatial alignment between the predicted detections and the ground-truth annotations. HOTA<sub>
<italic>&#x3b1;</italic>
</sub>, HOTA and LocA can be calculated using <xref ref-type="disp-formula" rid="e7">Eqs 7</xref>&#x2013;<xref ref-type="disp-formula" rid="e9">9</xref> respectively.<disp-formula id="e7">
<mml:math id="m7">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mi>Det</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m8">
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo>&#x222b;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x2248;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>19</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mmultiscripts>
<mml:mrow>
<mml:mn>0.05</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mn>0.95</mml:mn>
</mml:mrow>
<mml:none/>
<mml:none/>
</mml:mmultiscripts>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(8)</label>
</disp-formula>
<disp-formula id="e9">
<mml:math id="m9">
<mml:mi mathvariant="normal">L</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo>&#x222b;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:munder>
<mml:mi mathvariant="script">S</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mi>d</mml:mi>
<mml:mi>&#x3b1;</mml:mi>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Hierarchical diagram of the HOTA sub-metrics illustrating the different error types (<xref ref-type="bibr" rid="B19">Luiten et&#x20;al., 2020</xref>).</p>
</caption>
<graphic xlink:href="frsip-01-765006-g004.tif"/>
</fig>
</sec>
<sec id="s3-5">
<title>3.5 Datasets</title>
<p>Three datasets were used to train and test our solutions. First, MIO-TCD Localization (<xref ref-type="bibr" rid="B20">Luo et&#x20;al., 2018</xref>) allowed us to train our detectors for decoded frames. This public dataset acquired 140,000 annotated images captured at different times of the day and periods of the year by 8,000 traffic cameras deployed over North America. We trained a YOLOv4 and a tiny YOLOv4 detectors on MIO-TCD Localization and obtained an mAP of 80.4 and 71.5%, respectively, on the test set. The former is the state-of-the-art detector reported on the MIO-TCD localization challenge. Note that eleven mobile object classes were annotated in this challenge, but in this work all predictions were grouped into a single mobile object class to match the residual frames&#x2019; annotations. This has also been done to ease the detection task, given that this is a first exploration into deep learning object detection using residual frames. Another reason is because an implementation to evaluate multi-class dataset with HOTA metrics is not yet available.</p>
<p>Second, we applied our compression algorithm on video sequences from AICity Challenge 2021 Track 1 (<xref ref-type="bibr" rid="B24">Naphade et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B23">Naphade et&#x20;al., 2018</xref>) to obtain the residual frames dataset. We then trained a YOLOv4 and a tiny YOLOv4 detector on this dataset. For this purpose, we manually annotated 14,000 residual frames from six different points of view to make it appropriate for training. Indeed, there are some visibility discrepancies between decoded and residual frames. The latter representation relies on movement in the observed scene, leading to stationary vehicles often not being visible. Using an Nvidia GTX 1080Ti, it took approximately 16&#xa0;hours to train YOLOv4 and only 2&#xa0;hours to train the tiny version on this adapted dataset.</p>
<p>Finally, we used AICity Challenge 2021 Track 3, also known as CityFlowV2 (<xref ref-type="bibr" rid="B24">Naphade et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B31">Tang et&#x20;al., 2019</xref>). We selected four full HD (1,920 &#xd7; 1,080 pixels) video sequences (recorded at 10 frames per second), resulting in a total of 10,000 frames to test the performance of the four detector/tracker combinations on both representations. CityFlowV2 provides full ground-truths with vehicle IDs. It should be noted that the challenge targets multi-camera tracking. Therefore, only objects that travel across at least two cameras were annotated. Also, vehicles whose bounding boxes were smaller than 1,000 square pixels (smaller than 0.05% of the native resolution) were not annotated. To keep a fair comparison, predictions smaller than this threshold were also removed. Otherwise, detectors would have been wrongly penalized, since they are capable of detecting further objects resulting in false positives. Therefore, the appropriate test sequences were chosen to take the aforementioned constraints into account.</p>
</sec>
<sec id="s3-6">
<title>3.6 Experiments</title>
<p>In this paper, we set up two experiments on two different video surveillance scenarios to show that an object tracker trained and tested on residual frames can be just as effective as an object tracker trained and tested on decoded frames. Both experiments followed the same setup. The two experiments were tested on four detector/tracker combinations for both image representations (residual and decoded frames): YOLOv4/KIOU, YOLOv4/SORT, tiny YOLOv4/KIOU and tiny YOLOv4/SORT.</p>
<p>The scenario for Experiment One is shown in <xref ref-type="fig" rid="F5">Figure&#x20;5A</xref>. Three cameras observed the same intersection between a double-lane two-way street and a single-lane two-way street. This scenario is indeed very complex, with overlapping numbers of vehicles, and can be used for various video surveillance tasks such as traffic light violation detection, vehicle counting, and traffic jam monitoring. The goal of Experiment One was to show that a network trained and tested on residual frames could compete with a network trained and tested on decoded frames even in a highly complex scenario. The scenario for Experiment Two is shown in <xref ref-type="fig" rid="F5">Figure&#x20;5B</xref>. One camera observed a double-lane two-way street with uninterrupted traffic flow. This scenario is mainly used for vehicle counting and wrong-way driving violation detection. The goal of Experiment Two was to highlight the benefits of using residual frames in these types of scenarios, as the constant traffic flow should allow the block matching algorithm to generate motion vectors constantly. This would lead to uninterrupted and more visible residual frames in return.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Visualization of the two video surveillance scenarios. Experiment One <bold>(A)</bold> shows a complex scenario where an intersection is observed by three distinct cameras. Experiment Two <bold>(B)</bold> shows a simple scenario where a two-way street is observed by a single camera.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4 Results</title>
<sec id="s4-1">
<title>4.1 Experiment One</title>
<p>In this scenario, we tested the combinations on the intersection depicted in <xref ref-type="fig" rid="F5">Figure&#x20;5A</xref>. The total footage of the three cameras amounted to 6,000 frames recorded at 10 frames per second. We calculated the HOTA metric and sub-metrics for each detector/tracker combination on both representations. The results for Experiment One are shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Results of Experiment One (complex scenario where three distinct cameras observe an intersection) for residual versus decoded frames for each detector/tracker combination evaluated with HOTA metric and sub-metrics. Bold values highlight the overall HOTA scores and the underlined values show the best average scores for each metric between the two representations.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Representation</th>
<th align="center">Detector/Tracker</th>
<th align="center">HOTA</th>
<th align="center">DetA</th>
<th align="center">DetRe</th>
<th align="center">DetPr</th>
<th align="center">AssA</th>
<th align="center">AssRe</th>
<th align="center">AssPr</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">Residual</td>
<td align="left">YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>36.28</bold>
</td>
<td align="char" char=".">34.04</td>
<td align="char" char=".">44.24</td>
<td align="char" char=".">47.83</td>
<td align="char" char=".">39.95</td>
<td align="char" char=".">51.11</td>
<td align="char" char=".">60.89</td>
</tr>
<tr>
<td align="left">YOLOv4/SORT</td>
<td align="char" char=".">
<bold>37.47</bold>
</td>
<td align="char" char=".">32.96</td>
<td align="char" char=".">42.37</td>
<td align="char" char=".">47.43</td>
<td align="char" char=".">44.79</td>
<td align="char" char=".">57.14</td>
<td align="char" char=".">61.30</td>
</tr>
<tr>
<td align="left">Tiny YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>34.90</bold>
</td>
<td align="char" char=".">32.67</td>
<td align="char" char=".">44.73</td>
<td align="char" char=".">45.31</td>
<td align="char" char=".">38.55</td>
<td align="char" char=".">50.64</td>
<td align="char" char=".">60.26</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Tiny YOLOv4/SORT</td>
<td align="char" char=".">
<bold>35.02</bold>
</td>
<td align="char" char=".">30.91</td>
<td align="char" char=".">41.61</td>
<td align="char" char=".">43.79</td>
<td align="char" char=".">41.55</td>
<td align="char" char=".">54.84</td>
<td align="char" char=".">57.88</td>
</tr>
<tr>
<td colspan="2" align="left">Average scores</td>
<td align="char" char=".">
<bold>35.92</bold>
</td>
<td align="char" char=".">32.64</td>
<td align="char" char=".">43.24</td>
<td align="char" char=".">
<underline>46.09</underline>
</td>
<td align="char" char=".">41.21</td>
<td align="char" char=".">53.43</td>
<td align="char" char=".">60.08</td>
</tr>
<tr>
<td rowspan="3" align="left">Decoded</td>
<td align="left">YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>42.63</bold>
</td>
<td align="char" char=".">35.97</td>
<td align="char" char=".">61.26</td>
<td align="char" char=".">38.53</td>
<td align="char" char=".">52.94</td>
<td align="char" char=".">67.74</td>
<td align="char" char=".">61.37</td>
</tr>
<tr>
<td align="left">YOLOv4/SORT</td>
<td align="char" char=".">
<bold>43.21</bold>
</td>
<td align="char" char=".">35.29</td>
<td align="char" char=".">59.70</td>
<td align="char" char=".">38.21</td>
<td align="char" char=".">55.46</td>
<td align="char" char=".">69.44</td>
<td align="char" char=".">61.89</td>
</tr>
<tr>
<td align="left">Tiny YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>40.35</bold>
</td>
<td align="char" char=".">34.00</td>
<td align="char" char=".">62.77</td>
<td align="char" char=".">36.22</td>
<td align="char" char=".">49.75</td>
<td align="char" char=".">63.93</td>
<td align="char" char=".">60.78</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Tiny YOLOv4/SORT</td>
<td align="char" char=".">
<bold>41.29</bold>
</td>
<td align="char" char=".">33.38</td>
<td align="char" char=".">61.06</td>
<td align="char" char=".">35.92</td>
<td align="char" char=".">52.95</td>
<td align="char" char=".">66.86</td>
<td align="char" char=".">60.49</td>
</tr>
<tr>
<td colspan="2" align="left">Average scores</td>
<td align="char" char=".">
<underline>
<bold>41.87</bold>
</underline>
</td>
<td align="char" char=".">
<underline>34.66</underline>
</td>
<td align="char" char=".">
<underline>61.20</underline>
</td>
<td align="char" char=".">37.22</td>
<td align="char" char=".">
<underline>52.77</underline>
</td>
<td align="char" char=".">
<underline>66.99</underline>
</td>
<td align="char" char=".">
<underline>61.13</underline>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For the HOTA metric, we obtained an average score of 35.92% when residual frames were the input versus an average score of 41.87% when decoded frames were the input. Concerning the DetA sub-metric, we observed an average score of 32.64% when residual frames were the input versus an average score of 34.66% when decoded frames were the input. As for the AssA sub-metric, we derived an average score of 41.21% when residual frames were the input versus an average score of 52.77% when decoded frames were the&#x20;input.</p>
<p>On the whole, the HOTA score is 6% better on average when decoded frames are used. This difference mainly comes from AssA (11.5% difference on average) and more specifically from the association recall (AssRe) that measures how objects&#x2019; tracks are split into multiple tracks. This can be explained simply by the traffic light stop lines. When vehicles stand still in the residual frames, they temporary disappear, given that residual frames depend on motion vectors generated by block matching. These vehicles are then assigned new IDs when they start moving again. On the other hand, the DetA sub-score is not much affected by residual frames (only 2% difference on average). Likewise, the detection recall (DetRe) is strongly affected by the stop lines since it measures to what extent all detections are found, but is counterbalanced by the detection precision (DetPr), which is higher for residual frames detectors thanks to the background suppression it provides. This results in fewer false positive detections. That being said, both kinds of detectors locate the objects properly in the space as shown by <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Results on residual frames <bold>(A)</bold> versus decoded frames <bold>(B)</bold> in Experiment One. White bounding boxes represent the ground truth annotations, colored boxes represent the detections together with their respective IDs and confidence&#x20;score.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g006.tif"/>
</fig>
</sec>
<sec id="s4-2">
<title>4.2 Experiment Two</title>
<p>In this scenario, we tested the combinations on the street depicted in <xref ref-type="fig" rid="F5">Figure&#x20;5B</xref>. The total footage of the camera amounted to 4,000 frames recorded at 10 frames per second. We calculated the HOTA metric and sub-metrics for each detector/tracker combination on both representations. The results for Experiment Two are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Results of Experiment Two (simple scenario where a single camera observes a two-way street) for residual versus decoded frames for each detector/tracker combination evaluated with HOTA metric and sub-metrics. Bold values highlight the overall HOTA scores and the underlined values show the best average scores for each metric between the two representations.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Representation</th>
<th align="center">Detector/Tracker</th>
<th align="center">HOTA</th>
<th align="center">DetA</th>
<th align="center">DetRe</th>
<th align="center">DetPr</th>
<th align="center">AssA</th>
<th align="center">AssRe</th>
<th align="center">AssPr</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">Residual</td>
<td align="left">YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>35.57</bold>
</td>
<td align="char" char=".">37.96</td>
<td align="char" char=".">51.11</td>
<td align="char" char=".">46.60</td>
<td align="char" char=".">33.60</td>
<td align="char" char=".">41.15</td>
<td align="char" char=".">54.92</td>
</tr>
<tr>
<td align="left">YOLOv4/SORT</td>
<td align="char" char=".">
<bold>38.97</bold>
</td>
<td align="char" char=".">37.87</td>
<td align="char" char=".">50.67</td>
<td align="char" char=".">46.86</td>
<td align="char" char=".">40.33</td>
<td align="char" char=".">50.25</td>
<td align="char" char=".">55.18</td>
</tr>
<tr>
<td align="left">Tiny YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>34.92</bold>
</td>
<td align="char" char=".">39.74</td>
<td align="char" char=".">53.60</td>
<td align="char" char=".">46.86</td>
<td align="char" char=".">30.87</td>
<td align="char" char=".">39.58</td>
<td align="char" char=".">51.55</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Tiny YOLOv4/SORT</td>
<td align="char" char=".">
<bold>40.07</bold>
</td>
<td align="char" char=".">40.24</td>
<td align="char" char=".">53.63</td>
<td align="char" char=".">47.75</td>
<td align="char" char=".">40.17</td>
<td align="char" char=".">50.41</td>
<td align="char" char=".">53.89</td>
</tr>
<tr>
<td colspan="2" align="left">Average scores</td>
<td align="char" char=".">
<underline>
<bold>37.38</bold>
</underline>
</td>
<td align="char" char=".">
<underline>38.96</underline>
</td>
<td align="char" char=".">
<underline>52.25</underline>
</td>
<td align="char" char=".">
<underline>47.02</underline>
</td>
<td align="char" char=".">
<underline>36.24</underline>
</td>
<td align="char" char=".">
<underline>45.35</underline>
</td>
<td align="char" char=".">
<underline>53.89</underline>
</td>
</tr>
<tr>
<td rowspan="3" align="left">Decoded</td>
<td align="left">YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>31.72</bold>
</td>
<td align="char" char=".">32.35</td>
<td align="char" char=".">45.06</td>
<td align="char" char=".">39.96</td>
<td align="char" char=".">31.38</td>
<td align="char" char=".">39.76</td>
<td align="char" char=".">48.41</td>
</tr>
<tr>
<td align="left">YOLOv4/SORT</td>
<td align="char" char=".">
<bold>33.93</bold>
</td>
<td align="char" char=".">31.92</td>
<td align="char" char=".">44.25</td>
<td align="char" char=".">39.93</td>
<td align="char" char=".">36.46</td>
<td align="char" char=".">46.51</td>
<td align="char" char=".">48.80</td>
</tr>
<tr>
<td align="left">Tiny YOLOv4/KIOU</td>
<td align="char" char=".">
<bold>31.13</bold>
</td>
<td align="char" char=".">27.85</td>
<td align="char" char=".">44.93</td>
<td align="char" char=".">32.85</td>
<td align="char" char=".">34.95</td>
<td align="char" char=".">41.76</td>
<td align="char" char=".">49.12</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Tiny YOLOv4/SORT</td>
<td align="char" char=".">
<bold>32.23</bold>
</td>
<td align="char" char=".">27.69</td>
<td align="char" char=".">44.09</td>
<td align="char" char=".">32.93</td>
<td align="char" char=".">37.69</td>
<td align="char" char=".">45.29</td>
<td align="char" char=".">48.78</td>
</tr>
<tr>
<td colspan="2" align="left">Average scores</td>
<td align="char" char=".">
<bold>32.25</bold>
</td>
<td align="char" char=".">29.95</td>
<td align="char" char=".">44.58</td>
<td align="char" char=".">36.42</td>
<td align="char" char=".">35.12</td>
<td align="char" char=".">43.33</td>
<td align="char" char=".">48.78</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For the HOTA metric, we obtained an average score of 37.38% when residual frames were the input versus an average score of 32.25% when decoded frames were the input. Concerning the DetA sub-metric, we observed an average score of 38.96% when residual frames were the input versus an average score of 29.95% when decoded frames ere the. As for the AssA sub-metric, we derived an average score of 36.24% when residual frames were the input versus an average score of 35.12% when decoded frames were the&#x20;input.</p>
<p>Overall, the HOTA score is 5% higher on average for the residual representation in Experiment Two. This time, the main difference comes from the DetA sub-metric, where a 9% difference on average can be noticed. This is explained by a higher detection recall (DetRe) due to the vehicles&#x2019; constant movements causing them to appear in the residual frames. Also, fewer false positive detection results in a better precision (DetPr), similar to Experiment One. Less significantly, the association accuracy (AssA) is quite similar for both representations, with less than 1% difference on average. Those scores are closer since tracks are not split anymore in the case of residual frames. Furthermore, the camera is closer to the ground, making the distant vehicles less distinguishable for the trackers. Consequently, all the association scores are a bit lower than in Experiment One. With everything considered, both detectors still perform generally well, as shown by <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Results on residual frames <bold>(A)</bold> versus decoded frames <bold>(B)</bold> in Experiment Two. White bounding boxes represent the ground truth annotations, colored boxes represent the detections together with their respective IDs and confidence&#x20;score.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g007.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5 Discussion</title>
<p>In this section, we first discuss the benefits and the drawbacks of using the two frame representations. Secondly, we assess both residual and decoded frames on a privacy-friendly model. Thirdly, we position our research with respect to the state-of-the-art. Last, we expose the limitations of our&#x20;work.</p>
<sec id="s5-1">
<title>5.1 Decoded Frames Versus Residual Frames</title>
<p>The results of the two experiments yield two different observations. The first observation is seen in Experiment One (complex scenario where three distinct cameras observe an intersection) where the decoded frame representation outperforms the residual frame representation. This is due to the continuously interrupted traffic flow by the intersection&#x2019;s traffic lights. Given that compression codecs rely on movement to generate residual frames, standstill objects do not appear in the image, as shown in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>. Yet, residual frames are not completely limited in the scenario. For instance, it is still possible to count vehicles entering or exiting the intersection for statistical purposes.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Results on residual frames <bold>(A)</bold> versus decoded frames <bold>(B)</bold> showing false negatives in residual frames when cars are stopped. White bounding boxes represent the ground truth annotations, colored boxes represent the detections together with their respective IDs and confidence&#x20;score.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g008.tif"/>
</fig>
<p>The second observation is seen in Experiment Two (simple scenario where a single camera observes a two-way street) where the residual frames representation outperforms the decoded frame representation. <xref ref-type="fig" rid="F9">Figure&#x20;9</xref> shows a benefit and a drawback of using residual frames. On one hand, residual frames face possible detection mergers due to the uniform color distribution. On the other hand, the image smoothing offered by residual frames allows to get rid of false positives, which are sometimes predicted by deep learning techniques because of confusing shapes or colors. <xref ref-type="fig" rid="F10">Figure&#x20;10</xref> shows a second benefit to the use of the residual frame representation. It makes it possible to deal with backgrounds that contain objects that the model is able to detect but are not of interest. While it is possible to use a mask to perform detection only in a region of interest, this is not possible in this case, where there is a full parking lot in the background.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Results on residual frames <bold>(A)</bold> versus decoded frames <bold>(B)</bold> showing false positives in the background and foreground of the decoded frame and detection mergers in the residual frame. White bounding boxes represent the ground truth annotations, colored boxes represent the detections together with their respective IDs and confidence&#x20;score.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g009.tif"/>
</fig>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Results on residual frames <bold>(A)</bold> versus decoded frames <bold>(B)</bold> showing a parking with irrelevant cars in the background. Colored boxes represent the detections together with their respective IDs and confidence&#x20;score.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g010.tif"/>
</fig>
<p>In summary, in the complex scenario, the drawback caused by the high frequency of continuously interrupted targets outweighs the benefits of residual frames&#x2019; background subtraction. However, in the simple scenario, the uninterrupted traffic flow limits the drawback of residual frames and emphasizes the benefits of its background subtraction. Consequently, for object tracking purposes, the choice of the image representation may depend on the evaluated scenario. Yet, for storage purposes, it is obvious that we would rather choose lightweight residual frames over heavy decoded frames. Also, for privacy purposes, the choice is not straightforward. It remains an open-ended question whether decoded or residual frames are more privacy-friendly. This will be discussed in the next section.</p>
</sec>
<sec id="s5-2">
<title>5.2&#x20;Privacy-Friendly Model</title>
<p>In this paper, we not only want to address the two needs of data compression and data analysis but also tackle the need for data privacy. It is, however, very difficult to define a clear evaluation metric when measuring data privacy. Also, every country has different thresholds for the amount of information that may be leaked from video surveillance footage. In the European Union, the General Data Protection Regulation (GDPR) ensures the individual&#x2019;s right to ask for any information held about them, including but not limited to CCTV footage<xref ref-type="fn" rid="fn3">
<sup>2</sup>
</xref>. It is extremely hard to guarantee with 100% accuracy that one image, regardless of the representation used, does not reveal any private information about individuals in the visual scene. In our situation, we have to compromise between effective tracking and protecting the private information of individuals in the field. We can observe that the residual frames representation used is a sort of information filter on the entire image achieved by removing the background and distorting the foreground.</p>
<p>However, all the previously presented traditional methods of privacy modeling look at the explicit features for identification, such as visual text or facial features, only; they do not include implicit features such as location, time, and actions observed in the scene (<xref ref-type="bibr" rid="B10">Dufaux and Ebrahimi, 2006</xref>; <xref ref-type="bibr" rid="B6">Carrillo et&#x20;al., 2008</xref>). To have a global privacy loss measurement, we not only need to consider all the features involved in the observed frames but should also have a non-binary evaluation metric to reflect this trade-off. To this end, a global privacy loss metric (&#x393;) has been put forward by <xref ref-type="bibr" rid="B27">Saini et&#x20;al. (2010)</xref>. The metric takes into consideration the four key information features that can be associated to detected objects: Who, What, When, and Where. The Who information features represent the explicit features associated with identity. The What, When and Where information features represent the implicit features that, if combined with contextual knowledge of the scene and accumulated over several frames, can represent identity with a certain level of certainty. All four key information features have scores ranging from 0 to 1, with 0 indicating no evidence of privacy loss and 1 indicating sufficient evidence of privacy loss resulting in identification. The logistic function modeling the privacy loss is shown in <xref ref-type="disp-formula" rid="e10">Eq. 10</xref>:<disp-formula id="e10">
<mml:math id="m10">
<mml:mi mathvariant="normal">&#x393;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x2a;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(10)</label>
</disp-formula>where <italic>&#x3b1;</italic> is the scaling coefficient and <italic>&#x3b2;</italic> is the translation coefficient. <italic>I</italic>
<sub>
<italic>Who</italic>
</sub> and <italic>I</italic>
<sub>
<italic>What</italic>,<italic>When</italic>,<italic>Where</italic>
</sub> are the privacy leakage due to the explicit and implicit features, respectively. It should be noted that <italic>I</italic>
<sub>
<italic>Who</italic>
</sub> carries the highest weight of the four key information features.</p>
<p>In the initial model of using residual frames proposed by this paper, we would opt for encrypting the entire encoded bit-stream except for the residual frames, which would be publicly available. If we wanted to improve the privacy of our system further, we would need to apply a simple clustering filter to the residual frames. The filter would simply group together every YxZ pixel cluster by replacing them with their median value. A sample of the filtered residual frames is shown in <xref ref-type="fig" rid="F11">Figure&#x20;11</xref> with Y &#x3d; Z &#x3d; 7. Thus, to apply this clustering filter, we would need to tweak the compression codec used. This means that on top of having the residual frame (needed for decoding) in the encoded bit-stream, we would have an additional filtered residual frame. In that case we would encrypt the residual frame alongside the rest of the stream and keep only the filtered residual frame public. However, this gain in privacy comes at a higher encoding cost for the compression codec as well as a performance drop for the deep learning object detectors and trackers. In short, there will always be a trade-off between the three needs of compression, analysis, and privacy.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Comparison of explicit features leakage in <bold>(A)</bold> decoded frames, <bold>(B)</bold> residual frames and <bold>(C)</bold> filtered residual frames.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g011.tif"/>
</fig>
<p>We put the model into practice by enumerating the different explicit and implicit features in the case of traffic video surveillance and assigned them scores ranging from very low to very high evidence of privacy loss for each of the three proposed representations. This is depicted in <xref ref-type="fig" rid="F12">Figure&#x20;12</xref>. Among the explicit features (<italic>I</italic>
<sub>
<italic>Who</italic>
</sub>), the evidence of privacy loss drops for colors, since residual representations highlight the prediction error, which is mapped onto a grayscale image. The brand and model of a vehicle is less recognizable compared with decoded frames, in particular for filtered residual frames. License plates, for their part, can directly leak the identity in the case of the decoded frames, but are less likely to be readable in residual frames and are scrambled in filtered residual frames. The same goes for dents and other damage, which can be considered as a unique feature on someone&#x2019;s vehicle. As for implicit features, <italic>I</italic>
<sub>
<italic>Where</italic>
</sub> can be decomposed mainly into text and background features. The former are texts featured on traffic signs, for instance, while the latter can be any building or monument capable of revealing the location. As previously shown, the background filtering present in the residual representations addresses this concern. <italic>I</italic>
<sub>
<italic>When</italic>
</sub> considered mainly time related information. For example, the weather and the sunlight can communicate information on the day and time of the scene. Residual representations are mostly agnostic of <italic>I</italic>
<sub>
<italic>When</italic>
</sub> features. Some exceptions could occur in the case of severe weather conditions. <italic>I</italic>
<sub>
<italic>What</italic>
</sub> can be split into simple and complex detection tasks. Simple tasks such as vehicle counting or wrong-way driving violations can be carried just as effectively with residual representations or decoded frames. Complex tasks such as detecting emergency vehicles are easier to achieve when dealing with decoded frames rather than both residual representations. Overall, we observe that the global privacy loss is better for residual frames than for decoded frames and can also be improved by applying clustering filters to the residual frames.</p>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>Radar chart for both explicit <bold>(A)</bold> and implicit <bold>(B)</bold> features highlighting the reduction in information leakage when using residual and filtered residual frames over decoded frames.</p>
</caption>
<graphic xlink:href="frsip-01-765006-g012.tif"/>
</fig>
</sec>
<sec id="s5-3">
<title>5.3 Research Positioning</title>
<p>The proposed work is a new approach to object tracking based exclusively on the analysis of compressed domain residual frames. We therefore opted to position our paper not only on its object tracking results but also by highlighting its other benefits by observing key similarities and differences with the previous works mentioned in <xref ref-type="sec" rid="s2">Section 2</xref>. A comparison can be made with respect to the proposed work on CNN training using the ROI extracted by classical background subtraction (<xref ref-type="bibr" rid="B14">Kim et&#x20;al., 2018</xref>). Similarly to their work, we also take advantage of the residual frame&#x2019;s background subtraction by-product to obtain the changing ROI to train our network. However, contrarily to their proposal, the residual frame&#x2019;s background subtraction by-product is auto-generated by the already existing video compression codec and therefore does not require any supplementary computational power. In the same scope, we find further similarities of our work with the VCM proposal by MPEG (<xref ref-type="bibr" rid="B9">Duan et&#x20;al., 2020</xref>) as the two works aim to facilitate feature extraction. Another comparison can be made with regards to research propositions that have integrated compressed domain motion vectors for object detection (<xref ref-type="bibr" rid="B32">Ujiie et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B16">Liu et&#x20;al., 2019</xref>). Even though we have utilized the compressed domain residual frames in our case, we still differ from their work as we propose to train and test our network exclusively on the residual frames. This will ensure that we do not rely on the original key frames (also called I frames) nor on the motion vectors for the detection and tracking process. We therefore are not only able to store our data in the compressed format but we also show that this alternative could potentially provide a privacy-friendly solution to deep learning-based object tracking. We also find similarities with the privacy-friendly video surveillance scrambler proposal that distorts the ROI in the frame to hide critical information (<xref ref-type="bibr" rid="B10">Dufaux and Ebrahimi, 2006</xref>). In our paper, we use the clustering filter proposed in <xref ref-type="sec" rid="s5-2">Section 5.2</xref> to scramble the residual frame-generated ROI. Given that the ROI are auto-generated by the residual frames and that the clustering filtered is a basic median value filter, our algorithm would only require minor additional computational power compared to having to extract the ROI with complex algorithms.</p>
</sec>
<sec id="s5-4">
<title>5.4 Limitations</title>
<p>As for all research work, we reached some limitations that were either external or based on decisions made within our team. Firstly, we decided to chose the same parameters for all video sequences for the detectors and trackers. We decided to set parameters that worked well for all sequences because optimizing parameters for each one would have been arbitrary and could lead to biased results. For example, depending on the scenario, one could adjust a detector to favor false positives over false negatives, such as in intrusion detection systems. Conversely, urban planners would rather balance false positives and false negatives to obtain correct estimations.</p>
<p>Concerning the generic compression algorithm, it can be optimized in one of the two directions: either gain storage space and limit the amount of information disclosed at a cost of lower detection and tracking performance, or lose storage space and increase the amount of information leakage to improve the detection and tracking performance. In this study, we worked with fixed compression parameters for all chosen training and test sequences. The parameters have been chosen to balance storage space and tracking performance while maximizing privacy protection.</p>
<p>Regarding our detectors, they were trained on two different datasets. Nevertheless, they all proved to generalize well on other video sequences. Even though the detectors for residual frames were trained on AICity footage (different from those used for testing), there is less need variety in the observed scene to obtain a generic model given the simple appearance of the moving objects in the residual representation.</p>
<p>An important external factor that impacted our results was the AICity annotations. The fact that vehicles have to travel across at least two cameras to be annotated results in missing ground truths for vehicles passing in front of a single camera. Moreover, to ensure full coverage of the vehicles, these ground truths were annotated larger than normal. Furthermore, only annotations larger than 2/3 of the visible vehicle bodies were kept (<xref ref-type="bibr" rid="B24">Naphade et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B31">Tang et&#x20;al., 2019</xref>). All these factors do not really impact our comparison, as they are common to both kinds of detectors. However, the annotation restrictions lowered the HOTA percentages for all tested sequences.</p>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion and Future Work</title>
<p>In this work, we put forward an object tracking method which adapts both video compression and video analysis while reducing the amount of private information leakage in the video stream. This research addresses the three major needs created by the large surge in video data consumption. This was done by setting up two experiments based on two different video surveillance scenarios following the same experimental setup. The two experiments were tested on four detector/tracker combinations for both image representations (residual and decoded frames): YOLOv4/KIOU, YOLOv4/SORT, tiny YOLOv4/KIOU and tiny YOLOv4/SORT. Using the HOTA evaluation metric, we showed that using inexpensive compressed domain residual frames as an image representation can be just as effective as using decoded frames for deep learning-based object tracking. This research is to be seen as a positive result to encourage the use of compressed domain representations in deep learning-based video analysis. It is also a first step towards providing a publicly available data format for deep learning-based traffic monitoring.</p>
<p>Several future work propositions would extend the validation of our results. Further testing on other deep learning-based object detectors such as EfficientDet or other YOLO-family models (PP-YOLO, scaled YOLOv4, or YOLOR) should be done to consolidate our hypothesis (<xref ref-type="bibr" rid="B30">Tan et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B18">Long et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B35">Wang C. et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B36">Wang et&#x20;al., 2021</xref>). Feature-based trackers such as DeepSORT, JDETracker or TPM could be trained on residual frames to see if they are capable of capturing more information, albeit at a higher cost (<xref ref-type="bibr" rid="B25">Peng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B40">Zhang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B37">Wojke et&#x20;al., 2017</xref>). Another interesting path to explore is the use of residual frames in night-time object tracking, as it is much more robust to illumination and color changes than decoded frames. Finally, further exploration into the privacy evaluation metrics could be investigated with the goal of further validating our claim to providing a privacy-friendly solution to video surveillance object tracking.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>KE mainly worked on the adaptive block matching algorithm, the privacy-friendly model and put forward the proposed experimental setup. JS set up the datasets, metrics, and detection and tracking algorithms as well as conducted the experiments. BM supervised the entire project and was key in setting up the state of the art for object tracking and the privacy-friendly model. All authors contributed to the analysis of the results and discussion as well as the writing process.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>JS is supported by the APTITUDE project (&#x23;1910045) funded by the Walloon Region Win2Wal program and ACIC S.&#x20;A.</p>
</sec>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of Interest</title>
<p>This study received funding from ACIC S.A. The funder was not involved in the study design, collection, analysis, interpretation of data, the writing of this article or the decision to submit it for publication. All authors declare no other competing interests.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We would first like to thank our colleagues at the Pixels and interactions LAB (<ext-link ext-link-type="uri" xlink:href="https://pilab.be">PiLAB</ext-link>) for their support. We would also like to thank the participants of the (<ext-link ext-link-type="uri" xlink:href="https://trail.ac">TRAIL&#x2019;20</ext-link>) workshop for their initial contributions to the project.</p>
</ack>
<fn-group>
<fn id="fn2">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://twiki.cern.ch/twiki/pub/HEPIX/TechwatchNetwork/HtwNetworkDocuments/white-paper-c11-741490.pdf">https://twiki.cern.ch/twiki/pub/HEPIX/TechwatchNetwork/HtwNetworkDocuments/white-paper-c11-741490.pdf</ext-link>
</p>
</fn>
<fn id="fn3">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://webaim.org/">https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32016R0679&#x0026;from=EN</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barjatya</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Block Matching Algorithms for Motion Estimation</article-title>. <source>Final Project Paper for Spring 2004 Digital Image Processing Course at the Utah State Univ.</source> <comment>Available at <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/profile/Sp-Immanuel/publication/50235332">https://www.researchgate.net/profile/Sp-Immanuel/publication/50235332"</ext-link>
</comment> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernardin</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Stiefelhagen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Evaluating Multiple Object Tracking Performance: The clear Mot Metrics</article-title>. <source>EURASIP J.&#x20;Image Video Process.</source> <volume>2008</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1155/2008/246309</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bewley</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ge</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ott</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ramos</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Upcroft</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Simple Online and Realtime Tracking</article-title>,&#x201d; in <conf-name>2016 IEEE International Conference on Image Processing</conf-name>, <conf-loc>Phoenix, Arizona, United States</conf-loc>, <conf-date>September 25&#x2013;28, 2016</conf-date> (<publisher-name>ICIP</publisher-name>), <fpage>3464</fpage>&#x2013;<lpage>3468</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2016.7533003</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bochinski</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Eiselein</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Sikora</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>High-speed Tracking-By-Detection without Using Image Information</article-title>,&#x201d; in <conf-name>2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</conf-name>, <conf-loc>Lecce, Italy</conf-loc>, <conf-date>August 29&#x2013;September 1, 2017</conf-date>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/avss.2017.8078516</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bochkovskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>H. M.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Yolov4: Optimal Speed and Accuracy of Object Detection</source>. <comment>CoRR abs/2004</comment>, <fpage>10934</fpage>. </citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Carrillo</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kalva</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Magliveras</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Compression Independent Object Encryption for Ensuring Privacy in Video Surveillance</article-title>,&#x201d; in <conf-name>2008 IEEE International Conference on Multimedia and Expo</conf-name>, <conf-loc>Hannover, Germany</conf-loc>, <conf-date>June 23&#x2013;26, 2008</conf-date>, <fpage>273</fpage>&#x2013;<lpage>276</lpage>. <pub-id pub-id-type="doi">10.1109/ICME.2008.4607424</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chien</surname>
<given-names>W.-J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Winken</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>R.-L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Motion Vector Coding and Block Merging in Versatile Video Coding Standard</article-title>. <source>IEEE Trans. Circuits Syst. Video Tech.</source> <volume>1</volume>, <fpage>3848</fpage>&#x2013;<lpage>3861</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2021.3101212</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics</article-title>. <source>IEEE Trans. Image Process.</source> <volume>29</volume>, <fpage>8680</fpage>&#x2013;<lpage>8695</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2020.3016485</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dufaux</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ebrahimi</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Scrambling for Video Surveillance with Privacy</article-title>,&#x201d; in <conf-name>2006 Conference on Computer Vision and Pattern Recognition Workshop (CVPRW&#x2019;06)</conf-name>, <conf-loc>New York, United States</conf-loc>, <conf-date>June 17&#x2013;22, 2006</conf-date>, <fpage>160</fpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2006.184</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>37</volume>, <fpage>1904</fpage>&#x2013;<lpage>1916</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2389824</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.-W.</given-names>
</name>
<name>
<surname>Hsu</surname>
<given-names>C.-W.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C.-Y.</given-names>
</name>
<name>
<surname>Chuang</surname>
<given-names>T.-D.</given-names>
</name>
<name>
<surname>Hsiang</surname>
<given-names>S.-T.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C.-C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A Vvc Proposal with Quaternary Tree Plus Binary-Ternary Tree Coding Block Structure and Advanced Coding Techniques</article-title>. <source>IEEE Trans. Circuits Syst. Video Tech.</source> <volume>30</volume>, <fpage>1311</fpage>&#x2013;<lpage>1325</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2019.2945048</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>Y.-M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Hybrid Framework Combining Background Subtraction and Deep Neural Networks for Rapid Person Detection</article-title>. <source>J.&#x20;Big Data</source> <volume>5</volume>, <fpage>1</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-018-0131-x</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>T.-Y.</given-names>
</name>
<name>
<surname>Maire</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Belongie</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hays</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Perona</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ramanan</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Microsoft Coco: Common Objects in Context</article-title>. <source>Computer Vis. &#x2013; ECCV.</source> <volume>8693</volume>, <fpage>740</fpage>&#x2013;<lpage>755</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-10602-1_48</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Real-time Online Multi-Object Tracking in Compressed Domain</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>76489</fpage>&#x2013;<lpage>76499</lpage>. <pub-id pub-id-type="doi">10.1109/access.2019.2921975</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Qin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Path Aggregation Network for Instance Segmentation</article-title>,&#x201d; in <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Salt Lake City, Utah, United States</conf-loc>, <conf-date>June 18&#x2013;22, 2018</conf-date>, <fpage>8759</fpage>&#x2013;<lpage>8768</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00913</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Long</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Dang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <source>PP-YOLO: An Effective and Efficient Implementation of Object Detector</source>. <comment>CoRR abs/2007</comment>, <fpage>12099</fpage>. </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luiten</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Aljo&#x161;a</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Dendorfer</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Torr</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Geiger</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Leal-Taix&#xe9;</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Hota: A Higher Order Metric for Evaluating Multi-Object Tracking</article-title>. <source>Int. J.&#x20;Comp. Vis.</source> <volume>129</volume>, <fpage>548</fpage>&#x2013;<lpage>578</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-020-01375-2</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Branchaud-Charron</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lemaire</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Konrad</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mishra</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Mio-tcd: A New Benchmark Dataset for Vehicle Classification and Localization</article-title>. <source>IEEE Trans. Image Process.</source> <volume>27</volume>, <fpage>5129</fpage>&#x2013;<lpage>5141</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2018.2848705</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Misra</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Mish: A Self Regularized Non-monotonic Neural Activation Function</source>. <comment>CoRR abs/1908</comment>, <fpage>08681</fpage>. </citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mukherjee</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bankoski</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Grange</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Koleszar</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wilkins</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). &#x201c;<article-title>The Latest Open-Source Video Codec Vp9 - an Overview and Preliminary Results</article-title>,&#x201d; in <conf-name>2013 Picture Coding Symposium</conf-name>, <conf-loc>San Jose, California</conf-loc>, <conf-date>December 8&#x2013;11, 2013</conf-date> (<publisher-name>PCS</publisher-name>), <fpage>390</fpage>&#x2013;<lpage>393</lpage>. <pub-id pub-id-type="doi">10.1109/PCS.2013.6737765</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Naphade</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>M.-C.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Anastasiu</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Jagarlamudi</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Chakraborty</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). &#x201c;<article-title>The 2018 Nvidia Ai City challenge</article-title>,&#x201d; in <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>, <conf-loc>Salt Lake City, Utah, United States</conf-loc>, <conf-date>June 18&#x2013;22, 2018</conf-date>, <fpage>53</fpage>&#x2013;<lpage>537</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2018.00015</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Naphade</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Anastasiu</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>M.-C.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). &#x201c;<article-title>The 5th Ai City challenge</article-title>,&#x201d; in <conf-name>The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops</conf-name>, <conf-loc>Virtual</conf-loc>, <conf-date>June 19&#x2013;25, 2021</conf-date>. <pub-id pub-id-type="doi">10.1109/cvprw53098.2021.00482</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>See</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wen</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Tpm: Multiple Object Tracking with Tracklet-Plane Matching</article-title>. <source>Pattern Recognition</source> <volume>107</volume>, <fpage>107480</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2020.107480</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Tracing Boundary Contours in a Binary Image</article-title>. <source>Image Vis. Comput.</source> <volume>20</volume>, <fpage>125</fpage>&#x2013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1016/s0262-8856(01)00091-9</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Saini</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Atrey</surname>
<given-names>P. K.</given-names>
</name>
<name>
<surname>Mehrotra</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Emmanuel</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kankanhalli</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Privacy Modeling for Video Data Publication</article-title>,&#x201d; in <conf-name>2010 IEEE International Conference on Multimedia and Expo</conf-name>, <conf-loc>Singapore</conf-loc>, <conf-date>July 19&#x2013;23, 2010</conf-date>, <fpage>60</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1109/ICME.2010.5583334</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sullivan</surname>
<given-names>G. J.</given-names>
</name>
<name>
<surname>Ohm</surname>
<given-names>J.-R.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>W.-J.</given-names>
</name>
<name>
<surname>Wiegand</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Overview of the High Efficiency Video Coding (Hevc) Standard</article-title>. <source>IEEE Trans. Circuits Syst. Video Tech.</source> <volume>22</volume>, <fpage>1649</fpage>&#x2013;<lpage>1668</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2012.2221191</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzuki</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Abe</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Topological Structural Analysis of Digitized Binary Images by Border Following</article-title>. <source>Comp. Vis. Graphics, Image Process.</source> <volume>30</volume>, <fpage>32</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/0734-189X(85)90016-7</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>Q. V.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Efficientdet: Scalable and Efficient Object Detection</article-title>,&#x201d; in <conf-name>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Virtual</conf-loc>, <conf-date>July 14&#x2013;19, 2020</conf-date>, <fpage>10778</fpage>&#x2013;<lpage>10787</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01079</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Naphade</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M.-Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Birchfield</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Cityflow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-identification</article-title>,&#x201d; in <conf-name>In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Long Beach, California, USA</conf-loc>, <conf-date>June 16&#x2013;20, 2019</conf-date>, <fpage>8789</fpage>&#x2013;<lpage>8798</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00900</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ujiie</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hiromoto</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sato</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Interpolation-based Object Detection Using Motion Vectors for Embedded Real-Time Tracking Systems</article-title>,&#x201d; in <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>, <conf-loc>Salt Lake City, Utah, USA</conf-loc>, <conf-date>June 18&#x2013;22, 2018</conf-date>. <pub-id pub-id-type="doi">10.1109/cvprw.2018.00104</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vermaut</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Deville</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Marichal</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Macq</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>A Distributed Adaptive Block Matching Algorithm: Dis-Abma</article-title>. <source>Signal. Processing: Image Commun.</source> <volume>16</volume>, <fpage>431</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1016/s0923-5965(00)00008-4</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>C.-Y.</given-names>
</name>
<name>
<surname>Mark Liao</surname>
<given-names>H.-Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y.-H.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.-Y.</given-names>
</name>
<name>
<surname>Hsieh</surname>
<given-names>J.-W.</given-names>
</name>
<name>
<surname>Yeh</surname>
<given-names>I.-H.</given-names>
</name>
</person-group> (<year>2020b</year>). &#x201c;<article-title>Cspnet: A New Backbone that Can Enhance Learning Capability of Cnn</article-title>,&#x201d; in <conf-name>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>, <conf-loc>Virtual</conf-loc>, <conf-date>June 14&#x2013;19, 2020</conf-date>, <fpage>1571</fpage>&#x2013;<lpage>1580</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW50498.2020.00203</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bochkovskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>H. M.</given-names>
</name>
</person-group> (<year>2020a</year>). <source>Scaled-yolov4: Scaling Cross Stage Partial Network</source>. <comment>CoRR abs/2011</comment>, <fpage>08036</fpage>. </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yeh</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>H. M.</given-names>
</name>
</person-group> (<year>2021</year>). <source>You Only Learn One Representation: Unified Network for Multiple tasks</source>. <comment>CoRR Abs/2105</comment>, <fpage>04206</fpage>. </citation>
</ref>
<ref id="B37">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wojke</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bewley</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Paulus</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Simple Online and Realtime Tracking with a Deep Association Metric</article-title>,&#x201d; in <conf-name>2017 IEEE International Conference on Image Processing (ICIP)</conf-name>, <conf-loc>Beijing, China</conf-loc>, <conf-date>September 17&#x2013;20, 2017</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>3645</fpage>&#x2013;<lpage>3649</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2017.8296962</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yun</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chun</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Oh</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Yoo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Choe</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Cutmix: Regularization Strategy to Train strong Classifiers with Localizable Features</article-title>,&#x201d; in <conf-name>2019 IEEE/CVF International Conference on Computer Vision</conf-name>, <conf-loc>Seoul</conf-loc>, <conf-date>October 27&#x2013;November 2, 2019</conf-date> (<publisher-name>ICCV</publisher-name>), <fpage>6022</fpage>&#x2013;<lpage>6031</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00612</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Chuang</surname>
<given-names>H. C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>History-based Motion Vector Prediction in Versatile Video Coding</article-title>,&#x201d; in <conf-name>2019 Data Compression Conference</conf-name>, <conf-loc>Snowbird, Utah, USA</conf-loc>, <conf-date>March 26&#x2013;29, 2019</conf-date> (<publisher-name>DCC</publisher-name>), <fpage>43</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1109/DCC.2019.00012</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <source>A Simple Baseline for Multi-Object tracking</source>. <comment>CoRR Abs/2004</comment>, <fpage>01888</fpage>. </citation>
</ref>
</ref-list>
</back>
</article>