<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Sustain. Cities</journal-id>
<journal-title>Frontiers in Sustainable Cities</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Sustain. Cities</abbrev-journal-title>
<issn pub-type="epub">2624-9634</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frsc.2023.1197434</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Sustainable Cities</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Unsupervised video anomaly detection in UAVs: a new approach based on learning and inference</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Gang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Shu</surname> <given-names>Lisheng</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Yang</surname> <given-names>Yuhui</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2269303/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Jin</surname> <given-names>Chen</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2260013/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Research Department of Aeronautics, Zhejiang Scientific Research Institute of Transport</institution>, <addr-line>Hangzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Telecommunications Engineering, Xidian University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>College of Civil Aviation, Nanjing University of Aeronautics and Astronautics</institution>, <addr-line>Nanjing</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Nan Gao, Tsinghua University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Shiyi Liu, Arizona State University, United States; Shusen Jing, University of California, Davis, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Chen Jin <email>jinchen&#x00040;fastmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>06</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>5</volume>
<elocation-id>1197434</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>03</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>05</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Liu, Shu, Yang and Jin.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Liu, Shu, Yang and Jin</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>In this paper, an innovative approach to detecting anomalous occurrences in video data without supervision is introduced, leveraging contextual data derived from visual characteristics and effectively addressing the semantic discrepancy that exists between visual information and the interpretation of atypical incidents. Our work incorporates Unmanned Aerial Vehicles (UAVs) to capture video data from a different perspective and to provide a unique set of visual features. Specifically, we put forward a technique for discerning context through scene comprehension, which entails the construction of a spatio-temporal contextual graph to represent various aspects of visual information. These aspects encompass the manifestation of objects, their interrelations within the spatio-temporal domain, and the categorization of the scenes captured by UAVs. To encode context information, we utilize Transformer with message passing for updating the graph&#x00027;s nodes and edges. Furthermore, we have designed a graph-oriented deep Variational Autoencoder (VAE) approach for unsupervised categorization of scenes, enabling the extraction of the spatio-temporal context graph across diverse settings. In conclusion, by utilizing contextual data, we ascertain anomaly scores at the frame-level to identify atypical occurrences. We assessed the efficacy of the suggested approach by employing it on a trio of intricate data collections, specifically, the UCF-Crime, Avenue, and ShanghaiTech datasets, which provided substantial evidence of the method&#x00027;s successful performance.</p></abstract>
<kwd-group>
<kwd>drone video anomaly detection</kwd>
<kwd>spatio-temporal graph</kwd>
<kwd>unsupervised learning</kwd>
<kwd>Unmanned Aerial Vehicles</kwd>
<kwd>Variational Autoencoder</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="5"/>
<equation-count count="13"/>
<ref-count count="59"/>
<page-count count="13"/>
<word-count count="8797"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Smart Technologies and Cities</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The detection of abnormal events in videos, including those captured by Unmanned Aerial Vehicles (UAVs), poses a formidable obstacle as a result of the extensive spectrum of occurrences, coupled with the restricted accessibility of learning resources, and the contextualized definition of abnormal events (Li et al., <xref ref-type="bibr" rid="B19">2013</xref>; Ionescu et al., <xref ref-type="bibr" rid="B13">2019</xref>; Song et al., <xref ref-type="bibr" rid="B41">2019</xref>). UAVs, also known as drones, have seen a significant rise in popularity and usage in recent years due to their cost-effectiveness, versatility, and ability to access hard-to-reach areas. These advanced aerial systems have been increasingly utilized for various applications, such as surveillance in military and civilian contexts, search and rescue operations in disaster-stricken areas, environmental monitoring to track changes in ecosystems, infrastructure inspection, and even in the entertainment industry for aerial photography and filming.</p>
<p>As a result, detecting abnormal events in UAV-captured videos has become increasingly important for ensuring safety and security. Abnormal events in this context can refer to a wide range of occurrences, from intrusions and suspicious activities in surveillance scenarios to detecting signs of natural disasters or accidents in search and rescue operations. The challenge lies in the fact that these events are often context-dependent and can vary greatly in appearance, making it difficult for traditional computer vision algorithms to detect and classify them effectively.</p>
<p>In order to confront this particular issue related to Unmanned Aerial Vehicles (UAVs) or drones, a multitude of pre-existing approaches have been put forth to address challenges such as object detection, tracking, and anomaly recognition. These techniques endeavor to acquire customary spatial and temporal configurations pertaining to appearance and movement of UAVs in various environments. Consequently, they facilitate the identification of irregular occurrences, such as unauthorized UAVs entering restricted areas or deviating from their designated flight paths, by differentiating them from the established normative patterns (Feng et al., <xref ref-type="bibr" rid="B8">2016</xref>; Xu et al., <xref ref-type="bibr" rid="B51">2017a</xref>).</p>
<p>Typically, visual features are extracted from either an entire image (Hasan et al., <xref ref-type="bibr" rid="B10">2016</xref>; Chong and Tay, <xref ref-type="bibr" rid="B5">2017</xref>) or a specific zone of focal concern (Ionescu et al., <xref ref-type="bibr" rid="B13">2019</xref>) to acquire a comprehensive understanding of the inherent spatial and temporal configurations in nature. For instance, these techniques might analyze the shape, size, color, and texture of the UAVs in the imagery data. Additionally, the extracted features can be used to classify the UAVs into different categories, such as fixed-wing, rotary-wing, or hybrid designs, as well as to determine their speed, altitude, and flight patterns.</p>
<p>Investigations within the realm of psychological science have established that individuals possess the capability to accurately identify entities and environments through the employment of contextual visual data (Bar, <xref ref-type="bibr" rid="B2">2004</xref>; Tang et al., <xref ref-type="bibr" rid="B45">2020</xref>). Furthermore, context information has been shown to be beneficial for various computer vision tasks (Ionescu et al., <xref ref-type="bibr" rid="B14">2017</xref>; Hasan et al., <xref ref-type="bibr" rid="B11">2019</xref>; Sun et al., <xref ref-type="bibr" rid="B44">2019</xref>; Kang et al., <xref ref-type="bibr" rid="B15">2022</xref>), including those involving Unmanned Aerial Vehicles (UAVs) or drones. UAVs, which have rapidly advanced in recent years, offer a versatile and efficient solution for remote sensing, surveillance, and data collection in various domains, such as agriculture, disaster response, and law enforcement.</p>
<p>Hence, it is crucial to extract extensive contextual data that surpasses the scope of image-based and object-based characteristics for the precise detection of atypical occurrences within a video, as the visual context serves a pivotal function in ascertaining the existence of such anomalies. In the case of UAV-based surveillance systems, context information can enhance the system&#x00027;s ability to recognize and respond to unusual events or potential threats, which in turn leads to improved safety and security. Taking pedestrian behavior as an example, as depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>, strolling along a pedestrian pathway is perceived as a customary occurrence, whereas ambulating on a busy highway constitutes an atypical incident. Ignoring context information related to sidewalks and highways could lead to false detection in UAV-based surveillance systems.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Unmanned Aerial Vehicle (UAV) perspective for abnormal event detection. The <bold>(Left)</bold> subplot illustrates the same entity(the pedestrian) in different scenes (highway and sidewalk) generating an abnormal event and a common event, respectively. The <bold>(Right)</bold> subplot depicts the abnormal and normal events caused by different objects (vehicles and the pedestrian) in the same environment(the park on a university campus).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsc-05-1197434-g0001.tif"/>
</fig>
<p>In the realm of computer vision, context-based abnormal event detection has consistently garnered considerable scholarly attention, including UAV-captured videos. As the use of UAVs, commonly known as drones, has become more widespread in various industries such as agriculture, surveillance, disaster management, and aerial photography, there is an increased demand for reliable and accurate automated analysis of the captured data. Previous efforts in this area manually predefined the collections of context based on human experiences, such as relationship context and scene context. However, the correctness and completeness of these collections are difficult to guarantee, and it is impossible to consider all possible context information. To address these issues, we propose an automatic methodology of inferential reasoning within contextual framework that mines elevated contextual insights derived from the fundamental visual characteristics of information sources. In our approach, a solitary frame is employed to generate a spatial context graph, which facilitates the understanding of object characteristics and their spatial associations. Subsequently, a structural recurrent neural network incorporates these graphs to establish a spatio-temporal context graph. To reason about context, an iterative process involving the mean-field technique is utilized to modify the conditions of nodes and edges within the spatio-temporal graph, ultimately enabling the extraction of semantic context from visual attributes.</p>
<p>UAVs, also known as drones, have gained significant traction in various applications, such as monitoring system, and search and rescue operations. They are equipped with sophisticated sensors and cameras to capture high-resolution aerial imagery and video footage. Analyzing this data for the detection of abnormal events is crucial to ensuring the safety and efficiency of UAV operations. To account for the rarity of abnormal events, we present an unsupervised approach for inferring spatio-temporal context graphs through the implementation of a scene segmentation technique. By employing a deep Variational Autoencoder model grounded in graph theory, we effectively partition scenes into distinct clusters. Subsequently, event categorization into normal and anomalous occurrences is performed based on their respective cluster affiliations. The proposed methodology is rigorously evaluated using the UCF-Crime, Avenue, and ShanghaiTech datasets, including UAV-captured videos. Our approach substantially surpasses the performance of leading unsupervised techniques while yielding a noticeable enhancement in comparison to established supervised methodologies.</p>
<p>Although we mentioned earlier that visual contextual information is crucial for comprehensive object and scene recognition, it is especially beneficial for various computer vision tasks, including those involving UAVs. However, manually defining context sets is limited in the case of diverse, changing, and unpredictable context-related events (Pang et al., <xref ref-type="bibr" rid="B28">2020</xref>). Therefore, our method automatically learns contextual information from data instead of manually pre-defining contextual content. By performing a contextual reasoning approach, we establish a connection between the visual context and the interpretation of deviant occurrences, thus overcoming the semantic disparity that exists between the two. The methodology we propose is versatile, and relevant to an extensive variety of applications in which contextual factors are crucial, such as monitoring through Unmanned Aerial Vehicle systems.</p>
<p>To summarize, our work offers the subsequent contributions.</p>
<list list-type="bullet">
<list-item><p>An innovative method that utilizes scene-aware context reasoning: Our study presents a novel methodology for identifying anomalous incidents in videos, including those captured by UAVs, by leveraging scene-aware context reasoning, which helps bridge the disparity in meaning between the visual environment and anomalous occurrences. This is rare in methods of the field.</p></list-item>
<list-item><p>Development of a contextual graph for spatio-temporal data: We construct a spatio-temporal context graph that encodes and reasons about context information, thereby enhancing the precision of identifying anomalous events in various scenarios, including UAV-captured videos.</p></list-item>
<list-item><p>Introduction of deep Variational Autoencoder architecture that incorporates graph-based techniques for unsupervised visual environment clustering: Our method incorporates a graph-based deep Variational Autoencoders for unsupervised scenario clustering, which enables the identification of different scene types and the accurate detection of aberrations that are obscure and contextual in nature.</p></list-item>
<list-item><p>Enhanced discrimination between normal and those deemed abnormal events: By leveraging scene clustering, our approach can better discriminate between normal and abnormal events in various scenes, leading to more accurate detections, particularly in UAV-based surveillance systems.</p></list-item>
<list-item><p>Significant improvement in unsupervised abnormal event detection accuracy: Our proposed method demonstrates a substantial increase in the accuracy of unsupervised abnormal event detection when evaluated against cutting-edge approaches, including those applied to UAV-captured videos.</p></list-item>
</list>
</sec>
<sec id="s2">
<title>2. Related work</title>
<p>Over the last several years, a considerable number of scholars (Mehran et al., <xref ref-type="bibr" rid="B26">2009</xref>; Li et al., <xref ref-type="bibr" rid="B19">2013</xref>; Luo et al., <xref ref-type="bibr" rid="B23">2017</xref>; Ribeiro et al., <xref ref-type="bibr" rid="B33">2018</xref>; Feng et al., <xref ref-type="bibr" rid="B7">2021</xref>; Georgescu et al., <xref ref-type="bibr" rid="B9">2021</xref>) have conducted research on video anomaly recognition. Generally, the studies can be partitioned into three stages according to the specific technology used, which are the traditional machine learning (Mahadevan et al., <xref ref-type="bibr" rid="B24">2010</xref>; Anti&#x00107; and Ommer, <xref ref-type="bibr" rid="B1">2011</xref>; Li et al., <xref ref-type="bibr" rid="B19">2013</xref>; Lu et al., <xref ref-type="bibr" rid="B22">2013</xref>; Cheng et al., <xref ref-type="bibr" rid="B3">2015</xref>; Hasan et al., <xref ref-type="bibr" rid="B10">2016</xref>), the hybrid stage combining machine learning and deep learning (Hinami et al., <xref ref-type="bibr" rid="B12">2017</xref>; Luo et al., <xref ref-type="bibr" rid="B23">2017</xref>; Smeureanu et al., <xref ref-type="bibr" rid="B40">2017</xref>; Ravanbakhsh et al., <xref ref-type="bibr" rid="B31">2018</xref>; Sabokrou et al., <xref ref-type="bibr" rid="B36">2018a</xref>), and the deep learning stage (Sabokrou et al., <xref ref-type="bibr" rid="B34">2015</xref>, <xref ref-type="bibr" rid="B35">2017</xref>, <xref ref-type="bibr" rid="B37">2018b</xref>; Xu et al., <xref ref-type="bibr" rid="B50">2015</xref>; Chong and Tay, <xref ref-type="bibr" rid="B5">2017</xref>; Ionescu et al., <xref ref-type="bibr" rid="B13">2019</xref>).</p>
<p>The initial video anomaly detection studies mainly used manual features to construct feature spaces. These studies used traditional machine learning methods, such as basic methods to determine whether events obey the normal state distribution (Saligrama et al., <xref ref-type="bibr" rid="B38">2010</xref>), Gaussian mixture models (Kratz and Nishino, <xref ref-type="bibr" rid="B16">2009</xref>), and Markov models to infer anomalous features (Tipping and Bishop, <xref ref-type="bibr" rid="B46">1999</xref>; Leyva et al., <xref ref-type="bibr" rid="B18">2017</xref>), and sparse learning methods (Luo et al., <xref ref-type="bibr" rid="B23">2017</xref>) that are more effective than the former two (Lu et al., <xref ref-type="bibr" rid="B22">2013</xref>). However, these traditional machine learning methods have a certain degree of dependence on the selection of features and are often adapted to particular scenarios.</p>
<p>Subsequently, with the emergence of deep learning methods, artificial features are replaced by deep features, which can effectively monitor and analyze video semantic concepts. With the properties of automatically learning and extracting video features according to the environment, deep learning methods have led to a series of studies on unsupervised learning methods (Xu et al., <xref ref-type="bibr" rid="B50">2015</xref>; Hasan et al., <xref ref-type="bibr" rid="B10">2016</xref>; Luo et al., <xref ref-type="bibr" rid="B23">2017</xref>; Sabokrou et al., <xref ref-type="bibr" rid="B35">2017</xref>; Liu et al., <xref ref-type="bibr" rid="B20">2018</xref>; Wang et al., <xref ref-type="bibr" rid="B49">2018</xref>; Ye et al., <xref ref-type="bibr" rid="B54">2019</xref>). These works reconstruct the video so that the model gets a stronger response to anomalous frames during testing. Although such studies overcome the problem of feature dependency, they are only applicable to video types with few anomalous patterns and short time series, having the limitation of low generalization ability. In contrast to these approaches, Zhao et al. (<xref ref-type="bibr" rid="B56">2017</xref>) started to consider the use of video local spatio-temporal information for video reconstruction by 3D convolutional autoencoder, yet their study still has weak generalization ability for anomalous event diversity. Besides, due to redundancy on consecutive frames (Zhou et al., <xref ref-type="bibr" rid="B58">2018</xref>), using 3D convolutional kernels (Tran et al., <xref ref-type="bibr" rid="B47">2015</xref>) to extract features in dense RGB frame sequences can sometimes be computationally more expensive.</p>
<p>In addition, along with the objective social needs such as diversification and complexity of video anomaly detection, researchers have also been prompted to focus on the full and reasonable utilization of video multidimensional information. Sultani et al. (<xref ref-type="bibr" rid="B42">2018</xref>) adopted multiple instance learning methods to achieve better abnormality detection. This work is one of the early studies to provide new solutions for weakly supervised video abnormality detection. On the contrary, Zhu and Newsam (<xref ref-type="bibr" rid="B59">2019</xref>) focused on the influence of dynamic information on &#x0201C;anomalies&#x0201D; and introduced the attention mechanism to highlight the response of anomalous video feature segments. But the potential relationship between instances is still not exploited. Nevertheless, the overall unreasonable assumption of independent identical distribution between instances remains yet. Furthermore Zhang et al. (<xref ref-type="bibr" rid="B55">2019</xref>) analyzed the contrast between affirmative and opposing states of instances in multiple instance learning, proposed the concept of quantification of intra-packet loss, and used temporal convolutional network (TCN) for temporal information correlation. Although this work mostly achieved anomalous differentiation between normal and abnormal, it was not sensitive to other more neutral video segments within and the differentiation was not significant. In order to clean away label noise, enhancing sensitivity to neutral segments, Zhong et al. (<xref ref-type="bibr" rid="B57">2019</xref>) adopted a hybrid approach of weakly supervised and fully supervised to handle the detection task. They used graph convolution to transfer information to denoise the normal clips in the abnormal video, obtaining pseudo-labels to train 3D convolutional network (C3D) (Tran et al., <xref ref-type="bibr" rid="B47">2015</xref>) for abnormality recognition. With high training complexity in the denoising process, it is very likely that abnormal &#x0201C;miscleaning&#x0201D; will occur. The outcome of this phenomenon is a reduction in the precision of detecting and locating abnormal events.</p>
<p>Among the approaches above, reconstruction and prediction is the current mainstream video anomaly event detection method. Although it settled the problems of strong feature dependency and insufficient utilization of video multidimensional information in previous, the present investigation into the detection of anomalous video activity is faced with two major challenges. The first challenge is the considerable complexity inherent in the training process of the employed methods. The second challenge pertains to the burden of differentiating noise from redundant information in the features. Notably, some interesting studies using relational inference modules to enhance the efficacy of video anomalous event detection (Choi et al., <xref ref-type="bibr" rid="B4">2012</xref>; Park et al., <xref ref-type="bibr" rid="B29">2012</xref>; Leach et al., <xref ref-type="bibr" rid="B17">2014</xref>) have been proposed recently. Most of them have detected deviating anomalous events by defining the rule set of context for inference with the help of temporal relationships between video frames. While such methods alleviate the ambiguity of anomalous event definitions in redundant information, they in turn increase the scene-dependence of anomalous event definitions. Inspired by them, we adopt the unsupervised idea of scene-aware inference, and after encoding the contextual information into a spatio-temporal relationship graph, we attempt attention mechanisms to perform the next inference step and anomaly detection by updating the state in the graph. Differently, our method can achieve automatic mining of high-level features directly related to anomalous events based on the underlying visual information. That means it can be applied to detect anomalous occurrences in diverse settings while bridging the huge gap between the underlying data and the concepts of anomalous events.</p>
</sec>
<sec id="s3">
<title>3. Methodologies</title>
<p>During our study, we have evolved a spatio-temporal graph-based method for context-aware anomaly event inference in videos. As illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref>, we modeled graphical representation of spatial relationships using every individual frame of the video as an object. In this spatial graph model, we encapsulated the manifestation of entities and their spatial interconnections within each frame. The spatial graph was then input into a transformer to learn dynamic features of each object in the temporal dimension, ultimately developing the spatio-temporal graph model. Inspired by Shao et al. (<xref ref-type="bibr" rid="B39">2016</xref>), we employed unsupervised clustering to classify scenes and deduce the spatio-temporal scene graph model. Finally, we detected anomalous events utilizing the spatio-temporal scene graph model and the context features acquired through scene clustering.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Overview of our proposed framework. The framework is proposed that involves the use of a spatial scene graph (SSG) to analyze each frame <italic>I</italic> based on object bounding boxes. The SSG nodes are denoted by <italic>i</italic>, <italic>j</italic>, and <italic>k</italic>, with timestamps <italic>t</italic><sub>0</sub>, <italic>t</italic><sub>1</sub>, and <italic>t</italic><sub>2</sub>. A spatio-temporal scene graph (STSG) is then constructed by incorporating the temporal modeling of the SSG using a Transformer. Moreover, scene clustering is concurrently performed on the SSG to differentiate between different scene types. Finally, a reasoning model is established on top of the STSG to detect abnormal events across different scenes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsc-05-1197434-g0002.tif"/>
</fig>
<sec>
<title>3.1. Spatio-temporal scene graph</title>
<p>We represent video feature information as a Spatio-Temporal Scene Graph (STSG), where nodes encode object appearance, spatial graph model edges depict object relationships, and temporal edges model object dynamic features. This encoding scheme allows us to infer more semantic information about objects (nodes), object spatio-temporal relationships (edges), and event occurrence scenarios (entire graph model) compared to existing research. Consequently, we can detect single-point anomalies, relationship anomalies, and group anomalies. We achieve the goal of detecting various scenario anomalies in videos by constructing and inferring Spatio-Temporal Scene Graphs, implementing the inference through iterative updates of the node and edge states in the Spatio-Temporal Graph Model.</p>
<sec>
<title>3.1.1. Formulation</title>
<p>Assuming an input of a visual medium in the form of a video consisting of a total of <italic>N</italic> frames, denoted as <italic>V</italic> &#x0003D; [<italic>I</italic><sup>1</sup>, <italic>I</italic><sup>2</sup>, &#x022EF;&#x02009;, <italic>I</italic><sup><italic>N</italic></sup>], where <italic>I</italic> is a certain frame of the video. The Region Proposal Network (RPN) (Ren et al., <xref ref-type="bibr" rid="B32">2015</xref>) is employed to generate object bounding boxes for individual frames. The top-<italic>K</italic> bounding boxes in the <italic>n</italic>-th frame are selected as <italic>B</italic><sup><italic>n</italic></sup>, which includes the entire frame as an additional bounding box. To construct a spatial scene graph (SSG) for each frame, we use the image containing <italic>k</italic> enclosure boxes. In the SSG, each node <italic>v</italic><sub><italic>i</italic></sub> represents an object, while the edge <italic>e</italic><sub><italic>i,j</italic></sub> represents the association among the objects. To enable the inference function through iterative graph updates, we assign a &#x0201C;normal&#x0201D; or &#x0201C;abnormal&#x0201D; label to each node and edge of the SSG.</p>
<p>In the context of the <italic>n</italic>-th frame, the designation of the <italic>i</italic>-th object is denoted by the label <inline-formula><mml:math id="M1"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, while the label <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is assigned to the relationship between the <italic>i</italic>-th and <italic>j</italic>-th objects. We utilize binary categorization, whereby the label &#x02018;0&#x00027; denotes normal objects or inter-object relationships, while &#x02018;1&#x00027; indicates one-field objects or inter-object relations. To establish a comprehensive definition of anomalous labels, we define the set encompassing all such labels: <inline-formula><mml:math id="M3"><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. And then SSG model can be formalized as <inline-formula><mml:math id="M4"><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02223;</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Next, we use a Transformer to incorporate temporal information of multiple SSGs to generate the STSG. In the <italic>n</italic>-th frame, the node denoted as <italic>v</italic><sub><italic>i</italic></sub> is linked solely to its corresponding node <italic>v</italic><sub><italic>i</italic></sub> in the <italic>n</italic> &#x0002B; 1-th frame by means of the temporal edge denoted as <italic>e</italic><sub><italic>i,i</italic></sub>, which has a corresponding relation label denoted as <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Inspired by (Sun et al., <xref ref-type="bibr" rid="B43">2020</xref>), the conclusive probability distribution is determined with</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M8"><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> denotes collection of total exception tags of the input.</p>
</sec>
<sec>
<title>3.1.2. Inferencing on graphs</title>
<p>As previously mentioned, we perform contextual semantic inference in videos by generating states of nodes and edges in the model of graph. In this work, we adopt the mean-field graph model inference method (Xu et al., <xref ref-type="bibr" rid="B52">2017b</xref>; Qi et al., <xref ref-type="bibr" rid="B30">2019</xref>). The probability function <italic>P</italic>(<bold>y</bold>&#x02223;&#x000B7;) will be approximated as <italic>Q</italic>(<bold>y</bold>&#x02223;&#x000B7;). In particular, we utilized the states <inline-formula><mml:math id="M9"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to denote the present state of node <italic>i</italic> and the edge between node <italic>i,j</italic> in frame <italic>n</italic>, individually. <inline-formula><mml:math id="M10"><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mo>&#x000B7;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> of node <italic>i</italic> depends on the state <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of all nodes and edges.</p>
<p>The approximation <inline-formula><mml:math id="M12"><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mo>&#x000B7;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> depends only on the current state, i.e., <inline-formula><mml:math id="M13"><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mo>&#x000B7;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. The probability distribution approximation for the edge is also implemented using this idea.</p>
<p>Drawing inspiration from Vaswani et al. (<xref ref-type="bibr" rid="B48">2017</xref>); Xu et al. (<xref ref-type="bibr" rid="B53">2022</xref>), we leverage Transformers to calculate the state of <italic>Q</italic>. Depicting in <xref ref-type="fig" rid="F3">Figure 3</xref>, STSG model we proposed employs the nodeTransformer and edgeTransformer to model the hidden states of nodes and spatio-temporal edges, respectively. Both the nodeTransformer and edgeTransformer serve to update node and edge states, allowing for the inference of context semantics from visual features. The nodeTransformer computation is expressed as follows:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M14"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>K</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>V</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>M</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:mtext>softmax</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mfrac><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>K</mml:mi></mml:mstyle><mml:mo>&#x022A4;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo stretchy='false'>)</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>V</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mtext>LN</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mtext>MLP</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>M</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>E</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the node embedding matrix, <inline-formula><mml:math id="M16"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the hidden state matrix of nodeTransformer at time <italic>t</italic>, <inline-formula><mml:math id="M17"><mml:mstyle mathvariant="bold"><mml:mtext>Q</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>K</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>V</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are the query, key, and value matrices, respectively, <inline-formula><mml:math id="M18"><mml:mstyle mathvariant="bold"><mml:mtext>M</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the output of the self-attention mechanism, LN denotes layer normalization, and MLP represents a multi-layer perceptron that applies non-linear transformations to the concatenated input features. Here, <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters, <italic>d</italic><sub><italic>k</italic></sub> and <italic>d</italic><sub><italic>v</italic></sub> denote the dimensions of the key and value vectors, respectively, and MLP consists of two fully connected layers with ReLU activation and a skip connection.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Schematic representation of the STSG inference process. The delineation of object boundaries is accomplished through the utilization of a region proposal network (RPN), which processes raw frames. Subsequently, node spatial edge and temporal edge attributes are derived through specialized feature extraction modules. These features serve as the initial input for nodeTransformer and edgeTransformer in reasoning. The message-passing technique is employed to renew Transformers. After computation, nodeTransformer and edgeTransformer yield the predicted anomaly labels <italic>y</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsc-05-1197434-g0003.tif"/>
</fig>
<p>Similarly, the computation of edgeTransformer is formulated as</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M21"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>4</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>K</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>4</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>V</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>4</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>M</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:mtext>softmax</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mfrac><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>K</mml:mi></mml:mstyle><mml:mo>&#x022A4;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo stretchy='false'>)</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>V</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mtext>LN</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mtext>MLP</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>M</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>E</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the edge embedding matrix, <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the hidden state matrix of edgeTransformer at time <italic>t</italic>, and <bold>Q</bold>, <bold>K</bold>, <bold>V</bold>, <bold>M</bold> are defined similarly as in the node self-attention mechanism. Here, <inline-formula><mml:math id="M24"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters, and MLP consists of two fully connected layers with ReLU activation and a skip connection. In this formulation, the edge embeddings are fed into the edge self-attention mechanism, and the resulting representation is used to update the hidden states of edges through a multi-layer perceptron.</p>
<p>To enhance the inference process&#x00027;s efficiency, we utilize message-passing techniques during computation. The message-passing matrix <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> at time <italic>t</italic>&#x0002B;1 is computed using a simple linear transformation:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mtext>ReLU</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>node</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the message passing matrix at time <italic>t</italic>&#x0002B;1, <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>msg</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters, and ReLU denotes the rectified linear unit activation function.</p>
<p>In the aforementioned formulation, node and edge representations are integrated using a linear transformation. This can be perceived as a simplified version of message passing, where messages are computed as linear combinations of the hidden states of nodes and edges. Such a formulation can be more efficient than conventional message passing, particularly for large graphs with dense connections. The resulting model can encapsulate both spatial and temporal dependencies within the graph and learn to deduce context semantics from visual features.</p>
</sec>
</sec>
<sec>
<title>3.2. Scenario clustering</title>
<p>The identification of scene types is essential for comprehending abnormal events since it is typical for a standard event in one scene to be considered abnormal in another. To tackle this issue, we propose an unsupervised scene clustering approach to discern scene types and deduce the STSG. Given that humans can effortlessly differentiate The categorization of scenarios based on a solitary image, we cluster the SSG of static frames to distinguish between various scenes. By categorizing events into different scenes, each group can possess its own standard events, which can be utilized to deduce the contextual framework and identify anomalous occurrences.</p>
<p>To cluster the scenarios, we introduce a graphic tech-based Variational Autoencoder (VAE). Specifically, we consider a graph represented by an adjacency matrix <bold>A</bold> &#x02208; &#x0211D;<sup><italic>K</italic>&#x000D7;<italic>K</italic></sup>, where each node is associated with a feature vector X &#x02208; &#x0211D;<sup><italic>K</italic>&#x000D7;<italic>D</italic></sup>. Our objective is to learn a node representation Z<sup>(<italic>l</italic>)</sup> in the <italic>l</italic>-th layer of a Graph Convolutional Network (GCN), as depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>. The representation of a node in the <italic>l</italic>-th layer is provided by</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>Z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="true"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>A</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mtext>Z</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3;(&#x000B7;) represents an activation function, and <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the trainable weight matrix of the <italic>l</italic>-th layer. In the first layer, we initialize Z<sup>(0)</sup> &#x0003D; X. To generate the scene graph, we assume that each node is connected to all other nodes and thus set all elements of <bold>A</bold> to 1.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Depiction of the graph-based deep Variational Autoencoders (VAEs).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsc-05-1197434-g0004.tif"/>
</fig>
<p>To execute clustering in the latent space, the Soft K-Means algorithm can be employed with the subsequent equations:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M32"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x0007C;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo stretchy='false'>|</mml:mo><mml:mo stretchy='false'>|</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>|</mml:mo><mml:msup><mml:mo stretchy='false'>|</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:mi>exp</mml:mi></mml:mrow></mml:mstyle><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo stretchy='false'>|</mml:mo><mml:mo stretchy='false'>|</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>|</mml:mo><mml:msup><mml:mo stretchy='false'>|</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mover accent='true'><mml:mi>&#x003D5;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mover accent='true'><mml:mi>&#x003BC;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>z</italic><sub><italic>i</italic></sub> represents the latent variable of data point <italic>i</italic>, &#x003BC;<sub><italic>m</italic></sub> denotes the mean of the <italic>m</italic>-th cluster, &#x003B2; is the temperature parameter controlling the assignment&#x00027;s softness, and <inline-formula><mml:math id="M33"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:math></inline-formula> is the soft assignment of data point <italic>i</italic> to the <italic>m</italic>-th cluster.</p>
<p>The soft mixture-component membership prediction <inline-formula><mml:math id="M34"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> can be obtained using the softmax function as follows:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mo class="qopname">softmax</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo class="qopname">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M36"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the soft assignment matrix derived from the Soft K-Means algorithm and &#x003B1; is the temperature parameter.</p>
<p>To obtain the estimated parameters of each cluster, we can employ the following equation</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M37"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mover accent='true'><mml:mo>&#x003A3;</mml:mo><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>&#x003BC;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>&#x003BC;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M38"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003A3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the covariance matrix of the <italic>m</italic>-th cluster.</p>
<p>The energy function can be expressed as follows:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M39"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x003B2;</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo class="qopname">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mo class="qopname">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B2; is the temperature parameter and <inline-formula><mml:math id="M40"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the fraction of data points assigned to the <italic>j</italic>-th cluster.</p>
<p>The loss function for clustering using the VAE model can be formulated as follows:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M41"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>clu</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>E</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext>&#x003A3;</mml:mtext></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BB;<sub>1</sub> represents a trade-off factor, while <italic>d</italic> represents latent-space&#x00027;dimensionality.</p>
<p>During clustering, the number of clustering centers <italic>m</italic> denotes the preset number of scene types, and we set <italic>m</italic> to 10 to cover common scenes (e.g., campuses, highways, subways, etc). The training batch of scene clustering is set to 1,024, and we use the RMSprop optimizer with a 0.0001 learning rate to train the clustering model.</p>
<p>The VAE model learns a lower-dimensional representation of the graph utilized for clustering. The Soft K-Means algorithm is applied to cluster the nodes in the latent space, and the estimated parameters of each cluster are obtained using equations akin to those employed in the Gaussian Mixture Model (GMM). The loss function for clustering using the VAE model is adapted to encompass the reconstruction loss and the Kullback-Leibler (KL) divergence term, which serve to train the VAE model.</p>
</sec>
<sec>
<title>3.3. Model optimization</title>
<p>Utilizing the methodology of segmenting visual environments, the contextual scenarios in videos were partitioned into distinct groups, and the objects and relationships in these scenes were accurately labeled. This labeled data was then utilized in the STSG inference, wherein the network&#x00027;s nodes were trained to predict the abnormality of objects, while the edges were trained to predict the abnormality of the relationships between the objects. The architecture underwent comprehensive education through the utilization of the backpropagation method, employing a holistic approach from inception to completion, and all the learnable model parameters optimized simultaneously. To accomplish this, we utilized a cross-entropy loss function with regularization, the purpose of which was to augment the likelihood delineated before. For every cluster, the probability of normalcy or deviation was ascertained for each vertex <italic>v</italic><sub><italic>i</italic></sub> and corresponding edge <italic>e</italic><sub><italic>i,j</italic></sub> within the <italic>n</italic>-th frame, individually.</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M42"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">softmax</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo class="qopname">softmax</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the term MLN refers to a multi-layer neural network that is comprised of two fully-connected layers. In order to simplify the notation, we denote the anomaly probability of <inline-formula><mml:math id="M43"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> given <italic>V</italic> and <italic>B</italic> as <inline-formula><mml:math id="M45"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M46"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, respectively. The loss function of graph inference is given by</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M47"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi mathvariant="script">L</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>N</mml:mi><mml:mi>K</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>N</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:msub><mml:mtext>&#x003BB;</mml:mtext><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The loss function utilized in the training of the MLN classifier represents the logistic loss function, denoted as <inline-formula><mml:math id="M48"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x000B7;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Variables associated with the Multilayer Neural Network(MLN) are denoted as <bold>W</bold><sub><italic>m</italic></sub>. In this investigation, the regularization parameter, denoted as &#x003BB;<sub>2</sub>, is assigned a value of 0.0001. The exponent <italic>m</italic> in <inline-formula><mml:math id="M49"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> indicates that the classification model has undergone training with the <italic>m</italic>-th cluster. Within each cluster, all entities <inline-formula><mml:math id="M50"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and associations <inline-formula><mml:math id="M51"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> receive a designation as typical, while data points from other groups are randomly selected and labeled as simulated irregularities to facilitate the training of the model.</p>
</sec>
<sec>
<title>3.4. Irregularity score</title>
<p>In order to detect abnormal events in videos, independent classifiers are trained for objects and relationships within each group. These classifiers generate classification scores, which are then utilized to calculate the final anomaly scores. In the context of the <italic>m</italic>-th group, assessment scores for categorization, represented as <inline-formula><mml:math id="M52"><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M53"><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02223;</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, are derived from the examination of test data through the evaluation of Equation 12. The anomaly score is then determined as the minimal categorization metric observed across all scenario clusters.</p>
<p>Deviation metrics pertaining to entities and their interconnections are denoted as <inline-formula><mml:math id="M54"><mml:msubsup><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M55"><mml:msubsup><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, respectively, and are computed for each of the M groups. The former score reflects the detection of individual anomalies, while the latter reflects the detection of group anomalies. To achieve granular detection at the frame level, the highest score encompassing all entities and associations within a single frame is identified as the representative anomaly score for that particular frame. To ensure that the irregularity score varies smoothly across frames, we apply a Gaussian filter to enforce temporal smoothness of the final frame-level anomaly scores.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments</title>
<sec>
<title>4.1. Datasets</title>
<p>The effectiveness of our proposed method is evaluated on three benchmark datasets, namely UCF-Crime (Sultani et al., <xref ref-type="bibr" rid="B42">2018</xref>), Avenue (Lu et al., <xref ref-type="bibr" rid="B22">2013</xref>), and ShanghaiTech (Luo et al., <xref ref-type="bibr" rid="B23">2017</xref>). The UCF-Crime dataset is an extensive compilation of authentic surveillance footage encompassing 13 distinct categories of anomalous occurrences across various environments. This collection is comprised of 1,610 videos for training purposes and 290 videos designated for evaluation, all of which were utilized in our experiments. Avenue, on the other hand, contains 16 training and 21 testing videos with a total of 35,240 frames, each lasting about 2 minutes. The dataset includes four types of abnormal events: running, walking in opposite direction, throwing objects, and loitering. Lastly, the ShanghaiTech dataset includes 13 scenes with complex light conditions and various viewpoints, and consists of over 270,000 training frames and 130 abnormal events. We utilized all of these datasets to comprehensively evaluate the performance of our proposed method.</p>
</sec>
<sec>
<title>4.2. Evaluation metric</title>
<p>We assess the performance of our proposed method at the frame level by computing anomaly scores for each frame. The performance of the method is evaluated using the Receiver Operating Characteristic (ROC) curve (Fawcett, <xref ref-type="bibr" rid="B6">2006</xref>), which involves progressively adjusting the benchmark for irregularity values. The relevant evaluation metrics employed encompass the Area Under the Curve (AUC &#x02191;) and the Equal Error Rate (EER &#x02193;). Moreover, the false alarm rate serves as an assessment indicator for the likelihood of incorrect categorization. Enhanced performance of the anomaly detection technique is signified by an elevated AUC merit (Lobo et al., <xref ref-type="bibr" rid="B21">2008</xref>), a diminished EER merit, and other merits.</p>
</sec>
<sec>
<title>4.3. Comparisons</title>
<sec>
<title>4.3.1. Analysis on the UCF-crime dataset</title>
<p>The methodology we put forth undergoes assessment and juxtaposition with numerous prevalent unsupervised and supervised techniques, employing the UCF-Crime dataset for this comparative analysis. The performance of our method is reported in <xref ref-type="table" rid="T1">Table 1</xref> in terms of the AUC and false alarm rate, respectively. To ensure a fair comparison, we reconstructed the research conducted by Ionescu et al. (<xref ref-type="bibr" rid="B13">2019</xref>), substituting their employed detection mechanism with the Region Proposal Network (RPN) detector to enhance the methodology. The performances of other compared methods are taken from Sultani et al. (<xref ref-type="bibr" rid="B42">2018</xref>). The results show that our method outperforms cutting-edge unsupervised technique, with an improvement of 8.9 and 1.6% on the Area Under the Curve (AUC) and false alarm rate evaluations, respectively. Moreover, our approach exhibits similarity to the most advanced supervised technique (Sultani et al., <xref ref-type="bibr" rid="B42">2018</xref>) currently available in the field, achieving comparable AUC scores and false alarm rates without the need for video-level annotations. This demonstrates the effectiveness of our method in detecting unknown abnormal events in real-world applications. The ROC curves of our method are plotted in <xref ref-type="fig" rid="F5">Figure 5</xref>, which encompasses the contours of unsupervised methodologies and surpasses the study of Hasan et al. (<xref ref-type="bibr" rid="B10">2016</xref>); Ionescu et al. (<xref ref-type="bibr" rid="B13">2019</xref>); Lu et al. (<xref ref-type="bibr" rid="B22">2013</xref>) at diverse benchmarks. True positive of the proposed method marginally exceeds the research of Sultani et al. (<xref ref-type="bibr" rid="B42">2018</xref>) when a middle threshold is selected, indicating the effectiveness of our method.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>A comparative evaluation of abnormal event detection outcomes between unsupervised and supervised techniques.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Training</bold></th>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>AUC &#x02191;</bold></th>
<th valign="top" align="left"><bold>False alarm &#x02193;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="1">Unsupervised</td>
<td valign="top" align="left">Hasan et al. (<xref ref-type="bibr" rid="B10">2016</xref>)</td>
<td valign="top" align="center">49.8%</td>
<td valign="top" align="left">26.9%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ionescu et al. (<xref ref-type="bibr" rid="B13">2019</xref>)</td>
<td valign="top" align="center">62.1%</td>
<td valign="top" align="left">9.3%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Lu et al. (<xref ref-type="bibr" rid="B22">2013</xref>)</td>
<td valign="top" align="center">67.4%</td>
<td valign="top" align="left">3.8%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>76.3%</bold></td>
<td valign="top" align="left">2.2%</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="1">Supervised</td>
<td valign="top" align="left">SVM baseline</td>
<td valign="top" align="center">50%</td>
<td valign="top" align="left">&#x02013;</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Sultani et al. (<xref ref-type="bibr" rid="B42">2018</xref>)</td>
<td valign="top" align="center">69.2%</td>
<td valign="top" align="left"><bold>2.1%</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x02191; Indicates that advanced values correspond to superior performance, and &#x02193; signifies that lower scores are indicative of better results.</p>
<p>The bold values indicate the best value among all experimental results.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparison of ROC curves for various unsupervised and supervised approaches on the UCF-Crime dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsc-05-1197434-g0005.tif"/>
</fig>
</sec>
<sec>
<title>4.3.2. Analysis on the avenue dataset</title>
<p>On the Avenue dataset, our method outperforms all existing methods in terms of both the AUC and EER evaluations, as shown in <xref ref-type="table" rid="T2">Table 2</xref>. The cutting-edge research of Ye et al. (<xref ref-type="bibr" rid="B54">2019</xref>) achieved AUC values of 85.9%, while our approach gained an advancement of 4.0%, demonstrating the effectiveness and robustness of our method.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparative evaluation of abnormal event detection performance using AUC and EER metrics.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>AUC &#x02191;</bold></th>
<th valign="top" align="center"><bold>EER &#x02193;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Chong and Tay (<xref ref-type="bibr" rid="B5">2017</xref>)</td>
<td valign="top" align="center">78.2%</td>
<td valign="top" align="center">21.3%</td>
</tr>
<tr>
<td valign="top" align="left">Hasan et al. (<xref ref-type="bibr" rid="B10">2016</xref>)</td>
<td valign="top" align="center">69.4%</td>
<td valign="top" align="center">26.1%</td>
</tr>
<tr>
<td valign="top" align="left">Ionescu et al. (<xref ref-type="bibr" rid="B13">2019</xref>)</td>
<td valign="top" align="center">81.0%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Luo et al. (<xref ref-type="bibr" rid="B23">2017</xref>)</td>
<td valign="top" align="center">82.1%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Liu et al. (<xref ref-type="bibr" rid="B20">2018</xref>)</td>
<td valign="top" align="center">83.5%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Wang et al. (<xref ref-type="bibr" rid="B49">2018</xref>)</td>
<td valign="top" align="center">84.7%</td>
<td valign="top" align="center">22.9%</td>
</tr>
<tr>
<td valign="top" align="left">Morais et al. (<xref ref-type="bibr" rid="B27">2019</xref>)</td>
<td valign="top" align="center">85.6%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ye et al. (<xref ref-type="bibr" rid="B54">2019</xref>)</td>
<td valign="top" align="center">85.9%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>89.9%</bold></td>
<td valign="top" align="center"><bold>20.4%</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The bold values indicate the best value among all experimental results.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.3.3. Analysis on the ShanghaiTech dataset</title>
<p>Furthermore, we present the findings of our experimental evaluation conducted on the demanding ShanghaiTech dataset, which contains complex scenes and various actions. According to the information presented in <xref ref-type="table" rid="T3">Table 3</xref>, the proposed approach overtakes the leading-edge strategies on this dataset, demonstrating its effectiveness in detecting abnormal events in challenging settings.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparative evaluation of abnormal event detection performance using frame-level AUC and EER metrics.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>AUC &#x02191;</bold></th>
<th valign="top" align="center"><bold>EER &#x02193;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Chong and Tay (<xref ref-type="bibr" rid="B5">2017</xref>)</td>
<td valign="top" align="center">61.2%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Luo et al. (<xref ref-type="bibr" rid="B23">2017</xref>)</td>
<td valign="top" align="center">67.9%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Wang et al. (<xref ref-type="bibr" rid="B49">2018</xref>)</td>
<td valign="top" align="center">71.7%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Liu et al. (<xref ref-type="bibr" rid="B20">2018</xref>)</td>
<td valign="top" align="center">72.1%</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Ionescu et al. (<xref ref-type="bibr" rid="B13">2019</xref>)</td>
<td valign="top" align="center">72.9%</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>75.2%</bold></td>
<td valign="top" align="center"><bold>25.1%</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The bold values indicate the best value among all experimental results.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec>
<title>4.4. Ablation study</title>
<p><xref ref-type="table" rid="T4">Table 4</xref> showcases a comparative analysis of distinct constituents&#x00027; contributions within the proposed technique designed for detecting unsupervised abnormal events. The term &#x0201C;w/o spatial relationships&#x0201D; signifies the exclusion of associations in space dimension, wherein the STSGs are converted into object-oriented series over numerous frames, which are then simulated by Transformers. The term &#x0201C;w/o temporal relationships&#x0201D; implies its performance on the SSG inference disregarding any temporal connections, while &#x0201C;w/o relationships&#x0201D; employs a twin set of fully-connected layers to simulate individual objects in a standalone manner. We conducted identical scene clustering for the aforementioned three scenarios. The term &#x0201C;w/o scene clustering&#x0201D; denotes the exclusion of scenario clustering and solely relying on a one-class discriminator to differentiate between regular and aberrant occurrences. Referring to <xref ref-type="table" rid="T4">Table 4</xref>, it is observed that discarding spatial dependencies, temporal dependencies, or spatio-temporal dependencies decreases the AUC execution by 6.9%&#x02212;13.6%, indicating the relevance of information associations for distinguishing irregular occurrences. Scenario clustering significantly improves performance, and the exhibited performance in distinguishing diverse environments to detect anomalous incidents affirms the efficacy of this approach. Furthermore, the enhancement reinforces the effectiveness of the unsupervised scene clustering technique utilized during the training phase.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Performance evaluation of individual components of the proposed approach in terms of AUC and false alarm.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>AUC &#x02191;</bold></th>
<th valign="top" align="center"><bold>False alarm &#x02193;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">W/o temporal relationships</td>
<td valign="top" align="center">69.4%</td>
<td valign="top" align="center">2.9%</td>
</tr>
<tr>
<td valign="top" align="left">W/o spatial relationships</td>
<td valign="top" align="center">62.7%</td>
<td valign="top" align="center">5.2%</td>
</tr>
<tr>
<td valign="top" align="left">W/o relationships</td>
<td valign="top" align="center">62.8%</td>
<td valign="top" align="center">12.5%</td>
</tr>
<tr>
<td valign="top" align="left">W/o scene clustering</td>
<td valign="top" align="center">64.9%</td>
<td valign="top" align="center">7.1%</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>76.3%</bold></td>
<td valign="top" align="center"><bold>1.9%</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The bold values indicate the best value among all experimental results.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.5. Incident reckon</title>
<p>In order to detect and determine anomalous incidents based on irregularity values, we employ a method that involves choosing the local maxima among the chronological progression of irregularity values within a given video. To identify meaningful local maxima, we utilize the persistence1D algorithm, and then define a fixed time interval for the region. To group nearby expanded local maximum regions, we adopt the approach outlined in Luo et al. (<xref ref-type="bibr" rid="B23">2017</xref>), which results in the final abnormal temporal regions where anomalous incidents can be accurately determined.</p>
<p>On the Avenue dataset, the outcomes of the proposed approach are presented in <xref ref-type="table" rid="T5">Table 5</xref>, which exhibits the quantity of identified anomalous incidents and false alarms. The strategy we developed can reliably identify anomalous incidents in comparison to the approaches taken in Wang et al. (<xref ref-type="bibr" rid="B49">2018</xref>) and Luo et al. (<xref ref-type="bibr" rid="B23">2017</xref>). The false alarm rate of the proposed approach exceeds the work in Medel and Savakis (<xref ref-type="bibr" rid="B25">2016</xref>), primarily due to their use of minute benchmarks of irregularity values to identify anomalous incidents. However, the approach we proposed identifies 47 true anomalous incidents, compared to the 39 anomalous incidents identified by the strategy in Wang et al. (<xref ref-type="bibr" rid="B49">2018</xref>). These outcomes manifest the superior validity of the methodology we developed in verifying the time span of anomalous incidents, rendering it a more feasible choice for implementation in real-world scenarios.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Outlier event identification outcomes in terms of count of identified occurrences and false alarms.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>True positives &#x02191;</bold></th>
<th valign="top" align="center"><bold>False alarm &#x02193;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Wang et al. (<xref ref-type="bibr" rid="B49">2018</xref>)</td>
<td valign="top" align="center">39</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Luo et al. (<xref ref-type="bibr" rid="B23">2017</xref>)</td>
<td valign="top" align="center">44</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Morais et al. (<xref ref-type="bibr" rid="B27">2019</xref>)</td>
<td valign="top" align="center">42</td>
<td valign="top" align="center">13</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>47</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The bold values indicate the best value among all experimental results.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>Throughout this article, we propose a novel approach for unsupervised anomalous incidents identification in videos, particularly those acquired via Unmanned Aerial Vehicles (UAVs), which involves the utilization of a contextually responsive reasoning strategy. As UAVs are increasingly utilized in various applications such as surveillance, search and rescue, and environmental monitoring, anomalous incident detection in UAV-captured videos is pivotal for ensuring safety and security. Contextual inference overtly entails the extraction of high-level environmental knowledge from low-level vision-oriented characteristics. Our approach generates a spatiotemporal scenario graph to facilitate the explicit establishment of the vision-oriented environment, by embedding objects&#x00027; visual morphology and their spatiotemporal associations in graphic representations. This approach is particularly beneficial in UAV-captured videos, where the aerial perspective offers unique contextual information. Furthermore, we evolve a graph-based deep Variational Autoencoder model for scenario clustering that can capably ascertain scenario categories and deduce the spatio-temporal scenario graph within unsupervised. This enables our method to accurately detect aberrant occurrences with contextual dependencies and ambiguous sources in various environments, including those captured by UAVs. Our experiments on three datasets, including UAV-captured videos, exhibit the superiority of our approach over current unsupervised methodologies, while simultaneously highlighting its comparability with contemporary supervised techniques that represent the cutting-edge of the field. Subsequent research will endeavor to investigate more detailed contextual feature in order to expand the methodology from detecting anomalies at the frame dimension to the more precise pixel level, further enhancing the effectiveness of atypical occurrences identification in UAV-captured videos.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>GL proposed the idea and wrote this article with YY and LS. YY conducted the experiments. CJ organized the entire research, provided funding supports, handles the manuscript, and correspondence during the publication process. All authors contributed to manuscript writing, revision, read, and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This work was supported by the Zhejiang &#x02018;JIANBING&#x00027; R&#x00026;D Project (No. 2022C01055) and R&#x00026;D Project of Department of Transport of Zhejiang Province (No. 2021010).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anti&#x00107;</surname> <given-names>B.</given-names></name> <name><surname>Ommer</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Video parsing for abnormality detection.&#x0201D;</article-title> in 2011 <italic>International Conference on Computer Vision</italic> (Barcelona: IEEE), <fpage>2415</fpage>&#x02013;<lpage>2422</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2011.6126525</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bar</surname> <given-names>M.</given-names></name></person-group> (<year>2004</year>). <article-title>Visual objects in context</article-title>. <source>Nat. Rev. Neurosci</source>. <volume>5</volume>, <fpage>617</fpage>&#x02013;<lpage>629</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1476</pub-id><pub-id pub-id-type="pmid">15263892</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>K.-W.</given-names></name> <name><surname>Chen</surname> <given-names>Y.-T.</given-names></name> <name><surname>Fang</surname> <given-names>W.-H.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Video anomaly detection and localization using hierarchical feature representation and gaussian process regression,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2909</fpage>&#x02013;<lpage>2917</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298909</pub-id><pub-id pub-id-type="pmid">26394423</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>M. J.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Willsky</surname> <given-names>A. S.</given-names></name></person-group> (<year>2012</year>). <article-title>Context models and out-of-context objects</article-title>. <source>Pattern Recogn. Lett</source>. <volume>33</volume>, <fpage>853</fpage>&#x02013;<lpage>862</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2011.12.004</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chong</surname> <given-names>Y. S.</given-names></name> <name><surname>Tay</surname> <given-names>Y. H.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Abnormal event detection in videos using spatiotemporal autoencoder,&#x0201D;</article-title> in <source>Advances in Neural Networks-ISNN 2017: 14th International Symposium, ISNN 2017, Sapporo, Hakodate, and Muroran, Hokkaido, Japan, June 21-26, 2017, Proceedings, Part II 14</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>189</fpage>&#x02013;<lpage>196</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-59081-3_23</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fawcett</surname> <given-names>T.</given-names></name></person-group> (<year>2006</year>). <article-title>An introduction to roc analysis</article-title>. <source>Pattern Recogn. Lett</source>. <volume>27</volume>, <fpage>861</fpage>&#x02013;<lpage>874</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2005.10.010</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>J.-C.</given-names></name> <name><surname>Hong</surname> <given-names>F.-T.</given-names></name> <name><surname>Zheng</surname> <given-names>W.-S.</given-names></name></person-group> (<year>2021</year>). &#x0201C;Mist: multiple instance self-training framework for video anomaly detection,&#x0201D; <italic>in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</italic> (Nashville, TN: IEEE), <fpage>14009</fpage>&#x02013;<lpage>14018</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01379</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>Y.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep representation for abnormal event detection in crowded scenes,&#x0201D;</article-title> in <source>Acm on Multimedia Conference</source>, 591&#x02013;595. <pub-id pub-id-type="doi">10.1145/2964284.2967290</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Georgescu</surname> <given-names>M. I.</given-names></name> <name><surname>Barbalau</surname> <given-names>A.</given-names></name> <name><surname>Ionescu</surname> <given-names>R. T.</given-names></name> <name><surname>Khan</surname> <given-names>F. S.</given-names></name> <name><surname>Popescu</surname> <given-names>M.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Anomaly detection in video via self-supervised and multi-task learning,&#x0201D;</article-title> in <source>Computer Vision and Pattern Recognition</source> (<publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>IEEE</publisher-name>). <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01255</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasan</surname> <given-names>M.</given-names></name> <name><surname>Choi</surname> <given-names>J.</given-names></name> <name><surname>Neumann</surname> <given-names>J.</given-names></name> <name><surname>Roy-Chowdhury</surname> <given-names>A. K.</given-names></name> <name><surname>Davis</surname> <given-names>L. S.</given-names></name></person-group> (<year>2016</year>). &#x0201C;Learning temporal regularity in video sequences, in <italic>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic> (Las Vegas, NV, USA: IEEE). <pub-id pub-id-type="doi">10.1109/CVPR.2016.86</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasan</surname> <given-names>M.</given-names></name> <name><surname>Paul</surname> <given-names>S.</given-names></name> <name><surname>Mourikis</surname> <given-names>A. I.</given-names></name> <name><surname>Roy-Chowdhury</surname> <given-names>A. K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Context-aware query selection for active learning in event recognition,&#x0201D;</article-title> in <source>IEEE Transactions on Pattern Analysis</source> &#x00026; <italic>Machine Intelligence</italic> (IEEE), <fpage>1</fpage>.<pub-id pub-id-type="pmid">30387722</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinami</surname> <given-names>R.</given-names></name> <name><surname>Mei</surname> <given-names>T.</given-names></name> <name><surname>Satoh</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Joint detection and recounting of abnormal events by learning deep generic knowledge,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3619</fpage>&#x02013;<lpage>3627</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.391</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ionescu</surname> <given-names>R. T.</given-names></name> <name><surname>Khan</surname> <given-names>F. S.</given-names></name> <name><surname>Georgescu</surname> <given-names>M.-I.</given-names></name> <name><surname>Shao</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Object-centric auto-encoders and dummy anomalies for abnormal event detection in video,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>7842</fpage>&#x02013;<lpage>7851</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00803</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ionescu</surname> <given-names>R. T.</given-names></name> <name><surname>Smeureanu</surname> <given-names>S.</given-names></name> <name><surname>Alexe</surname> <given-names>B.</given-names></name> <name><surname>Popescu</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Unmasking the abnormal events in video</article-title>, in <italic>2017 IEEE International Conference on Computer Vision (ICCV)</italic> (Venice, Italy: IEEE). <pub-id pub-id-type="doi">10.1109/ICCV.2017.315</pub-id><pub-id pub-id-type="pmid">17884759</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>Y.</given-names></name> <name><surname>Rahaman</surname> <given-names>M. S.</given-names></name> <name><surname>Ren</surname> <given-names>Y.</given-names></name> <name><surname>Sanderson</surname> <given-names>M.</given-names></name> <name><surname>White</surname> <given-names>R. W.</given-names></name> <name><surname>Salim</surname> <given-names>F. D.</given-names></name></person-group> (<year>2022</year>). <article-title>App usage on-the-move: context-and commute-aware next app prediction</article-title>. <source>Pervasive Mobile Comput</source>. 87, 101704. <pub-id pub-id-type="doi">10.1016/j.pmcj.2022.101704</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kratz</surname> <given-names>L.</given-names></name> <name><surname>Nishino</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Anomaly detection in extremely crowded scenes using spatio-temporal motion pattern models,&#x0201D;</article-title> in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1446</fpage>&#x02013;<lpage>1453</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206771</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leach</surname> <given-names>M. J.</given-names></name> <name><surname>Sparks</surname> <given-names>E. P.</given-names></name> <name><surname>Robertson</surname> <given-names>N. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Contextual anomaly detection in crowded surveillance scenes</article-title>. <source>Pattern Recogn. Lett</source>. <volume>44</volume>, <fpage>71</fpage>&#x02013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2013.11.018</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leyva</surname> <given-names>R.</given-names></name> <name><surname>Sanchez</surname> <given-names>V.</given-names></name> <name><surname>Li</surname> <given-names>C. T.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Video anomaly detection with compact feature sets for online performance,&#x0201D;</article-title> in <source>IEEE Transactions on Image Processing</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>3463</fpage>&#x02013;<lpage>3478</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2695105</pub-id><pub-id pub-id-type="pmid">28436865</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Mahadevan</surname> <given-names>V.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>N.</given-names></name></person-group> (<year>2013</year>). <article-title>Anomaly detection and localization in crowded scenes</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intellig</source>. <volume>36</volume>, <fpage>18</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2013.111</pub-id><pub-id pub-id-type="pmid">28221995</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Lian</surname> <given-names>D.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Future frame prediction for anomaly detection-a new baseline,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6536</fpage>&#x02013;<lpage>6545</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00684</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lobo</surname> <given-names>J. M.</given-names></name> <name><surname>Jim&#x000E9;nez-Valverde</surname> <given-names>A.</given-names></name> <name><surname>Real</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Auc: a misleading measure of the performance of predictive distribution models</article-title>. <source>Glob. Ecol. Biogeogr</source>. <volume>17</volume>, <fpage>145</fpage>&#x02013;<lpage>151</lpage>. <pub-id pub-id-type="doi">10.1111/j.1466-8238.2007.00358.x</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>C.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Abnormal event detection at 150 fps in matlab,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2720</fpage>&#x02013;<lpage>2727</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2013.338</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;A revisit of sparse coding based anomaly detection in stacked rnn framework,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>341</fpage>&#x02013;<lpage>349</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.45</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mahadevan</surname> <given-names>V.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Bhalodia</surname> <given-names>V.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>N.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Anomaly detection in crowded scenes,&#x0201D;</article-title> in <source>2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1975</fpage>&#x02013;<lpage>1981</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2010.5539872</pub-id><pub-id pub-id-type="pmid">28221995</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Medel</surname> <given-names>J. R.</given-names></name> <name><surname>Savakis</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Anomaly detection in video using predictive convolutional long short-term memory networks</article-title>. <source>arXiv [Preprint]</source>. arXiv:1612.00390.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehran</surname> <given-names>R.</given-names></name> <name><surname>Oyama</surname> <given-names>A.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Abnormal crowd behavior detection using social force model,&#x0201D;</article-title> in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>935</fpage>&#x02013;<lpage>942</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206641</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morais</surname> <given-names>R.</given-names></name> <name><surname>Le</surname> <given-names>V.</given-names></name> <name><surname>Tran</surname> <given-names>T.</given-names></name> <name><surname>Saha</surname> <given-names>B.</given-names></name> <name><surname>Mansour</surname> <given-names>M.</given-names></name> <name><surname>Venkatesh</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Learning regularity in skeleton trajectories for anomaly detection in videos,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>11996</fpage>&#x02013;<lpage>12004</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01227</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>G.</given-names></name> <name><surname>Yan</surname> <given-names>C.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Hengel</surname> <given-names>A. v. d</given-names></name>  <name><surname>Bai</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Self-trained deep ordinal regression for end-to-end video anomaly detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>12173</fpage>&#x02013;<lpage>12182</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01219</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>W.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Abnormal object detection by canonical scene-based contextual model,&#x0201D;</article-title> in <source>Computer Vision-ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part III 12</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>651</fpage>&#x02013;<lpage>664</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-33712-3_47</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qi</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Qin</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>A.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>Stagnet: an attentive semantic rnn for group activity and individual action recognition</article-title>. <source>IEEE Trans. Circuits Syst. Video Technol</source>. <volume>30</volume>, <fpage>549</fpage>&#x02013;<lpage>565</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2019.2894161</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ravanbakhsh</surname> <given-names>M.</given-names></name> <name><surname>Nabi</surname> <given-names>M.</given-names></name> <name><surname>Mousavi</surname> <given-names>H.</given-names></name> <name><surname>Sangineto</surname> <given-names>E.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Plug-and-play cnn for crowd motion analysis: an application in abnormal event detection,&#x0201D;</article-title> in <source>2018 IEEE Winter Conference on Applications of Computer Vision (WACV)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1689</fpage>&#x02013;<lpage>1698</lpage>. <pub-id pub-id-type="doi">10.1109/WACV.2018.00188</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Faster R-CNN: Towards real-time object detection with region proposal networks,&#x0201D;</article-title> in <source>Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1</source> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>91</fpage>&#x02013;<lpage>99</lpage>.<pub-id pub-id-type="pmid">27295650</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ribeiro</surname> <given-names>M.</given-names></name> <name><surname>Lazzaretti</surname> <given-names>A. E.</given-names></name> <name><surname>Lopes</surname> <given-names>H. S.</given-names></name></person-group> (<year>2018</year>). <article-title>A study of deep convolutional auto-encoders for anomaly detection in videos</article-title>. <source>Pattern Recogn. Lett</source>. <volume>105</volume>:<fpage>13</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2017.07.016</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sabokrou</surname> <given-names>M.</given-names></name> <name><surname>Fathy</surname> <given-names>M.</given-names></name> <name><surname>Hoseini</surname> <given-names>M.</given-names></name> <name><surname>Klette</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Real-time anomaly detection and localization in crowded scenes,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>56</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2015.7301284</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sabokrou</surname> <given-names>M.</given-names></name> <name><surname>Fayyaz</surname> <given-names>M.</given-names></name> <name><surname>Fathy</surname> <given-names>M.</given-names></name> <name><surname>Klette</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep-cascade: cascading 3d deep neural networks for fast anomaly detection and localization in crowded scenes</article-title>. <source>IEEE Trans. Image Process</source>. <volume>26</volume>, <fpage>1992</fpage>&#x02013;<lpage>2004</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2670780</pub-id><pub-id pub-id-type="pmid">28221995</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sabokrou</surname> <given-names>M.</given-names></name> <name><surname>Fayyaz</surname> <given-names>M.</given-names></name> <name><surname>Fathy</surname> <given-names>M.</given-names></name> <name><surname>Moayed</surname> <given-names>Z.</given-names></name> <name><surname>Klette</surname> <given-names>R.</given-names></name></person-group> (<year>2018a</year>). <article-title>Deep-anomaly: fully convolutional neural network for fast anomaly detection in crowded scenes</article-title>. <source>Comput. Vis. Image Understanding</source> <volume>172</volume>, <fpage>88</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2018.02.006</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sabokrou</surname> <given-names>M.</given-names></name> <name><surname>Khalooei</surname> <given-names>M.</given-names></name> <name><surname>Fathy</surname> <given-names>M.</given-names></name> <name><surname>Adeli</surname> <given-names>E.</given-names></name></person-group> (<year>2018b</year>). <article-title>&#x0201C;Adversarially learned one-class classifier for novelty detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3379</fpage>&#x02013;<lpage>3388</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00356</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saligrama</surname> <given-names>V.</given-names></name> <name><surname>Konrad</surname> <given-names>J.</given-names></name> <name><surname>Jodoin</surname> <given-names>P.-M.</given-names></name></person-group> (<year>2010</year>). <article-title>Video anomaly identification</article-title>. <source>IEEE Signal Process. Magaz</source>. <volume>27</volume>, <fpage>18</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1109/MSP.2010.937393</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shao</surname> <given-names>W.</given-names></name> <name><surname>Salim</surname> <given-names>F. D.</given-names></name> <name><surname>Song</surname> <given-names>A.</given-names></name> <name><surname>Bouguettaya</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Clustering big spatiotemporal-interval data</article-title>. <source>IEEE Trans. Big Data</source> <volume>2</volume>, <fpage>190</fpage>&#x02013;<lpage>203</lpage>. <pub-id pub-id-type="doi">10.1109/TBDATA.2016.2599923</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smeureanu</surname> <given-names>S.</given-names></name> <name><surname>Ionescu</surname> <given-names>R. T.</given-names></name> <name><surname>Popescu</surname> <given-names>M.</given-names></name> <name><surname>Alexe</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Deep appearance features for abnormal behavior detection in video,&#x0201D;</article-title> in <source>Image Analysis and Processing-ICIAP 2017: 19th International Conference, Catania, Italy, September 11-15, 2017, Proceedings, Part II 19</source> (<publisher-loc>Catania</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>779</fpage>&#x02013;<lpage>789</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-68548-9_70</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>H.</given-names></name> <name><surname>Sun</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Learning normal patterns via adversarial attention-based autoencoder for abnormal event detection in videos,&#x0201D;</article-title> in <source>IEEE Transactions on Multimedia</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1</fpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sultani</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). &#x0201C;Real-world anomaly detection in surveillance videos,&#x0201D; <italic>in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</italic> (Salt Lake City, UT: IEEE), <fpage>6479</fpage>&#x02013;<lpage>6488</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00678</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>C.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Scene-aware context reasoning for unsupervised abnormal event detection in videos,&#x0201D;</article-title> in <source>Proceedings of the 28th ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>184</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1145/3394171.3413887</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>C.</given-names></name> <name><surname>Song</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Learning weighted video segments for temporal action localization,&#x0201D;</article-title> in <source>Pattern Recognition and Computer Vision: Second Chinese Conference, PRCV 2019, Xi&#x00027;an, China, November 8-11, 2019, Proceedings, Part I 2</source> (<publisher-loc>Xian</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>181</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-31654-9_16</pub-id><pub-id pub-id-type="pmid">35994544</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>B.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Learning to compose dynamic tree structures for visual contexts,&#x0201D;</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>). <pub-id pub-id-type="doi">10.1109/CVPR.2019.00678</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tipping</surname> <given-names>M. E.</given-names></name> <name><surname>Bishop</surname> <given-names>C. M.</given-names></name></person-group> (<year>1999</year>). <article-title>Mixtures of probabilistic principal component analyzers</article-title>. <source>Neural Comput</source>. <volume>11</volume>, <fpage>443</fpage>&#x02013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1162/089976699300016728</pub-id><pub-id pub-id-type="pmid">9950739</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tran</surname> <given-names>D.</given-names></name> <name><surname>Bourdev</surname> <given-names>L.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>Torresani</surname> <given-names>L.</given-names></name> <name><surname>Paluri</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Learning spatiotemporal features with 3d convolutional networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4489</fpage>&#x02013;<lpage>4497</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.510</pub-id><pub-id pub-id-type="pmid">30530363</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmer</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Attention is all you need,&#x0201D;</article-title> in <source>Proceedings of the 31st International Conference on Neural Information Processing Systems</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates</publisher-name>), <fpage>6000</fpage>&#x02013;<lpage>6010</lpage>.</citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name> <name><surname>Tan</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Abnormal event detection in videos using hybrid spatio-temporal autoencoder,&#x0201D;</article-title> in <source>2018 25th IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>2276</fpage>&#x02013;<lpage>2280</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2018.8451070</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Ricci</surname> <given-names>E.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>J.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning deep representations of appearance and motion for anomalous event detection</article-title>. <source>arXiv [Preprint]</source>. arXiv:1510.01553. <pub-id pub-id-type="doi">10.5244/C.29.8</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Ricci</surname> <given-names>E.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2017a</year>). <article-title>Detecting anomalous events in videos by learning deep representations of appearance and motion</article-title>. <source>Comput. Vis. Image Understanding</source> <volume>156</volume>, <fpage>117</fpage>&#x02013;<lpage>127</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2016.10.010</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Choy</surname> <given-names>C. B.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2017b</year>). <article-title>&#x0201C;Scene graph generation by iterative message passing,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5410</fpage>&#x02013;<lpage>5419</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.330</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>R.</given-names></name> <name><surname>Tu</surname> <given-names>Z.</given-names></name> <name><surname>Xiang</surname> <given-names>H.</given-names></name> <name><surname>Shao</surname> <given-names>W.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Ma</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Cobevt: Cooperative bird&#x00027;s eye view semantic segmentation with sparse transformers</article-title>. <source>arXiv [Preprint]</source>. arXiv:2207.02202.</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ye</surname> <given-names>M.</given-names></name> <name><surname>Peng</surname> <given-names>X.</given-names></name> <name><surname>Gan</surname> <given-names>W.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Qiao</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Anopcn: Video anomaly detection via deep predictive coding network,&#x0201D;</article-title> in <source>Proceedings of the 27th ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1805</fpage>&#x02013;<lpage>1813</lpage>. <pub-id pub-id-type="doi">10.1145/3343031.3350899</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Qing</surname> <given-names>L.</given-names></name> <name><surname>Miao</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Temporal convolutional network with complementary inner bag loss for weakly supervised anomaly detection,&#x0201D;</article-title> in <source>2019 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Taipei</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4030</fpage>&#x02013;<lpage>4034</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2019.8803657</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Deng</surname> <given-names>B.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Hua</surname> <given-names>X.-S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Spatio-temporal autoencoder for video anomaly detection,&#x0201D;</article-title> in <source>Proceedings of the 25th ACM international conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1933</fpage>&#x02013;<lpage>1941</lpage>. <pub-id pub-id-type="doi">10.1145/3123266.3123451</pub-id><pub-id pub-id-type="pmid">36681841</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>J.-X.</given-names></name> <name><surname>Li</surname> <given-names>N.</given-names></name> <name><surname>Kong</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>T. H.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Graph convolutional label noise cleaner: train a plug-and-play action classifier for anomaly detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1237</fpage>&#x02013;<lpage>1246</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00133</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Andonian</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Temporal relational reasoning in videos,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>803</fpage>&#x02013;<lpage>818</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01246-5_49</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Newsam</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Motion-aware feature for improved video anomaly detection</article-title>. <source>arXiv [Preprint]</source>. arXiv:1907.10211.</citation>
</ref>
</ref-list> 
</back>
</article>