<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Sci.</journal-id>
<journal-title>Frontiers in Computer Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Sci.</abbrev-journal-title>
<issn pub-type="epub">2624-9898</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcomp.2022.876846</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Computer Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Ani-GIFs: A benchmark dataset for domain generalization of action recognition from GIFs</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Majumdar</surname> <given-names>Shoumik Sovan</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Jain</surname> <given-names>Shubhangi</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Tourni</surname> <given-names>Isidora Chara</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1712296/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Mustafin</surname> <given-names>Arsenii</given-names></name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Lteif</surname> <given-names>Diala</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1556567/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sclaroff</surname> <given-names>Stan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/415023/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Saenko</surname> <given-names>Kate</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Bargal</surname> <given-names>Sarah Adel</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1319231/overview"/>
</contrib>
</contrib-group>
<aff><institution>Department of Computer Science, Boston University</institution>, <addr-line>Boston, MA</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Elisa Ricci, University of Trento, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yang Wang, University of Manitoba, Canada; Massimiliano Mancini, University of T&#x000FC;bingen, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Diala Lteif <email>dlteif&#x00040;bu.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Computer Vision, a section of the journal Frontiers in Computer Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>26</day>
<month>09</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>4</volume>
<elocation-id>876846</elocation-id>
<history>
<date date-type="received">
<day>15</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>08</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Majumdar, Jain, Tourni, Mustafin, Lteif, Sclaroff, Saenko and Bargal.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Majumdar, Jain, Tourni, Mustafin, Lteif, Sclaroff, Saenko and Bargal</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Deep learning models perform remarkably well for the same task under the assumption that data is always coming from the same distribution. However, this is generally violated in practice, mainly due to the differences in data acquisition techniques and the lack of information about the underlying source of new data. Domain generalization targets the ability to generalize to test data of an unseen domain; while this problem is well-studied for images, such studies are significantly lacking in spatiotemporal visual content&#x02014;videos and GIFs. This is due to (1) the challenging nature of misalignment of temporal features and the varying appearance/motion of actors and actions in different domains, and (2) spatiotemporal datasets being laborious to collect and annotate for multiple domains. We collect and present the first synthetic video dataset of Animated GIFs for domain generalization, <italic>Ani-GIFs</italic>, that is used to study the domain gap of videos vs. GIFs, and animated vs. real GIFs, for the task of action recognition. We provide a training and testing setting for <italic>Ani-GIFs</italic>, and extend two domain generalization baseline approaches, based on data augmentation and explainability, to the spatiotemporal domain to catalyze research in this direction.</p></abstract>
<kwd-group>
<kwd>domain generalization</kwd>
<kwd>domain adaptation</kwd>
<kwd>video action recognition</kwd>
<kwd>GIFs</kwd>
<kwd>transfer learning</kwd>
<kwd>explainability</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="79"/>
<page-count count="12"/>
<word-count count="9125"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Deep neural networks allow us to learn representations for a variety of computer vision tasks when large amounts of labeled data are available, but are susceptible to a <italic>domain shift</italic>, when applied to unseen data of new domains at test time. Solutions such as further fine-tuning the network on new data, are not always efficient or trivial, and data collection and annotation are expensive and time-consuming processes, setting obstacles to the application and generalization of the existing models to other domains.</p>
<p>Domain adaptation attempts to address these shortcomings, by training a network on labeled data from a single (Pan et al., <xref ref-type="bibr" rid="B51">2010</xref>; Baktashmotlagh et al., <xref ref-type="bibr" rid="B1">2016</xref>; Long et al., <xref ref-type="bibr" rid="B42">2016</xref>) or multiple (Duan et al., <xref ref-type="bibr" rid="B16">2012</xref>; Jhuo et al., <xref ref-type="bibr" rid="B25">2012</xref>; Yang and Hospedales, <xref ref-type="bibr" rid="B74">2014</xref>; Liu et al., <xref ref-type="bibr" rid="B39">2016</xref>; Xu et al., <xref ref-type="bibr" rid="B73">2018</xref>) source domains, and on a related but different target domain, to learn more transferable representations. Since labeled data are often limited and hard to obtain, unsupervised domain adaptation (Long et al., <xref ref-type="bibr" rid="B44">2015</xref>, <xref ref-type="bibr" rid="B41">2018</xref>; Ganin et al., <xref ref-type="bibr" rid="B19">2016</xref>; Sun and Saenko, <xref ref-type="bibr" rid="B64">2016</xref>; Wilson and Cook, <xref ref-type="bibr" rid="B72">2020</xref>) is of most interest, aiming to leverage the few or no labeled samples. A more complex problem is deep domain generalization (Muandet et al., <xref ref-type="bibr" rid="B48">2013</xref>; Ghifary et al., <xref ref-type="bibr" rid="B20">2015</xref>; Li et al., <xref ref-type="bibr" rid="B34">2017</xref>, <xref ref-type="bibr" rid="B37">2018b</xref>), in which the model is completely unaware of the target domain, and does not see any samples from the target distribution during training. These methods have been widely explored for images, but the scarcity of work and applications in videos serves as a motivation for our current approach.</p>
<p>Our paper comes to address the crucial need to build high-quality benchmark video datasets, in multiple domains, to objectively measure the performance of these techniques. This is because well-defined, rich in features, labeled datasets allow for a universal evaluation of the different methods (Ponce et al., <xref ref-type="bibr" rid="B53">2006</xref>; Torralba and Efros, <xref ref-type="bibr" rid="B66">2011</xref>; Russakovsky et al., <xref ref-type="bibr" rid="B57">2015</xref>; Beery et al., <xref ref-type="bibr" rid="B2">2018</xref>; Recht et al., <xref ref-type="bibr" rid="B54">2019</xref>). Given the arduity of real-world data collection and labeling, synthetic data have grown in popularity, as they can be generated in abundance, introducing a substantial domain gap when compared to other domains (Ros et al., <xref ref-type="bibr" rid="B55">2016</xref>; R&#x000F6;ssler et al., <xref ref-type="bibr" rid="B56">2018</xref>; Cruz et al., <xref ref-type="bibr" rid="B12">2020</xref>; Kong et al., <xref ref-type="bibr" rid="B30">2020</xref>; Scheck et al., <xref ref-type="bibr" rid="B60">2020</xref>).</p>
<p>Our focus is on videos, and, more specifically, on Animated GIFs (Eppink, <xref ref-type="bibr" rid="B17">2014</xref>), in which this gap is identified in both space and time (unlike in images, which suffer only from spatial domain shift). Temporal features can be misaligned between domains, which makes the problem more challenging, and significantly under-explored. GIFs are videos that are short in duration, designed to repeat (or re-play), and do not include audio. They typically illustrate a certain action, and have the ability to express a broad spectrum of emotions, aiming at performance of affect and conveyance of cultural knowledge (Miltner and Highfield, <xref ref-type="bibr" rid="B45">2017</xref>). GIFs are created by sampling frames from a video and are extensively used nowadays on the internet, especially in social networks and online communication (Tolins and Samermit, <xref ref-type="bibr" rid="B65">2016</xref>; Jiang et al., <xref ref-type="bibr" rid="B26">2018</xref>). Animated GIFs are synthetically generated and tend to exaggerate or emphasize action motion. In this work, we aim to answer the following questions: <italic>How large is the domain gap between (1) videos and GIFs, and (2) animated and real GIFs?</italic></p>
<p>We propose the first synthetic domain generalization Animated GIFs dataset, <italic>Ani-GIFs</italic>, designed for the task of action recognition in videos. To our knowledge, no other synthetic GIFs dataset exists designed explicitly for spatiotemporal domain generalization, as depicted in <xref ref-type="table" rid="T1">Table 1</xref>. <xref ref-type="fig" rid="F1">Figure 1</xref> presents sample examples from <italic>Ani-GIFs</italic>, and contrasts it with GIFs of the real domain from the Kinetics GIFs dataset. We evaluate domain generalization baselines on <italic>Ani-GIFs</italic> using an I3D action recognition model (Carreira and Zisserman, <xref ref-type="bibr" rid="B5">2018</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Comparing our proposed benchmark to existing ones for spatiotemporal action recognition.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="left"><bold>Domain</bold></th>
<th valign="top" align="center"><bold>Dataset size</bold></th>
<th valign="top" align="center"><bold>Number of classes</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Kinetics-600</td>
<td valign="top" align="left">Real videos</td>
<td valign="top" align="center">70,901</td>
<td valign="top" align="center">600</td>
</tr>
<tr>
<td valign="top" align="left">HMDB-51</td>
<td valign="top" align="left">Real videos</td>
<td valign="top" align="center">6,849</td>
<td valign="top" align="center">51</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="left">Synthetic GIFs</td>
<td valign="top" align="center">17,095</td>
<td valign="top" align="center">536</td>
</tr>
<tr>
<td/>
<td valign="top" align="left"><italic>(animated, cartoon, graphics)</italic></td>
<td/>
</tr>
</tbody>
</table><table-wrap-foot><p>Dataset size represents the number of samples, videos or GIFs, in each dataset. Ani-GIFs is the first synthetic GIFs dataset to our knowledge designed explicitly to study the domain gap between videos vs. GIFs, and animated vs. real GIFs for the task of domain generalization.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>This figure highlights the spatiotemporal domain gap between <italic>Ani-GIFs</italic>, our proposed benchmark dataset, and GIFs of the real domain&#x02014;from the Kinetics dataset (Kay et al., <xref ref-type="bibr" rid="B28">2017</xref>)&#x02014;for three classes: <italic>Bench Pressing, Brushing Teeth</italic>, and <italic>Break Dancing</italic>. This illustrates the domain gap between real vs. animated frames.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-876846-g0001.tif"/>
</fig>
<p>In order to verify the model robustness on our benchmark and the suitability of the dataset for testing domain adaptation and domain generalization methods, we employ the data augmentation approach proposed by Volpi and Murino (<xref ref-type="bibr" rid="B68">2019</xref>) for images and extend it to GIF (video) frames. We define a series of content-preserving frame transformations (e.g.,contrast enhancement, sharpness/color adjustment), which do not alter the content of the frames, but only the way it is presented. Starting with the identity transformation, we apply a set of concatenated data transformations, given as tuples of a specific size, to the training data, in an alternating process of augmenting the samples with a uniformly selected tuple from the set, and training the model to choose the one among those applied which maximizes the model loss, using a random-search algorithm for selection, so as to strengthen our model.</p>
<p>We also extend an explainability-based domain generalization technique initially proposed for images (Zunino et al., <xref ref-type="bibr" rid="B79">2020</xref>) to the spatiotemporal domain. Explainability, i.e.,using the correct evidence for prediction, is utilized to bridge the gap between the real and the synthetic domains. The black-box nature of deep neural network models creates highly non-linear feature representations that make it difficult to understand what causes models to make certain classification decisions. We use the extended saliency-based explainability approach to identify regions in the image that contribute the most to the model&#x00027;s predictions. We leverage these spatiotemporal saliency tubes to guide the model in focusing on image regions, where a particular action is being performed, as opposed to focusing on domain-specific details, that do not necessarily generalize across domains.</p>
<p>To summarize, our contributions are: providing (1) a spatiotemporal dataset, (2) a training and testing setting, (3) a spatiotemporal baseline, (4) an augmentation-based spatiotemporal training strategy, and (5) an explainability-based spatiotemporal training strategy, to enable research addressing the challenging domain generalization problem.</p>
<p>Our paper is organized as follows: First, we discuss the related work on GIF and video datasets, state-of-the-art methods for domain generalization, domain adaptation, data augmentation (Section 2), and explainability. Second, we describe our dataset and the processes of collection and annotation (Section 3). We then analyze the selected baseline methods for the task of action recognition (Section 4) and evaluate the performance presenting the experimental results of our approach (Section 5), before concluding our work (Section 6). Our dataset and baseline implementations will be made publicly available upon acceptance.</p>
</sec>
<sec id="s2">
<title>2. Related work</title>
<sec>
<title>2.1. Domain adaptation</title>
<p>Domain adaptation tackles the problem of domain shift between one or more source domains to a different but related target domain. The case where unlabeled data from the target are available for training is addressed by Unsupervised Domain Adaptation (UDA). UDA methods can be categorized into: divergence-based (Long et al., <xref ref-type="bibr" rid="B43">2017</xref>; Saito et al., <xref ref-type="bibr" rid="B58">2018</xref>), adversarial-based (Ganin et al., <xref ref-type="bibr" rid="B19">2016</xref>; Tzeng et al., <xref ref-type="bibr" rid="B67">2017</xref>; Liu et al., <xref ref-type="bibr" rid="B40">2021</xref>), and reconsrtuction-based (image-level translation) methods (Hoffman et al., <xref ref-type="bibr" rid="B23">2018</xref>; Murez et al., <xref ref-type="bibr" rid="B49">2018</xref>). Divergence-based methods focus on minimizing a divergence criterion between the source and target distributions, like the Maximum Mean Discrepancy (MMD) (Long et al., <xref ref-type="bibr" rid="B43">2017</xref>). Adversarial-based approaches focus on making features from different domains indistinguishable. Semi-Supervised Domain Adaptation (SSDA) addresses the other case where a few target labels are provided. In addition, other image domain adaptation methods can be applied to cross-domain tasks, like domain generalization, UDA, and SSDA (Nam et al., <xref ref-type="bibr" rid="B50">2021</xref>).</p>
</sec>
<sec>
<title>2.2. Video domain adaptation</title>
<p>The problem of domain adaptation in video action recognition is still under-explored, despite the extensive work in this area for image classification and object recognition. Two approaches are introduced by Jamal et al. (<xref ref-type="bibr" rid="B24">2018</xref>), Action Modeling on Latent Subspace (AMLS), which models the videos as points or sequences of points in a latent space, and uses adaptive kernels to learn from source domain points to target domain point sequences, and Deep Adversarial Action Adaptation (DAAA), an adversarial learning framework built to minimize the domain shift. In a most recent work, Chen et al. (<xref ref-type="bibr" rid="B8">2019</xref>), a variety of alignment and learning techniques are proposed or extended to minimize domain discrepancy in videos along the spatial and temporal directions. In Chen et al. (<xref ref-type="bibr" rid="B7">2020</xref>), the authors propose a generative adversarial network, VideoGAN, which uses an X-shape generator to preserve the intra-video consistency during translation of video data across different domains, and a color-based loss, to tune the color distribution of each translated frame and bridge the domain gap.</p>
</sec>
<sec>
<title>2.3. Domain generalization</title>
<p>In domain generalization methods, a relaxed approach is adopted in learning distributions of source domains to generalize to unseen domains, without prior knowledge of the target distribution. Common DG methods can be categorized into: domain agnostic/invariant model learning (Muandet et al., <xref ref-type="bibr" rid="B48">2013</xref>; Ghifary et al., <xref ref-type="bibr" rid="B20">2015</xref>; Dou et al., <xref ref-type="bibr" rid="B15">2019</xref>), self-supervision based (Kim et al., <xref ref-type="bibr" rid="B29">2021</xref>), data-augmentation based (Volpi et al., <xref ref-type="bibr" rid="B69">2018</xref>; Yao et al., <xref ref-type="bibr" rid="B76">2019</xref>), and feature-augmentation based DG methods (Li et al., <xref ref-type="bibr" rid="B32">2021</xref>).</p>
</sec>
<sec>
<title>2.4. Video domain generalization</title>
<p>Several techniques have been introduced to solve this problem with deep models (Muandet et al., <xref ref-type="bibr" rid="B48">2013</xref>; Li et al., <xref ref-type="bibr" rid="B34">2017</xref>, <xref ref-type="bibr" rid="B35">2018a</xref>; Motiian et al., <xref ref-type="bibr" rid="B47">2017</xref>), and with important results for a variety of datasets and data types, but the area is significantly under-explored with respect to video datasets, due to the complexity of entangling spatial and temporal domain shifts. In Yao et al. (<xref ref-type="bibr" rid="B76">2019</xref>, <xref ref-type="bibr" rid="B75">2021</xref>), the only recent prominent work in this area, the authors present the Adversarial Pyramid Network (APN), a network capturing the videos&#x00027; local-, global-, and multi-layer cross-relation features. They also extend an adversarial data augmentation method in Volpi et al. (<xref ref-type="bibr" rid="B69">2018</xref>), ADA, to videos. Their improved approach, namely Robust Adaptive Data Augmentation (RADA), uses robust regularization to improve the robustness of APN to various adversarial perturbations derived from the relational features at multiple levels. Given the reliability of RADA on those relational features, it is intimately coupled with the proposed APN architecture and does not perform as well on other non-hierarchical models, contrary to other model-agnostic domain generalization approaches.</p>
</sec>
<sec>
<title>2.5. Video domain adaptation/generalization datasets</title>
<p>Several existing datasets built for Video analysis tasks are or could be extended to solve the problem of domain shift in action videos, but few new video datasets have been introduced exclusively for the task of domain adaptation or generalization for video action recognition, and are all depicting real actions. The Gameplay dataset (Chen et al., <xref ref-type="bibr" rid="B8">2019</xref>) is a collection of videos of length 1&#x02013;10 s in 91 categories from two video games. Selecting 30 overlapping categories between Gameplay and Kinetics (Kay et al., <xref ref-type="bibr" rid="B28">2017</xref>; Carreira et al., <xref ref-type="bibr" rid="B6">2018</xref>), the authors create the Kinetics-Gameplay dataset, observing a significant domain shift in the distributions of virtual and real data. In the same work, all relevant and overlapping categories between existing video datasets UCF101 (Soomro et al., <xref ref-type="bibr" rid="B63">2012</xref>) and HMDB51 (Kuehne et al., <xref ref-type="bibr" rid="B31">2011</xref>) are combined in UCF-HMDB<sub>full</sub>, a large-scale collection of videos of length 1&#x02013;33 s in 12 classes, used in evaluating several state-of-the-art video domain adaptation methods (Ganin and Lempitsky, <xref ref-type="bibr" rid="B18">2014</xref>; Long et al., <xref ref-type="bibr" rid="B43">2017</xref>; Li et al., <xref ref-type="bibr" rid="B38">2018c</xref>; Saito et al., <xref ref-type="bibr" rid="B58">2018</xref>). For domain generalization, Yao et al. (<xref ref-type="bibr" rid="B76">2019</xref>, <xref ref-type="bibr" rid="B75">2021</xref>) propose four video domain generalization benchmarks, UCF-HMDB, Something-Something, PKU-MMD, and NTU, built from existing action recognition videos, in which they divide the source and target domains according to different datasets, consequences of actions, and camera views, to test their method&#x00027;s performance. In parallel, datasets with a focus on more specific tasks such as autonomous driving (Yu et al., <xref ref-type="bibr" rid="B77">2018</xref>) and medical diagnosis (Cheplygina et al., <xref ref-type="bibr" rid="B10">2017</xref>) have been introduced, allowing for domain adaptation evaluation in a variety of sub-domains.</p>
</sec>
<sec>
<title>2.6. GIF datasets and analysis techniques</title>
<p>There is an abundance of GIF datasets collected and available in the literature. TGIF (Li et al., <xref ref-type="bibr" rid="B36">2016</xref>) is a dataset of 100K animated GIFs from Tumblr and 120K natural language descriptions obtained <italic>via</italic> crowdsourcing, serving as a benchmark for the task of visual content captioning, namely in the generation of natural language descriptions for animated GIFs or video clips. In Vid2GIF (Gygli et al., <xref ref-type="bibr" rid="B22">2016</xref>), a robust framework, RankNet, is proposed, to learn the content in videos most frequently selected for creating popular animated GIFs, and produce a ranked list of segments according to their suitability, generalizing this ability to other tasks such as video highlight detection. To this purpose, a dataset of 120K user-generated animated GIFs with their corresponding video sources is collected, that is one to two orders of magnitude larger than existing datasets in video highlight detection. GIF Super-Resolution (Wang et al.) is an approach proposed to tackle the problem of slow download speed of GIFs, by using the first and last high-resolution frames of a GIF and a low-resolution representation of it, to reconstruct a GIF easier to process. To this purpose, the authors create GIFSR, a dataset of 1,000 GIFs in 5 categories: Emotion, Action, Scene, Animation, and Animal. In GIFGIF&#x0002B; (Chen et al., <xref ref-type="bibr" rid="B9">2017</xref>), an emotions GIF dataset is introduced, consisting of 23,544 GIFs over 17 emotion categories, as the authors propose a novel method for animated GIFs collection, to explore the problem of automatic analysis of emotions in GIFs. Similarly, in Jou et al. (<xref ref-type="bibr" rid="B27">2014</xref>), 4,000 GIFs are collected, with scores for 17 discrete emotions, and are used in a computational analysis and evaluation of emotion prediction on animated GIFs. However, all these datasets were designed to be used for tasks other than domain adaptation or generalization.</p>
</sec>
<sec>
<title>2.7. Data augmentation</title>
<p>Data augmentation is widely used as a model domain generalization improvement technique in computer vision, to obtain more information from the training dataset, and reduce the gap between this and the unseen validation set, preventing the model from performing poorly in evaluation (Shorten and Khoshgoftaar, <xref ref-type="bibr" rid="B62">2019</xref>). When applied on image datasets, data augmentation techniques exploit the spatial properties of the data, and can range from image manipulations, such as geometric or color transformations, rotation, or blurring (Ciregan et al., <xref ref-type="bibr" rid="B11">2012</xref>; Wan et al., <xref ref-type="bibr" rid="B70">2013</xref>; Sato et al., <xref ref-type="bibr" rid="B59">2015</xref>), to feature space augmentation (DeVries and Taylor, <xref ref-type="bibr" rid="B13">2017</xref>), adversarial training techniques (Moosavi-Dezfooli et al., <xref ref-type="bibr" rid="B46">2016</xref>; Volpi et al., <xref ref-type="bibr" rid="B69">2018</xref>; Zajac et al., <xref ref-type="bibr" rid="B78">2019</xref>), and GAN-based approaches (Bowles et al., <xref ref-type="bibr" rid="B3">2018</xref>). Expanding the objective to videos, the proposed methods augment the dataset in both spatial and temporal dimensions, in domain generalization approaches for tasks such as semantic segmentation (Budvytis et al., <xref ref-type="bibr" rid="B4">2017</xref>) and video action recognition (Yao et al., <xref ref-type="bibr" rid="B76">2019</xref>, <xref ref-type="bibr" rid="B75">2021</xref>).</p>
</sec>
<sec>
<title>2.8. Explainability</title>
<p>Explainability techniques were initially developed as a diagnostic tool to visualize and explain a model&#x00027;s behavior. GradCAM (Selvaraju et al., <xref ref-type="bibr" rid="B61">2017</xref>) is a gradient-based approach that uses gradients flowing into a target layer to compute coarse localization maps at that layer. In recent work on explainability, Zunino et al. (<xref ref-type="bibr" rid="B79">2020</xref>) use an explainability-based training strategy on images to boost model performance. We extend this to the spatiotemporal domain by computing saliency tubes using GradCAM (Selvaraju et al., <xref ref-type="bibr" rid="B61">2017</xref>) in space and time.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Our dataset: <italic>Ani-GIFs</italic></title>
<p>In this section, we introduce our benchmark dataset together with conducted collection and filtration procedures. Our dataset focuses on actions occurring in Animated GIFs, in mirror classes of the Kinetics-600 dataset.</p>
<p>We propose <italic>Ani-GIFs</italic> as a domain generalization benchmark, acting as the target domain in a domain generalization approach from a <bold>real</bold> source domain of actions performed by human characters, to a <bold>synthetic</bold> target domain of actions performed by animated/cartoon/graphical characters. As the real domain dataset, we are using the GIFs from the existing Kinetics dataset (Kay et al., <xref ref-type="bibr" rid="B28">2017</xref>), and we collect the GIFs in the synthetic domain, forming the proposed dataset, <italic>Ani-GIFs</italic>.</p>
<sec>
<title>3.1. Data collection</title>
<p>We created the <italic>Ani-GIFs</italic> dataset by collecting animated GIFs using the Bing search engine. For each action class in the Kinetics-600 dataset, we set up an automated script to search and download datapoints. Three search keywords were used, the first being &#x0201C;animated&#x0201D; or &#x0201C;cartoon&#x0201D; or &#x0201C;graphics&#x0201D;, the second being the action class, and the third being &#x0201C;GIF&#x0201D;. For example, for the action class &#x0201C;Applauding&#x0201D;, we performed three separate searches: &#x0201C;animated Applauding gif&#x0201D;, &#x0201C;cartoon Applauding gif&#x0201D;, and &#x0201C;graphics Applauding gif&#x0201D;. The keywords &#x0201C;animated&#x0201D;, &#x0201C;graphics&#x0201D;, and &#x0201C;cartoon&#x0201D; were used synonymously as means of maximizing the number of retrievals from the search engine. We then collected GIFs from each separately. Each of the three collection processes, for all 600 classes, took approximately 100 h to complete.</p>
</sec>
<sec>
<title>3.2. Filtration and annotation</title>
<p>After collecting the animated GIFs, we performed extensive filtering. The first stage of filtering was combining search results of animated, cartoons, and graphics and removing duplicates. The second stage was performed manually by four graduate students. This stage involved ensuring a downloaded video was indeed: (1) a GIF, (2) performed by an animated, cartoon or graphics character, and (3) depicting the exact class action in Kinetics-600. Annotation was performed based on the action classes only, and not on the type of synthetic domain. <xref ref-type="fig" rid="F2">Figure 2</xref> provides examples of collected animated GIFs which were rejected or accepted during the filtering process.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>In this figure, we can see that first GIFs, in <italic>Blowing out candles</italic> and <italic>Bending metal</italic> classes, were rejected as the actions are not performed by any character. We also rejected GIFs in the <italic>Shopping</italic> class, as the action was not relevant to the class (i.e.,no shopping action was observed).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-876846-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.3. Correspondence with Kinetics-600</title>
<p><italic>Ani-GIFs</italic> is designed to have one-to-one correspondence with the classes of Kinetics-600, to act as a domain generalization benchmark. 60 classes from Kinetics-600 did not have corresponding animated GIFs after filtration. Examples for such classes that do not typically have associated animated GIFs, are: <italic>Arranging Flowers, Changing Oil, Curling Hair, Feeding Goats, Making Jewellery, Sharpening Knives, Putting On Sari</italic>. Therefore, the resulting <italic>Ani-GIFs</italic> dataset has 536 classes, and 17,095 animated GIFs in total, all intersecting with Kinetics-600. <xref ref-type="fig" rid="F3">Figure 3</xref> shows the number of GIF samples per class in the <italic>Ani-GIFs</italic> dataset for the forty top-frequency classes.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The 40 most frequent classes after filtration of the <italic>Ani-GIFs</italic> dataset. Classes that have the highest frequency belong to actions with a large number of associated GIFs, e.g.,common actions and emotions. We choose these forty classes as a subset for GIFs domain adaptation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-876846-g0003.tif"/>
</fig>
</sec>
<sec>
<title>3.4. Subset for domain adaptation</title>
<p>While our dataset is designed for the task of GIF domain generalization, we identify a subset of <italic>Ani-GIFs</italic> for the task of GIF domain adaptation for action recognition. The subset consists of the forty classes having the highest frequency. This would allow for standard testing of domain adaptation, i.e.,from Real to Animated GIFs and vice versa.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Spatiotemporal domain generalization</title>
<p>In this work we address the challenging problem of single-source domain generalization for spatiotemporal GIFs. At training time, we only have access to a single source domain, and at test time we have access to a different target domain that is unseen at training time. We focus on the real videos/GIFs source domain and the animated GIFs target domain. While the problem of attributing an action to an animated spatiotemporal progression is trivial for humans, it is a significantly challenging task for machine learning models that have only been trained on real video data. The gap between the two domains in this problem setting is large. The two domains exhibit significant variations in color templates, as animated GIFs tend to only have a few colors in all frames, while real videos or GIFs have a significantly richer color template. Moreover, animated GIFs tend to have a smaller level of detail, in contrast to real videos or GIFs. At the same time, animated GIFs exhibit a faster speed for actions than real videos or GIFs, i.e., while the difference in motion between subsequent frames in real videos is usually small even after sub-sampling, the difference between subsequent frames in GIFs is significantly larger. We demonstrate how large this domain gap is experimentally in Section 5.</p>
<p>To reduce this huge domain gap, we use a GIF version of the Kinetics dataset&#x02014;Kinetics GIFs (Kay et al., <xref ref-type="bibr" rid="B28">2017</xref>)&#x02014;as the source domain in our data augmentation baseline experiments. Samples in Kinetics GIFs are GIFs produced from original Kinetics videos, which have a fixed length of 40 frames and a significantly smaller resolution, typically of 400 by 400 pixels. After training the model on Kinetics GIFs we evaluate it on <italic>Ani-GIFs</italic> to obtain a baseline performance, that is then compared to applying domain generalization techniques.</p>
<p>We also use the AVA-Kinetics Localized Human Actions Video Dataset (Li et al., <xref ref-type="bibr" rid="B33">2020</xref>) to extend the explainable training strategy of Zunino et al. (<xref ref-type="bibr" rid="B79">2020</xref>) on images to the spatiotemporal domain to achieve better evidence for domain generalization. The dataset is an extension of the Kinetics dataset with AVA-style bounding boxes and atomic actions, which makes it suitable as a train set in our explainability-based approach. AVA-Kinetics has more than 230k clips labeled with one of 80 AVA action classes, which are manually mapped to their corresponding top related Kinetics classes.</p>
<sec>
<title>4.1. Data augmentation approach</title>
<p>We extend the work of Volpi and Murino (<xref ref-type="bibr" rid="B68">2019</xref>) on images and develop a spatiotemporal data augmentation approach for animated GIFs. Data augmentation is a very powerful technique to create additional representations and increase the generalization ability of a model to domains that are unseen at training time. We artificially inflate the dataset by applying transformations in space and time. Following Volpi and Murino (<xref ref-type="bibr" rid="B68">2019</xref>), we apply a set of image transformations <inline-formula><mml:math id="M1"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> from the Python library Pillow, to compute the augmented versions of each GIF. We consider transformation tuples of length four, i.e., four transformations are applied concurrently to a GIF of the training set for every augmentation. The pool of transformations is (intensity in parenthesis): auto-contrast (20), sharpness (20), brightness (20), color (20), contrast (20), gray scale conversion (1), R-channel enhancer (30), G-channel enhancer (30), B-channel enhancer (30), solarize (20).</p>
<p>Starting with a model pre-trained on the Kinetics-600 dataset, and the transformations set <inline-formula><mml:math id="M2"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> containing only identity transformations, we perform a fine-tuning process to identify the tuple of transformations that the model is most vulnerable to. Vulnerability of the model is defined to be the tuple of transformations that leads to the highest value of cross-entropy loss when applied to the input batches. At every iteration of the training process, we randomly sample a tuple from our set of vulnerable transformations <inline-formula><mml:math id="M3"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, and apply those to our training batches with their associated intensity values. We train our model using Stochastic Gradient Descent to minimize the cross-entropy loss. The transformations set is updated every 200 training iterations, using random search.</p>
<p>The identification process targets adding one tuple of transformations to the set of known vulnerable transformations <inline-formula><mml:math id="M4"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> using a random search approach. At every iteration of the random search, four transformations are randomly sampled with repetition, from the pool of transformations at random intensity values to create a tuple. While extending the augmentation approach to account for temporal shifts, all four transformations of the tuple are applied to all frames of the input batches. This ensures that the same transformation is performed on all frames of the video to obtain a single augmented instance. The vulnerability of the model to this tuple of transformation is then determined by evaluating the cross-entropy loss. At the end of 50 iterations of the searching process, the tuple of transformations that led to the highest cross-entropy loss is identified and added to our set of vulnerable transformations <inline-formula><mml:math id="M5"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, along with its intensity value. In subsequent iterations of the standard training process, this identified tuple of transformations is available to be randomly sampled from our set <inline-formula><mml:math id="M6"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> and applied to the training batches for training with the Adam optimizer. In <xref ref-type="fig" rid="F4">Figure 4</xref>, we show images after different tuples of transformations applied to frames that were equally sampled from a video in class &#x0201C;Yoga&#x0201D; taken from the Kinetics GIFs dataset. Transformations are applied after Batch Normalization.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>This figure presents frames sampled equally from a video in class &#x0201C;Yoga&#x0201D;. The first set of frames belongs to the Kinetics GIFs dataset, and is followed by the frames after Batch Normalization is applied. The subsequent set of images depicts the frames after tuples of transformations, chosen by the random search approach, are applied to them.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-876846-g0004.tif"/>
</fig>
</sec>
<sec>
<title>4.2. Explainability approach</title>
<p>We extend and apply a saliency-based spatiotemporal explainability approach (Zunino et al., <xref ref-type="bibr" rid="B79">2020</xref>) on our dataset. At training time, saliency maps for the ground-truth class are periodically computed as saliency tubes in space and time. As training progresses, we have access to these regions as bounding box co-ordinates for the input batch. Saliency maps are computed using the GradCAM (Selvaraju et al., <xref ref-type="bibr" rid="B61">2017</xref>) algorithm after the last block of the feature extactor layer <italic>l</italic> of the model. We estimate saliency on the last spatial layer as it models higher level spatial patterns, that are most correlated with the target label. If the peak saliency does not fall within the ground-truth region, we enforce that by utilizing a multiplicative binary 3D-mask (saliency tube) that is applied to the forward activations of layer <italic>l</italic>. This mask contains a value of 1 for pixels that lie within the spatiotemporal region of interest and 0 otherwise. We run the saliency estimation periodically every 200 batches, and train using the Adam optimizer.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Experiments</title>
<p>In this section, we start by experimentally demonstrating the huge domain gap between real videos vs. GIFs of the same videos, and real videos vs. animated GIFs. We then demonstrate how spatiotemporal domain generalization can reduce the gap in the latter scenario. Results of this first study are shown in <xref ref-type="table" rid="T2">Table 2</xref>. Furthermore, we conduct a second study to emphasize the effectiveness of our spatiotemporal and explainability-based approaches over baseline domain generalization and report results in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Study 1 results. Top-1 and top-5 test accuracies of our baseline algorithm are given, from various training on different testing domains.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center"><bold>Source (train) domain</bold></th>
<th valign="top" align="center"><bold>Target (test) domain</bold></th>
<th valign="top" align="center"><bold>Spatiotemporal augmentation</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Test accuracy (%)</bold></th>
</tr>
<tr>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td/>
<td/>
<td valign="top" align="center"><bold>Top-1</bold></td>
<td valign="top" align="center"><bold>Top-5</bold></td>
</tr> <tr>
<td valign="top" align="left">Kinetics</td>
<td valign="top" align="left">Kinetics</td>
<td valign="top" align="center">&#x02718;</td>
<td valign="top" align="center">71.70</td>
<td valign="top" align="center">90.40</td>
</tr>
<tr>
<td valign="top" align="left">Kinetics</td>
<td valign="top" align="left">Kinetics GIFs</td>
<td valign="top" align="center">&#x02718;</td>
<td valign="top" align="center">21.12</td>
<td valign="top" align="center">40.86</td>
</tr>
<tr>
<td valign="top" align="left">Kinetics GIFs</td>
<td valign="top" align="left">Kinetics GIFs</td>
<td valign="top" align="center">&#x02718;</td>
<td valign="top" align="center">23.10</td>
<td valign="top" align="center">46.28</td>
</tr>
<tr>
<td valign="top" align="left">Kinetics GIFs</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="center">&#x02718;</td>
<td valign="top" align="center">1.95</td>
<td valign="top" align="center">6.09</td>
</tr>
<tr>
<td valign="top" align="left">Kinetics GIFs</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="center">&#x02714;</td>
<td valign="top" align="center">2.91</td>
<td valign="top" align="center">8.44</td>
</tr>
</tbody>
</table><table-wrap-foot><p>The difference in the reported accuracies between rows 1 and 2 demonstrates the existing domain gap from Kinetics to Kinetics GIFs, and in rows 3 and 4 the domain shift between Kinetics GIFs and <italic>Ani-GIFs</italic>, with the latter dataset used in its entirety for measuring accuracy while testing. The increase from row 4 to 5 shows the gain in accuracy yielded by extending and applying the spatiotemporal data augmentation algorithm for domain generalization on the training dataset, Kinetics GIFs.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Study 2 results. Top-1 and top-5 test accuracies of our baseline algorithms are given, from fine tuning on 12 classes of the AVA-Kinetics dataset, on the Ani-GIFs dataset, the testing domain.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Source (train) domain</bold></th>
<th valign="top" align="left"><bold>Target (test) domain</bold></th>
<th valign="top" align="left"><bold>Approach</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Test accuracy (%)</bold></th>
</tr>
<tr>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td/>
<td/>
<td valign="top" align="center"><bold>Top-1</bold></td>
<td valign="top" align="center"><bold>Top-5</bold></td>
</tr> <tr>
<td valign="top" align="left">AVA-Kinetics</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="left">Random Augmentation</td>
<td valign="top" align="center">8.56</td>
<td valign="top" align="center">34.47</td>
</tr>
<tr>
<td valign="top" align="left">AVA-Kinetics</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="left">Spatiotemporal Augmentation</td>
<td valign="top" align="center">11.88</td>
<td valign="top" align="center">42.74</td>
</tr>
<tr>
<td valign="top" align="left">AVA-Kinetics</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="left">RADA</td>
<td valign="top" align="center">12.55</td>
<td valign="top" align="center">34.63</td>
</tr>
<tr>
<td valign="top" align="left">AVA-Kinetics</td>
<td valign="top" align="left"><italic>Ani-GIFs</italic></td>
<td valign="top" align="left">Explainability</td>
<td valign="top" align="center">17.31</td>
<td valign="top" align="center">58.78</td>
</tr>
</tbody>
</table><table-wrap-foot><p>The difference in the reported accuracies between rows 1 and 2 shows that our spatiotemporal augmentation approach outperforms random augmentation on the training dataset, AVA-Kinetics. Row 3 shows that spatiotemporal augmentation outperforms Robust Adaptive Data Augmentation (RADA) in the top-5 test accuracy. Furthermore, we show that our explainability approach in row 4 outperforms all reported augmentation baselines, verifying the effectiveness of explainability for domain generalization.</p>
</table-wrap-foot>
</table-wrap>
<sec>
<title>5.1. Datasets</title>
<p>In the domain gap study, which we call Study 1 (<xref ref-type="table" rid="T2">Table 2</xref>), we use the Kinetics-600 dataset as the source domain for both training and fine-tuning, and the (<italic>real</italic>) Kinetics GIFs vs. <italic>Ani-GIFs</italic> datasets as target domains. As for the second study (<xref ref-type="table" rid="T3">Table 3</xref>), we use AVA-Kinetics as the source domain to fine-tune a baseline model pre-trained on Kinetics-600 using our spatiotemporal domain generalization algorithms. We then use <italic>Ani-GIFs</italic> as the target domain to compare the performance of our algorithms with baseline domain generalization approaches.</p>
</sec>
<sec>
<title>5.2. Experimental setup</title>
<p>We choose the I3D model architecture (Carreira and Zisserman, <xref ref-type="bibr" rid="B5">2018</xref>) as the first baseline for the spatiotemporal training and testing of our videos and animated GIFs, because of its increased transferability and its ability to capture a fine-grained temporal structure of actions. This model builds upon state-of-the-art image classification architectures, expanding their filters and pooling kernels (and optionally their parameters) into 3D, hence learning seamless spatiotemporal features from videos while leveraging successful ImageNet architecture designs and their parameters. More specifically, starting from a 2D architecture, all the filters and pooling kernels are inflated with an additional temporal dimension. The model is trained on 64-frame video snippets at 25 frames per second, processing all video frames at test time, and learning high temporal resolution features.</p>
<p>While training, we perform certain preprocessing on the input frames that aims to improve quality by suppressing unwanted noise in the frames, and enhancing important features. Animated GIFs are preprocessed frame-wise&#x02014;each frame was rescaled such that its shorter side has length of 224 pixels. Realignment was followed by center cropping, resulting in a frame of size 224 by 224. Hence, during training, each training sample has a fixed size of (40, 224, 224, 3). The number of frames in <italic>Ani-GIFs</italic> samples may though vary, so we upsampled frames for animated GIFs that had less than 9 frames, and subsampled frames of animated GIFs that had more than 60 frames, such that the chosen frames have equal spacing in time. All values were rescaled to the [-1, 1] interval.</p>
<p>The models were trained on four <italic>Nvidia TITAN V</italic> GPUs for 60 epochs with a batch size of 32 samples. We start with an I3D model that is pre-trained on Kinetics videos (Piergiovanni, <xref ref-type="bibr" rid="B52">2018</xref>).</p>
<p>In Study 1, we start with the model trained on Kinetics GIFs using the I3D model architecture, and fine-tune it further with the random search approach and Adam optimizer (Diederik P. Kingma, <xref ref-type="bibr" rid="B14">2014</xref>). The fine-tuning process was performed on three <italic>Nvidia TITAN V</italic> GPUs, in batches of eight animated GIFs. The model was tuned for 600 random search iterations. We used the same upsampling and subsampling criteria as in the training process, which resulted in every animated GIF having a fixed shape of (40, 224, 224, 3). Every frame was similarly preprocessed with realignment, center cropping and rescaling. In order to augment the animated GIFs, we made sure the same transformations are applied to the entire batch of input GIFs, resulting in a batch with a shape of (8*40, 224, 224, 3).</p>
<p>The second study uses AVA-Kinetics for fine-tuning, to compare our spatiotemporal augmentation and explainability algorithms against a baseline with random data augmentations and Yao et al. (<xref ref-type="bibr" rid="B75">2021</xref>)&#x00027;s Robust Adaptive Data Augmentation (RADA). While incorporating the saliency-based approach for our training, we start with a model that is pre-trained on the Kinetics dataset. This pre-trained model uses the same I3D architecture as in Study 1 and is further tuned on the AVA Kinetics dataset using the GradCAM saliency algorithm. We filter the AVA Kinetics dataset and use only 12 classes that are a one-to-one mapping to classes in Ani-GIFs. To ensure a balanced training setting, we sample data from these 12 classes such that every class contains 2000 training datapoints. The fine-tuning process was performed on three <italic>Nvidia TITAN V</italic> GPUs, in batches of 8 videos. We run the saliency estimation every 200 batches. We also apply RADA (Yao et al., <xref ref-type="bibr" rid="B75">2021</xref>) at the same iteration frequency to ensure a fair comparison, and follow the authors&#x00027; implementation which is available online. We maintain the same data preprocessing as in our augmentation experiment and make sure that the entire batch undergoes the same preprocessing which results in batches with the shape (8*40,224,224,3). All models were trained and fine-tuned using the Adam optimizer with the following hyperparameters: learning rate = 10<sup>&#x02212;4</sup>, &#x003B2;<sub>1</sub> &#x0003D; 0.9, and &#x003B2;<sub>2</sub> &#x0003D; 0.999.</p>
</sec>
<sec>
<title>5.3. Experimental results</title>
<p>The results of our two studies are presented in <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref>. In <xref ref-type="table" rid="T2">Table 2</xref>, we begin with two experiments demonstrating the domain gap within videos, and also between videos and GIFs, both from the same (real) domain. The first row of the table reports the results of training and testing processes on Kinetics-600 real videos (Kay et al., <xref ref-type="bibr" rid="B28">2017</xref>), with a 71.7% top-1 accuracy, and the second row reports the outcome of testing the same model on the GIFs version of the Kinetics-600 dataset (Gituma, <xref ref-type="bibr" rid="B21">2019</xref>), similarly in the real domain.</p>
<p>We mark the significant accuracy drop, to a 21.12% top-1 accuracy, which we can attribute to the frame sampling process in GIFs, or the difference in GIF frames&#x00027; speed, in comparison to videos, between the source and target domains, in the second variation of the model application. We then train a model on the Kinetics GIFs dataset (Gituma, <xref ref-type="bibr" rid="B21">2019</xref>) and test on GIFs from the same dataset and, hence, domain. This, as we can observe, increases the model performance to a higher top-1 accuracy of 23.1%, compared to the previous experiment, as expected when training and testing within the same domain. This result is given in row 3 of <xref ref-type="table" rid="T2">Table 2</xref>, while rows 4 and 5 show how our domain generalization baseline performs, when trained on the Kinetics GIFs dataset and tested on the <italic>Ani-GIFs</italic> dataset, with and without data augmentation. We can see how our proposed data augmentation approach gives an absolute improvement of 0.96% in the top-1 accuracy and 2.35% in the top-5 accuracy and can serve as an initial baseline for <italic>Ani-GIFs</italic>.</p>
<p>We further compare our spatiotemporal augmentation approach against random and adversarial data augmentation and demonstrate the results in <xref ref-type="table" rid="T3">Table 3</xref>. We use the same I3D model architecture as in Study 1, pre-trained on Kinetics-600 and fine-tuned on AVA-Kinetics using the same fine tuning process as earlier. We show in rows 1 and 2 of <xref ref-type="table" rid="T3">Table 3</xref> that our spatiotemporal approach outperforms random augmentation by 3.32% and 8.27% in the top-1 and top-5 accuracies respectively. In rows 2 and 3, while RADA (Yao et al., <xref ref-type="bibr" rid="B75">2021</xref>) gets a slight top-1 accuracy improvement of less than 1% over our approach, we show that the latter none-the-less outperforms it by 8.11% in the top-5 accuracy. Furthermore, our proposed explainability-based approach outperforms all three augmentation baselines in both the top-1 and top-5 accuracies, which emphasizes the effectiveness of explainability in domain generalization.</p>
</sec>
<sec>
<title>5.4. Explainability for spatiotemporal domain generalization</title>
<p>We utilize explainability as a visualization tool for evaluating the generalization capability of models for domain generalization on spatiotemporal data. We show that a model is able to generalize an action across various domains, in our case real vs. animated GIFs for the task of action recognition.</p>
<p>Typically, classification accuracy is reported to summarize the recognition capability of models on classification datasets. However, it alone is not indicative as to whether the models have learnt to generalize an action across the source and target domains. For example, it may be that the model is correctly classifying a sample based on the wrong cues. <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates examples of poor generalization ability of the baseline model from the source, AVA-Kinetics, to the target domain, <italic>Ani-GIFs</italic>, compared against the saliency model trained with domain adaptation using the explainability approach. We use GradCAM to visualize saliency on different GIFs from the <italic>Ani-GIFs</italic> dataset. In addition, we report evaluation results in <xref ref-type="table" rid="T3">Table 3</xref> of our explainability approach on the Ani-GIFs dataset, for an I3D model pre-trained on Kinetics and fine-tuned on 12 classes of AVA-Kinetics using saliency. The results show how using explainability for domain generalization outperforms all three augmentation baselines by a maximum of 8.75% in the top-1 accuracy and 24.31% in the top-5 accuracy.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>This figure presents the visualizations of predictions of four different GIFs from the AniGIFs dataset (GIF frames equally sampled) demonstrating evidence of the baseline model without domain generalization and the model trained with the explainability-based domain generalization approach (Section 4.2). The top two examples show that for some GIFs the explainability approach boosts both the model accuracy and generalization ability. And the bottom two examples show that even when making a correct prediction, the baseline model does not use discriminative cues to make that prediction. In contrast, the model trained with the explainability-based domain generalization approach accurately highlights the correct action-specific cues.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-876846-g0005.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusion</title>
<p>We introduce the first domain generalization GIFs Dataset, <italic>Ani-GIFs</italic>, designed for the task of video action recognition in a synthetic domain, which consists of 536 classes, mirroring the classes in the real domain of the Kinetics GIFs dataset. We discuss the collection and filtration process, provide the results of evaluating a domain generalization baseline, trained on Kinetics GIFs, and an explainability-based domain generalization model, trained on the AVA-Kinetics Localized Human Actions Video Dataset, and also evaluate the baselines after extending and applying an existing image data augmentation technique. Our results show that it is evident that the domain gap in the temporal space is a great challenge. Current domain generalization techniques for images, when extended to Videos/GIFs, showcase a performance improvement, although small enough to highlight the need for better methods tailored toward the temporal dimension. Our dataset serves as a benchmark to catalyze the development and testing of state-of-the-art domain generalization techniques tailored for videos and animated GIFs, and as a motivation for further exploration and enrichment of the existing GIF datasets, to span different domains for the tasks of domain adaptation and domain generalization.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The urls of the full set of unfiltered images of our dataset supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s8">
<title>Author contributions</title>
<p>SS, KS, and SB: study conception and design. SM, SJ, IT, and AM: data collection and filtration. SM, SJ, IT, AM, and DL: implementation and model training. AM and DL: analysis and interpretation of results. SM, SJ, IT, AM, DL, SS, KS, and SB: draft manuscript preparation. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>

<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baktashmotlagh</surname> <given-names>M.</given-names></name> <name><surname>Harandi</surname> <given-names>M.</given-names></name> <name><surname>Salzmann</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Distribution-matching embedding for visual domain adaptation</article-title>. <source>J. Mach. Learn. Res</source>. <volume>17</volume>, <fpage>3760</fpage>&#x02013;<lpage>3789</lpage>. <pub-id pub-id-type="doi">10.5555/2946645.3007061</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Beery</surname> <given-names>S.</given-names></name> <name><surname>Van Horn</surname> <given-names>G.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Recognition in terra incognita,</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>456</fpage>&#x02013;<lpage>473</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01270-0_28</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bowles</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Guerrero</surname> <given-names>R.</given-names></name> <name><surname>Bentley</surname> <given-names>P.</given-names></name> <name><surname>Gunn</surname> <given-names>R.</given-names></name> <name><surname>Hammers</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Gan augmentation: Augmenting training data using generative adversarial networks</article-title>. <source>arXiv preprint arXiv:1810.10863</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1810.10863</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Budvytis</surname> <given-names>I.</given-names></name> <name><surname>Sauer</surname> <given-names>P.</given-names></name> <name><surname>Roddick</surname> <given-names>T.</given-names></name> <name><surname>Breen</surname> <given-names>K.</given-names></name> <name><surname>Cipolla</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Large scale labelled video data augmentation for semantic segmentation in driving scenarios,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision Workshops</source> (<publisher-loc>Venice</publisher-loc>), <fpage>230</fpage>&#x02013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1109/ICCVW.2017.36</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Carreira</surname> <given-names>J.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Quo vadis, action recognition? A new model and the kinetics dataset</article-title>. <source>arXiv preprint arXiv:1705.07750</source>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.502</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Carreira</surname> <given-names>J.</given-names></name> <name><surname>Noland</surname> <given-names>E.</given-names></name> <name><surname>Banki-Horvath</surname> <given-names>A.</given-names></name> <name><surname>Hillier</surname> <given-names>C.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>A short note about kinetics-600</article-title>. <source>arXiv preprint arXiv:1808.01340</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1808.01340</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Ma</surname> <given-names>K.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Generative adversarial networks for video-to-video domain adaptation</article-title>. <source>arXiv preprint arXiv:2004.08058</source>. <pub-id pub-id-type="doi">10.1609/aaai.v34i04.5750</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>M.-H.</given-names></name> <name><surname>Kira</surname> <given-names>Z.</given-names></name> <name><surname>AlRegib</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>Temporal attentive alignment for video domain adaptation</article-title>. <source>arXiv preprint arXiv:1905.10861</source>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00642</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Rudovic</surname> <given-names>O. O.</given-names></name> <name><surname>Picard</surname> <given-names>R. W.</given-names></name></person-group> (<year>2017</year>). <article-title>Gifgif&#x0002B;: Collecting emotional animated gifs with clustered multi-task learning,</article-title> in <source>2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>510</fpage>&#x02013;<lpage>517</lpage>. <pub-id pub-id-type="doi">10.1109/ACII.2017.8273647</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheplygina</surname> <given-names>V.</given-names></name> <name><surname>Pena</surname> <given-names>I. P.</given-names></name> <name><surname>Pedersen</surname> <given-names>J. H.</given-names></name> <name><surname>Lynch</surname> <given-names>D. A.</given-names></name> <name><surname>S&#x000F8;rensen</surname> <given-names>L.</given-names></name> <name><surname>de Bruijne</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Transfer learning for multicenter classification of chronic obstructive pulmonary disease</article-title>. <source>IEEE J. Biomed. Health Inform</source>. <volume>22</volume>, <fpage>1486</fpage>&#x02013;<lpage>1496</lpage>. <pub-id pub-id-type="doi">10.1109/JBHI.2017.2769800</pub-id><pub-id pub-id-type="pmid">29990220</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ciregan</surname> <given-names>D.</given-names></name> <name><surname>Meier</surname> <given-names>U.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Multi-column deep neural networks for image classification,</article-title> in <source>2012 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Providence, RI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3642</fpage>&#x02013;<lpage>3649</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2012.6248110</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cruz</surname> <given-names>S. D. D.</given-names></name> <name><surname>Wasenmuller</surname> <given-names>O.</given-names></name> <name><surname>Beise</surname> <given-names>H.-P.</given-names></name> <name><surname>Stifter</surname> <given-names>T.</given-names></name> <name><surname>Stricker</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>SVIRO: synthetic vehicle interior rear seat occupancy dataset and benchmark,</article-title> in <source>The IEEE Winter Conference on Applications of Computer Vision</source> (<publisher-loc>Snowmass, CO</publisher-loc>), <fpage>973</fpage>&#x02013;<lpage>982</lpage>. <pub-id pub-id-type="doi">10.1109/WACV45572.2020.9093315</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>DeVries</surname> <given-names>T.</given-names></name> <name><surname>Taylor</surname> <given-names>G. W.</given-names></name></person-group> (<year>2017</year>). <article-title>Dataset augmentation in feature space</article-title>. <source>arXiv preprint arXiv:1702.05538</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1702.05538</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Diederik</surname> <given-names>P.</given-names></name> <name><surname>Kingma</surname> <given-names>J. B.</given-names></name></person-group> (<year>2014</year>). <article-title>ADAM: a method for stochastic optimization</article-title>. <source>arXiv preprint arXiv</source>:1412.6980. <pub-id pub-id-type="doi">10.48550/arXiv.1412.6980</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dou</surname> <given-names>Q.</given-names></name> <name><surname>Coelho de Castro</surname> <given-names>D.</given-names></name> <name><surname>Kamnitsas</surname> <given-names>K.</given-names></name> <name><surname>Glocker</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>Domain generalization via model-agnostic learning of semantic features,</article-title> in <source>Annual Conference on Neural Information Processing Systems, Vol. 32</source>, eds <person-group person-group-type="editor"><name><surname>Wallach</surname> <given-names>H.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Beygelzimer</surname> <given-names>A.</given-names></name> <name><surname>d&#x00027;Alch&#x000E9;-Buc</surname> <given-names>F.</given-names></name> <name><surname>Fox</surname> <given-names>E.</given-names></name> <name><surname>Garnett</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>Curran Associates</publisher-name>), <fpage>6447</fpage>&#x02013;<lpage>6458</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>L.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Tsang</surname> <given-names>I. W.-H.</given-names></name></person-group> (<year>2012</year>). <article-title>Domain adaptation from multiple sources: a domain-dependent regularization approach</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>23</volume>, <fpage>504</fpage>&#x02013;<lpage>518</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2011.2178556</pub-id><pub-id pub-id-type="pmid">24808555</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eppink</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>A brief history of the gif (so far)</article-title>. <source>J. Visual Cult</source>. <volume>13</volume>, <fpage>298</fpage>&#x02013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1177/1470412914553365</pub-id><pub-id pub-id-type="pmid">15088643</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ganin</surname> <given-names>Y.</given-names></name> <name><surname>Lempitsky</surname> <given-names>V.</given-names></name></person-group> (<year>2014</year>). <article-title>Unsupervised domain adaptation by backpropagation</article-title>. <source>arXiv preprint arXiv:1409.7495</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1409.7495</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ganin</surname> <given-names>Y.</given-names></name> <name><surname>Ustinova</surname> <given-names>E.</given-names></name> <name><surname>Ajakan</surname> <given-names>H.</given-names></name> <name><surname>Germain</surname> <given-names>P.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Laviolette</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Domain-adversarial training of neural networks</article-title>. <source>J. Mach. Learn. Res</source>. <volume>17</volume>, <fpage>2096</fpage>&#x02013;<lpage>2030</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-58347-1_10</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ghifary</surname> <given-names>M.</given-names></name> <name><surname>Bastiaan Kleijn</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Balduzzi</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>Domain generalization for object recognition with multi-task autoencoders,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Santiago</publisher-loc>), <fpage>2551</fpage>&#x02013;<lpage>2559</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.293</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Gituma</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <source>The Kinetics Dataset Explorer Using GIFs</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://towardsdatascience.com/the-kinetics-dataset-explorer-using-gifs-8ceeebcbdaba">https://towardsdatascience.com/the-kinetics-dataset-explorer-using-gifs-8ceeebcbdaba</ext-link> (accessed February 24, 2019).</citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gygli</surname> <given-names>M.</given-names></name> <name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Video2gif: automatic generation of animated gifs from video,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>1001</fpage>&#x02013;<lpage>1009</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.114</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hoffman</surname> <given-names>J.</given-names></name> <name><surname>Tzeng</surname> <given-names>E.</given-names></name> <name><surname>Park</surname> <given-names>T.</given-names></name> <name><surname>Zhu</surname> <given-names>J.-Y.</given-names></name> <name><surname>Isola</surname> <given-names>P.</given-names></name> <name><surname>Saenko</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>CyCADA: cycle-consistent adversarial domain adaptation,</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>PMLR</publisher-loc>) (Stockholm), <fpage>1989</fpage>&#x02013;<lpage>1998</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jamal</surname> <given-names>A.</given-names></name> <name><surname>Namboodiri</surname> <given-names>V. P.</given-names></name> <name><surname>Deodhare</surname> <given-names>D.</given-names></name> <name><surname>Venkatesh</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep domain adaptation in action space,</article-title> in <source>BMVC</source> (<publisher-loc>Newcastle</publisher-loc>), <fpage>264</fpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jhuo</surname> <given-names>I.-H.</given-names></name> <name><surname>Liu</surname> <given-names>D.</given-names></name> <name><surname>Lee</surname> <given-names>D.</given-names></name> <name><surname>Chang</surname> <given-names>S.-F.</given-names></name></person-group> (<year>2012</year>). <article-title>Robust visual domain adaptation with low-rank reconstruction,</article-title> in <source>2012 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Providence, RI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2168</fpage>&#x02013;<lpage>2175</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2012.6247924</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J. A.</given-names></name> <name><surname>Fiesler</surname> <given-names>C.</given-names></name> <name><surname>Brubaker</surname> <given-names>J. R.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x02018;The perfect one? understanding communication practices and challenges with animated gifs</article-title>. <source>Proc. ACM Hum. Comput. Interact</source>. <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1145/3274349</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jou</surname> <given-names>B.</given-names></name> <name><surname>Bhattacharya</surname> <given-names>S.</given-names></name> <name><surname>Chang</surname> <given-names>S.-F.</given-names></name></person-group> (<year>2014</year>). <article-title>Predicting viewer perceived emotions in animated GIFs,</article-title> in <source>Proceedings of the 22nd ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>213</fpage>&#x02013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1145/2647868.2656408</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kay</surname> <given-names>W.</given-names></name> <name><surname>Carreira</surname> <given-names>J.</given-names></name> <name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Hillier</surname> <given-names>C.</given-names></name> <name><surname>Vijayanarasimhan</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>The kinetics human action video dataset</article-title>. <source>arXiv preprint arXiv:1705.06950</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1705.06950</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Yoo</surname> <given-names>Y.</given-names></name> <name><surname>Park</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Selfreg: self-supervised contrastive regularization for domain generalization,</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>9619</fpage>&#x02013;<lpage>9628</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00948</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kong</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>B.</given-names></name> <name><surname>Bradbury</surname> <given-names>K.</given-names></name> <name><surname>Malof</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>The synthinel-1 dataset: a collection of high resolution synthetic overhead imagery for building segmentation,</article-title> in <source>The IEEE Winter Conference on Applications of Computer Vision</source> (<publisher-loc>Snowmass, CO</publisher-loc>), <fpage>1814</fpage>&#x02013;<lpage>1823</lpage>. <pub-id pub-id-type="doi">10.1109/WACV45572.2020.9093339</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuehne</surname> <given-names>H.</given-names></name> <name><surname>Jhuang</surname> <given-names>H.</given-names></name> <name><surname>Garrote</surname> <given-names>E.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name> <name><surname>Serre</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>HMDB: a large video database for human motion recognition,</article-title> in <source>2011 International Conference on Computer Vision</source> (<publisher-loc>Barcelona</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2556</fpage>&#x02013;<lpage>2563</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2011.6126543</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Gong</surname> <given-names>S.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name> <name><surname>Hospedales</surname> <given-names>T. M.</given-names></name></person-group> (<year>2021</year>). <article-title>A simple feature augmentation for domain generalization,</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>8886</fpage>&#x02013;<lpage>8895</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00876</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>A.</given-names></name> <name><surname>Thotakuri</surname> <given-names>M.</given-names></name> <name><surname>Ross</surname> <given-names>D. A.</given-names></name> <name><surname>Carreira</surname> <given-names>J.</given-names></name> <name><surname>Vostrikov</surname> <given-names>A.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>The AVA-kinetics localized human actions video dataset</article-title>. <source>arXiv preprint arXiv:2005.00214</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2005.00214</pub-id><pub-id pub-id-type="pmid">24051735</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>Y.-Z.</given-names></name> <name><surname>Hospedales</surname> <given-names>T. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Deeper, broader and artier domain generalization,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source>, 5542&#x02013;5550. <pub-id pub-id-type="doi">10.1109/ICCV.2017.591</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>Y.-Z.</given-names></name> <name><surname>Hospedales</surname> <given-names>T. M.</given-names></name></person-group> (<year>2018a</year>). <article-title>Learning to generalize: meta-learning for domain generalization,</article-title> in <source>Thirty-Second AAAI Conference on Artificial Intelligence</source> (<publisher-loc>New Orleans, LA</publisher-loc>). <pub-id pub-id-type="doi">10.1609/aaai.v32i1.11596</pub-id><pub-id pub-id-type="pmid">35724300</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name> <name><surname>Tetreault</surname> <given-names>J.</given-names></name> <name><surname>Goldberg</surname> <given-names>L.</given-names></name> <name><surname>Jaimes</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>TGIF: a new dataset and benchmark on animated gif description,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>4641</fpage>&#x02013;<lpage>4650</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.502</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Tian</surname> <given-names>X.</given-names></name> <name><surname>Gong</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2018b</year>). <article-title>Deep domain generalization via conditional invariant adversarial networks,</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source>, 624&#x02013;639. <pub-id pub-id-type="doi">10.1007/978-3-030-01267-0_38</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>N.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Hou</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name></person-group> (<year>2018c</year>). <article-title>Adaptive batch normalization for practical domain adaptation</article-title>. <source>Pattern Recogn</source>. <volume>80</volume>, <fpage>109</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2018.03.005</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Shao</surname> <given-names>M.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Structure-preserved multi-source domain adaptation,</article-title> in <source>2016 IEEE 16th International Conference on Data Mining (ICDM)</source> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1059</fpage>&#x02013;<lpage>1064</lpage>. <pub-id pub-id-type="doi">10.1109/ICDM.2016.0136</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Guo</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name> <name><surname>Kuo</surname> <given-names>C.-C. J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Adversarial unsupervised domain adaptation with conditional and label shift: infer, align and iterate,</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</source>, 10367&#x02013;10376. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01020</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2018</year>). <article-title>Conditional adversarial domain adaptation,</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 31</source>, eds <person-group person-group-type="editor"><name><surname>Bengio</surname> <given-names>S.</given-names></name> <name><surname>Wallach</surname> <given-names>H.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Grauman</surname> <given-names>K.</given-names></name> <name><surname>Cesa-Bianchi</surname> <given-names>N.</given-names></name> <name><surname>Garnett</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>1640</fpage>&#x02013;<lpage>1650</lpage>.<pub-id pub-id-type="pmid">34487497</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2016</year>). <article-title>Unsupervised domain adaptation with residual transfer networks,</article-title> in <source>Advances in Neural Information Processing System, Vol. 29</source>, eds <person-group person-group-type="editor"><name><surname>Lee</surname> <given-names>D.</given-names></name> <name><surname>Sugiyama</surname> <given-names>M.</given-names></name> <name><surname>Luxburg</surname> <given-names>U.</given-names></name> <name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Garnett</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>136</fpage>&#x02013;<lpage>144</lpage>.<pub-id pub-id-type="pmid">32361548</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep transfer learning with joint adaptation networks,</article-title> in <source>Proceedings of the 34th International Conference on Machine Learning</source>, eds <person-group person-group-type="editor"><name><surname>Precup</surname> <given-names>D.</given-names></name> <name><surname>Teh</surname> <given-names>Y. W.</given-names></name></person-group> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>2208</fpage>&#x02013;<lpage>2217</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jordan</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning transferable features with deep adaptation networks,</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>Lille</publisher-loc>), <fpage>97</fpage>&#x02013;<lpage>105</lpage>.<pub-id pub-id-type="pmid">30188813</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Miltner</surname> <given-names>K. M.</given-names></name> <name><surname>Highfield</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Never gonna gif you up: analyzing the cultural significance of the animated gif</article-title>. <source>Soc. Media Soc</source>. 3, 2056305117725223. <pub-id pub-id-type="doi">10.1177/2056305117725223</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moosavi-Dezfooli</surname> <given-names>S.-M.</given-names></name> <name><surname>Fawzi</surname> <given-names>A.</given-names></name> <name><surname>Frossard</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Deepfool: a simple and accurate method to fool deep neural networks,</article-title> in <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>). <pub-id pub-id-type="doi">10.1109/CVPR.2016.282</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Motiian</surname> <given-names>S.</given-names></name> <name><surname>Piccirilli</surname> <given-names>M.</given-names></name> <name><surname>Adjeroh</surname> <given-names>D. A.</given-names></name> <name><surname>Doretto</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Unified deep supervised domain adaptation and generalization,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>5715</fpage>&#x02013;<lpage>5725</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.609</pub-id><pub-id pub-id-type="pmid">35009752</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Muandet</surname> <given-names>K.</given-names></name> <name><surname>Balduzzi</surname> <given-names>D.</given-names></name> <name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>Domain generalization via invariant feature representation,&#x0201D;</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>Atlanta, GA</publisher-loc>), <fpage>10</fpage>&#x02013;<lpage>18</lpage>.<pub-id pub-id-type="pmid">32377643</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Murez</surname> <given-names>Z.</given-names></name> <name><surname>Kolouri</surname> <given-names>S.</given-names></name> <name><surname>Kriegman</surname> <given-names>D.</given-names></name> <name><surname>Ramamoorthi</surname> <given-names>R.</given-names></name> <name><surname>Kim</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Image to image translation for domain adaptation,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>4500</fpage>&#x02013;<lpage>4509</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00473</pub-id><pub-id pub-id-type="pmid">35301702</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nam</surname> <given-names>H.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Yoon</surname> <given-names>W.</given-names></name> <name><surname>Yoo</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Reducing domain gap by reducing style bias,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Nashville, TN</publisher-loc>), <fpage>8690</fpage>&#x02013;<lpage>8699</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00858</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Tsang</surname> <given-names>I. W.</given-names></name> <name><surname>Kwok</surname> <given-names>J. T.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2010</year>). <article-title>Domain adaptation via transfer component analysis</article-title>. <source>IEEE Trans. Neural Netw</source>. <volume>22</volume>, <fpage>199</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1109/TNN.2010.2091281</pub-id><pub-id pub-id-type="pmid">21095864</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Piergiovanni</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <source>3D models trained on Kinetics.</source> Available online at: <ext-link ext-link-type="uri" xlink:href="https://github.com/piergiaj/pytorch-i3d">https://github.com/piergiaj/pytorch-i3d</ext-link> (accessed June 28, 2018).</citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ponce</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Dataset issues in object recognition,&#x00027;</article-title> in <source>Toward Category-Level Object Recognition. Lecture Notes in Computer Science, Vol 4170</source>, eds <person-group person-group-type="editor"><name><surname>Ponce</surname> <given-names>J.</given-names></name> <name><surname>Hebert</surname> <given-names>M.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>29</fpage>&#x02013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1007/11957959_2</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Recht</surname> <given-names>B.</given-names></name> <name><surname>Roelofs</surname> <given-names>R.</given-names></name> <name><surname>Schmidt</surname> <given-names>L.</given-names></name> <name><surname>Shankar</surname> <given-names>V.</given-names></name></person-group> (<year>2019</year>). <article-title>Do imagenet classifiers generalize to imagenet?</article-title> <source>arXiv preprint arXiv:1902.10811</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1902.10811</pub-id><pub-id pub-id-type="pmid">35111210</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ros</surname> <given-names>G.</given-names></name> <name><surname>Sellart</surname> <given-names>L.</given-names></name> <name><surname>Materzynska</surname> <given-names>J.</given-names></name> <name><surname>Vazquez</surname> <given-names>D.</given-names></name> <name><surname>Lopez</surname> <given-names>A. M.</given-names></name></person-group> (<year>2016</year>). <article-title>The synthia dataset: a large collection of synthetic images for semantic segmentation of urban scenes,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>3234</fpage>&#x02013;<lpage>3243</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.352</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>R&#x000F6;ssler</surname> <given-names>A.</given-names></name> <name><surname>Cozzolino</surname> <given-names>D.</given-names></name> <name><surname>Verdoliva</surname> <given-names>L.</given-names></name> <name><surname>Riess</surname> <given-names>C.</given-names></name> <name><surname>Thies</surname> <given-names>J.</given-names></name> <name><surname>Nie&#x000DF;ner</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Faceforensics: a large-scale video dataset for forgery detection in human faces</article-title>. <source>arXiv preprint arXiv:1803.09179</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1803.09179</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russakovsky</surname> <given-names>O.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Satheesh</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Imagenet large scale visual recognition challenge</article-title>. <source>Int. J. Comput. Vision</source> <volume>115</volume>, <fpage>211</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Saito</surname> <given-names>K.</given-names></name> <name><surname>Watanabe</surname> <given-names>K.</given-names></name> <name><surname>Ushiku</surname> <given-names>Y.</given-names></name> <name><surname>Harada</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Maximum classifier discrepancy for unsupervised domain adaptation,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>3723</fpage>&#x02013;<lpage>3732</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00392</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sato</surname> <given-names>I.</given-names></name> <name><surname>Nishimura</surname> <given-names>H.</given-names></name> <name><surname>Yokoi</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>APAC: augmented pattern classification with neural networks</article-title>. <source>arXiv preprint arXiv:1505.03229</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1505.03229</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Scheck</surname> <given-names>T.</given-names></name> <name><surname>Seidel</surname> <given-names>R.</given-names></name> <name><surname>Hirtz</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Learning from theodore: a synthetic omnidirectional top-view indoor dataset for deep transfer learning,</article-title> in <source>The IEEE Winter Conference on Applications of Computer Vision</source> (<publisher-loc>Snowmass, CO</publisher-loc>), <fpage>943</fpage>&#x02013;<lpage>952</lpage>. <pub-id pub-id-type="doi">10.1109/WACV45572.2020.9093563</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Selvaraju</surname> <given-names>R. R.</given-names></name> <name><surname>Cogswell</surname> <given-names>M.</given-names></name> <name><surname>Das</surname> <given-names>A.</given-names></name> <name><surname>Vedantam</surname> <given-names>R.</given-names></name> <name><surname>Parikh</surname> <given-names>D.</given-names></name> <name><surname>Batra</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Grad-CAM: visual explanations from deep networks via gradient-based localization,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>618</fpage>&#x02013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.74</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shorten</surname> <given-names>C.</given-names></name> <name><surname>Khoshgoftaar</surname> <given-names>T. M.</given-names></name></person-group> (<year>2019</year>). <article-title>A survey on image data augmentation for deep learning</article-title>. <source>J. Big Data</source> <volume>6</volume>, <fpage>60</fpage>. <pub-id pub-id-type="doi">10.1186/s40537-019-0197-0</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Soomro</surname> <given-names>K.</given-names></name> <name><surname>Zamir</surname> <given-names>A. R.</given-names></name> <name><surname>Shah</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <source>A Dataset of 101 Human Action Classes from Videos in the Wild</source>. Center for Research in Computer Vision.</citation>
</ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>B.</given-names></name> <name><surname>Saenko</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep coral: correlation alignment for deep domain adaptation,</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>443</fpage>&#x02013;<lpage>450</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-49409-8_35</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tolins</surname> <given-names>J.</given-names></name> <name><surname>Samermit</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>GIFs as embodied enactments in text-mediated conversation</article-title>. <source>Res. Lang. Soc. Interact</source>. <volume>49</volume>, <fpage>75</fpage>&#x02013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1080/08351813.2016.1164391</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Unbiased look at dataset bias,</article-title> in <source>CVPR 2011</source> (<publisher-loc>Colorado Springs, CO</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1521</fpage>&#x02013;<lpage>1528</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2011.5995347</pub-id></citation>
</ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tzeng</surname> <given-names>E.</given-names></name> <name><surname>Hoffman</surname> <given-names>J.</given-names></name> <name><surname>Saenko</surname> <given-names>K.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Adversarial discriminative domain adaptation,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>7167</fpage>&#x02013;<lpage>7176</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.316</pub-id><pub-id pub-id-type="pmid">32635540</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Volpi</surname> <given-names>R.</given-names></name> <name><surname>Murino</surname> <given-names>V.</given-names></name></person-group> (<year>2019</year>). <article-title>Addressing model vulnerability to distributional shifts over image transformation sets,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>7980</fpage>&#x02013;<lpage>7989</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00807</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Volpi</surname> <given-names>R.</given-names></name> <name><surname>Namkoong</surname> <given-names>H.</given-names></name> <name><surname>Sener</surname> <given-names>O.</given-names></name> <name><surname>Duchi</surname> <given-names>J. C.</given-names></name> <name><surname>Murino</surname> <given-names>V.</given-names></name> <name><surname>Savarese</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Generalizing to unseen domains via adversarial data augmentation,</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 31</source>, eds <person-group person-group-type="editor"><name><surname>Bengio</surname> <given-names>S.</given-names></name> <name><surname>Wallach</surname> <given-names>H.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Grauman</surname> <given-names>K.</given-names></name> <name><surname>Cesa-Bianchi</surname> <given-names>N.</given-names></name> <name><surname>Garnett</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>), <fpage>5334</fpage>&#x02013;<lpage>5344</lpage>.</citation>
</ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wan</surname> <given-names>L.</given-names></name> <name><surname>Zeiler</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Le Cun</surname> <given-names>Y.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>Regularization of neural networks using dropconnect,</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>Atlanta, GA</publisher-loc>), <fpage>1058</fpage>&#x02013;<lpage>1066</lpage>.</citation>
</ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name> <name><surname>Hellovera</surname> <given-names>A.</given-names></name></person-group> <source>Gif super-resolution.</source></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilson</surname> <given-names>G.</given-names></name> <name><surname>Cook</surname> <given-names>D. J.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of unsupervised deep domain adaptation</article-title>. <source>ACM Trans. Intell. Syst. Technol</source>. <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1145/3400066</pub-id><pub-id pub-id-type="pmid">35898010</pub-id></citation></ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>R.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep cocktail network: multi-source unsupervised domain adaptation with category shift,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>3964</fpage>&#x02013;<lpage>3973</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00417</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Hospedales</surname> <given-names>T. M.</given-names></name></person-group> (<year>2014</year>). <article-title>A unified perspective on multi-domain and multi-task learning</article-title>. <source>arXiv preprint arXiv:1412.7489</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1412.7489</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>P.</given-names></name> <name><surname>Long</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>VideoDG: generalizing temporal relations in videos to novel domains</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <pub-id pub-id-type="doi">10.1109/TPAMI.2021.3116945</pub-id><pub-id pub-id-type="pmid">34596532</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Du</surname> <given-names>X.</given-names></name> <name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Adversarial pyramid network for video domain generalization</article-title>. <source>arXiv preprint arXiv:1912.03716</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1912.03716</pub-id></citation>
</ref>
<ref id="B77">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>F.</given-names></name> <name><surname>Xian</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Liao</surname> <given-names>M.</given-names></name> <name><surname>Madhavan</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>BDD100K: a diverse driving video database with scalable annotation tooling</article-title>. <source>arXiv preprint arXiv:1805.04687</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1805.04687</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zajac</surname> <given-names>M.</given-names></name> <name><surname>Zo&#x00142;na</surname> <given-names>k.</given-names></name> <name><surname>Rostamzadeh</surname> <given-names>N.</given-names></name> <name><surname>Pinheiro</surname> <given-names>P. O.</given-names></name></person-group> (<year>2019</year>). <article-title>Adversarial framing for image and video classification,</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>10077</fpage>&#x02013;<lpage>10078</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.330110077</pub-id></citation>
</ref>
<ref id="B79">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zunino</surname> <given-names>A.</given-names></name> <name><surname>Bargal</surname> <given-names>S. A.</given-names></name> <name><surname>Volpi</surname> <given-names>R.</given-names></name> <name><surname>Sameki</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Sclaroff</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Explainable deep classification models for domain generalization</article-title>. <source>arXiv preprint arXiv:2003.06498</source>. <pub-id pub-id-type="doi">10.1109/CVPRW53098.2021.00361</pub-id><pub-id pub-id-type="pmid">34409286</pub-id></citation></ref>
</ref-list> 
</back>
</article>