<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">752831</article-id>
<article-id pub-id-type="doi">10.3389/frai.2021.752831</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Improving Adversarial Robustness via Attention and Adversarial Logit Pairing</article-title>
<alt-title alt-title-type="left-running-head">Li et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Adversarial Training With AT &#x2b; ALP</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Xingjian</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1424388/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Goodman</surname>
<given-names>Dou</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Liu</surname>
<given-names>Ji</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1227109/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wei</surname>
<given-names>Tao</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dou</surname>
<given-names>Dejing</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Big Data Lab, Baidu Research</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>X-Lab, Baidu Inc.</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1242884/overview">Yang Zhou</ext-link>, Auburn University, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1375376/overview">Jinghui Chen</ext-link>, The Pennsylvania State University (PSU), United&#x20;States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1485993/overview">Dingqi Yang</ext-link>, University of Macau, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1486003/overview">Mohamed Reda Bouadjenek</ext-link>, Deakin University, Australia</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Xingjian Li, <email>lixj04@gmail.com</email>; Dou Goodman, <email>liu.yan@baidu.com</email>
</corresp>
<fn fn-type="equal" id="fn1">
<label>
<sup>&#x2020;</sup>
</label>
<p>These authors have contributed equally to this work</p>
</fn>
<fn fn-type="other">
<p>This article was submitted to Machine Learning and Artificial Intelligence, a section of the journal Frontiers in Artificial Intelligence</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>4</volume>
<elocation-id>752831</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Li, Goodman, Liu, Wei and Dou.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Li, Goodman, Liu, Wei and Dou</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Though deep neural networks have achieved the state of the art performance in visual classification, recent studies have shown that they are all vulnerable to the attack of adversarial examples. In this paper, we develop improved techniques for defending against adversarial examples. First, we propose an enhanced defense technique denoted <bold>Attention and Adversarial Logit Pairing (AT &#x2b; ALP)</bold>, which encourages both attention map and logit for the pairs of examples to be similar. When being applied to clean examples and their adversarial counterparts, <bold>AT &#x2b; ALP</bold> improves accuracy on adversarial examples over adversarial training. We show that <bold>AT &#x2b; ALP</bold> can effectively increase the average activations of adversarial examples in the key area and demonstrate that it focuses on discriminate features to improve the robustness of the model. Finally, we conduct extensive experiments using a wide range of datasets and the experiment results show that our <bold>AT &#x2b; ALP</bold> achieves <bold>the state of the art</bold> defense performance. For example, on <bold>17 Flower Category Database</bold>, under strong 200-iteration Projected Gradient Descent (PGD) gray-box and black-box attacks where prior art has 34 and 39% accuracy, our method achieves <bold>50</bold> and <bold>51%</bold>. Compared with previous work, our work is evaluated under highly challenging PGD attack: the maximum perturbation <italic>&#x3f5;</italic> &#x2208; {0.25, 0.5} i.e. <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> &#x2208; {0.25, 0.5} with 10&#x2013;200 attack iterations. To the best of our knowledge, such a strong attack has not been previously explored on a wide range of datasets.</p>
</abstract>
<kwd-group>
<kwd>adversarial training</kwd>
<kwd>attention</kwd>
<kwd>adversarial robustness</kwd>
<kwd>adversarial example</kwd>
<kwd>deep learning</kwd>
<kwd>deep neural network</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>In recent years, deep neural networks have been extensively deployed for computer vision tasks, particularly for visual classification problems, where new algorithms have been reported to achieve even better performance than human beings <xref ref-type="bibr" rid="B15">Krizhevsky et&#x20;al. (2012)</xref>, <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref>, <xref ref-type="bibr" rid="B16">Li et&#x20;al. (2019a)</xref>. The success of deep neural networks has led to an explosion in demand. However, recent studies have shown that they are all vulnerable to the attack of adversarial examples <xref ref-type="bibr" rid="B29">Szegedy et&#x20;al. (2013)</xref>; <xref ref-type="bibr" rid="B5">Carlini and Wagner (2016)</xref>; <xref ref-type="bibr" rid="B20">Moosavi-Dezfooli et&#x20;al. (2016)</xref>; <xref ref-type="bibr" rid="B3">Bose and Aarabi (2018)</xref>. Small and often imperceptible perturbations to the input images are sufficient to fool the most powerful deep neural networks.</p>
<p>In <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>, we visualize the spatial attention map of a flower and its corresponding adversarial image on ResNet-50 <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref> pretrained on ImageNet <xref ref-type="bibr" rid="B26">Russakovsky et&#x20;al. (2015)</xref>. The figure suggests that adversarial perturbations, while small in the pixel space, lead to very substantial &#x201c;noise&#x201d; in the attention map of the network. Whereas the features for the clean image appear to focus primarily on semantically informative content in the image, the attention map for the adversarial image are activated across semantically irrelevant regions as well. The state of the art adversarial training methods only encourage hard labels <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B30">Tram&#xe8;r et&#x20;al. (2017)</xref> or logit <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref> for pairs of clean examples and adversarial counterparts to be similar. In our opinion, it is not enough to align the difference between the clean examples and adversarial counterparts only at the end part of the whole network, i.e.,&#x20;hard labels or logit, and we need to align the attention maps for important parts of the whole network. Motivated by this observation, we explore <bold>Attention and Adversarial Logit Pairing(AT &#x2b; ALP)</bold>, a method that encourages both attention map and logit for pairs of examples to be similar. When being applied to clean examples and their adversarial counterparts, <bold>AT &#x2b; ALP</bold> improves accuracy on adversarial examples over adversarial training.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>
<bold>(A)</bold> is original image and <bold>(B)</bold> is corresponding spatial attention map of ResNet-50 <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref> pretrained on ImageNet <xref ref-type="bibr" rid="B26">Russakovsky et&#x20;al. (2015)</xref> which shows where the network focuses in order to classify the given image. <bold>(C)</bold> is adversarial image of <bold>(A)</bold>, <bold>(D)</bold> is corresponding spatial attention&#x20;map.</p>
</caption>
<graphic xlink:href="frai-04-752831-g001.tif"/>
</fig>
<p>The contributions of this paper are summarized as follows:<list list-type="simple">
<list-item>
<p>&#x2022; We introduce enhanced adversarial training using a technique we call <bold>Attention and Adversarial Logit Pairing(AT &#x2b; ALP)</bold>, which encourages both attention map and logit for pairs of examples to be similar. When being applied to clean examples and their adversarial counterparts, <bold>AT &#x2b; ALP</bold> improves accuracy on adversarial examples over adversarial training.</p>
</list-item>
<list-item>
<p>&#x2022; We show that our <bold>AT &#x2b; ALP</bold> can effectively increase the average activations of adversarial examples in the key area and demonstrate that it focuses on more discriminate features to improve the robustness of the&#x20;model.</p>
</list-item>
<list-item>
<p>&#x2022; We show that our <bold>AT &#x2b; ALP</bold> achieves <bold>the state of the art</bold> defense on a wide range of datasets against strong <bold>PGD</bold> gray-box and black-box attacks. Compared with previous work, our work is evaluated under highly challenging PGD attack: the maximum perturbation <italic>&#x3f5;</italic> &#x2208; {0.25, 0.5}, i.e.,&#x20;<italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> &#x2208; {0.25, 0.5} with 10&#x2013;200 attack iterations. To the best of our knowledge, such a strong attack has not been previously explored on a wide range of datasets.</p>
</list-item>
</list>
</p>
<p>The rest of the paper is organized as follows: in <xref ref-type="sec" rid="s2">Section 2</xref>, we present the related works; in <xref ref-type="sec" rid="s3">Section 3</xref>, we introduce definitions and threat models; in <xref ref-type="sec" rid="s4">Section 4</xref> we propose our <bold>Attention and Adversarial Logit Pairing(AT &#x2b; ALP)</bold> method; in <xref ref-type="sec" rid="s5">Section 5</xref>, we show extensive experimental results; and <xref ref-type="sec" rid="s6">Section 6</xref> concludes.</p>
</sec>
<sec id="s2">
<title>2 Related Work</title>
<p>
<xref ref-type="bibr" rid="B2">Athalye et&#x20;al. (2018)</xref> evaluate the robustness of nine papers <xref ref-type="bibr" rid="B4">Buckman et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B18">Ma et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B11">Guo et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B8">Dhillon et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B31">Xie et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B28">Song et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B27">Samangouei et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B21">Na et&#x20;al. (2017)</xref> accepted by ICLR 2018 as non-certified white-box-secure defenses to adversarial examples. They find that seven of the nine defenses use obfuscated gradients, a kind of gradient masking, as a phenomenon that leads to a false sense of security in defenses against adversarial examples. Obfuscated gradients provide a limited increase in robustness and can be broken by improved attack techniques they develop. The only defense they observe that significantly increases robustness to adversarial examples within the threat model proposed is <bold>adversarial training</bold> <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>.</p>
<p>Adversarial training <xref ref-type="bibr" rid="B10">Goodfellow et&#x20;al. (2015)</xref>; <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B30">Tram&#xe8;r et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B23">Pang et&#x20;al. (2019)</xref> defends against adversarial perturbations by training networks on adversarial images that are generated on-the-fly during training. For adversarial training, the most relevant work to our study is <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>, which introduce a technique they call <bold>Adversarial Logit Pairing (ALP)</bold>. This method encourages logits for pairs of examples to be similar. Our <bold>AT &#x2b; ALP</bold> encourages both attention map and logit for pairs of examples to be similar. When being applied to clean examples and their adversarial counterparts, <bold>AT &#x2b; ALP</bold> improves accuracy on adversarial examples over adversarial training. <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref> adds random noise at training and inference time, <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref> adds denoising blocks to the model to increase adversarial robustness, while neither of the above approaches focuses on the attention&#x20;map.</p>
<p>In terms of methodologies, our work is also related to deep transfer learning and knowledge distillation problems, and the most relevant work to our study is <xref ref-type="bibr" rid="B33">Zagoruyko and Komodakis (2016)</xref>; <xref ref-type="bibr" rid="B17">Li et&#x20;al. (2019b)</xref>, which constrain the <italic>L</italic>
<sub>2</sub>-norm of the difference between their behaviors (i.e.,&#x20;the feature maps of outer layer outputs in the source/target networks). Our <bold>AT &#x2b; ALP</bold> constrains attention map and logit for pairs of clean examples and their adversarial counterparts to be similar.</p>
</sec>
<sec id="s3">
<title>3 Definitions and Threat Models</title>
<p>In this paper, we always assume the attacker is capable of forming attacks that consist of perturbations of limited <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub>-norm. This is a simplified task chosen because it is more amenable to benchmark evaluations. We consider two different threat models characterizing amounts of information the adversary can have:<list list-type="simple">
<list-item>
<p>&#x2022; <bold>Gray-box Attack</bold> We focus on defense against gray-box attacks in this paper. In a gray-back attack, the attacker knows both the original network and the defense algorithm. Only the parameters of the defense model are hidden from the attacker. This is also a standard setting assumed in many security systems and applications <xref ref-type="bibr" rid="B24">Pfleeger and Pfleeger (2004)</xref>.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Black-box Attack</bold> The attacker has no information about the model&#x2019;s architecture or parameters, and no ability to send queries to the model to gather more information.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s4">
<title>4 Methods</title>
<sec id="s4-1">
<title>4.1 Architecture</title>
<p>
<xref ref-type="fig" rid="F2">Figure&#x20;2</xref> represents architecture of <bold>Attention and Adversarial Logit Pairing (AT &#x2b; ALP)</bold>: a baseline model is adversarial trained so as, not only to make similar logits, but to also have similar spatial attention maps to those of original image and adversarial&#x20;image.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Schematic representation of <bold>Attention and Adversarial Logit Pairing (AT &#x2b; ALP)</bold>: a baseline model is trained so as, not only to make similar logits, but to also have similar spatial attention maps to those of original image and adversarial&#x20;image.</p>
</caption>
<graphic xlink:href="frai-04-752831-g002.tif"/>
</fig>
</sec>
<sec id="s4-2">
<title>4.2 Adversarial Training</title>
<p>We use adversarial training with <bold>Projected Gradient Descent (PGD)</bold> <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref> as the underlying basis for our methods:<disp-formula id="e1">
<mml:math id="m1">
<mml:munder>
<mml:mrow>
<mml:mi>arg min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>&#x2009;data&#x2009;</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>&#x2009;data&#x2009;</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the underlying training data distribution, <italic>L</italic> (<italic>&#x3b8;</italic>, <italic>x</italic>&#x20;&#x2b; <italic>&#x3b4;</italic>, <italic>y</italic>) is a loss function at data point <italic>x</italic> which has true class <italic>y</italic> for a model with parameters <italic>&#x3b8;</italic>, and the maximization with respect to <italic>&#x3b4;</italic> is approximated using PGD. In this paper, the loss is defined as:<disp-formula id="e2">
<mml:math id="m3">
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>L</italic>
<sub>
<italic>CE</italic>
</sub> is cross entropy, <italic>&#x3b1;</italic> and <italic>&#x3b2;</italic> are hyperparameters.</p>
</sec>
<sec id="s4-3">
<title>4.3 Adversarial Logit Pairing</title>
<p>We also use <bold>Adversarial Logit Pairing (ALP)</bold> to encourage the logits from clean examples and their adversarial counterparts to be similar to each other. For a model that takes inputs <italic>x</italic> and computes a vector of logit <italic>z</italic>&#x20;&#x3d; <italic>f</italic>(<italic>x</italic>), logit pairing adds a loss:<disp-formula id="e3">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>In this paper we use <italic>L</italic>
<sub>2</sub> loss for&#x20;<italic>L</italic>
<sub>
<italic>a</italic>
</sub>.</p>
</sec>
<sec id="s4-4">
<title>4.4 Attention Map</title>
<p>We use <bold>Attention Map (AT)</bold> to encourage the attention map from clean examples and their adversarial counterparts to be similar to each other. Let <italic>I</italic> denote the indices of all activation layer pairs, for which we want to pay attention. Then, we can define the following total loss:<disp-formula id="e4">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>Let <italic>O</italic>, <italic>ADV</italic> denote clean examples and their adversarial counterparts, where <inline-formula id="inf2">
<mml:math id="m6">
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mo movablelimits="false" form="prefix">vec</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> and <inline-formula id="inf3">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mo movablelimits="false" form="prefix">vec</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> are respectively the <italic>j</italic>th pair of clean examples and their adversarial counterparts attention maps in vectorized form, and <italic>p</italic> refers to norm type (in the experiments we use <italic>p</italic>&#x20;&#x3d;&#x20;2).</p>
</sec>
<sec id="s4-5">
<title>4.5 Experiments: White-Box Settings</title>
<p>White-box attack is the most challenging task for evaluating a model&#x2019;s adversarial robustness. In white-box settings, attackers are assumed to know all details about the model, including its architecture and parameters. We conduct white-box experiments following common practices <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>; <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>. Specifically, we use ResNet-18 <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref> trained with CIFAR-10 <xref ref-type="bibr" rid="B14">Krizhevsky and Hinton (2009)</xref>.</p>
<p>We use Fast Gradient Sign (FGS) <xref ref-type="bibr" rid="B10">Goodfellow et&#x20;al. (2015)</xref>, Projected Gradient Descent (PGD) <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>, AutoAttack <xref ref-type="bibr" rid="B7">Croce and Hein (2020)</xref> and RayS <xref ref-type="bibr" rid="B6">Chen and Gu (2020)</xref> to perform white-box attacks towards evaluated models. We consider untargeted attack, which is more challenging for defense than targeted attack. Adversarial perturbations are measured by <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> norm (i.e.,&#x20;maximum perturbation for each pixel), with an allowed maximum value of <italic>&#x3f5;</italic> &#x3d; 8/255.</p>
</sec>
<sec id="s4-6">
<title>4.6 Image Database</title>
<p>The CIFAR-10 <xref ref-type="bibr" rid="B14">Krizhevsky and Hinton (2009)</xref> dataset contains 50,000 training samples and 10,000 test samples, uniformly distributed across 10 classes. Each sample is a 32&#x20;&#xd7; 32 color image. Though with a low image resolution, CIFAR-10 is a popular benchmark to evaluate the adversarial robustness of a&#x20;model.</p>
</sec>
<sec id="s4-7">
<title>4.7 Experimental Setup</title>
<p>For white-box settings, we use ResNet-18 <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref> as the model architecture. Models are first trained on CIFAR-10 with different adversarial training methods, including PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>, ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>, TRADES <xref ref-type="bibr" rid="B34">Zhang et&#x20;al. (2019)</xref> and our proposed Attention Map (<bold>AT</bold>). We train all models with 100 epochs following practices suggested by TRADES <xref ref-type="bibr" rid="B34">Zhang et&#x20;al. (2019)</xref>. For adversarial attacks, we adopt 1-step FSG attack <xref ref-type="bibr" rid="B10">Goodfellow et&#x20;al. (2015)</xref>, 7-iteration PGD attack <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref> and AutoAttack <xref ref-type="bibr" rid="B7">Croce and Hein (2020)</xref> with the common used perturbation magnitude of <italic>&#x3f5;</italic> &#x3d; 8/255 under <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> norm. We also evaluate them with RayS <xref ref-type="bibr" rid="B6">Chen and Gu (2020)</xref>, which is a gradient-free adversarial attack requiring only the target model&#x2019;s hard-label output. We run each experiments three times and report the average top-1 accuracy. We also report the training time of each method for a more comprehensive comparison. Our experiments are run on Nvidia Tesla V100-SXM2&#x20;GPUs.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s5">
<title>5 Results and Discussion</title>
<p>We present results of the white-box experiment in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. We compare the proposed Attention adversarial training (<bold>AT</bold>) against relevant methods including PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>, ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref> and TRADES <xref ref-type="bibr" rid="B34">Zhang et&#x20;al. (2019)</xref>. As seen in <xref ref-type="table" rid="T1">Table&#x20;1</xref>, all of these methods show certain degree of robustness, even under the advanced adversarial attacks such as AutoAttack. Specifically, our <bold>AT</bold> is superior to baseline methods PAT and ALP, with higher clean accuracy, robust accuracy under FSG, PGD and AutoAttack. TRADES <xref ref-type="bibr" rid="B34">Zhang et&#x20;al. (2019)</xref> improves ALP by involving an inner maximization to generate a most <italic>different</italic> counterpart for the clean example. Therefore, TRADES achieves higher adversarial accuracy than other methods. However, the drawback lies in its efficiency, i.e. TRADES is slower than other adversarial training methods by about %46. This is because TRADES needs 10 adversarial steps per batch to achieve good performance, while seven steps are enough for ALP and AT. Moreover, the proposed <bold>AT</bold> achieves the highest clean accuracy among all these adversarial training methods.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Defense against white-box attack on CIFAR-10. The adversarial perturbations were produced using Fast Gradient Sign (FGS) <xref ref-type="bibr" rid="B10">Goodfellow et&#x20;al. (2015)</xref>, Projected Gradient Descent (PGD) <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>, AutoAttack (AA) <xref ref-type="bibr" rid="B7">Croce and Hein (2020)</xref> and RayS <xref ref-type="bibr" rid="B6">Chen and Gu (2020)</xref>. The perturbation magnitude is <italic>&#x3f5;</italic> &#x3d; 8/255 under <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub>&#x20;norm.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Defense on CIFAR-10 database</th>
<th align="center">Clean</th>
<th align="center">FGS</th>
<th align="center">PGD</th>
<th align="center">AA</th>
<th align="center">RayS</th>
<th align="center">Time (hours)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">No Defence</td>
<td align="char" char=".">95.3</td>
<td align="center">
<inline-formula id="inf4">
<mml:math id="m8">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula>1</td>
<td align="center">
<inline-formula id="inf5">
<mml:math id="m9">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula>1</td>
<td align="center">
<inline-formula id="inf6">
<mml:math id="m10">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula>1</td>
<td align="center">
<inline-formula id="inf7">
<mml:math id="m11">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula>1</td>
<td align="char" char=".">0.4</td>
</tr>
<tr>
<td align="left">PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>
</td>
<td align="char" char=".">83.2</td>
<td align="center">55.7</td>
<td align="center">51.6</td>
<td align="center">46.1</td>
<td align="center">57.3</td>
<td align="char" char=".">2.6</td>
</tr>
<tr>
<td align="left">ALP <xref ref-type="bibr" rid="B13">Kannan, Kurakin, and Goodfellow (2018)</xref>
</td>
<td align="char" char=".">82.7</td>
<td align="center">56.4</td>
<td align="center">52.7</td>
<td align="center">46.8</td>
<td align="center">59.4</td>
<td align="char" char=".">2.6</td>
</tr>
<tr>
<td align="left">Our <bold>AT</bold>
</td>
<td align="char" char=".">83.5</td>
<td align="center">56.9</td>
<td align="center">53.0</td>
<td align="center">48.4</td>
<td align="center">59.2</td>
<td align="char" char=".">2.6</td>
</tr>
<tr>
<td align="left">TRADES <xref ref-type="bibr" rid="B34">Zhang et&#x20;al. (2019)</xref>
</td>
<td align="char" char=".">82.1</td>
<td align="center">58.1</td>
<td align="center">54.6</td>
<td align="center">49.0</td>
<td align="center">58.9</td>
<td align="char" char=".">3.7</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>RayS <xref ref-type="bibr" rid="B6">Chen and Gu (2020)</xref> performs adversarial attack from a different perspective. As RayS is gradient-free and independent of certain adversarial losses, it can be used to detect possible falsely robust models, especially those may overfit to specific types of gradient-based attacks and adversarial losses. As seen in <xref ref-type="table" rid="T1">Table&#x20;1</xref>, all advanced adversarial training methods including AL, AT and TRADES, show higher robustness under RayS attack. Our results are consistent with those reported in RayS <xref ref-type="bibr" rid="B6">Chen and Gu (2020)</xref> that, when evaluated on really robust models, the robust accuracy of RayS is usually higher than that of standard&#x20;PGD.</p>
<sec id="s5-1">
<title>5.1 Experiments: Gray and Black-Box Settings</title>
<p>To evaluate the effectiveness of our defense strategy, we performed a series of image-classification experiments on <bold>17 Flower Category Database</bold> <xref ref-type="bibr" rid="B22">Nilsback and Zisserman (2006)</xref>, <bold>Part of ImageNet Database</bold> and <bold>Dogs-vs.-Cats Database</bold>. Following <xref ref-type="bibr" rid="B2">Athalye et&#x20;al. (2018)</xref>; <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref>, we assume an adversary that uses the state of the art PGD adversarial attack method.</p>
<p>We consider untargeted attacks when evaluating under the gray and black-box settings; untargeted attacks are also used in our adversarial training. We evaluate top-1 classification accuracy on validation images that are adversarially perturbed by the attacker. In this paper, adversarial perturbation is considered under <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> norm. The value of <italic>&#x3f5;</italic> is relative to the pixel intensity scale of 256, we use <italic>&#x3f5;</italic> &#x3d; 64/256 &#x3d; 0.25 and <italic>&#x3f5;</italic> &#x3d; 128/256 &#x3d; 0.5. PGD attacker with 10&#x2013;200 attack iterations and step size <italic>&#x3b1;</italic> &#x3d; 1.0/256 &#x3d; 0.0039. Our baselines are ResNet-101/152. There are four groups of convolutional structures in the baseline model, group-0 extracts of low-level features, group-1 and group-2 extract of mid-level features, group-3 extracts of high-level features <xref ref-type="bibr" rid="B33">Zagoruyko and Komodakis (2016)</xref>, which are described as <italic>conv</italic>2_<italic>x</italic>, <italic>conv</italic>3_<italic>x</italic>, <italic>conv</italic>4_<italic>x</italic> and <italic>conv</italic>5_<italic>x</italic> in <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref>.</p>
</sec>
<sec id="s5-2">
<title>5.2 Image Database</title>
<p>We performed a series of image-classification experiments on a wide range of datasets.<list list-type="simple">
<list-item>
<p>&#x2022; <bold>17 Flower Category Database</bold> <xref ref-type="bibr" rid="B22">Nilsback and Zisserman (2006)</xref> contains images of flowers belonging to 17 different categories. The images were acquired by searching the web and taking pictures. There are 80 images for each category.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Part of ImageNet Database</bold> contains images of four objects. These four objects are randomly selected from the ImageNet Database <xref ref-type="bibr" rid="B26">Russakovsky et&#x20;al. (2015)</xref>. In this experiment, they are tench, goldfish, white shark and dog. Each object contains 1,300 training images and 50 test images.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Dogs-vs.-Cats Database</bold>
<xref ref-type="fn" rid="FN1">
<sup>1</sup>
</xref> contains 8,000 images of dogs and cats in the train dataset and 2,000 in the test val dataset.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s5-3">
<title>5.3 Experimental Setup</title>
<p>To perform image classification, we use ResNet-101/152 that were trained on the <bold>17 Flower Category Database</bold>, <bold>Part of ImageNet Database</bold> and <bold>Dogs-vs.-Cats Database</bold> training set. We consider two different attack settings: 1) a gray-box attack setting in which the model used to generate the adversarial images is the same as the image-classification model, viz. the ResNet-101; and 2) a black-box attack setting in which the adversarial images are generated using the ResNet-152 model; The backend prediction model of gray-box and black-box is ResNet-101 with different implementations of the state of the art defense methods, such as IGR <xref ref-type="bibr" rid="B25">Ross and Doshi-Velez (2017)</xref>, PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>, RAT <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref>,Randomization <xref ref-type="bibr" rid="B31">Xie et&#x20;al. (2017)</xref>, ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>, FD <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref> and ADP <xref ref-type="bibr" rid="B23">Pang et&#x20;al. (2019)</xref>.</p>
</sec>
<sec id="s5-4">
<title>5.4 Results and Discussion</title>
<p>Here, we first present results with <bold>AT &#x2b; ALP</bold> on <bold>17 Flower Category Database</bold>. Compared with previous work, <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref> was evaluated under 10-iteration PGD attack and <italic>&#x3f5;</italic> &#x3d; 0.0625, our work are evaluated under highly challenging PGD attack:the maximum perturbation <italic>&#x3f5;</italic> &#x2208; {0.25, 0.5}, i.e.,&#x20;<italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> &#x2208; {0.25, 0.5} with 10&#x2013;200 attack iterations. The bigger the value of <italic>&#x3f5;</italic>, the bigger the disturbance, the more significant the adversarial image effect is. To the best of our knowledge, such a strong attack has not been previously explored on a wide range of datasets. As shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> that <bold>our AT &#x2b; ALP outperform the state-of-the-art in adversarial robustness against highly challenging gray-box and black-box PGD attacks</bold>. For example, under strong 200-iteration <bold>PGD</bold> gray-box and black-box attacks where prior art has 34 and 39% accuracy, our method achieves <bold>50</bold> and&#x20;<bold>51%</bold>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Defense against gray-box and black-box attacks on 17 Flower Category Database. <bold>(A,C)</bold> shows results against a gray-box PGD attacker with 10&#x2013;200 attack iterations. <bold>(B,D)</bold> shows results against a black-box PGD attacker with 10&#x2013;200 attack iterations. The maximum perturbation is <italic>&#x3f5;</italic> &#x2208; {0.25, 0.5}, i.e.,&#x20;<italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub> &#x2208; {0.25, 0.5}.Our <bold>AT &#x2b; ALP</bold> (purple line) outperform the state-of-the-art in adversarial robustness against highly challenging gray-box and black-box PGD attacks.</p>
</caption>
<graphic xlink:href="frai-04-752831-g003.tif"/>
</fig>
<p>
<xref ref-type="table" rid="T2">Table&#x20;2</xref> shows <bold>Main Result</bold> of our work: under strong 200-iteration PGD gray-box and black-box attacks, <bold>our AT &#x2b; ALP outperform the state-of-the-art in adversarial robustness on all these databases</bold>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Defense against gray-box and black-box attacks on 17 Flower Category Database, Part of ImageNet Database and Dogs-vs.-Cats Database. The adversarial perturbation were produced using PGD with step size <italic>&#x3b1;</italic> &#x3d; 1.0/256 &#x3d; 0.0039 and 200 attack iterations. As shown in this table, <bold>AT &#x2b; ALP got the highest Top-1 Accuracy on all these database</bold>.</p>
</caption>
<table>
<thead>
<tr>
<td align="left">17 flower category database</td>
<td colspan="2" align="center">Gray-box</td>
<td colspan="2" align="center">Black-box</td>
</tr>
<tr>
<td align="left">
<italic>&#x03B5;</italic> &#x3d; <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub>
</td>
<td align="center">0.25</td>
<td align="center">0.5</td>
<td align="center">0.25</td>
<td align="center">0.5</td>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">No Defence</td>
<td align="center">0</td>
<td align="center">0</td>
<td align="center">15</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">IGR <xref ref-type="bibr" rid="B25">Ross and Doshi-Velez (2017)</xref>
</td>
<td align="center">10</td>
<td align="center">3</td>
<td align="center">17</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>
</td>
<td align="center">55</td>
<td align="center">34</td>
<td align="center">57</td>
<td align="center">39</td>
</tr>
<tr>
<td align="left">RAT <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref>
</td>
<td align="center">54</td>
<td align="center">30</td>
<td align="center">57</td>
<td align="center">32</td>
</tr>
<tr>
<td align="left">Randomization <xref ref-type="bibr" rid="B31">Xie et&#x20;al. (2017)</xref>
</td>
<td align="center">12</td>
<td align="center">6</td>
<td align="center">27</td>
<td align="center">16</td>
</tr>
<tr>
<td align="left">ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>
</td>
<td align="center">47</td>
<td align="center">23</td>
<td align="center">49</td>
<td align="center">25</td>
</tr>
<tr>
<td align="left">FD <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref>
</td>
<td align="center">33</td>
<td align="center">10</td>
<td align="center">33</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">ADP <xref ref-type="bibr" rid="B23">Pang et&#x20;al. (2019)</xref>
</td>
<td align="center">22</td>
<td align="center">8</td>
<td align="center">23</td>
<td align="center">8</td>
</tr>
<tr>
<td align="left">Our <bold>AT</bold>
</td>
<td align="center">41</td>
<td align="center">24</td>
<td align="center">45</td>
<td align="center">29</td>
</tr>
<tr>
<td align="left">Our <bold>AT &#x2b; ALP</bold>
</td>
<td align="center">68</td>
<td align="center">50</td>
<td align="center">70</td>
<td align="center">51</td>
</tr>
</tbody>
</table>
<table>
<thead>
<tr>
<th align="left">
<bold>Part of ImageNet database</bold>
</th>
<th colspan="2" align="center">
<bold>Gray-box</bold>
</th>
<th colspan="2" align="center">
<bold>Black-box</bold>
</th>
</tr>
<tr>
<th align="left">
<bold>
<italic>&#x03B5;</italic>
</bold> <bold>&#x3d;</bold> <bold>
<italic>L</italic>
</bold>
<sub>
<bold>
<italic>&#x221e;</italic>
</bold>
</sub>
</th>
<th align="center">
<bold>0.25</bold>
</th>
<th align="center">
<bold>0.5</bold>
</th>
<th align="center">
<bold>0.25</bold>
</th>
<th align="center">
<bold>0.5</bold>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">No Defence</td>
<td align="center">2</td>
<td align="center">3</td>
<td align="center">52</td>
<td align="center">50</td>
</tr>
<tr>
<td align="left">IGR <xref ref-type="bibr" rid="B25">Ross and Doshi-Velez (2017)</xref>
</td>
<td align="center">32</td>
<td align="center">32</td>
<td align="center">34</td>
<td align="center">34</td>
</tr>
<tr>
<td align="left">PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>
</td>
<td align="center">76</td>
<td align="center">76</td>
<td align="center">77</td>
<td align="center">77</td>
</tr>
<tr>
<td align="left">RAT <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref>
</td>
<td align="center">76</td>
<td align="center">76</td>
<td align="center">77</td>
<td align="center">76</td>
</tr>
<tr>
<td align="left">Randomization <xref ref-type="bibr" rid="B31">Xie et&#x20;al. (2017)</xref>
</td>
<td align="center">40</td>
<td align="center">41</td>
<td align="center">62</td>
<td align="center">59</td>
</tr>
<tr>
<td align="left">ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>
</td>
<td align="center">54</td>
<td align="center">54</td>
<td align="center">55</td>
<td align="center">55</td>
</tr>
<tr>
<td align="left">FD <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref>
</td>
<td align="center">60</td>
<td align="center">61</td>
<td align="center">61</td>
<td align="center">61</td>
</tr>
<tr>
<td align="left">ADP <xref ref-type="bibr" rid="B23">Pang et&#x20;al. (2019)</xref>
</td>
<td align="center">42</td>
<td align="center">44</td>
<td align="center">43</td>
<td align="center">44</td>
</tr>
<tr>
<td align="left">Our <bold>AT</bold>
</td>
<td align="center">76</td>
<td align="center">76</td>
<td align="center">77</td>
<td align="center">76</td>
</tr>
<tr>
<td align="left">Our <bold>AT &#x2b; ALP</bold>
</td>
<td align="center">82</td>
<td align="center">82</td>
<td align="center">82</td>
<td align="center">82</td>
</tr>
</tbody>
</table>
<table>
<thead>
<tr>
<th align="left">
<bold>Dogs-vs.-Cats Database</bold>
</th>
<th colspan="2" align="center">
<bold>Gray-box</bold>
</th>
<th colspan="2" align="center">
<bold>Black-box</bold>
</th>
</tr>
<tr>
<th align="left">
<bold>
<italic>&#x03B5;</italic>
</bold> <bold>&#x3d;</bold> <bold>
<italic>L</italic>
</bold>
<sub>
<bold>
<italic>&#x221e;</italic>
</bold>
</sub>
</th>
<th align="center">
<bold>0.25</bold>
</th>
<th align="center">
<bold>0.5</bold>
</th>
<th align="center">
<bold>0.25</bold>
</th>
<th align="center">
<bold>0.5</bold>
</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">No Defence</td>
<td align="center">1</td>
<td align="center">1</td>
<td align="center">52</td>
<td align="center">53</td>
</tr>
<tr>
<td align="left">IGR <xref ref-type="bibr" rid="B25">Ross and Doshi-Velez (2017)</xref>
</td>
<td align="center">57</td>
<td align="center">60</td>
<td align="center">51</td>
<td align="center">52</td>
</tr>
<tr>
<td align="left">PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>
</td>
<td align="center">51</td>
<td align="center">51</td>
<td align="center">52</td>
<td align="center">52</td>
</tr>
<tr>
<td align="left">RAT <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref>
</td>
<td align="center">49</td>
<td align="center">49</td>
<td align="center">50</td>
<td align="center">50</td>
</tr>
<tr>
<td align="left">Randomization <xref ref-type="bibr" rid="B31">Xie et&#x20;al. (2017)</xref>
</td>
<td align="center">10</td>
<td align="center">8</td>
<td align="center">55</td>
<td align="center">54</td>
</tr>
<tr>
<td align="left">ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>
</td>
<td align="center">57</td>
<td align="center">56</td>
<td align="center">57</td>
<td align="center">57</td>
</tr>
<tr>
<td align="left">FD <xref ref-type="bibr" rid="B32">Xie et&#x20;al. (2018)</xref>
</td>
<td align="center">57</td>
<td align="center">57</td>
<td align="center">57</td>
<td align="center">57</td>
</tr>
<tr>
<td align="left">ADP <xref ref-type="bibr" rid="B23">Pang et&#x20;al. (2019)</xref>
</td>
<td align="center">50</td>
<td align="center">50</td>
<td align="center">50</td>
<td align="center">50</td>
</tr>
<tr>
<td align="left">Our <bold>AT</bold>
</td>
<td align="center">50</td>
<td align="center">50</td>
<td align="center">50</td>
<td align="center">50</td>
</tr>
<tr>
<td align="left">Our <bold>AT &#x2b; ALP</bold>
</td>
<td align="center">67</td>
<td align="center">67</td>
<td align="center">71</td>
<td align="center">71</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We visualized activation attention maps for defense against PGD attacks. Baseline model is ResNet-101 <xref ref-type="bibr" rid="B12">He et&#x20;al. (2015)</xref>, which is pre-trained on <bold>ImageNet</bold> <xref ref-type="bibr" rid="B26">Russakovsky et&#x20;al. (2015)</xref> and fine-tuned on <bold>17 Flower Category Database</bold> <xref ref-type="bibr" rid="B22">Nilsback and Zisserman (2006)</xref>, group-0 to group-3 represent the activation attention maps of four groups of convolutional structures in the baseline model, i.e.,&#x20;<italic>conv</italic>2_<italic>x</italic>, <italic>conv</italic>3_<italic>x</italic>, <italic>conv</italic>4_<italic>x</italic> and <italic>conv</italic>5_<italic>x</italic> of ResNet-101, group-0 extracts of low-level features, group-1 and group-2 extract of mid-level features, group-3 extracts of high-level features <xref ref-type="bibr" rid="B33">Zagoruyko and Komodakis (2016)</xref>;. We found from <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> that group-0 of <bold>AT &#x2b; ALP</bold> can extract the outline and texture of flowers more accurately, and group-3 has a higher level of activation on the whole flower, compared with other defense methods, only <bold>AT &#x2b; ALP</bold> makes accurate prediction.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Activation attention maps for defense against gray-box PGD attacks (<italic>&#x3f5;</italic> &#x3d; 0.25) on 17 Flower Category Database. <bold>(A)</bold> is original image and <bold>(B)</bold> is corresponding adversarial image. <bold>(C)</bold> and <bold>(D)</bold> are activation attention maps of baseline model for original image and adversarial image, <bold>(E,F,G)</bold> are activation attention maps of <bold>ALP</bold>, <bold>AT</bold> and <bold>AT &#x2b; ALP</bold> for adversarial image. Group-0 to group-3 represent the activation attention maps of four groups of convolutional structures in the baseline model, group-0 extracts of low-level features, group-1 and group-2 extract of mid-level features, group-3 extracts of high-level features <xref ref-type="bibr" rid="B33">Zagoruyko and Komodakis (2016)</xref>. It can be clearly found that group-0 of AT &#x2b; ALP can extract the outline and texture of flowers more accurately, and group-3 has a higher level of activation on the whole flower, compared with other defense methods, only it makes accurate prediction.</p>
</caption>
<graphic xlink:href="frai-04-752831-g004.tif"/>
</fig>
<p>We compared average activations on discriminate parts of <bold>17 Flower Category Database</bold> for different defense methods. <bold>17 Flower Category Database</bold> defined discriminative parts of flowers. See <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> for an illustrative example. These discriminative parts are annotated by humans, according to their contributions to recognize a target. In other words, they are crucial features for the classification. For example, the head and feather should be discriminative parts to recognize a species of bird. Using all testing examples of <bold>17 Flower Category Database</bold>, we calculated normalized activations on these key regions of these different defense methods. As shown in <xref ref-type="table" rid="T3">Table&#x20;3</xref>, <bold>AT &#x2b; ALP</bold> got the highest average activations on those key regions, demonstrating that <bold>AT &#x2b; ALP</bold> focused on more discriminate features for flowers recognition. We also demonstrate in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> that <bold>AT &#x2b; ALP</bold> shows smoother loss landscapes, which further verifies its effectiveness.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>
<bold>(A)</bold> is original image and <bold>(B)</bold> is corresponding discriminative parts. <bold>17 Flower Category Database</bold> defined discriminative parts of flowers. So for each image, we got several key regions which are very important to discriminate its category.</p>
</caption>
<graphic xlink:href="frai-04-752831-g005.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Comparing <bold>average activations</bold> on discriminate parts of <bold>17 Flower Category Database</bold> for different defense methods. In addition, we included new statistical results of activations on part locations of <bold>17 Flower Category Database</bold> supporting the above qualitative cases. The <bold>17 Flower Category Database</bold> defined discriminative parts of flowers. So for each image, we got several key regions which are very important to discriminate its category. Using all testing examples of <bold>17 Flower Category Database</bold>, we calculated normalized activations on these key regions of these different defense methods. As shown in this table, <bold>AT &#x2b; ALP</bold> got the highest average activations on those key regions, demonstrating that <bold>AT &#x2b; ALP</bold> focused on more discriminate features for flowers recognition.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Defense</th>
<th colspan="2" align="center">Black-box</th>
<th colspan="2" align="center">Gray-box</th>
</tr>
<tr>
<td align="left">
<italic>&#x03B5;</italic> &#x3d; <italic>L</italic>
<sub>
<italic>&#x221e;</italic>
</sub>
</td>
<td align="char" char=".">0.25</td>
<td align="char" char=".">0.5</td>
<td align="char" char=".">0.25</td>
<td align="char" char=".">0.5</td>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">No Defense</td>
<td align="char" char=".">0.41</td>
<td align="char" char=".">0.41</td>
<td align="char" char=".">0.21</td>
<td align="char" char=".">0.21</td>
</tr>
<tr>
<td align="left">ALP <xref ref-type="bibr" rid="B13">Kannan et&#x20;al. (2018)</xref>
</td>
<td align="char" char=".">0.16</td>
<td align="char" char=".">0.16</td>
<td align="char" char=".">0.15</td>
<td align="char" char=".">0.15</td>
</tr>
<tr>
<td align="left">IGR <xref ref-type="bibr" rid="B25">Ross and Doshi-Velez (2017)</xref>
</td>
<td align="char" char=".">0.37</td>
<td align="char" char=".">0.37</td>
<td align="char" char=".">0.33</td>
<td align="char" char=".">0.33</td>
</tr>
<tr>
<td align="left">PAT <xref ref-type="bibr" rid="B19">Madry et&#x20;al. (2017)</xref>
</td>
<td align="char" char=".">0.42</td>
<td align="char" char=".">0.42</td>
<td align="char" char=".">0.44</td>
<td align="char" char=".">0.44</td>
</tr>
<tr>
<td align="left">RAT <xref ref-type="bibr" rid="B1">Araujo et&#x20;al. (2019)</xref>
</td>
<td align="char" char=".">0.40</td>
<td align="char" char=".">0.40</td>
<td align="char" char=".">0.41</td>
<td align="char" char=".">0.41</td>
</tr>
<tr>
<td align="left">Our <bold>AT</bold>
</td>
<td align="char" char=".">0.55</td>
<td align="char" char=".">0.54</td>
<td align="char" char=".">0.56</td>
<td align="char" char=".">0.56</td>
</tr>
<tr>
<td align="left">Our <bold>AT &#x2b; ALP</bold>
</td>
<td align="char" char=".">0.98</td>
<td align="char" char=".">0.98</td>
<td align="char" char=".">0.96</td>
<td align="char" char=".">0.96</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Comparison of loss landscapes generated by No defence <bold>(A)</bold>, ALP <bold>(B)</bold>, AT <bold>(C)</bold> and AT&#x002B;ALP <bold>(D)</bold>. We see that <bold>ALP</bold> and <bold>AT</bold> sometimes induces decreased loss near the input locally, and gives a &#x201c;bumpier&#x201d; optimization landscape, our <bold>AT &#x2b; ALP</bold> has better robustness. The <italic>z</italic> axis represents the loss. If <italic>x</italic> is the original input, then we plot the loss varying along the space determined by two vectors: <inline-formula id="inf8">
<mml:math id="m12">
<mml:mi>r</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>n</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x25bd;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> and <italic>r</italic>2 &#x223c; <italic>Rademacher</italic> (0.5). We thus plot the following function: <italic>z</italic>&#x20;&#x3d; <italic>loss</italic> (<italic>x</italic>&#x20;&#x22c5; <italic>r</italic>1 &#x2b; <italic>y</italic>&#x20;&#x22c5; <italic>r</italic>2).</p>
</caption>
<graphic xlink:href="frai-04-752831-g006.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion</title>
<p>In this paper, we introduced enhanced defense using a technique we called <bold>Attention and Adversarial Logit Pairing (AT &#x2b; ALP)</bold>, a method that encouraged both attention map and logit for pairs of examples to be similar. When being applied to clean examples and their adversarial counterparts, <bold>AT &#x2b; ALP</bold> improved accuracy on adversarial examples over adversarial training. Our <bold>AT &#x2b; ALP</bold> achieves <bold>the state of the art</bold> defense on a wide range of datasets against <bold>PGD</bold> gray-box and black-box attacks. Compared with other defense methods, our <bold>AT &#x2b; ALP</bold> is simple and effective, without modifying the model structure, and without adding additional image preprocessing&#x20;steps.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>XL and DG conducted the experiments and initial writing. JL, TW and DD helped revised the paper.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>XL, DG, JL, TW and DD were employed by the company Baidu Inc.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="FN1">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/chetankv/dogs-cats-images">https://www.kaggle.com/chetankv/dogs-cats-images</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Araujo</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pinot</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Negrevergne</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Meunier</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Atif</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Robust Neural Networks Using Randomized Adversarial Training</article-title>. <comment>arXiv preprint arXiv:1903.10219</comment>. </citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Athalye</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Carlini</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wagner</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples</article-title>,&#x201d; in <conf-name>In Proceedings of the 35th International Conference on Machine Learning</conf-name>, <conf-loc>Stockholmsm&#x00E4;ssan, Stockholm, Sweden</conf-loc>, <conf-date>July 10&#x2013;15, 2018</conf-date> (<publisher-name>ICML 2018</publisher-name>). </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bose</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Aarabi</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Adversarial Attacks on Face Detectors Using Neural Net Based Constrained Optimization</article-title>,&#x201d; in <conf-name>2018 IEEE 20th International Workshop on Multimedia Signal Processing (MMSP)</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>, <conf-date>August 29&#x2013;31, 2018</conf-date>. <pub-id pub-id-type="doi">10.1109/mmsp.2018.8547128</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Buckman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Roy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Raffel</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Thermometer Encoding: One Hot Way to Resist Adversarial Examples</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations</conf-name>, <conf-loc>Vancouver Convention Center, Vancouver, Canada</conf-loc>, <conf-date>April 30&#x2013;May 3, 2018</conf-date>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carlini</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wagner</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Towards Evaluating the Robustness of Neural Networks</article-title>. <comment>arXiv preprint arXiv:1608.04644</comment>. </citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Rays: A ray Searching Method for Hard-Label Adversarial Attack</article-title>,&#x201d; in <conf-name>Proceedings of the 26rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</conf-name>, <conf-date>August 23&#x2013;27, 2020</conf-date>. <pub-id pub-id-type="doi">10.1145/3394486.3403225</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Croce</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Hein</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Reliable Evaluation of Adversarial Robustness with an Ensemble of Diverse Parameter-free Attacks</article-title>,&#x201d; in <conf-name>International conference on machine learning</conf-name>, <conf-date>July 13&#x2013;18, 2020</conf-date> (<publisher-name>PMLR</publisher-name>), <fpage>2206</fpage>&#x2013;<lpage>2216</lpage>. </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dhillon</surname>
<given-names>G. S.</given-names>
</name>
<name>
<surname>Azizzadenesheli</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lipton</surname>
<given-names>Z. C.</given-names>
</name>
<name>
<surname>Bernstein</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kossaifi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Khanna</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Stochastic Activation Pruning for Robust Adversarial Defense</article-title>. <comment>arXiv preprint arXiv:1803.01442</comment>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I. J.</given-names>
</name>
<name>
<surname>Shlens</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Explaining and Harnessing Adversarial Examples</article-title>. <conf-name>3rd International Conference on Learning Representations, (ICLR) 2015</conf-name>, <conf-loc>San Diego, CA</conf-loc>, <conf-date>May 7&#x2013;9, 2015</conf-date>. <comment>arxiv.org/abs/1412.6572</comment>. </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rana</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cisse</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Van Der Maaten</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Countering Adversarial Images Using Input Transformations</article-title>. <comment>arXiv preprint arXiv:1711.00117</comment>. </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep Residual Learning for Image Recognition</article-title>. <comment>CoRR abs/1512.03385</comment>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kannan</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kurakin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Goodfellow</surname>
<given-names>I. J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Adversarial Logit Pairing</article-title>. <comment>CoRR abs/1803.06373</comment>. </citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2009</year>). <source>Learning Multiple Layers of Features from Tiny Images</source>. </citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Imagenet Classification with Deep Convolutional Neural Networks</article-title>,&#x201d; in <conf-name>International Conference on Neural Information Processing Systems</conf-name>, <conf-loc>Harrahs and Harveys, Lake Tahoe</conf-loc>, <conf-date>Dec 3&#x2013;8, 2012</conf-date>. </citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Qian</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). &#x201c;<article-title>Dsfd: Dual Shot Face Detector</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Long Beach, CA</conf-loc>, <conf-date>June 16&#x2013;20, 2019</conf-date>, <fpage>5060</fpage>&#x2013;<lpage>5069</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2019.00520</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Rao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huan</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019b</year>). <article-title>Delta: Deep Learning Transfer Using Feature Map with Attention for Convolutional Networks</article-title>. <comment>arXiv preprint arXiv:1901.09229</comment>. </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Erfani</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Wijewickrema</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schoenebeck</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality</article-title>. <comment>arXiv preprint arXiv:1801.02613</comment>. </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Madry</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Makelov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Schmidt</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tsipras</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Vladu</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Towards Deep Learning Models Resistant to Adversarial Attacks</article-title>. <comment>arXiv preprint arXiv:1706.06083</comment>. </citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Moosavi-Dezfooli</surname>
<given-names>S.-M.</given-names>
</name>
<name>
<surname>Fawzi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Frossard</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deepfool: a Simple and Accurate Method to Fool Deep Neural Networks</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <conf-loc>Las Vegas</conf-loc>, <conf-date>June 26&#x2013;July 1, 2016</conf-date>, <fpage>2574</fpage>&#x2013;<lpage>2582</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2016.282</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Na</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ko</surname>
<given-names>J.&#x20;H.</given-names>
</name>
<name>
<surname>Mukhopadhyay</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Cascade Adversarial Machine Learning Regularized with a Unified Embedding</article-title>. <comment>arXiv preprint arXiv:1708.02582</comment>. </citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Nilsback</surname>
<given-names>M.-E.</given-names>
</name>
<name>
<surname>Zisserman</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>A Visual Vocabulary for Flower Classification</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>New York, NY</conf-loc>, <conf-date>June 17&#x2013;22, 2006</conf-date> <volume>2</volume>, <fpage>1447</fpage>&#x2013;<lpage>1454</lpage>. </citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Improving Adversarial Robustness via Promoting Ensemble Diversity</article-title>. In<source>International Conference on Machine Learning</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>9</fpage>&#x2013;<lpage>15</lpage>. </citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pfleeger</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Pfleeger</surname>
<given-names>S. L.</given-names>
</name>
</person-group> (<year>2004</year>). <source>Security in Computing</source>. <edition>3rd Edn</edition>. <publisher-name>Prentice Hall PTR</publisher-name>. </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ross</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Doshi-Velez</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing Their Input Gradients</source>. <conf-name>Thirty-second AAAI conference on artificial intelligence 2018</conf-name>, <conf-loc>New Orleans, LA</conf-loc>, <conf-date>February 2&#x2013;7, 2018</conf-date>. </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Russakovsky</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Krause</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Satheesh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>ImageNet Large Scale Visual Recognition Challenge</article-title>. <source>Int. J.&#x20;Comput. Vis.</source> <volume>115</volume> (<issue>3</issue>), <fpage>211</fpage>&#x2013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Samangouei</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kabkab</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chellappa</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Defense-gan: Protecting Classifiers against Adversarial Attacks Using Generative Models</article-title>. <comment>arXiv preprint arXiv:1805.06605</comment>. </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Nowozin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ermon</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kushman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Pixeldefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples</article-title>. <comment>arXiv preprint arXiv:1710.10766</comment>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zaremba</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bruna</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Erhan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Intriguing Properties of Neural Networks</article-title>. <comment>arXiv preprint arXiv:1312.6199</comment>. </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tram&#xe8;r</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kurakin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Papernot</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Boneh</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>McDaniel</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Ensemble Adversarial Training: Attacks and Defenses</article-title>. <comment>arXiv preprint arXiv:1705.07204</comment>. </citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yuille</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Mitigating Adversarial Effects through Randomization</source>. <conf-name>International Conference on Learning Representations. 2018</conf-name>, <conf-loc>Vancouver Convention Center, Vancouver, Canada</conf-loc>, <conf-date>April 30&#x2013;May 3, 2018</conf-date>. </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Maaten</surname>
<given-names>L. V. D.</given-names>
</name>
<name>
<surname>Yuille</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Feature Denoising for Improving Adversarial Robustness</source>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Long Beach, CA</conf-loc>, <conf-date>June 16&#x2013;20, 2019</conf-date>. </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zagoruyko</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Komodakis</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer</article-title>. <comment>CoRR abs/1612.03928</comment>. </citation>
</ref>
<ref id="B34">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jiao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>E. P.</given-names>
</name>
<name>
<surname>Ghaoui</surname>
<given-names>L. E.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>M. I.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Theoretically Principled Trade-Off between Robustness and Accuracy</article-title>,&#x201d; in <conf-name>International Conference on Machine Learning</conf-name>, <conf-loc>Long Beach, CA</conf-loc>, <conf-date>June 16&#x2013;20, 2019</conf-date>. </citation>
</ref>
</ref-list>
</back>
</article>