<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2021.785808</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>ShapeEditor: A StyleGAN Encoder for Stable and High Fidelity Face Swapping</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Yang</surname> <given-names>Shuai</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1497932/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Qiao</surname> <given-names>Kai</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/527442/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Qin</surname> <given-names>Ruoxi</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Xie</surname> <given-names>Pengfei</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1487301/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Shi</surname> <given-names>Shuhao</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1478134/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liang</surname> <given-names>Ningning</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Linyuan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/572226/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Jian</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/612406/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Hu</surname> <given-names>Guoen</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1400316/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Yan</surname> <given-names>Bin</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/572228/overview"/>
</contrib>
</contrib-group>
<aff><institution>Henan Key Laboratory of Imaging and Intelligent Processing, People&#x00027;s Liberation Army (PLA) Strategy Support Force Information Engineering University</institution>, <addr-line>Zhengzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Xin Jin, Yunnan University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Durai Raj Vincent P.M, VIT University, India; Yanan Guo, Beijing Information Science and Technology University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Bin Yan <email>ybspace&#x00040;hotmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>785808</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Yang, Qiao, Qin, Xie, Shi, Liang, Wang, Chen, Hu and Yan.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yang, Qiao, Qin, Xie, Shi, Liang, Wang, Chen, Hu and Yan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>With the continuous development of deep-learning technology, ever more advanced face-swapping methods are being proposed. Recently, face-swapping methods based on generative adversarial networks (GANs) have realized many-to-many face exchanges with few samples, which advances the development of this field. However, the images generated by previous GAN-based methods often show instability. The fundamental reason is that the GAN in these frameworks is difficult to converge to the distribution of face space in training completely. To solve this problem, we propose a novel face-swapping method based on pretrained StyleGAN generator with a stronger ability of high-quality face image generation. The critical issue is how to control StyleGAN to generate swapped images accurately. We design the control strategy of the generator based on the idea of encoding and decoding and propose an encoder called ShapeEditor to complete this task. ShapeEditor is a two-step encoder used to generate a set of coding vectors that integrate the identity and attribute of the input faces. In the first step, we extract the identity vector of the source image and the attribute vector of the target image; in the second step, we map the concatenation of the identity vector and attribute vector onto the potential internal space of StyleGAN. Extensive experiments on the test dataset show that the results of the proposed method are not only superior in clarity and authenticity than other state-of-the-art methods but also sufficiently integrate identity and attribute.</p></abstract>
<kwd-group>
<kwd>face swapping</kwd>
<kwd>generative adversarial network</kwd>
<kwd>disentanglement</kwd>
<kwd>style transfer</kwd>
<kwd>deepfake</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="3"/>
<equation-count count="5"/>
<ref-count count="37"/>
<page-count count="10"/>
<word-count count="6788"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>As one of the main contents of deepfake, face swapping declares to the world today that seeing is not always believing. Face swapping refers to transferring the identity of a source image to the face of another target image while keeping unchanged the illumination, head posture, expression, dress, background, and other attribute information of the target image. Face swapping has received widespread attention since its birth, catering to the affluent needs of social life, such as hairstyle simulation, film and television shooting, privacy protection, and so on (Ross and Othman, <xref ref-type="bibr" rid="B26">2010</xref>).</p>
<p>Face swapping is accompanied not only by its interesting and operational application prospects but also by various challenges between reality and vision. The early face-swapping methods (Bitouk et al., <xref ref-type="bibr" rid="B4">2008</xref>; Korshunova et al., <xref ref-type="bibr" rid="B13">2017</xref>) require many images of source and target characters to provide sufficient facial information. Otherwise, the models would not have a suitable reference basis to produce good results. Some three-dimensional-based (3D-based) methods (Olszewski et al., <xref ref-type="bibr" rid="B23">2017</xref>; Nirkin et al., <xref ref-type="bibr" rid="B21">2018</xref>; Sun et al., <xref ref-type="bibr" rid="B30">2018</xref>) make use of the advantage of fitting 3D face models to deal with the problems of large angle and small samples. At the same time, due to the limited accuracy of 3D face models, it is impossible to generate works with better details and higher fidelity. Recently, with the continuous tapping of the potential of generative adversarial networks (GANs) (Nandhini Abirami et al., <xref ref-type="bibr" rid="B17">2021</xref>), some face-swapping methods based on GANs (Bao et al., <xref ref-type="bibr" rid="B3">2018</xref>; Natsume et al., <xref ref-type="bibr" rid="B18">2018a</xref>,<xref ref-type="bibr" rid="B19">b</xref>; Li et al., <xref ref-type="bibr" rid="B15">2019</xref>; Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>) can achieve a good fusion of identity and attribute information with only a small number of samples, reflecting the effect of great creativity. Unfortunately, the surprising creativity of these methods does not offset the adverse impacts of their frequent artifacts and low-resolution limitation.</p>
<p>On another track, the most advanced face image generation methods have generated facial images with high resolution and realistic texture. Most notably, StyleGAN (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>) can randomly generate a variety of clear faces with a resolution of up to 1024 &#x000D7; 1024. StyleGAN has three potential spaces: initial potential space <inline-formula><mml:math id="M1"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>Z</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula>, intermediate potential space <inline-formula><mml:math id="M2"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula>, and extended potential space <inline-formula><mml:math id="M3"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula>. (Abdal et al., <xref ref-type="bibr" rid="B1">2019</xref>) proved that the concatenation of 18 different 512-dimensional vectors is the easiest way to embed an image and obtain a reasonable result. On this basis, various works (Gu et al., <xref ref-type="bibr" rid="B6">2020</xref>; H&#x000E4;rk&#x000F6;nen et al., <xref ref-type="bibr" rid="B8">2020</xref>; Richardson et al., <xref ref-type="bibr" rid="B25">2020</xref>; Zhu et al., <xref ref-type="bibr" rid="B37">2020</xref>) explore in detail the StyleGAN potential vector space: some (Shen and Zhou, <xref ref-type="bibr" rid="B29">2020</xref>; Shen et al., <xref ref-type="bibr" rid="B28">2020</xref>; Tewari et al., <xref ref-type="bibr" rid="B31">2020</xref>) find a linear direction to control the change of a single facial attribute, some (Nitzan et al., <xref ref-type="bibr" rid="B22">2020</xref>) control facial expression and posture in the original StyleGAN image domain, and others (Richardson et al., <xref ref-type="bibr" rid="B25">2020</xref>; Wang et al., <xref ref-type="bibr" rid="B32">2021</xref>) deal well with the difficult task of facial super-resolution.</p>
<p>In contrast with other face-swapping methods, the first criterion we pursue is that the images after face swapping have both higher clarity and better authenticity. We propose a many-to-many face-swapping method based on the pretrained StyleGAN model (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>), which strives to ensure the clarity and fidelity of the results while fusing identity and attribute information. Given the inherent ability of the pretrained StyleGAN model to generate random high-quality face images, the difficulty of this task is how to accurately render the corresponding latent vectors. To achieve this goal, we first designed an encoder, ShapeEditor, to find the corresponding codes in the <inline-formula><mml:math id="M4"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> vector space. The workflow of the encoder was divided into two stages, the first being the respective extraction of identity and attribute codes, and the second being to map the combination of two-channel codes into the potential input vector domain of the pretrained model. Moreover, we designed a set of loss functions with a strong monitoring ability to urge ShapeEditor to update parameters to learn to map step by step onto the latent space of StyleGAN. As verification, we made numerous qualitative and quantitative experimental comparisons with the existing face-swapping methods, which show the unique advantages of the proposed method.</p>
</sec>
<sec id="s2">
<title>2. Related Works</title>
<p>Recently, the GAN-based face-swapping methods have shown better performance, thus attracting more extensive research and attention. Although integrate attributes and identity information well, these methods generally have the common problem of poor clarity and authenticity. On the other hand, as GAN with better image quality has been proposed, many works are devoted to manipulating GAN&#x00027;s semantic space to generate clear and stable images. We creatively combine the advantages of the above two fields to improve the performance of face swapping, and make possible the more complex control of GAN&#x00027;s potential space.</p>
<sec>
<title>2.1. GAN-Based Face Swapping</title>
<p>Olszewski et al. (<xref ref-type="bibr" rid="B23">2017</xref>) fit the 3D face model of the source face and used a conditional generator of the coder-decoder structure to infer the converted face texture. Too simple generator network structure and training strategy make this method unable to separate identity and attribute information to further complete many-to-many identity exchange. Sun et al. (<xref ref-type="bibr" rid="B30">2018</xref>) trained a convolutional neural network to regress the parameters of a 3D model of the input face, replaced the identity parameters, and combined the region around the head to generate a realistic face-swapped image. Limited to the accuracy of the model reconstruction, 3D-based face-swapping methods are unsatisfactory in terms of attribute and identity fidelity. Face Swapping GAN (FSGAN) (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>) used sparse landmarks to track facial expression, and designed GANs with different functions for the three stages of face swapping. This method realized subject agnostic face swapping, while being limited by the resolution of the input image and the complexity of expression. Bao et al. (<xref ref-type="bibr" rid="B3">2018</xref>) implemented this task using a more concise coder-decoder architecture, in which two independent coders separate the identity and attributes of human faces. This method used an asymmetric training strategy to promote a large number of unlabeled faces to contribute to the training. Following the basic network framework and asymmetric training strategy of Bao et al. (<xref ref-type="bibr" rid="B3">2018</xref>), FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>) has done meaningful work on embedding multi-level information in the generator and handling occlusion more robustly. The generator leverages denormalizations for feature integration in multiple feature levels, showing a better representation of identity and attribute. However, the clarity and stability of the image generated by FaceShifter are not always ideal. As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, the eyebrows of the result in the first line appear ghosting, and the nose of the result in the second line appear artifact. These examples show that the most advanced GAN-based face-swapping method is still insufficient in authenticity.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Some abnormal results generated by FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-785808-g0001.tif"/>
</fig>
</sec>
<sec>
<title>2.2. The Potential and Challenge of Pretrained GAN Manipulation</title>
<p>While a lot of works have been done on how to control GAN to perform complex image operations, such as face swapping, others focus on improving the quality of images. Through carefully designed style-based network structure and layer-by-layer training, StyleGAN (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>) realized high-definition and high-quality face image generation. With the help of pretrained StyleGAN, image quality is easier to be improved. The manipulation of StyleGAN is a difficult task, and most early works are limited to understanding and reproducing the potential space of GAN. The inversion task of StyleGAN is to find the potential vector that best matches the given image. Abdal et al. (<xref ref-type="bibr" rid="B1">2019</xref>) took several minutes to embed a face into the StyleGAN image domain. Richardson et al. (<xref ref-type="bibr" rid="B25">2020</xref>), Zhu et al. (<xref ref-type="bibr" rid="B37">2020</xref>), and Gu et al. (<xref ref-type="bibr" rid="B6">2020</xref>) tried to improve efficiency using encoder structure, but the inversion results of wild images in their methods are unsatisfactory. Later, some more complex works appeared, such as changing individual attributes (smile, age, facial angle, etc.) (H&#x000E4;rk&#x000F6;nen et al., <xref ref-type="bibr" rid="B8">2020</xref>; Shen and Zhou, <xref ref-type="bibr" rid="B29">2020</xref>; Shen et al., <xref ref-type="bibr" rid="B28">2020</xref>), establishing relationship between 3D semantic parameters and genuine facial expressions (Tewari et al., <xref ref-type="bibr" rid="B31">2020</xref>), and super-resolution of low-quality facial images (Wang et al., <xref ref-type="bibr" rid="B32">2021</xref>). To the best of our knowledge, there is no face-swapping method based on StyleGAN. This task requires more complex semantic manipulation, and the current controllers are not competent. Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) did closely related work to control expression through latent space mapping. However, working in the <inline-formula><mml:math id="M5"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> space led to the failure of embedding wild images into potential space. In addition, the single vector of the attribute is too plain to carry the information of background, posture, expression, etc.</p>
</sec>
<sec>
<title>2.3. The Inheritance and Transcendence</title>
<p>We propose a StyleGAN encoder, called ShapeEditor, for stable and high-fidelity face swapping. As the combination of face swapping and pretrained GAN manipulation, ShapeEditor inherits and surpasses the latest ideas in the two fields.</p>
<p>We use an asymmetric training strategy similar to that in FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>) to realize the training process without labeled data, so as to ensure solid constraints and reduce data processing costs. Moreover, the well-designed coder-decoder structure of our framework can firmly guarantee image quality, which is the weakest aspect of FaceShifter. Inspired by SPADE (Park et al., <xref ref-type="bibr" rid="B24">2019</xref>) and AdaIN (Huang and Belongie, <xref ref-type="bibr" rid="B9">2017</xref>), the FaceShifter generator designs AAD layer-level denormalization for feature integration in multiple feature levels. By comparison, the internal mapper of ShapeEditor is composed of lightweight Multilayer Perceptrons (MLP) to generate feature vectors embedded in StyleGAN <inline-formula><mml:math id="M6"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space, which reduces the burden of model training.</p>
<p>Our method and Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) both use the decoupling framework to extract attribute and identity code through attribute extractor and identity extractor, respectively. The codes are then mapped into the latent space of the employed pretrained generator. Our key difference is that we select <inline-formula><mml:math id="M7"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> potential space as the mapping space, which is the premise of realizing the complex semantic operation of face swapping. In addition, in order to recover the attribute information more finely, we use multi-level feature mapping instead of a single output as attribute code like Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) did. The ablation study proves that our pertinent designs make a significant contribution to better semantic manipulation.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3. Methods</title>
<p>Our method requires two images as input: <italic>I</italic><sub>attr</sub> and <italic>I</italic><sub>id</sub>. We expect the output of the model to reflect the identity of <italic>I</italic><sub>id</sub> and the facial expression, head posture, hairstyle, lighting, and other attribute information of <italic>I</italic><sub>attr</sub>. Therefore, the main challenge of this work is to obtain the StyleGAN potential vectors that are consistent with the <inline-formula><mml:math id="M8"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> spatial distribution and better integrate attributes and identity. To solve this problem, we designed a two-step coding process. As shown in <xref ref-type="fig" rid="F2">Figure 2A</xref>, the entire mapping process is divided into two phases: ID-ATTR encoding and latent-space encoding. In the first stage, <italic>E</italic><sub>id</sub> extracts the identity vector of <italic>I</italic><sub>id</sub>, and <italic>E</italic><sub>attr</sub> extracts the attribute vector of <italic>I</italic><sub>attr</sub>. As shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>, inspired by pSp (Richardson et al., <xref ref-type="bibr" rid="B25">2020</xref>), <italic>E</italic><sub>attr</sub> consists of a pyramid-shaped three-layer feature map extraction structure and a set of convolutional mappers (CM). In the second stage, we input the concatenation of <italic>E</italic><sub>id</sub>(<italic>I</italic><sub>id</sub>) and <italic>E</italic><sub>attr</sub>(<italic>I</italic><sub>attr</sub>) into the multilayer perceptron (MLP) of each layer and map the vectors containing identity and attribute information directly to the <inline-formula><mml:math id="M9"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> potential vector space. In summary, the whole image conversion process can be represented as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">out</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo><mml:mo stretchy="true">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>G</italic>(&#x000B7;) is the pretrained StyleGAN model, <italic>MLP</italic>(&#x000B7;) is the multilayer perceptron, and [&#x000B7;, &#x000B7;] is the concatenation of two vectors.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The overall structure and data flow of the proposed model. <bold>(A)</bold> is the flow of our method. <bold>(B)</bold> is the structure of <italic>E</italic><sub>attr</sub>. <bold>(C)</bold> is the structure of Multilayer Perceptron (MLP).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-785808-g0002.tif"/>
</fig>
<sec>
<title>3.1. Network Architecture</title>
<p><italic>E</italic><sub>id</sub> is pretrained ArcFace (Deng et al., <xref ref-type="bibr" rid="B5">2019</xref>) model. We use ResNet-IR (Deng et al., <xref ref-type="bibr" rid="B5">2019</xref>) for Feature Extractor (<italic>FE</italic>), in which the feature output layers are 27, 30, and 44. The CM is a fully convolutional network that compresses the tensor of 8 &#x000D7; 8 &#x000D7; 512 dimensions into 1 &#x000D7; 1 &#x000D7; 512 dimensions through three convolution operations with a step size of two. As shown in <xref ref-type="fig" rid="F2">Figure 2C</xref>, <italic>MLP</italic> is a five-layer fully connected network. The StyleGAN generator is a pretrained model trained on FlickrFaces-HQ (FFHQ) (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>).</p>
<p>We mainly use convolution to reduce the dimensions of image encoding and use deconvolution to decode <inline-formula><mml:math id="M11"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> vectors. <italic>E</italic><sub>attr</sub> and <italic>E</italic><sub>id</sub> achieve the data-dimension reduction from image to vector through convolution and other network operations. The identity vector and attribute vector dimensions are both 1 &#x000D7; 512. The splicing of identity and attribute vectors is then input into a set of MLP to convert the face style and map the low-dimensional information to <inline-formula><mml:math id="M12"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space. The deconvolution process is mainly reflected in StyleGAN, which changes from vectors in <inline-formula><mml:math id="M13"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space to images. Note that we do not change any structure of StyleGAN but hope to use its powerful image-generation capabilities to make our face-changing images more stable and clear.</p>
</sec>
<sec>
<title>3.2. Training and Loss Functions</title>
<p>The advanced face-recognition model accurately identifies the face, so we believe that it can extract face-feature information and take the feature vector extracted by the pretrained ArcFace (Deng et al., <xref ref-type="bibr" rid="B5">2019</xref>) as the identity information. To ensure that the identity of <italic>I</italic><sub>out</sub> is consistent with <italic>I</italic><sub>id</sub>, we introduce the identity loss</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x02225;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">out</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>E</italic><sub>id</sub>(&#x000B7;) is the pretrained ArcFace model.</p>
<p>Similarly, we adopt certain restrictions to ensure that the attribute information of <italic>I</italic><sub>out</sub> is consistent with that of <italic>I</italic><sub>attr</sub>. Given that the three-layer feature map extraction structure should gradually have the ability to extract attribute information with the training process, we define the attribute loss function as</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M28"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x02225;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">out</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>P</italic>(&#x000B7;) is the extraction structure.</p>
<p>Note that the attribute information of <italic>I</italic><sub>attr</sub> and the identity information of <italic>I</italic><sub>id</sub> should not only exist in <italic>I</italic><sub>out</sub> but should also be well integrated. Based on this idea, we define the reconstruction loss as</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02225;</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">out</mml:mtext></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x02225;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">out</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">if&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">otherwise</mml:mtext><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>F</italic>(&#x000B7;) is the perceptual feature extractor in the loss of learned perceptual image patch similarity (Zhang et al., <xref ref-type="bibr" rid="B35">2018</xref>), which extracts the perceptual information of the image at the high-dimensional level. <inline-formula><mml:math id="M30"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> loss measures the difference between the two images at the pixel level. Note that <inline-formula><mml:math id="M31"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula> has a positive value only when <italic>I</italic><sub>id</sub> and <italic>I</italic><sub>attr</sub> are the same because only in this case should <italic>I</italic><sub>out</sub> and <italic>I</italic><sub>id</sub> (or <italic>I</italic><sub>attr</sub>) be so consistent that they are exactly the same; otherwise, we cannot expect a similar comparison between the two images. Overall, our total training loss is the weighted sum of all the losses mentioned above:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M32"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">total</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Based on the loss functions and model structure proposed above, we train the ShapeEditor encoder according to <xref ref-type="table" rid="T4">Algorithm 1</xref>.</p>
<table-wrap position="float" id="T4">
<label>Algorithm 1</label>
<caption><p>Training ShapeEditor using gradient descent.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td><bold>Input:</bold> <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold><italic>I</italic></bold><sub>attr</sub>: Image containing attribute information <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold><italic>I</italic></bold><sub>id</sub>: Image containing identity information<break/> &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>P</italic>: Identity-attribute image pair space <break/><bold>Functions:</bold> <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>Encoder ShapeEditor</bold>:<italic>P</italic> <inline-formula><mml:math id="M14"><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>Generator G</bold><inline-formula><mml:math id="M15"><mml:mo>:</mml:mo><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>Loss</bold> &#x02190; <inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>:</mml:mo></mml:math></inline-formula> Calculate the identity loss between <italic>I</italic><sub>id</sub> and <italic>I</italic><sub>out</sub>. <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>Loss</bold> &#x02190; <inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>:</mml:mo></mml:math></inline-formula> Calculate the attribute loss between <italic>I</italic><sub>attr</sub> and <italic>I</italic><sub>out</sub>. <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>Loss</bold> &#x02190; <inline-formula><mml:math id="M18"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>:</mml:mo></mml:math></inline-formula> Calculate the reconstruction loss between <italic>I</italic><sub>id</sub>(<italic>I</italic><sub>attr</sub>) and <italic>I</italic><sub>out</sub>. <break/><bold>Output:</bold> <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>I</italic>: Image space <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M19"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mo>:</mml:mo></mml:math></inline-formula> Potential vector space of StyleGAN <break/>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>I</italic><sub>out</sub>: Synthesized face-swapping image</td>
</tr>
<tr><td align="left" valign="top"> 1: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>for</bold> number of training iterations <bold>do:</bold> </td></tr>
<tr><td align="left" valign="top"> 2: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>for</bold> <italic>I</italic><sub>id</sub>, <italic>I</italic><sub>attr</sub> randomly selected in training dataset <bold>do:</bold> </td></tr>
<tr><td align="left" valign="top"> 3: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Generate the <inline-formula><mml:math id="M20"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo mathvariant="bold">&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space vector using [<italic>I</italic><sub>id</sub>, <italic>I</italic><sub>attr</sub>] </td></tr>
<tr><td align="left" valign="top"> 4: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;ShapeEditor:<italic>P</italic> <inline-formula><mml:math id="M21"><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> </td></tr>
<tr><td align="left" valign="top"> 5: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Generate the face-swapping image <italic>I</italic><sub>out</sub> using the <inline-formula><mml:math id="M22"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space vector </td></tr>
<tr><td align="left" valign="top"> 6: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;G<inline-formula><mml:math id="M23"><mml:mo>:</mml:mo><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> </td></tr>
<tr><td align="left" valign="top"> 7: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Calculate the identity loss <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">id</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula>, the attribute loss <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">attr</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula>, and the reconstruction loss <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>L</mml:mi></mml:mstyle></mml:mrow></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">rec</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula> </td></tr>
<tr><td align="left" valign="top"> 8: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Update ShapeEditor with loss </td></tr>
<tr><td align="left" valign="top"> 9: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;end </td></tr>
<tr><td align="left" valign="top"> 10: &#x000A0;end.</td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments</title>
<p><bold>Implementation Details:</bold> We use the FFHQ (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>) dataset as the training set, and the value of loss weights is set to &#x003BB;<sub>id</sub> &#x0003D; 0.5, &#x003BB;<sub>attr</sub> &#x0003D; 0.1, &#x003BB;<sub>rec</sub> &#x0003D; 1. The ratio of the training data with <italic>I</italic><sub>id</sub> &#x0003D; <italic>I</italic><sub>attr</sub> to that with <italic>I</italic><sub>id</sub>&#x02260;<italic>I</italic><sub>attr</sub> is set to 2:1. During the training, the network parameters of <italic>E</italic><sub>id</sub> and the StyleGAN generator remain unchanged, and the weights of the rest are updated with iterations. To compare with other methods, we train the model with images of 256 &#x000D7; 256 resolution in this section. This model was trained on a single NVIDIA TITAN RTX for about 2 days with a Ranger optimizer (Richardson et al., <xref ref-type="bibr" rid="B25">2020</xref>), with a batch size set to eight and a learning rate set to 0.0001.</p>
<sec>
<title>4.1. Qualitative Comparison With Previous Methods</title>
<p>We compare the proposed method with FSGAN (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>), FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>; Nitzan et al., <xref ref-type="bibr" rid="B22">2020</xref>) on the CelebAMask-HQ (Lee et al., <xref ref-type="bibr" rid="B14">2020</xref>) test dataset. <xref ref-type="fig" rid="F3">Figure 3</xref> shows, as expected because the proposed method is based on a pretrained StyleGAN (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>) with high-quality face-generation capabilities, that all the generation results (<xref ref-type="fig" rid="F3">Figure 3</xref>, column 6) are stable and clear enough that there are no errors such as artifacts and abnormal illumination.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Qualitative comparison with FSGAN (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>), FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>; Nitzan et al., <xref ref-type="bibr" rid="B22">2020</xref>) on the CelebAMask-HQ (Lee et al., <xref ref-type="bibr" rid="B14">2020</xref>) test dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-785808-g0003.tif"/>
</fig>
<p>Almost every output image (<xref ref-type="fig" rid="F3">Figure 3</xref>, column 3) of FSGAN (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>) shows unnatural lighting transition and lack of facial details, the abnormal region of the face is caused by directly extracting and filling the internal area of the face (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 3, column 4), which is completely avoided in the proposed method.</p>
<p>Because there is no pretrained model as the backbone, it is difficult for FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>) to avoid facial blur, some results even show facial illumination confusion (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 3, column 4) and eye ghosting (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 7, column 4), showing that its authenticity is significantly inferior to that of the proposed method.</p>
<p>Similar to the proposed method, Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) use StyleGAN (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>) as the backbone. However, it cannot accurately integrate identity and attribute information because of its simple encoder structure and the constraint of <inline-formula><mml:math id="M33"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> potential space. Therefore, although it can generate high-quality images (<xref ref-type="fig" rid="F3">Figure 3</xref>, column 6), it is not as good as the proposed method for fusing semantic information, which is reflected in the attributes of the target image, such as hairstyle and background, that are not contained.</p>
<p>In addition to the excellent performance in terms of authenticity and fidelity, the proposed method also deals with extreme lighting conditions (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 2, column 6) and even keeps the sense of age (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 3, column 6). Thanks to that, we use the facial recognition module to extract the identity vector instead of directly using the pixels in the facial area. We can extract the identity information very well even if the source image has facial occlusion (<xref ref-type="fig" rid="F3">Figure 3</xref>, row 4, column 6). The proposed model understands whether its output should have glasses (<xref ref-type="fig" rid="F3">Figure 3</xref>, column 6, rows 5 and 6), which is embedded in the potential space of the pretrained StyleGAN model (Karras et al., <xref ref-type="bibr" rid="B12">2019</xref>).</p>
</sec>
<sec>
<title>4.2. Quantitative Comparison With Previous Methods</title>
<p>As mentioned in the section 2.3, our method mainly inherits the ideas of latent space manipulation of pretrained models and GAN-based face swapping. To show the advantages, we compare the proposed method with other related. In the field of latent space manipulation, Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) is the most similar to our work, which is about controlling facial attributes with StyleGAN. In the field of GAN-based face swapping, DeepFakes (R&#x000F6;ssler et al., <xref ref-type="bibr" rid="B27">2019</xref>), FSGAN (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>), and FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>) occupy earlier positions and have achieved remarkable face exchange. To show the robustness of our method, we compare the proposed method with them quantitatively.</p>
<sec>
<title>4.2.1. Comparison With Nitzan et al.</title>
<p>Our method and Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) both make use of the image generation ability of pretrained StyleGAN, and make efforts to achieve adequate control of the human face. But we are different in the choice of mapping space and framework design. To show the significance of our improvement in semantic control, we quantitatively compare our method with Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) in terms of identity, pose, expression, and mood consistency on CelebAMask-HQ (Lee et al., <xref ref-type="bibr" rid="B14">2020</xref>) dataset.</p>
<p>The face swapping model not only needs to ensure the image quality but also needs to fuse the identity and attribute information to the greatest extent. We propose four indicators to measure these aspects. To calculate the identity information in the test stage, we use another advanced method called CurricularFace (Huang et al., <xref ref-type="bibr" rid="B10">2020</xref>) as the face-recognition module to extract the identity vectors of source faces and face-swapping results, then use L2 distance to calculate the difference between them to get the identity error. To ensure that the conversion results are consistent with the target image in attribute, we use 3DDFA-V2 (Guo et al., <xref ref-type="bibr" rid="B7">2020</xref>) to estimate the key face points and the head angle. For normalization, we use the two-dimensional (2D) coordinate information instead of 3D coordinate information to reduce the error impact of key-point estimation as much as possible, and calculate the average position of key points in each image, and then obtain the relative position of each point so as to establish a unified expression coordinate system. Based on the above, we take the difference between the target image and the resulting image in angle as pose error, in key face points as expression error. In addition to pose and expression, mood embodies the high-level semantics of face attribute. Inspired by Abirami and Vincent (<xref ref-type="bibr" rid="B2">2021</xref>), we use the emotion recognition model (Zhao et al., <xref ref-type="bibr" rid="B36">2021</xref>) to detect the ability of face-swapping methods to transmit emotional information. Specifically, we recognize the moods of the swapped images and calculate the consistency of the mood recognition results before and after face exchange.</p>
<p>We randomly extract images from the CelebAMask-HQ dataset as source faces and take the remaining images as target faces to form one-to-one corresponding face combinations as the test dataset. As shown in <xref ref-type="table" rid="T1">Table 1</xref>, our method is superior to Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>) in pose error, expression error, and mood consistency, which shows our advantages in attribute information transfer. Our identity error is slightly higher than Nitzan, that is because face swapping brings more changes in head area than expression manipulation. Our advantages in most indicators demonstrate that we have realized better work in latent space manipulation.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Quantitative comparison with Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>). Our method performs better in most indicators.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Identity Error</bold> <bold>&#x02193;</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Pose Error</bold> <bold>&#x02193;</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Expression Error</bold> <bold>&#x02193;</bold></th>
<th valign="top" align="center"><bold>Mood Consistency &#x02191;</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>Avg</bold>.</th>
<th valign="top" align="center"><bold>Std</bold>.</th>
<th valign="top" align="center"><bold>Avg</bold>.</th>
<th valign="top" align="center"><bold>Std</bold>.</th>
<th valign="top" align="center"><bold>Avg</bold>.</th>
<th valign="top" align="center"><bold>Std</bold>.</th>
<th valign="top" align="center"><bold>Acc. (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Nitzan et al. (<xref ref-type="bibr" rid="B22">2020</xref>)</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">5.99</td>
<td valign="top" align="center">7.16</td>
<td valign="top" align="center">10.13</td>
<td valign="top" align="center">5.38</td>
<td valign="top" align="center">65.35</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center">1.30</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center"><bold>3.82</bold></td>
<td valign="top" align="center">6.88</td>
<td valign="top" align="center"><bold>5.93</bold></td>
<td valign="top" align="center">3.63</td>
<td valign="top" align="center"><bold>75.38</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold values represent the best. &#x02191; represents that the larger the value, the better. &#x02193; represents that the smaller the value, the better</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.2.2. Comparison With Face Swapping Methods</title>
<p>To comprehensively show the face-swapping ability of our method, we conduct quantitative comparisons in transformation consistency and image quality with DeepFakes, FSGAN, and FaceShifter. Our work, FSGAN, and FaceShifter rely on a single reference or few references and are many-to-many approaches. At the same time, DeepFakes have to be supported by multi-images or videos to transfer faces in to two specific identities. Therefore, in order to ensure the effectiveness and efficiency of comparison, we extract DeepFakes conversion results from R&#x000F6;ssler et al. (<xref ref-type="bibr" rid="B27">2019</xref>) dataset. The calculations of identity error, pose error, expression error, and mood consistency is the same as in section 4.2.1, which represent transformation consistency evaluation. Following the work of Yao et al. (<xref ref-type="bibr" rid="B34">2020</xref>), we employ peak signal-to-noise ratio (PSNR) (Huynh-Thu and Ghanbari, <xref ref-type="bibr" rid="B11">2008</xref>) and structural similarity index (SSIM) (Wang et al., <xref ref-type="bibr" rid="B33">2004</xref>) to measure the image reconstruction similarity between the target face and swapped face. Last but not least, to evaluate the clarity and authenticity of images, we use Li and Lyu (<xref ref-type="bibr" rid="B16">2018</xref>), which can effectively capture the artifacts in the forged images, to identify fake faces according to the resolution of the generated images. Specifically, we calculate the Forgery Detection Rate (FDR) of the output images. In the analysis of section 4.1, we know that the problems of low-quality images are mainly reflected in insufficient resolution and abnormal artifact areas. Therefore, the method of Li and Lyu (<xref ref-type="bibr" rid="B16">2018</xref>) can evaluate the quality of face images to a certain extent.</p>
<p><xref ref-type="table" rid="T2">Table 2</xref> lists the comparison results of different face-swapping methods. Notably, our method performs best in SSIM, indicating that our method retains the brightness, contrast, and structure of the original images to the greatest extent. Besides, our method outperforms others in PSNR, which demonstrates that our method can better preserve the global similarity than others. Also, our method has the least scores in FDR under different thresholds, which implies that our method can generate images with more sufficient resolution and less abnormal artifact areas. Finally, it is worth noting that our method has the second-best or the same level scores in identity error, pose error, expression error, and mood consistency, indicating that our method is comparable to others in identity and attribute, while being superior to them in terms of image quality and stability.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Quantitative assessment with DeepFakes (R&#x000F6;ssler et al., <xref ref-type="bibr" rid="B27">2019</xref>), FSGAN (Nirkin et al., <xref ref-type="bibr" rid="B20">2019</xref>), and FaceShifter (Li et al., <xref ref-type="bibr" rid="B15">2019</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left" colspan="2"></th>
<th valign="top" align="center"><bold>DeepFakes</bold></th>
<th valign="top" align="center"><bold>FSGAN</bold></th>
<th valign="top" align="center"><bold>FaceShifter</bold></th>
<th valign="top" align="center"><bold>Ours</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Identity Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">1.35</td>
<td valign="top" align="center">1.51</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center">1.30</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.33</td>
</tr>
<tr>
<td valign="top" align="left">Pose Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">3.79</td>
<td valign="top" align="center"><bold>2.81</bold></td>
<td valign="top" align="center">3.04</td>
<td valign="top" align="center">3.82</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">1.99</td>
<td valign="top" align="center">4.41</td>
<td valign="top" align="center">6.70</td>
<td valign="top" align="center">6.88</td>
</tr>
<tr>
<td valign="top" align="left">Expression Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">8.82</td>
<td valign="top" align="center">5.03</td>
<td valign="top" align="center"><bold>4.53</bold></td>
<td valign="top" align="center">5.93</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">3.30</td>
<td valign="top" align="center">2.17</td>
<td valign="top" align="center">2.83</td>
<td valign="top" align="center">3.63</td>
</tr>
<tr>
<td valign="top" align="left">Mood Consistency &#x02191;</td>
<td valign="top" align="center">Acc. (%)</td>
<td valign="top" align="center">39.80</td>
<td valign="top" align="center">72.77</td>
<td valign="top" align="center"><bold>77.94</bold></td>
<td valign="top" align="center">75.38</td>
</tr>
<tr>
<td valign="top" align="left">SSIM &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.75</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.08</td>
</tr>
<tr>
<td valign="top" align="left">PSNR &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">20.54</td>
<td valign="top" align="center">23.76</td>
<td valign="top" align="center">28.17</td>
<td valign="top" align="center"><bold>20.22</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">2.60</td>
<td valign="top" align="center">2.30</td>
<td valign="top" align="center">1.92</td>
<td valign="top" align="center">1.62</td>
</tr>
<tr>
<td valign="top" align="left">FDR &#x02193;</td>
<td valign="top" align="center">Tsd.=0.01</td>
<td valign="top" align="center">91.42</td>
<td valign="top" align="center">76.59</td>
<td valign="top" align="center">37.67</td>
<td valign="top" align="center"><bold>15.18</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Tsd.=0.05</td>
<td valign="top" align="center">83.83</td>
<td valign="top" align="center">48.99</td>
<td valign="top" align="center">11.66</td>
<td valign="top" align="center"><bold>2.67</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Tsd.=0.1</td>
<td valign="top" align="center">77.45</td>
<td valign="top" align="center">35.86</td>
<td valign="top" align="center">6.05</td>
<td valign="top" align="center"><bold>1.09</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Tsd.=0.2</td>
<td valign="top" align="center">70.86</td>
<td valign="top" align="center">24.22</td>
<td valign="top" align="center">2.86</td>
<td valign="top" align="center"><bold>0.32</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN1"><p><italic>Tsd. represents the threshold, which is set to judge whether samples are forged or not. Bold values represent the best. &#x02191; represents that the larger the value, the better. &#x02193; represents that the smaller the value, the better</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec>
<title>4.3. Ablation Study</title>
<p>To verify the effectiveness of each component of the proposed method, we do the ablation study by evaluating the following degenerate models of our method:</p>
<list list-type="bullet">
<list-item><p><italic>Random StyleGAN</italic>. Using randomly initialized StyleGAN instead of pretrained generator.</p></list-item>
<list-item><p><italic>Single attribute vector</italic>. This variant uses a single output layer of Feature Extractor (<italic>FE</italic>), while the original uses multi-layer attribute information.</p></list-item>
<list-item><p><inline-formula><mml:math id="M34"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> <italic>space</italic>. Using <inline-formula><mml:math id="M35"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> potential space instead of <inline-formula><mml:math id="M36"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item><p><italic>Random</italic> <italic>E</italic><sub>id</sub>. Using randomly initialized <italic>E</italic><sub>id</sub> instead of pretrained face recognition model, with weight updating.</p></list-item>
</list>
<p>We report the qualitative results of the variants of our method in <xref ref-type="fig" rid="F4">Figure 4</xref>. We can see that our original model has better face-swapping results. The results of <italic>Random StyleGAN</italic> are too vague to recognize, indicating that the pretrained StyleGAN can help to generate clear and vivid faces. The results of <italic>Single attribute vector</italic> lose details of hair, wrinkles, and beard compared with ours, showing that multi-layer <italic>FE</italic> can deliver more attribute information. The results of <inline-formula><mml:math id="M37"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> <italic>space</italic> leak identity information and add unnecessary details like glasses, showing that <inline-formula><mml:math id="M38"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> potential space can more strictly embed wild faces into StyleGAN semantic space. The results of <italic>Random</italic> <italic>E</italic><sub>id</sub> leak identity information, which implies that using pretrained identity recognition model is of great significance.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Qualitative ablation study on different variants. Our original model performs better than others.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-785808-g0004.tif"/>
</fig>
<p><xref ref-type="table" rid="T3">Table 3</xref> shows the quantitative results of the variants of our method on the randomly selected data from Lee et al. (<xref ref-type="bibr" rid="B14">2020</xref>) dataset. With the help of <inline-formula><mml:math id="M40"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space and pretrained <italic>E</italic><sub>id</sub>, ours and <italic>Single attribute vector</italic> obtain lower identity error. The results of <inline-formula><mml:math id="M41"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> <italic>space</italic> are much inferior compared to ours in pose error and expression error, revealing the importance of the reasonable space choice. Also, we can see that <inline-formula><mml:math id="M42"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> <italic>space</italic> performs best in PSNR and SSIM, that is because face swapping in <inline-formula><mml:math id="M43"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> space tends to map a wild face to a most similar face in the StyleGAN face domain, which is a more natural result with better image quality. Thanks to the help of StyleGAN, every model in <xref ref-type="table" rid="T3">Table 3</xref> surpasses the existing face-swapping methods in PSNR and SSIM.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Quantitative ablation study on different variants for face swapping.</p></caption>
<table frame="hsides" rules="groups">
<tbody><tr>
<td/>
<td/>
<td valign="top" align="center"><bold><italic>Single attribute vector</italic></bold></td>
<td valign="top" align="center"><inline-formula><mml:math id="M39"><mml:mstyle mathvariant="bold"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle></mml:mrow></mml:mstyle></mml:math></inline-formula> <bold><italic>space</italic></bold></td>
<td valign="top" align="center"><bold><italic>Random E</italic><sub>id</sub></bold></td>
<td valign="top" align="center"><bold>Ours</bold></td>
</tr>
<tr>
<td valign="top" align="left">Identity Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center"><bold>1.29</bold></td>
<td valign="top" align="center">1.33</td>
<td valign="top" align="center">1.37</td>
<td valign="top" align="center"><bold>1.29</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.33</td>
</tr>
<tr>
<td valign="top" align="left">Pose Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">3.94</td>
<td valign="top" align="center">4.53</td>
<td valign="top" align="center">4.05</td>
<td valign="top" align="center"><bold>3.64</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">5.59</td>
<td valign="top" align="center">5.74</td>
<td valign="top" align="center">5.75</td>
<td valign="top" align="center">5.55</td>
</tr>
<tr>
<td valign="top" align="left">Expression Error &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">6.63</td>
<td valign="top" align="center">7.43</td>
<td valign="top" align="center">6.48</td>
<td valign="top" align="center"><bold>5.96</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">3.45</td>
<td valign="top" align="center">4.20</td>
<td valign="top" align="center">3.87</td>
<td valign="top" align="center">3.22</td>
</tr>
<tr>
<td valign="top" align="left">PSNR &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">19.38</td>
<td valign="top" align="center"><bold>18.42</bold></td>
<td valign="top" align="center">19.94</td>
<td valign="top" align="center">20.22</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">1.51</td>
<td valign="top" align="center">1.58</td>
<td valign="top" align="center">1.59</td>
<td valign="top" align="center">1.62</td>
</tr>
<tr>
<td valign="top" align="left">SSIM &#x02193;</td>
<td valign="top" align="center">Avg.</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center"><bold>0.70</bold></td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.75</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.07</td>
<td valign="top" align="center">0.08</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold values represent the best. &#x02193; represents that the smaller the value, the better</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.4. Discussion</title>
<p>The core of the proposed model is to use StyleGAN as the face decoder, which reduces the burden of face spatial feature learning and dramatically reduces the possibility of artifacts in the conversion results. However, the proposed method also has some defects. As shown in <xref ref-type="fig" rid="F5">Figure 5A</xref>, the letters in the background of the target image become blurred in the resulting image, which shows that the proposed model is not good at restoring the background. Although the pretrained model we use learns the potential features of face space, it does not learn well how to separate the head from the background. To deal with this problem, we will separate the head and background in the next step through image segmentation and then combine the background of the target image with the head of the resulting image. At the same time, <xref ref-type="fig" rid="F5">Figure 5B</xref> shows that the resulting image lacks Asian characteristics similar to those in the source image, which reflects the problem of insufficient potential vectors in the StyleGAN face space and is caused by the relative lack of Asian faces in the training dataset. Therefore, adding more types of faces to the pretrained model and selecting a better-pretrained model should also be a focus in future work.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Shortcomings of the proposed model. The problem in panel <bold>(A)</bold> is that the background of the conversion result is blurred. The problem in panel <bold>(B)</bold> is that the swapped face lacks Asian characteristics.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-785808-g0005.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>This article proposes a new face-swapping framework that includes ShapeEditor and a pretrained StyleGAN model. The pretrained model gives the proposed framework the potential to generate clear and realistic faces. The ShapeEditor encoder effectively extracts and integrates the attribute and identity information of the input images, then accurately maps them onto the <inline-formula><mml:math id="M44"><mml:mrow><mml:mstyle mathvariant="-tex-caligraphic"><mml:mi>W</mml:mi></mml:mstyle><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:math></inline-formula> space, thus controlling StyleGAN to output the appropriate results. Extensive experiments show that the proposed method performs better than existing frameworks in terms of clarity and authenticity, with sufficiently integrating identity and attribute.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>SY is responsible for code writing and thesis writing. KQ, RQ, PX, and SS are responsible for the inspiration of ideas. NL, LW, JC, GH, and BY put forward their opinions on the revision of the paper. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abdal</surname> <given-names>R.</given-names></name> <name><surname>Qin</surname> <given-names>Y.</given-names></name> <name><surname>Wonka</surname> <given-names>P.</given-names></name></person-group> (<year>2019</year>). <article-title>Image2stylegan: how to embed images into the stylegan latent space?,</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>4432</fpage>&#x02013;<lpage>4441</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abirami</surname> <given-names>R. N.</given-names></name> <name><surname>Vincent</surname> <given-names>P. D. R.</given-names></name></person-group> (<year>2021</year>). <article-title>Identity preserving multi-pose facial expression recognition using fine tuned vgg on the latent space vector of generative adversarial network</article-title>. <source>Math. Biosci. Eng.</source> <volume>18</volume>, <fpage>3699</fpage>&#x02013;<lpage>3717</lpage>. <pub-id pub-id-type="doi">10.3934/mbe.2021186</pub-id><pub-id pub-id-type="pmid">34198408</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Wen</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Hua</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Towards open-set identity preserving face synthesis,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>) <fpage>6713</fpage>&#x02013;<lpage>6722</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bitouk</surname> <given-names>D.</given-names></name> <name><surname>Kumar</surname> <given-names>N.</given-names></name> <name><surname>Dhillon</surname> <given-names>S.</given-names></name> <name><surname>Belhumeur</surname> <given-names>P.</given-names></name> <name><surname>Nayar</surname> <given-names>S. K.</given-names></name></person-group> (<year>2008</year>). <article-title>Face swapping: automatically replacing faces in photographs</article-title>. <source>ACM Trans. Graph.</source> <volume>27</volume>, <fpage>39</fpage>. <pub-id pub-id-type="doi">10.1145/1360612.1360638</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Xue</surname> <given-names>N.</given-names></name> <name><surname>Zafeiriou</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Arcface: additive angular margin loss for deep face recognition,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>4690</fpage>&#x02013;<lpage>4699</lpage>.<pub-id pub-id-type="pmid">34106845</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Image processing using multi-code gan prior,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>3012</fpage>&#x02013;<lpage>3021</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>F.</given-names></name> <name><surname>Lei</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>S. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Towards fast, accurate and stable 3d dense face aliganment,</article-title> in <source>Computer Vision&#x02013;ECCV 2020: 16th European Conference, Glasgow, UK, August 23&#x02013;28, 2020, Proceedings, Part XIX 16</source> (<publisher-loc>Glasgow</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>152</fpage>&#x02013;<lpage>168</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>H&#x000E4;rk&#x000F6;nen</surname> <given-names>E.</given-names></name> <name><surname>Hertzmann</surname> <given-names>A.</given-names></name> <name><surname>Lehtinen</surname> <given-names>J.</given-names></name> <name><surname>Paris</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Ganspace: aiscovering interpretable gan controls</article-title>. <source>arXiv preprint</source> arXiv:2004.02546.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Arbitrary style transfer in real-time with adaptive instance normalization,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>1501</fpage>&#x02013;<lpage>1510</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Tai</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Shen</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). Curricularface: adaptive curriculum learning loss for deep face recognition. in <italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</italic> (Seattle, WA), <fpage>5901</fpage>&#x02013;<lpage>5910</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huynh-Thu</surname> <given-names>Q.</given-names></name> <name><surname>Ghanbari</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Scope of validity of psnr in image/video quality assessment</article-title>. <source>Electron. Lett.</source> <volume>44</volume>, <fpage>800</fpage>&#x02013;<lpage>801</lpage>. <pub-id pub-id-type="doi">10.1049/el:20080522</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Karras</surname> <given-names>T.</given-names></name> <name><surname>Laine</surname> <given-names>S.</given-names></name> <name><surname>Aila</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>A style-based generator architecture for generative adversarial networks,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>4401</fpage>&#x02013;<lpage>4410</lpage>.<pub-id pub-id-type="pmid">32012000</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Korshunova</surname> <given-names>I.</given-names></name> <name><surname>Shi</surname> <given-names>W.</given-names></name> <name><surname>Dambre</surname> <given-names>J.</given-names></name> <name><surname>Theis</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Fast face-swap using convolutional neural networks,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>3677</fpage>&#x02013;<lpage>3685</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>C.-H.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>L.</given-names></name> <name><surname>Luo</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Maskgan: Towards diverse and interactive facial image manipulation,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>5549</fpage>&#x02013;<lpage>5558</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Bao</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Wen</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>Faceshifter: Towards high fidelity and occlusion aware face swapping</article-title>. <source>arXiv preprint</source> arXiv:1912.13457.</citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Lyu</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Exposing deepfake videos by detecting face warping artifacts</article-title>. <source>arXiv preprint</source> arXiv:1811.00656.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nandhini Abirami</surname> <given-names>R.</given-names></name> <name><surname>Durai Raj Vincent</surname> <given-names>P.</given-names></name> <name><surname>Srinivasan</surname> <given-names>K.</given-names></name> <name><surname>Tariq</surname> <given-names>U.</given-names></name> <name><surname>Chang</surname> <given-names>C.-Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep cnn and deep gan in computational visual perception-driven image analysis</article-title>. <source>Complexity</source> <volume>2021</volume>:<fpage>5541134</fpage>. <pub-id pub-id-type="doi">10.1155/2021/5541134</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Natsume</surname> <given-names>R.</given-names></name> <name><surname>Yatagawa</surname> <given-names>T.</given-names></name> <name><surname>Morishima</surname> <given-names>S.</given-names></name></person-group> (<year>2018a</year>). <article-title>Fsnet: an identity-aware generative model for image-based face swapping,</article-title> in <source>Asian Conference on Computer Vision</source> (<publisher-loc>Perth</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>117</fpage>&#x02013;<lpage>132</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Natsume</surname> <given-names>R.</given-names></name> <name><surname>Yatagawa</surname> <given-names>T.</given-names></name> <name><surname>Morishima</surname> <given-names>S.</given-names></name></person-group> (<year>2018b</year>). <article-title>Rsgan: face swapping and editing using face and hair representation in latent spaces</article-title>. <source>arXiv preprint</source> arXiv:1804.03447.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nirkin</surname> <given-names>Y.</given-names></name> <name><surname>Keller</surname> <given-names>Y.</given-names></name> <name><surname>Hassner</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Fsgan: subject agnostic face swapping and reenactment,</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>7184</fpage>&#x02013;<lpage>7193</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nirkin</surname> <given-names>Y.</given-names></name> <name><surname>Masi</surname> <given-names>I.</given-names></name> <name><surname>Tuan</surname> <given-names>A. T.</given-names></name> <name><surname>Hassner</surname> <given-names>T.</given-names></name> <name><surname>Medioni</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>On face segmentation, face swapping, and face perception,</article-title> in <source>2018 13th IEEE International Conference on Automatic Face &#x00026; Gesture Recognition (FG 2018)</source> (<publisher-loc>Xi&#x00027;an</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>98</fpage>&#x02013;<lpage>105</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nitzan</surname> <given-names>Y.</given-names></name> <name><surname>Bermano</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Cohen-Or</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Face identity disentanglement via latent space mapping</article-title>. <source>ACM Trans. Graph.</source> <volume>39</volume>, <fpage>1</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/3414685.3417826</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Olszewski</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>R.</given-names></name> <name><surname>Huang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Realistic dynamic facial textures from a single image using gans,</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>5429</fpage>&#x02013;<lpage>5438</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>M.-Y.</given-names></name> <name><surname>Wang</surname> <given-names>T.-C.</given-names></name> <name><surname>Zhu</surname> <given-names>J.-Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Semantic image synthesis with spatially-adaptive normalization,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>2337</fpage>&#x02013;<lpage>2346</lpage>.<pub-id pub-id-type="pmid">33914680</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Richardson</surname> <given-names>E.</given-names></name> <name><surname>Alaluf</surname> <given-names>Y.</given-names></name> <name><surname>Patashnik</surname> <given-names>O.</given-names></name> <name><surname>Nitzan</surname> <given-names>Y.</given-names></name> <name><surname>Azar</surname> <given-names>Y.</given-names></name> <name><surname>Shapiro</surname> <given-names>S.</given-names></name> <name><surname>Cohen-Or</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Encoding in style: a stylegan encoder for image-to-image translation</article-title>. <source>arXiv preprint</source> arXiv:2008.00951.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ross</surname> <given-names>A.</given-names></name> <name><surname>Othman</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Visual cryptography for biometric privacy</article-title>. <source>IEEE Trans. Inform. Forensics Secur.</source> <volume>6</volume>, <fpage>70</fpage>&#x02013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2010.2097252</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>R&#x000F6;ssler</surname> <given-names>A.</given-names></name> <name><surname>Cozzolino</surname> <given-names>D.</given-names></name> <name><surname>Verdoliva</surname> <given-names>L.</given-names></name> <name><surname>Riess</surname> <given-names>C.</given-names></name> <name><surname>Thies</surname> <given-names>J.</given-names></name> <name><surname>Nie&#x000DF;ner</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>FaceForensics&#x0002B;&#x0002B;: learning to detect manipulated facial images,</article-title> in <source>International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>).</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Interfacegan: interpreting the disentangled face representation learned by gans</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <pub-id pub-id-type="doi">10.1109/TPAMI.2020.3034267</pub-id><pub-id pub-id-type="pmid">33108282</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Closed-form factorization of latent semantics in gans</article-title>. <source>arXiv preprint</source> arXiv:2007.06600.</citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Q.</given-names></name> <name><surname>Tewari</surname> <given-names>A.</given-names></name> <name><surname>Xu</surname> <given-names>W.</given-names></name> <name><surname>Fritz</surname> <given-names>M.</given-names></name> <name><surname>Theobalt</surname> <given-names>C.</given-names></name> <name><surname>Schiele</surname> <given-names>B.</given-names></name></person-group> (<year>2018</year>). <article-title>A hybrid model for identity obfuscation by face replacement,</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>553</fpage>&#x02013;<lpage>569</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tewari</surname> <given-names>A.</given-names></name> <name><surname>Elgharib</surname> <given-names>M.</given-names></name> <name><surname>Bharaj</surname> <given-names>G.</given-names></name> <name><surname>Bernard</surname> <given-names>F.</given-names></name> <name><surname>Seidel</surname> <given-names>H.-P.</given-names></name> <name><surname>P&#x000E9;rez</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Stylerig: Rigging stylegan for 3d control over portrait images,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>6142</fpage>&#x02013;<lpage>6151</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Shan</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Towards real-world blind face restoration with generative facial prior</article-title>. <source>arXiv preprint</source> arXiv:2101.04061.</citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name> <name><surname>Sheikh</surname> <given-names>H. R.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E. P.</given-names></name></person-group> (<year>2004</year>). <article-title>Image quality assessment: from error visibility to structural similarity</article-title>. <source>IEEE Trans. Image Process.</source> <volume>13</volume>, <fpage>600</fpage>&#x02013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2003.819861</pub-id><pub-id pub-id-type="pmid">15376593</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>G.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Shao</surname> <given-names>T.</given-names></name> <name><surname>Zhou</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Mesh guided one-shot face reenactment using graph convolutional networks,</article-title> in <source>Proceedings of the 28th ACM International Conference on Multimedia</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>1773</fpage>&#x02013;<lpage>1781</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>R.</given-names></name> <name><surname>Isola</surname> <given-names>P.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name> <name><surname>Shechtman</surname> <given-names>E.</given-names></name> <name><surname>Wang</surname> <given-names>O.</given-names></name></person-group> (<year>2018</year>). <article-title>The unreasonable effectiveness of deep features as a perceptual metric,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>586</fpage>&#x02013;<lpage>595</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Zhou</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Robust lightweight facial expression recognition network with label distribution training,</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, Vol. <volume>35</volume>, <fpage>3510</fpage>&#x02013;<lpage>3519</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>D.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>In-domain gan inversion for real image editing,</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Glasgow</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>592</fpage>&#x02013;<lpage>608</lpage>.</citation>
</ref>
</ref-list> 
</back>
</article> 