<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Virtual Real.</journal-id>
<journal-title>Frontiers in Virtual Reality</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Virtual Real.</abbrev-journal-title>
<issn pub-type="epub">2673-4192</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">739010</article-id>
<article-id pub-id-type="doi">10.3389/frvir.2021.739010</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Virtual Reality</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep4D: A Compact Generative Representation for Volumetric Video</article-title>
<alt-title alt-title-type="left-running-head">Regateiro et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Deep4D</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Regateiro</surname>
<given-names>Jo&#xe3;o</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1366406/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Volino</surname>
<given-names>Marco</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1373883/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Hilton</surname>
<given-names>Adrian</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Centre for Vision Speech and Signal Processing, University of Surrey, <addr-line>Guildford</addr-line>, <country>United&#x20;Kingdom</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP (Institute of Engineering Univ. Grenoble Alpes), LJK, <addr-line>Grenoble</addr-line>, <country>France</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/490345/overview">Fabien Danieau</ext-link>, InterDigital, France</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1412299/overview">Anna Hilsmann</ext-link>, Heinrich Hertz Institute (FHG), Germany</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/851728/overview">Weiya Chen</ext-link>, Huazhong University of Science and Technology, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jo&#xe3;o Regateiro, <email>j.regateiro@inria.fr</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Technologies for VR, a section of the journal Frontiers in Virtual Reality</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>2</volume>
<elocation-id>739010</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Regateiro, Volino and Hilton.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Regateiro, Volino and Hilton</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>This paper introduces Deep4D a compact generative representation of shape and appearance from captured 4D volumetric video sequences of people. 4D volumetric video achieves highly realistic reproduction, replay and free-viewpoint rendering of actor performance from multiple view video acquisition systems. A deep generative network is trained on 4D video sequences of an actor performing multiple motions to learn a generative model of the dynamic shape and appearance. We demonstrate the proposed generative model can provide a compact encoded representation capable of high-quality synthesis of 4D volumetric video with two orders of magnitude compression. A variational encoder-decoder network is employed to learn an encoded latent space that maps from 3D skeletal pose to 4D shape and appearance. This enables high-quality 4D volumetric video synthesis to be driven by skeletal motion, including skeletal motion capture data. This encoded latent space supports the representation of multiple sequences with dynamic interpolation to transition between motions. Therefore we introduce Deep4D motion graphs, a direct application of the proposed generative representation. Deep4D motion graphs allow real-tiome interactive character animation whilst preserving the plausible realism of movement and appearance from the captured volumetric video. Deep4D motion graphs implicitly combine multiple captured motions from a unified representation for character animation from volumetric video, allowing novel character movements to be generated with dynamic shape and appearance detail.</p>
</abstract>
<kwd-group>
<kwd>volumetric video</kwd>
<kwd>generative networks</kwd>
<kwd>motion graphs</kwd>
<kwd>animation</kwd>
<kwd>performance capture</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Volumetric video is an emerging media that allows free-viewpoint rendering and replay of dynamic scenes with the visual quality approaching that of the of captured video. This has the potential to allow highly-realistic content production for immersive virtual and augmented reality experiences. Volumetric video is produced from multiple camera performance capture studios that generally consist of synchronised cameras that simultaneously record a performance (<xref ref-type="bibr" rid="B16">Collet et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B46">Starck and Hilton, 2007</xref>; <xref ref-type="bibr" rid="B17">de Aguiar et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B10">Carranza et&#x20;al., 2003</xref>). The generated content usually consists of 4D dynamic mesh and texture sequences that represent the visual features of the scene, for example, shape, motion and appearance. This allows replay of the performance from any viewpoint and moment in time, although it requires a huge computational effort to process and store. Volumetric video capture is currently limited to replay of the captured performance and does not support animation to modify, combine or generate novel movement sequences. Previous work has introduced methods for animation from volumetric video based on re-sampling and concatenation of volumetric sequences (<xref ref-type="bibr" rid="B26">Huang et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B40">Prada et&#x20;al., 2016</xref>).</p>
<p>Rendering realistic human appearance is a particularly challenging problem. Humans are social animals that have evolved to read emotions through body language and facial expressions (<xref ref-type="bibr" rid="B19">Ekman, 1980</xref>). As a result, humans are extremely sensitive to movement and rendering artefacts, which gives rise to the well-known uncanny valley in photo-realistic rendering of human appearance. Recently there has been significant progress using deep generative models to synthesise highly realistic images (<xref ref-type="bibr" rid="B21">Goodfellow et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B31">Kingma and Welling, 2013</xref>; <xref ref-type="bibr" rid="B58">Zhu et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B27">Isola et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B52">Ulyanov et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B36">Ma et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B45">Siarohin et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B39">Paier et&#x20;al., 2020</xref>) and videos (<xref ref-type="bibr" rid="B54">Vondrick et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B51">Tulyakov et&#x20;al., 2017</xref>) of scenes, which is important for applications such as image manipulation, video animation and rendering of virtual environments. Human avatars are typically rendered using detailed, explicit 3D models, which consist of meshes and textures, and animated using tailored motion models to simulate human behaviour and activity.</p>
<p>Recent work <xref ref-type="bibr" rid="B23">Holden et&#x20;al. (2017)</xref> has shown that it is possible to learn and animate natural human behaviour (e.g. walking, jumping, etc.) from human skeletal motion capture data (MoCap) of actor performance. On the other hand, designing a realistic 3D model of a person is still a laborious process. Given the tremendous success of deep generative models (<xref ref-type="bibr" rid="B21">Goodfellow et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B31">Kingma and Welling, 2013</xref>; <xref ref-type="bibr" rid="B57">Zhu et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B29">Karras et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B27">Isola et&#x20;al., 2016</xref>), the question arises, why not also learn to generate realistic rendering of a person? By conditioning the image generation process of a generative model on additional input data, mappings between different data domains are learned (<xref ref-type="bibr" rid="B58">Zhu et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B27">Isola et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B28">Johnson et&#x20;al., 2016</xref>), which, for instance, allows for controlling and manipulating object shape, turning sketches into images and images into paintings. Generative methods have improved recently on the resolution and quality of images produced (<xref ref-type="bibr" rid="B29">Karras et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B37">Miyato et&#x20;al., 2018</xref> <xref ref-type="bibr" rid="B6">Brock et&#x20;al., 2018</xref>). Yet generators continue to operate as black boxes, and despite recent efforts, the understanding of various aspects of the image synthesis process is unknown. The properties of the latent space are also poorly understood, and the commonly demonstrated latent space interpolation (<xref ref-type="bibr" rid="B18">Dosovitskiy et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B43">Sainburg et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B34">Laine, 2018</xref>) provide no quantitative way to compare different generators against each other. Motivated by recent advances in generative networks (<xref ref-type="bibr" rid="B30">Karras et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B29">Karras et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B21">Goodfellow et&#x20;al., 2014</xref>) we propose an architecture for learning to generate dynamic 4D shape and high resolution appearance that exposes ways to control image synthesis. Our appearance generator starts from a learned motion space and adjusts the resolution of the image at each convolution layer based on the latent motion code, therefore directly controlling the strength of image features at different scales.</p>
<p>This work proposes Deep4D, a deep generative representation of dynamic shape and appearance from 4D volumetric video of a human character. The proposed approach learns an efficient compressed latent space representation and generative model from 4D volumetric video sequences of a person performing multiple motions. Compact latent space representation is achieved using a variational encoder-decoder to learn the mapping from 3D skeletal motion to the corresponding full 4D volumetric shape, motion and appearance. The encoded latent space supports interpolation of dynamic shape and appearance to seamlessly transition between captured 4D volumetric video sequences. This work presents Deep4D motion graphs, which exploit generative representation of multiple 4D volumetric video sequences in the learnt latent space to enable interactive animation with optimal transition between motions. The primary novel contributions of this paper are:<list list-type="simple">
<list-item>
<p>&#x2022; Deep4D, a generative shape and appearance representation for 4D volumetric video that enables compact storage and real-time interactive animation.</p>
</list-item>
<list-item>
<p>&#x2022; Mapping of skeletal motion to 4D volumetric video to synthesise dynamic shape and appearance.</p>
</list-item>
<list-item>
<p>&#x2022; Deep4D motion graphs, an animation framework built on top of the Deep4D representation that allows high-level of 4D characters enabling synthesis of novel motions and real-time user interaction.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Related Work</title>
<p>
<bold>4D Volumetric Video:</bold> has been an active area of research (<xref ref-type="bibr" rid="B46">Starck and Hilton, 2007</xref>; <xref ref-type="bibr" rid="B16">Collet et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B10">Carranza et&#x20;al., 2003</xref>; <xref ref-type="bibr" rid="B17">de Aguiar et&#x20;al., 2008</xref>), that has emerged to address the increasing demand for realistic content of human performance. Recently, <xref ref-type="bibr" rid="B16">Collet et&#x20;al. (2015)</xref> presented a full pipeline to capture, reconstruct and replay high-quality volumetric video. The system uses approximately 100 synchronised cameras that simultaneously capture the volume from multiple viewpoints. Volumetric video captures the dynamic surface geometry and photo-realistic appearance of a subject. This unlocks enormous creative potential for highly realistic animated content production based on the captured performance. Recent research provides frameworks to ease the manipulation of this content (<xref ref-type="bibr" rid="B26">Huang et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B40">Prada et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B50">Tejera and Hilton, 2013</xref>; <xref ref-type="bibr" rid="B7">Budd et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B8">Cagniart et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B53">Vlasic et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B14">Casas et&#x20;al., 2014</xref>), allowing an artist to perform manual adjustments on 4D dynamic geometry and combine multiple sequences in a motion graph. However, use of 4D volumetric video in content production remains limited due to the challenge of manipulation, animation and rendering of shape sequences whilst maintaining the realism of appearance and clothing dynamics.</p>
<p>
<bold>Learnt Mesh Sequence Representations:</bold> <xref ref-type="bibr" rid="B50">Tejera and Hilton (2013)</xref> proposed a part-based spatio-temporal mesh sequence editing technique that learns surface deformation models in Laplacian coordinates. This approach constrains the mesh deformation to plausible surface shapes learnt from a set of examples. Part-based learning of surface deformation allows local manipulation of the mesh and achieves greater animation flexibility, allowing the generation of novel posed meshes. <xref ref-type="bibr" rid="B48">Tan et&#x20;al. (2018)</xref> use a variational autoencoder (VAE) to learn a representation of parameterised dynamic shapes. Their network trains on a pre-processed feature space of the training data, demonstrating very low reconstruction error for the ground truth shapes. <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al. (2018)</xref> proposed a learnt model of shape and appearance conditioned on viewpoint allowing recovery of view-dependent texture detail. This network demonstrates the ability to learn 3D dynamic shapes from vertices, avoiding the need to pre-process information. This demonstrates the real-time capabilities of VAEs, being able to decode shape and appearance in less than 5 milliseconds. Recently, <xref ref-type="bibr" rid="B41">Regateiro et&#x20;al. (2019)</xref> demonstrated the capabilities of learning 3D dynamic shapes to produce realistic animation using a VAE to learn the geometric space of a human character and re-use the decoder in real-time to synthesise 3D geometry.</p>
<p>
<bold>Learnt Representation of Appearance:</bold> Recently, <xref ref-type="bibr" rid="B20">Esser et&#x20;al. (2019)</xref> presented an approach towards a holistic learning framework for rendering human behaviour trained from skeletal motion capture data for realistic control and rendering. They learn a mapping from an abstract pose representation to target images conditioned on a latent representation of a VAE for appearance. <xref ref-type="bibr" rid="B29">Karras et&#x20;al. (2017)</xref> propose a novel training methodology for generative networks. that progressively grows both the generator and discriminator, starting from a low image resolution and ending at the original image resolution. They demonstrate that the model increasingly learns fine details as the training progresses, hence improving training speed and stability, and producing high-quality images. Although photorealism is a hard problem to solve, this approach is a step towards recreating high quality images that are indistinguishable from real images. More recently, <xref ref-type="bibr" rid="B30">Karras et&#x20;al. (2018)</xref> redefine the architecture of generative networks for style-based transfer. Using a similar approach to <xref ref-type="bibr" rid="B29">Karras et&#x20;al., 2017</xref>, they have demonstrated high quality images results, for example, the ability to learn the exact placement of hair, stubble, freckles, or skin pores. This demonstrates the potential to synthesise high resolution images of humans, whilst preserving natural details that are essential for perception of realism.</p>
<p>
<bold>4D Volumetric Video Animation:</bold> Motion graphs for character animation from skeletal motion capture sequences (<xref ref-type="bibr" rid="B1">Arikan et&#x20;al., 2003</xref>; <xref ref-type="bibr" rid="B33">Kovar et&#x20;al., 2002</xref>; <xref ref-type="bibr" rid="B49">Tanco and Hilton, 2000</xref>) use a structured graph representation to enable interactive control. The skeletal motion graphs are constructed using a frame-to-frame similarity metric which identifies similar poses and motion. The concept of motion graphs has been applied to volumetric video using both unstructured meshes (<xref ref-type="bibr" rid="B47">Starck et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B25">Huang et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B26">Hunag et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B40">Prada et&#x20;al., 2016</xref>) and temporally consistent structured meshes (<xref ref-type="bibr" rid="B14">Casas et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B3">Boukhayma and Boyer, 2017</xref>; <xref ref-type="bibr" rid="B22">Hilsmann et&#x20;al., 2020</xref>). Initial approaches (<xref ref-type="bibr" rid="B47">Starck et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B25">Huang et&#x20;al., 2009</xref>) concatenate unstructured dynamic mesh sequences without temporal consistency of the mesh connectivity based on shape and motion similarity. <xref ref-type="bibr" rid="B40">Prada et&#x20;al. (2016)</xref> instead performs mesh and texture alignment at defined transitions points to ensure smooth blending. This overcomes the challenging problem of global mesh alignment and only considers alignment of geometry and texture where necessary. In contrast, <xref ref-type="bibr" rid="B3">Boukhayma and Boyer (2017)</xref> and <xref ref-type="bibr" rid="B14">Casas et&#x20;al. (2014)</xref> leverage global alignment of the mesh sequence to obtain temporally consistent mesh connectivity from the volumetric video. This allows 4D motion graphs with mesh blending for high-level parametric control of the motion and smooth transitions between motions.</p>
<p>In this paper we introduce Deep4D, a learnt generative representation of volumetric video sequences, presented in <xref ref-type="sec" rid="s3">Section 3</xref>. Deep4D provides compact representation, which overcomes the memory and computation requirement of previous approaches to explicitly represent all captured sequences at run-time through the learnt parameters of the network. In <xref ref-type="sec" rid="s4">Section 4</xref> we present Deep4D motion graphs, a direct application of the proposed generative network to produce seamless animations of both dynamic shape and appearance between learnt captured motion sequences. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> presents a quantitative and qualitative evaluation of the proposed method.</p>
</sec>
<sec id="s3">
<title>3 Deep4D Representation</title>
<p>The work presents a step forward to allow control and synthesis of 4D volumetric video, while preserving the realism of dynamic shape and appearance. This section introduces the use of a generative network to represent 4D volumetric video content from performance capture data efficiently. Pre-processing of the captured volumetric video into a form suitable for neural networks is first presented. The generative network for the learning of 4D shape from captured volumetric sequences is described, together with the use of a variational encoder-decoder to ensure a compact latent space representation mapping from 3D skeletal pose to corresponding 4D dynamic shape. Finally, we present a generative network for 4D video appearance that learns to synthesise high-resolution dynamic texture appearance from the compact latent space representation, <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>. Enforcing a compact latent space representation enables interpolation between skeletal poses to generate plausible intermediate mesh shape and appearance. These sections individually describe the contribution of the generative network, illustrated in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>. Deep4D generative representation enables the generation of realistic renderings of human characters, with the ability to re-target new skeletal motion information.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The generative network is driven from 3D skeletal motion to synthesise 4D volumetric video interactively.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g001.tif"/>
</fig>
<sec id="s3-1">
<title>3.1 Volumetric Video Pre-processing</title>
<p>In the context of this work, 4D volumetric video represents 4D mesh sequences <inline-formula id="inf1">
<mml:math id="m1">
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, 2D textures <inline-formula id="inf2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and 3D skeletal motion <inline-formula id="inf3">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> computed from multiple view video capture. A 4D volumetric video dataset consists of <italic>N</italic>
<sub>
<italic>S</italic>
</sub> sequences <italic>s</italic>&#x20;&#x3d; [1 &#x2026; <italic>N</italic>
<sub>
<italic>S</italic>
</sub>] and each sequence consists of <inline-formula id="inf4">
<mml:math id="m4">
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> frames at a time instance <inline-formula id="inf5">
<mml:math id="m5">
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2026;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>State-of-the-art volumetric performance capture of people with loose clothing and hair (<xref ref-type="bibr" rid="B16">Collet et&#x20;al., 2015</xref>) results in high resolution reconstructed shape and texture appearance. Raw volumetric video typically results in an unstructured mesh sequence where both the mesh shape and connectivity changes from frame-to-frame (<xref ref-type="bibr" rid="B40">Prada et&#x20;al., 2016</xref>). Several approaches have been introduced for temporal alignment over short subsequences to compress the storage requirements (<xref ref-type="bibr" rid="B16">Collet et&#x20;al., 2015</xref>) or global alignment across complete sequences (<xref ref-type="bibr" rid="B24">Huang et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B8">Cagniart et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>).</p>
<p>In this work we employ the skeleton-driven volumetric surface alignment framework (<xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>) to pre-process captured 4D volumetric video of people to obtain a temporally coherent mesh structure across multiple sequences. This framework receives as input synchronised multiple view video from calibrated cameras and returns 3D skeletal joints and temporally consistent 3D meshes with the same mesh connectivity at every frame. The texture appearance is retrieved by re-mapping the original multiple view camera images onto the temporally consistent 3D meshes providing a dynamic texture map with consistent coordinates for all captured frames. The input to the deep network presented in the following sections consists of centred 4D temporally consistent mesh sequences with the corresponding 2D texture maps and 3D skeletal joint locations.</p>
</sec>
<sec id="s3-2">
<title>3.2 Deep4D: Pose2Shape Network</title>
<p>Variational networks have become a popular approach to learn a compact latent space representation which can integrate with deep neural networks. In this section, we employ a variational encoder-decoder to learn a compact latent space mapping between 3D skeletal pose and the corresponding 4D shape, illustrated in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. The generative network architecture maximises the probability distribution of the 3D skeletal joint positions <inline-formula id="inf6">
<mml:math id="m6">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, encoded in the latent space <inline-formula id="inf7">
<mml:math id="m7">
<mml:mi>z</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and learns the generative mapping of the decoder to the corresponding 4D mesh <inline-formula id="inf8">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. While we define input <italic>p</italic> as 3D skeletal joint positions, it can be replaced with other pose representations consisting of 3D landmarks, e.g. facial keypoints.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Pose2Shape network overview. The input and the output of the encoder and decoder is 3D skeletal motion and 4D shape respectively.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g002.tif"/>
</fig>
<p>Generative networks learn dependencies from the input data and capture them in a low-dimensional latent vector <inline-formula id="inf9">
<mml:math id="m9">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, creating compact representations <inline-formula id="inf10">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, where <italic>d</italic> is the latent space dimension (128 dimensions throughout this work). The probability density function <italic>P</italic>(<italic>p</italic>) for the skeletal pose is given by:<disp-formula id="e1">
<mml:math id="m11">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x222b;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.28em"/>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>z</mml:mi>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>The distribution <italic>P</italic>(<italic>p</italic>&#x7c;<italic>z</italic>) denotes the maximum likelihood estimation of dependencies of <italic>p</italic> over the latent vector <italic>z</italic>, and <italic>P</italic> (<italic>z</italic>) is the prior probability distribution of a latent vector <italic>z</italic>. To ensure a compact representation <italic>P</italic>(<italic>p</italic>&#x7c;<italic>z</italic>) is modelled as a Gaussian distribution with mean <italic>&#x3bc;</italic>(<italic>z</italic>) and diagonal co-variance <italic>&#x3c3;</italic>(<italic>z</italic>) multiplied by the identity <bold>
<italic>I</italic>
</bold>, which implicitly assumes independence between the dimensions of <italic>z</italic>.<disp-formula id="e2">
<mml:math id="m12">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="italic">N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>&#x3bc;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>&#x3c3;</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2a;</mml:mo>
<mml:mi mathvariant="bold-italic">I</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>The Pose2Shape network architecture is composed of an encoder, which receives 3D skeletal joint positions as input, and a decoder, see supplementary material network details, that generates high resolution 3D meshes. The encoder is trained to map the posterior distribution of data samples <italic>p</italic> to the latent space <italic>z</italic>, meanwhile forcing the latent variables <italic>z</italic> to comply with the prior distribution of <italic>P</italic>(<italic>z</italic>). However, both the posterior distribution <italic>P</italic>(<italic>z</italic>&#x7c;<italic>p</italic>) and <italic>P</italic>(<italic>p</italic>) are unknown. Therefore, variational networks give the solution that the posterior distribution is a variational distribution <inline-formula id="inf11">
<mml:math id="m13">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. In order to make <inline-formula id="inf12">
<mml:math id="m14">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> consistent with the distribution <italic>P</italic>(<italic>z</italic>), we use the Kullback-Leiber (KL) divergence (<xref ref-type="bibr" rid="B31">Kingma and Welling, 2013</xref>):<disp-formula id="e3">
<mml:math id="m15">
<mml:mi>K</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>The decoder is trained to regress from any latent vector <inline-formula id="inf13">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> in the learnt space <italic>z</italic> to a 4D mesh representation <inline-formula id="inf14">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. <xref ref-type="disp-formula" rid="e4">Eq. (4)</xref> defines the loss function minimised by the network to achieve a compact latent space representation and generative network output.<disp-formula id="e4">
<mml:math id="m18">
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>K</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>This is an optimal approximation of the true samples <inline-formula id="inf15">
<mml:math id="m19">
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, where <italic>&#x3c9;</italic> weighs the importance of the KL divergence, and <inline-formula id="inf16">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the ground truth 4D mesh for the 3D skeletal pose <inline-formula id="inf17">
<mml:math id="m21">
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> of sequence <italic>s</italic> at time&#x20;<italic>t</italic>.</p>
<sec id="s3-2-1">
<title>3.2.1 Training Details</title>
<p>The network architecture used to regress 3D skeletal pose to 4D mesh shape is summarised in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. The network was empirically found to learn a good latent space distribution with accurate 4D shape generation using a training cycle of 10<sup>4</sup> epochs, which is optimised through validation data to avoid over-fitting with a learning rate of 0.001. The datasets are split by randomly selecting frames from each motion sequence with &#x2248;80<italic>%</italic> used for training and &#x2248;20<italic>%</italic> used for validation. We set the prior probability over latent variables to be a Gaussian distribution with zero mean and unit standard variation, <italic>p</italic> (<italic>z</italic>) &#x3d; <italic>N</italic> (<italic>z</italic>; 0, <bold>
<italic>I</italic>
</bold>). We use Adam optimisation (<xref ref-type="bibr" rid="B31">Kingma and Welling, 2013</xref>) with a momentum of 0.9 to optimise <xref ref-type="disp-formula" rid="e4">Eq. (4)</xref> between the reconstructed and ground truth mesh vertices, and simultaneously the KL divergence of the 3D skeletal pose distribution. Evaluation of the performance of the network for shape representation from skeletal pose is given in <xref ref-type="sec" rid="s5">Section&#x20;5</xref>.</p>
</sec>
</sec>
<sec id="s3-3">
<title>3.3 Deep4D: Pose2Appearance Network</title>
<p>In this section, we propose the use of a Pose2Appearance network for the synthesis of high-resolution dynamic mesh texture maps from the encoded skeletal pose latent space representation. A similar approach described as the progressive growing of GANs was first introduced by <xref ref-type="bibr" rid="B29">Karras et&#x20;al. (2017)</xref> to improve image synthesis quality and training stability of Generative Adversarial Networks (GANs) (<xref ref-type="bibr" rid="B21">Goodfellow et&#x20;al., 2014</xref>).</p>
<p>A GAN consists of two networks, a generator and a discriminator. The generator produces images from a latent code, and the distribution of these images should be indistinguishable from the training distribution. The discriminator evaluates the quality of the images produced by the generator, forcing the generator to learn how to produce high-quality images so that the discriminator cannot tell the difference. A progressive generator generally consists of a network where the training begins with a low-resolution image and progressively increases the resolution until it reaches a target resolution. This incremental multi-resolution approach allows the training first to discover the large-scale structure of the distribution of the images and then shifts the attention to finer-scale details, whereas in traditional GAN architectures, all scales are learned simultaneously.</p>
<p>In this section, we adapt the generator from the progressive growing of GANs (<xref ref-type="bibr" rid="B29">Karras et&#x20;al., 2017</xref>) to learn how to synthesise high-resolution texture appearance from the latent probability distribution learned from 3D skeletal motion, <xref ref-type="sec" rid="s3-2">Section 3.2</xref>. The proposed Pose2Appearance for high-resolution texture map synthesis from the latent space vector is illustrated in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Progressive appearance network overview. The input is learnt latent vectors from 3D skeletal motion, and the output is a high-resolution 2D texture appearance.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g003.tif"/>
</fig>
<p>The Pose2Appearance initially starts with a small feed-forward network, see supplementary for details, which consists of four fully-connected layers, where the input consists of learnt latent vector <inline-formula id="inf18">
<mml:math id="m22">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> of dimension 128, which corresponds to the dimensions of the latent space learned from the Pose2Shape network, and the output dimension of the fourth layer is 512, to match the input size requirements of the first convolutional layer, as illustrated in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>. The convolutional layers consist of nine blocks, where each block represents a different resolution, and its output is a high-resolution texture <inline-formula id="inf19">
<mml:math id="m23">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>.</p>
<p>We also experimented with a VAE network for appearance synthesis. This experiment was found to result in significant blur and loss of detail. The VAE assumes the same input and output, hence not being a suitable architecture for the problem. For this reason, a more sophisticated network approach is required, <xref ref-type="sec" rid="s5-3">Section 5.3</xref> for comparison with state-of-the-art methods.</p>
<sec id="s3-3-1">
<title>3.3.1 Training Details</title>
<p>The Pose2Appearance training starts with a 4 &#xd7; 4 resolution and progressively grows the network layers until it reaches 1,024 &#xd7; 1,024 resolution. The network progresses through the training by adding new layers with double the size. There are two stages for training the growing process (<xref ref-type="fig" rid="F4">Figure&#x20;4</xref>), the first stage is when a new layer is added a fading stage begins where the new layer will be smoothly added to the network. This new layer will operate as a residual block, whose weight <italic>&#x3c4;</italic> increases linearly from 0 to 1. When the fading stage is over the second stage is initiated, the stabilising stage, where the new layer is fully integrated with the network, and it iterates over another training cycle. This training pattern repeats until it reaches the full resolution of 1,024 &#xd7; 1,024. For every stage, we gradually decrease the minibatch size, vary the stabiliser number of training iterations and vary the convergence tolerance. These parameters are necessary to avoid exceeding the available memory budget and decrease the training&#x20;time.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Progressive appearance training overview. Convolutional layers training stages, Stabilising stage with 16 &#xd7; 16 resolution, and Fading stage to the next 32 &#xd7; 32 resolution&#x20;layer.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g004.tif"/>
</fig>
<p>The generator network is trained using Adam (<xref ref-type="bibr" rid="B31">Kingma and Welling, 2013</xref>), with a constant learning rate of 0.001 across the full training. We use leaky ReLU (<xref ref-type="bibr" rid="B48">Tan et&#x20;al., 2018</xref>) with a leakiness value of 0.2, equalised learning rate for all layers, except the last layer that uses linear activation, and pixel normalisation of the feature vector after each Conv 3 &#xd7; 3 layer. All weights of the convolutional, fully-connected and affine transform layers are initialised using a Gaussian distribution with zero mean and unit standard variation, <italic>p</italic> (<italic>z</italic>) &#x3d; <italic>N</italic> (<italic>z</italic>; 0, <bold>
<italic>I</italic>
</bold>). Stochastic gradient descent with a momentum of 0.9 is used to minimise the mean squared error (MSE) loss between reconstructed image <inline-formula id="inf20">
<mml:math id="m24">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and the ground truth samples&#x20;<inline-formula id="inf21">
<mml:math id="m25">
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>.</p>
</sec>
</sec>
<sec id="s3-4">
<title>3.4 4D Volumetric Video Synthesis</title>
<p>The latent space of the learnt motion allows the pre-trained generators for shape <inline-formula id="inf22">
<mml:math id="m26">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and texture <inline-formula id="inf23">
<mml:math id="m27">
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> to interpolate between the captured 4D volumetric video shape and appearance sequences. Because the variational encoder-decoder produces a compact latent space it is possible to generate novel content by sampling from the learned space or interpolation of sampled latent vectors. Sampling of the latent space allows reproduction of the original 4D volumetric video sequences with a low reconstruction error. The sampling can be performed in two ways: random walk in the latent space that fits in the Gaussian distribution learned; or through 3D joint positions given as input to the networks. In this work, sampling is performed through 3D joint position as input. Interpolation in the learnt latent space allows transitions between observed sequences to create plausible novel motions. Interpolating the latent space is only possible because of the compact space representation produced by the generative network. This is performed, firstly, by sampling latent vectors using 3D joint position as input to the network. Once the latent vector is computed for two motion frames then interpolation is performed according to&#x20;<xref ref-type="disp-formula" rid="e5">Eq. 5</xref>.<disp-formula id="e5">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
<label>(5)</label>
</disp-formula>where <italic>z</italic>
<sub>
<italic>i</italic>
</sub> is the interpolated latent vector, <italic>&#x3b1;</italic> defines a normalised weighting [0&#x2025;1] between latent vectors <inline-formula id="inf24">
<mml:math id="m29">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf25">
<mml:math id="m30">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. Intermediate 4D shape and texture frames are synthesised to qualitatively evaluate how well the network is representing the 4D shape and appearance, <xref ref-type="sec" rid="s5-2">Section&#x20;5.2</xref>.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Deep4D Motion Graphs</title>
<p>The following section introduces Deep4D motion graphs, a novel approach to generate motion graphs (<xref ref-type="bibr" rid="B11">Casas et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B12">Casas et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B5">Boukhayma and Boyer, 2015</xref>; <xref ref-type="bibr" rid="B4">Boukhayma and Boyer, 2019</xref>) from the Deep4D representation introduced in <xref ref-type="sec" rid="s3">Section 3</xref>. The motion graphs, animation and rendering blocks presented in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> are discussed in detail to demonstrate the steps taken to generate motion graphs capable of animating learnt characters from 4D volumetric video datasets. The goal is to merge the popular deep learning research field with traditional animation pipelines to begin a new era for computer graphics, creating novel mechanisms to produce realistic human animations.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Animation framework pipeline. The image illustrates the online block that generates high-resolution 4D volumetric video. The 4D Motion Graph contains the learnt motion representation allowing the decoders in real-time to synthesise 4D volumetric video sequences consisting of 4D mesh, 2D texture appearance.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g005.tif"/>
</fig>
<p>Firstly, we discuss the input data to the animation framework along with the pre-requisites for initialisation. Secondly, the generation of motion graphs for learnt 4D volumetric video is presented along with a discussion of the metrics chosen to evaluate similarity and transition costs between motion frames. Finally, a real-time motion synthesis approach to generate 4D video sequences with interactive animation control by concatenating and blending between the captured motion sequences is presented.</p>
<sec id="s4-1">
<title>4.1 Input Data</title>
<p>The framework receives as input, skeletal motion data from 4D volumetric video estimated using a Skeleton Driven Surface Registration (SDSR) framework (<xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>) and latent vectors <inline-formula id="inf26">
<mml:math id="m31">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> of each motion sequence <inline-formula id="inf27">
<mml:math id="m32">
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> learned in <xref ref-type="sec" rid="s3">Section 3</xref> for 4D shape and appearance learnt from a skeletal&#x20;pose.</p>
<p>In the context of this section, a sequence of motion frames <inline-formula id="inf28">
<mml:math id="m33">
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> refers to collections of frames which contain representative latent vectors, and skeletal structures given by the SDSR framework as follows, <inline-formula id="inf29">
<mml:math id="m34">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf30">
<mml:math id="m35">
<mml:msubsup>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is a skeletal structure from a motion sequence, which contains <inline-formula id="inf31">
<mml:math id="m36">
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> number of frames <inline-formula id="inf32">
<mml:math id="m37">
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2026;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, representative of the original motion dataset. Lastly, it is necessary to utilise the pre-trained mesh generator <inline-formula id="inf33">
<mml:math id="m38">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and the appearance generator <inline-formula id="inf34">
<mml:math id="m39">
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> from <xref ref-type="sec" rid="s3">Section 3</xref> to interpret each latent vectors <inline-formula id="inf35">
<mml:math id="m40">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> stored as a motion frame in latent motion sequence&#x20;<inline-formula id="inf36">
<mml:math id="m41">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>.</p>
<p>The generative networks synthesise <inline-formula id="inf37">
<mml:math id="m42">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> meshes and <inline-formula id="inf38">
<mml:math id="m43">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> texture maps for every <inline-formula id="inf39">
<mml:math id="m44">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, which represents a temporally consistent 4D mesh and appearance, i.e. the topology, vertex connectivity and texture coordinates are constant across all frames and sequences. The construction of a motion graph is independent of the learnt model, allowing the framework to generalise its application to other types of models. A motion graph is interpreted as a directed weighted graph structure built from captured 4D volumetric video sequences, where graph nodes represent frames that contain latent vectors which hold information about shape, motion and appearance, and edges link nodes together to represent motion pathways between frames.</p>
</sec>
<sec id="s4-2">
<title>4.2&#x20;Pre-processing</title>
<p>The data is required to be pre-processed; this offline process starts with training the generative networks described in <xref ref-type="sec" rid="s3">Section 3</xref> for a skeleton motion sequences of a human character. Once training is complete the generators <inline-formula id="inf40">
<mml:math id="m45">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf41">
<mml:math id="m46">
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> are used to recover the 3D meshes and 2D textures represented by each latent vector <inline-formula id="inf42">
<mml:math id="m47">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, to allow the pre-processing step to be automated. The first step in the pre-processing stage is to connect frames within the same sequences automatically, and if possible create loops for cyclic motions, consequently a sequence can infinitely repeat itself. Loops are generated via searching on a similarity matrix <inline-formula id="inf43">
<mml:math id="m48">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> for all pairs of frames in the same sequence to automatically choose the minimum cost, <xref ref-type="sec" rid="s4-3">Section 4.3</xref> and <xref ref-type="sec" rid="s4-4">Section 4.4</xref>. Transitions within the same sequence should produce the most natural motion; hence the shape and motion cost should be&#x20;small.</p>
<p>The next step is to fully connect the graph by adding all possible transition combinations between sequences to allow better path estimations to be found for all frames. This step will generate a fully connected graph with appropriated edge weights using shape, motion and dynamic time warping metrics, as detailed in the following sections. Lastly, the graph is optimised using Dijkstra&#x2019;s algorithm to minimise the number of transition in the final motion graph, as detailed in <xref ref-type="sec" rid="s4-5">Section&#x20;4.5</xref>.</p>
</sec>
<sec id="s4-3">
<title>4.3 Shape Similarity Metric</title>
<p>Similarity is computed for every pair of frames in the input 4D volumetric video sequences <inline-formula id="inf44">
<mml:math id="m49">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf45">
<mml:math id="m50">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is a frame <italic>t</italic>
<sub>
<italic>u</italic>
</sub> from the <italic>i</italic>th sequence <inline-formula id="inf46">
<mml:math id="m51">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, comprising meshes <inline-formula id="inf47">
<mml:math id="m52">
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and textures <inline-formula id="inf48">
<mml:math id="m53">
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, where <italic>i</italic>&#x20;&#x3d; [1 &#x2026; <italic>N</italic>
<sub>
<italic>S</italic>
</sub>]. For a given latent vector <inline-formula id="inf49">
<mml:math id="m54">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> the decoder <inline-formula id="inf50">
<mml:math id="m55">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> reconstructs temporally consistent geometry, and the appearance generator <inline-formula id="inf51">
<mml:math id="m56">
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> reconstructs the 2D texture appearance of generated frame. The shape, motion and appearance similarity is computed for every pair of source <inline-formula id="inf52">
<mml:math id="m57">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and target <inline-formula id="inf53">
<mml:math id="m58">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>frames, having <inline-formula id="inf54">
<mml:math id="m59">
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and <italic>t</italic>
<sub>
<italic>v</italic>
</sub> &#x2208; [1, <italic>N</italic>
<sub>
<italic>T</italic>
</sub>] frames for all sequences <italic>i</italic>, <italic>j</italic>&#x20;&#x2208; [1, <italic>N</italic>
<sub>
<italic>S</italic>
</sub>].<disp-formula id="e6">
<mml:math id="m60">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">A</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>Where <italic>&#x3b8;</italic> weights the relative importance of shape and appearance similarity, giving a complete similarity matrix <inline-formula id="inf55">
<mml:math id="m61">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> for all frames generated by the learnt 4D volumetric video representation. To measure shape similarity we use the Euclidean distances and velocities between mesh vertices as illustrated in <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>.<disp-formula id="e7">
<mml:math id="m62">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>Where vertex velocity <inline-formula id="inf56">
<mml:math id="m63">
<mml:msubsup>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and <italic>N</italic>
<sub>
<italic>V</italic>
</sub> is the number of vertices. The appearance similarity uses the average absolute difference of the 2D texture appearance between two frames as illustrated in <xref ref-type="disp-formula" rid="e8">Eq. 8</xref>.<disp-formula id="e8">
<mml:math id="m64">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">A</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:math>
<label>(8)</label>
</disp-formula>Where <italic>N</italic>
<sub>
<italic>X</italic>
</sub> is the number of pixels. The similarities are normalised to the range (0,1) as follows:<disp-formula id="e9">
<mml:math id="m65">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(9)</label>
</disp-formula>Where <bold>
<italic>SIMQ</italic>
</bold> (&#x22c5;) is either <bold>
<italic>SIM</italic>
</bold>
<sub>
<bold>
<italic>M</italic>
</bold>
</sub> (&#x22c5;) or <bold>
<italic>SIM</italic>
</bold>
<sub>
<bold>
<italic>A</italic>
</bold>
</sub> (&#x22c5;) similarity metrics for shape and appearance. The pre-computed similarity matrix <inline-formula id="inf57">
<mml:math id="m66">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> for all frames allows to evaluate in real-time the similarity cost between any source and target meshes.</p>
</sec>
<sec id="s4-4">
<title>4.4 Transition Edge Cost</title>
<p>An edge in a motion graph represents a transition between two frames, where for clarity frames will be described as nodes. For every edge, we associate a weight to represent the similarity of shape transitions between nodes quantitatively. Realistic transitions should require little change in shape and appearance corresponding to a small similarity score. Hence the metric used takes into account the optimal surface interpolation cost between any pair of nodes (<xref ref-type="bibr" rid="B3">Boukhayma and Boyer, 2017</xref>). The cost of transitioning is the sum of intermediate poses between source node <italic>u</italic> and destination node <italic>v</italic> weighted by the similarity score for each intermediate&#x20;frame.</p>
<p>In order to smoothly blend source node <italic>u</italic> from a 3D mesh sequence to destination node <italic>v</italic> from another sequence, it is necessary to consider a blend window of length <italic>b</italic>. This window represents a successive number of nodes <italic>b</italic>
<sub>
<italic>u</italic>
</sub>, on the source sequence it begins at node <italic>u</italic> and ends at node <italic>u</italic>&#x20;&#x2b; <italic>b</italic>
<sub>
<italic>u</italic>
</sub> &#x2212; 1, in the destination sequence a window <italic>b</italic>
<sub>
<italic>v</italic>
</sub> ending at node <italic>v</italic> and starting at node <italic>v</italic>&#x20;&#x2212; <italic>b</italic>
<sub>
<italic>v</italic>
</sub> &#x2b; 1. Once, the window frame is initialised between source and destination sequence, it is necessary to extrapolate the nodes that gradually blend both sequences, generating smooth realistic transitions. To extract the optimal nodes from source and destination sequences we use a variant of dynamic time warping (DTW) (<xref ref-type="bibr" rid="B38">Muller, 2007</xref>; <xref ref-type="bibr" rid="B56">Witkin and Popovic, 1995</xref> <xref ref-type="bibr" rid="B55">Wang and Bodenheimer, 2008</xref>; <xref ref-type="bibr" rid="B12">Casas et&#x20;al., 2013</xref>) to estimate the best temporal warps <italic>w</italic>
<sub>
<italic>u</italic>
</sub> and <italic>w</italic>
<sub>
<italic>v</italic>
</sub> respectively with respect to the similarity metric defined in <xref ref-type="disp-formula" rid="e6">Eq. (6)</xref>. DTW was first introduced by <xref ref-type="bibr" rid="B44">Sakoe and Chiba (1990)</xref> for signal time alignment, it was used in conjunction with dynamic programming techniques for the recognition of isolated words and it had been widely used since then mainly for recognition tasks. The transition duration varies within a third of a second and 2&#xa0;s (<xref ref-type="bibr" rid="B55">Wang and Bodenheimer, 2008</xref>), hence we allow the length <italic>b</italic>
<sub>
<italic>u</italic>
</sub> and <italic>b</italic>
<sub>
<italic>v</italic>
</sub> to vary between boundaries <italic>b</italic>
<sub>min</sub> and <italic>b</italic>
<sub>
<italic>max</italic>
</sub>. The optimal transitions with minimal total similarity cost <italic>D</italic>(<italic>u</italic>, <italic>v</italic>) through the path generated from the DTW algorithm.<disp-formula id="e10">
<mml:math id="m67">
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>where <inline-formula id="inf58">
<mml:math id="m68">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the shape similarity cost defined in <xref ref-type="sec" rid="s4-3">Section 4.3</xref>, and <italic>D</italic>
<sub>
<italic>l</italic>
</sub> is the length of the path found by the DTW algorithm considered as the transition duration, see supplementary material for illustration. The optimisation above finds the following optimal parameters (<italic>b</italic>
<sub>
<italic>u</italic>
</sub>, <italic>b</italic>
<sub>
<italic>v</italic>
</sub>, <italic>w</italic>
<sub>
<italic>u</italic>
</sub>, <italic>w</italic>
<sub>
<italic>v</italic>
</sub>, <italic>D</italic>
<sub>
<italic>l</italic>
</sub>, <italic>D</italic>
<sub>
<italic>l</italic>
</sub>), which are considered later for motions synthesis. Similar to <xref ref-type="sec" rid="s4-3">Section 4.3</xref>, we define the edge weight between nodes to be the surface deformation cost <italic>D</italic> (<italic>u</italic>, <italic>v</italic>) and its interpolated duration cost <italic>D</italic>
<sub>
<italic>l</italic>
</sub>(<italic>u</italic>, <italic>v</italic>). <xref ref-type="disp-formula" rid="e11">Eq. (11)</xref> summarises the definition for the edge cost between nodes <italic>u</italic> and <italic>v</italic>.<disp-formula id="e11">
<mml:math id="m69">
<mml:msup>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mspace width="0.28em"/>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
<p>For the case nodes <italic>u</italic> and <italic>v</italic> are from the same sequence the surface deformation should be minimal. To control the tolerance between surface deformation and transition duration we add weight <italic>&#x3b1;</italic>.</p>
<p>This process will create a fully connected digraph where edges are weighted for the shape similarity and transition cost between nodes, in the following <xref ref-type="sec" rid="s4-5">Section 4.5</xref> we will discuss how to prune and optimise the connectivity of the complete digraph.</p>
</sec>
<sec id="s4-5">
<title>4.5 Motion Graph Optimisation</title>
<p>The last stage in the framework aims to find a globally optimal solution to minimise the number of transitions between nodes. Plausible transitions can be achieved by selecting the minimum cost transition from the similarity matrix between sequences, to generate a motion graph. A fully connected digraph was generated from <xref ref-type="sec" rid="s4-4">Section 4.4</xref>, which connects every pair of nodes for all existing motion sequences. Therefore selecting the minimum cost transition for every node would maintain dense connectivity in the&#x20;graph.</p>
<p>We have implemented a globally optimal strategy that extracts and maintains only the best paths between every pair of nodes (<xref ref-type="bibr" rid="B25">Huang et&#x20;al., 2009)</xref> <xref ref-type="bibr" rid="B13">Casas et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B12">Casas et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B5">Boukhayma and Boyer, 2015</xref>; <xref ref-type="bibr" rid="B3">Boukhayma and Boyer, 2017</xref>). This strategy corresponds to extracting the essential sub-graph from the complete digraph induced from the input sequences (<xref ref-type="bibr" rid="B2">Bordino et&#x20;al., 2008</xref>). This method ensures the existence of at least one transition between any two nodes in the graph, which potentially yields a better use of the original data with less dead ends. Given the fully connected digraph, we use the Dijkstra algorithm on every pair of nodes to extract the shortest paths between source and target nodes. Once this process is completed, we remove all edges that do not belong to the new generated paths, giving a connected digraph that contains only the necessary least cost transitions. The resulting structure is also referred to as the union of shortest-path trees rooted at every graph node. This solution will guarantee the minimal difference when transitioning from frames of different sequences.</p>
</sec>
<sec id="s4-6">
<title>4.6 4D Volumetric Video Animation</title>
<p>This section demonstrates generation of 4D volumetric video using the Deep4D motion graphs. To generate a continuous stream of animation between motion sequences it is necessary to calculate the least costly transition path between a source frame <inline-formula id="inf59">
<mml:math id="m70">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and a target frame <inline-formula id="inf60">
<mml:math id="m71">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> from different motion sequences. As discussed previously, the least costly transition should be a transition within the same motion sequence, consequently if the animation remains unchanged by the user the framework will play the same motion in a loop. If the user requests the character change to a new motion state, the animation framework computes the minimum transition cost <inline-formula id="inf61">
<mml:math id="m72">
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> from the current motion frame <inline-formula id="inf62">
<mml:math id="m73">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> to the selected motion sequence <inline-formula id="inf63">
<mml:math id="m74">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and returns the following parameters (<italic>b</italic>
<sub>
<italic>u</italic>
</sub>, <italic>b</italic>
<sub>
<italic>v</italic>
</sub>, <italic>w</italic>
<sub>
<italic>u</italic>
</sub>, <italic>w</italic>
<sub>
<italic>v</italic>
</sub>, <italic>D</italic>
<sub>
<italic>l</italic>
</sub>), <xref ref-type="sec" rid="s4-4">Section 4.4</xref>. These parameters allow interpolation of the intermediate frames between frame <italic>u</italic> and <italic>v</italic> with a transition length of <italic>D</italic>
<sub>
<italic>l</italic>
</sub>, creating a seamless transition in real-time between different motion sequences. The approach presented in <xref ref-type="sec" rid="s4-4">Section 4.4</xref> finds the corresponding pair of frames by computing the shortest path on the warps (<italic>w</italic>
<sub>
<italic>u</italic>
</sub>, <italic>w</italic>
<sub>
<italic>v</italic>
</sub>). The following sub-sections discuss how to synthesise 4D volumetric video and how intermediate frames are generated using generative networks.</p>
<sec id="s4-6-1">
<title>4.6.1 4D Motion Synthesis</title>
<p>For every node in the motion graph we store the latent vector <inline-formula id="inf64">
<mml:math id="m75">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> that corresponds to a particular frame of a motion sequence. This allows for the pre-trained generator <inline-formula id="inf65">
<mml:math id="m76">
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf66">
<mml:math id="m77">
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> from the generative networks to reconstruct 3D mesh and 2D texture appearance for any given latent vector. At run-time the framework provides a latent vector <inline-formula id="inf67">
<mml:math id="m78">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> of the current frame and generates the corresponding dynamic mesh shape and texture appearance to synthesise the 4D volumetric video. <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> illustrates synthesised 4D volumetric video sequences. The motion graph representation generates seamless transitions to enable interactive character animation. The world coordinates of each frame are given by the root of the original 3D skeletal motion information which is used to transform the 3D mesh content given by the generators, allowing it to reproduce the original physical motion translations.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Latent space interpolation, two frames with distinct motions are selected, on the left surrounded with a green box is the source, on the right surrounded with a red box is the target, in between is the interpolated results for 4D shape and appearance.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g006.tif"/>
</fig>
</sec>
<sec id="s4-6-2">
<title>4.6.2 Motion Frames Interpolation</title>
<p>Edges in the motion graph represent transitions between frames take into account the shape, motion and appearance similarity. It is necessary to create intermediate blend frames to smoothly transition between different sequences. As seen in <xref ref-type="sec" rid="s3-4">Section 3.4</xref>, the generative network allows synthesis of frames via interpolation of the latent vectors. Therefore, we perform a linear interpolation of the latent vectors for the given transition parameters, see <xref ref-type="sec" rid="s4-4">Section 4.4</xref>, to create smooth human character animation. <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>, <xref ref-type="fig" rid="F8">Figure&#x20;8</xref> illustrate interpolation between distinct body and face poses generating plausible intermediate mesh and texture.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Thomas appearance synthesis qualitative evaluation of <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al. (2018)</xref>.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g007.tif"/>
</fig>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Interpolation between two frames from the original motion sequence surrounded with a green and orange box. The top row represents the original sequence composed of five consequent frames. The middle row is the result of interpolating the learnt latent vector of the proposed network. The bottom row is the result of using linear blend skinning.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g008.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5 Results and Evaluation</title>
<p>This section presents results and evaluation for the proposed Pose2Shape network, the Pose2Appearance network from motion, and their applicability using Deep4D motion graphs to generated realistic animations, introduced in <xref ref-type="sec" rid="s3-2">Section 3.2</xref> and <xref ref-type="sec" rid="s3-3">Section 3.3</xref>. To evaluate the 4D animation framework we use publicly available volumetric video datasets for whole body and facial performance. The SurfCap dataset, JP and Roxanne characters, and Dan character (<xref ref-type="bibr" rid="B14">Casas et&#x20;al., 2014</xref>) are reconstructed using multi-view stereo (<xref ref-type="bibr" rid="B46">Starck and Hilton, 2007</xref>) and temporally aligned with SDSR (<xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>) which allows for surface pose manipulation. Martin dataset (<xref ref-type="bibr" rid="B32">Klaudiny and Hilton, 2012</xref>) consists of one sequence of temporally aligned geometry and texture appearance of a human face, and 3D facial key-points given by OpenPose (<xref ref-type="bibr" rid="B9">Cao et&#x20;al., 2021</xref>). Thomas dataset (<xref ref-type="bibr" rid="B5">Boukhayma and Boyer, 2015</xref>) consists of four sequences of temporally aligned meshes and texture appearance. An overview of dataset properties is shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. Examples of character animation using Deep4D motion graphs are shown in <xref ref-type="fig" rid="F9">Figures&#x20;9</xref>,<xref ref-type="fig" rid="F6">6</xref>. Results demonstrate that the proposed generative representation allows interactive character animation with seamless transitions between sequences based on interpolation of the latent space. The meshes are coloured to illustrate different motion sequences and interpolation between them when performing a blend transition. The learned generative model for shape and appearance synthesises animation with a quality similar to the input 4D&#x20;video.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Dan and Roxanne characters performing several animations, top-left: transition from jump to walk to reach; top-right: from walk to stand; bottom-left: from jump short to jump long; bottom-right: from jump short to jump high sequence. Mesh colours indicate motion sequences and the generated transitions.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g009.tif"/>
</fig>
<sec id="s5-1">
<title>5.1 Quantitative Results</title>
<p>The variational encoder-decoder uses <xref ref-type="disp-formula" rid="e4">Eq. (4)</xref> as a metric to predict plausible shape reconstructions from skeletal pose. The Pose2Appearance network uses the mean squared error (MSE) as loss function between generated images and ground truth as a metric to predict plausible high resolution textures. The comparison was performed between the training data, to ensure minimum error when sampling the original sequences, and validation data to guarantee a plausible result when generating unseen&#x20;mesh.</p>
<p>We compare generated 3D meshes with ground truth geometry acquired from multiple view stereo reconstruction (<xref ref-type="bibr" rid="B46">Starck and Hilton, 2007</xref>). 3D mesh evaluation is performed using Hausdorff distance defined as <italic>d</italic>
<sub>
<italic>H</italic>
</sub>(<italic>A</italic>, <italic>B</italic>) &#x3d; <italic>max</italic>{&#x2009;sup<sub>
<italic>a</italic>&#x2208;<italic>A</italic>
</sub>
<italic>d</italic>(<italic>a</italic>, <italic>B</italic>),&#x2009;sup<sub>
<italic>b</italic>&#x2208;<italic>B</italic>
</sub>
<italic>d</italic>(<italic>b</italic>, <italic>A</italic>)}, where <italic>d</italic>(<italic>a</italic>, <italic>B</italic>) and <italic>d</italic>(<italic>b</italic>, <italic>A</italic>) is the distance from a point <italic>a</italic> to a set <italic>B</italic> and from a point <italic>b</italic> to a set <italic>A</italic>, which has been shown to be a good measurement between 3D meshes. The comparison contains training and validation data for all sequences, <xref ref-type="table" rid="T1">Table&#x20;1</xref>. The appearance is evaluated using three metrics that are commonly used to assess image quality: mean squared distance (MSE); multi-scaled structural similarity (MS-SSIM); peak signal to noise ratio (PSNR), <xref ref-type="table" rid="T1">Table&#x20;1</xref> for results.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison of error metrics used for evaluation of 3D mesh and 2D texture appearance. The values represent the average error across the all motion sequence for different datasets.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Dataset</th>
<th colspan="2" align="center">Mesh</th>
<th colspan="3" align="center">Appearance</th>
</tr>
<tr>
<th align="center">RMSE (m)</th>
<th align="center">STDDV</th>
<th align="center">MSE</th>
<th align="center">SSIM</th>
<th align="center">PSNR</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Dan <xref ref-type="bibr" rid="B14">Casas et&#x20;al. (2014)</xref>
</td>
<td align="char" char=".">0.0158</td>
<td align="char" char=".">0.0156</td>
<td align="char" char=".">0.0008</td>
<td align="char" char=".">0.8417</td>
<td align="char" char=".">30.7327</td>
</tr>
<tr>
<td align="left">JP <xref ref-type="bibr" rid="B46">Starck and Hilton (2007)</xref>
</td>
<td align="char" char=".">0.0266</td>
<td align="char" char=".">0.0257</td>
<td align="char" char=".">0.0007</td>
<td align="char" char=".">0.9610</td>
<td align="char" char=".">31.1675</td>
</tr>
<tr>
<td align="left">Martin <xref ref-type="bibr" rid="B32">Klaudiny and Hilton (2012)</xref>
</td>
<td align="char" char=".">0.0027</td>
<td align="char" char=".">0.0015</td>
<td align="char" char=".">0.0001</td>
<td align="char" char=".">0.9813</td>
<td align="char" char=".">38.6342</td>
</tr>
<tr>
<td align="left">Roxanne <xref ref-type="bibr" rid="B46">Starck and Hilton (2007)</xref>
</td>
<td align="char" char=".">0.0166</td>
<td align="char" char=".">0.0161</td>
<td align="char" char=".">0.0002</td>
<td align="char" char=".">0.9804</td>
<td align="char" char=".">36.0430</td>
</tr>
<tr>
<td align="left">Thomas <xref ref-type="bibr" rid="B5">Boukhayma and Boyer (2015)</xref>
</td>
<td align="char" char=".">0.0125</td>
<td align="char" char=".">0.0122</td>
<td align="char" char=".">0.0002</td>
<td align="char" char=".">0.9889</td>
<td align="char" char=".">35.9946</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5-2">
<title>5.2 Qualitative Evaluation</title>
<p>We compare our network generated results to rendered images of the original textured model and synthesised 4D volumetric content, <xref ref-type="fig" rid="F10">Figure&#x20;10</xref> and supplementary material for more results. Our network is able to capture dynamic shape detail and high frequency appearance details such as wrinkles and hair movement, <xref ref-type="fig" rid="F10">Figure&#x20;10</xref>. The network is also capable of interpolating the existing data to generate novel geometry and appearance within the learned space. To test the interpolation performance of the network, the mesh and appearance of two encoded frames were selected and intermediate frames synthesised. <xref ref-type="fig" rid="F7">Figures&#x20;7</xref>,<xref ref-type="fig" rid="F8">8</xref> shows a more challenging example for two randomly selected frames with large differences in shape and appearance, note that the method is able to produce a natural transition between frames.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>The images represent the same motion frame, the top images were rendered using the original geometry and texture appearance, and the bottom images were rendered using the generative networks. The error image is an error histogram between the original and resulted images. It is visible from the reconstructed that it preserves original details such as, reflection on the eye ball, clothing patterns and wrinkles.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g010.tif"/>
</fig>
<p>The proposed generative network maps 3D skeletal pose to 4D volumetric video sequences consisting of shape and appearance. To evaluate this capability we use existing public skeletal motion capture sequences (<xref ref-type="bibr" rid="B15">CMU Graphics Lab, 2001</xref>) to synthesise novel 4D animations. To drive the generative network, we use the 3D skeletal joint positions <inline-formula id="inf68">
<mml:math id="m79">
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> to obtain the encoded latent vectors <inline-formula id="inf69">
<mml:math id="m80">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, sampling from the learnt distribution <italic>P</italic> (<italic>p</italic>&#x7c;<italic>z</italic>). <xref ref-type="fig" rid="F9">Figure&#x20;9</xref> shows three characters driven using a novel motion capture sequence This demonstrates the potential to generate novel plausible 4D shape and appearance sequences from MoCap input for similar motions.</p>
</sec>
<sec id="s5-3">
<title>5.3 Appearance Synthesis Evaluation</title>
<p>In this section, we evaluate the performance of the progressive appearance generator against the state-of-the-art method proposed by <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al. (2018)</xref> for facial image synthesis. The variant network architecture was chosen to allow for appearance synthesis only, as we intend to evaluate the texture synthesis quality, see supplementary for network illustration. Therefore we have removed the mesh and view-point conditioning from the original network architecture. We trained this network on 2D textures from the Thomas (<xref ref-type="bibr" rid="B5">Boukhayma and Boyer, 2015</xref>) and Martin (<xref ref-type="bibr" rid="B32">Klaudiny and Hilton, 2012</xref> datasets, where the training took approximately 10&#xa0;days for 10<sup>4</sup> training cycles, with a mini-batch size of 64. This network minimises the MSE error and the KL-divergence simultaneously, similar to the proposed approach.</p>
<p>
<xref ref-type="fig" rid="F11">Figure&#x20;11</xref> illustrates qualitative evaluation for this experiment. We have chosen one random sample from the training dataset to evaluate the quality of the texture synthesis given a seen example. <xref ref-type="fig" rid="F11">Figure&#x20;11A</xref> presents heat-map images to compare the synthesised result against the ground-truth for the proposed and Lombardi networks. It is visible that the proposed network outperforms the <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al., 2018</xref> approach, this is more visible on the close-up <xref ref-type="fig" rid="F11">Figure&#x20;11B</xref>, where the details on the t-shirt have been lost when using the <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al., 2018</xref> network. The proposed network is capable of preserving the printed image on the t-shirt along with wrinkles present in the original image. The lack of detail and the presence of blurred results from state-of-the-art <xref ref-type="bibr" rid="B35">Lombardi et&#x20;al., 2018</xref> network has led to the network presented in <xref ref-type="sec" rid="s3-3">Section 3.3</xref>. The proposed approach is a more sophisticated network, capable of preserving fine details and complex structures, and achieves faster training given limited computational hardware.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Latent space interpolation, two frames with distinct 3D motion landmarks are selected, on the left surrounded with a green box is the source, on the right surrounded with a red box is the target, in between is the interpolated results for shape and appearance.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g011.tif"/>
</fig>
</sec>
<sec id="s5-4">
<title>5.4 Linear Blend Skinning Comparison</title>
<p>This section includes a comparison of the proposed Pose2Shape network against linear blend skinning (LBS) techniques demonstrating the benefits of using the proposed network. LBS is a widely used approach in real-time character animation for deforming a surface mesh according to an underlying bone structure, where every bone contains a transformation matrix that affects a group of vertices. This relation is given by a weighting attribute that weights the contribution of a bone transformation on a vertex. LBS is computationally efficient and commonly used in animation frameworks, allowing real-time character animation by manipulation of surface geometry using a low-dimensional skeletal structure. Although, it does not allow propagation of non-linear surface deformation, and it can cause artefacts on the mesh surface. To understand if the proposed Pose2Shape model is capable of learning non-linear attributes from the input data instead of only learning a linear mapping, we compare the results against LBS. For this comparison, we present two experiments; the first experiment evaluates the interpolation performance against LBS. The second experiment compares the synthesis of a mesh sequence against using LBS to animate the same motion sequence, please see supplementary material for second experiment. To compare the meshes, we use the Hausdorff distance metrics, discussed in <xref ref-type="sec" rid="s5-1">Section&#x20;5.1</xref>.</p>
<p>
<xref ref-type="fig" rid="F12">Figure&#x20;12</xref> illustrates the results for the first experiment using the Thomas (<xref ref-type="bibr" rid="B5">Boukhayma and Boyer, 2015</xref>) dataset. The top row represents the original sequence of walking motion, the source and target frames surround by green and orange boxes, respectively, represent the frames used for interpolation. The middle row shows the results of interpolating the latent vectors representative of the source and target frames. Latent vectors were generated by encoding the respective skeletons of the source and target frames. As a consequence, we can synthesise intermediate poses following <xref ref-type="disp-formula" rid="e5">Eq. (5)</xref>. The bottom row shows the LBS results for source and target frames. LBS is achieved using the animation capabilities of the SDSR framework (<xref ref-type="bibr" rid="B42">Regateiro et&#x20;al., 2018</xref>), which allows mesh manipulation through skeletal animation. Therefore, given the original skeletal motion frames, we map the source frame onto the target frame whilst generating the intermediate frames, as illustrated in the bottom&#x20;row.</p>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>The top row is a motion capture sequence which is used to synthesise the bottom rows. The bottom rows are 4D volumetric content of Roxanne. Dan and Thomas datasets. It is visible that is able to generate 4D volumetric video from skeletal motion for different subjects.</p>
</caption>
<graphic xlink:href="frvir-02-739010-g012.tif"/>
</fig>
<p>This experiments demonstrates the ability to generate a more accurate reconstruction of the original mesh compared to LBS. To support these figures, <xref ref-type="table" rid="T1">Table&#x20;1</xref>, shows quantitative evaluation for all the datasets between LBS and the proposed results.</p>
</sec>
<sec id="s5-5">
<title>5.5 Compression</title>
<p>
<xref ref-type="table" rid="T2">Table&#x20;2</xref> demonstrates the proposed approach is capable of compressing 4D volumetric video through a deep learnt representation. The latent space representation achieves up to two orders of magnitude reduction in the size of the captured 4D volumetric video depending on sequence length. The decoders have an approximate size 105&#xa0;<italic>MB</italic> with the texture encoder size constant, 94&#xa0;<italic>MB</italic>, due to the fixed texture image resolution and the mesh encoder dependent on the mesh resolution, 10&#x2013;18&#xa0;<italic>MB</italic>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The table illustrates the total amount of disk space occupied in Megabytes (MB). The original column represents 3D mesh and 2D textures of the original dataset, and the latent space and decoder columns represent the required memory to synthesise 3D meshes and 2D texture appearance.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Dataset</th>
<th align="center">Vertices</th>
<th align="center">Frames</th>
<th align="center">Original (MB)</th>
<th align="center">Latent space (MB)</th>
<th align="center">Decoder (MB)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Dan <xref ref-type="bibr" rid="B14">Casas et&#x20;al. (2014)</xref>
</td>
<td align="center">2,667</td>
<td align="center">1,447</td>
<td align="char" char=".">768.2</td>
<td align="char" char=".">2.6</td>
<td align="char" char=".">104</td>
</tr>
<tr>
<td align="left">JP <xref ref-type="bibr" rid="B46">Starck and Hilton (2007)</xref>
</td>
<td align="center">3,463</td>
<td align="center">1788</td>
<td align="char" char=".">1,272.7</td>
<td align="char" char=".">4.7</td>
<td align="char" char=".">106.8</td>
</tr>
<tr>
<td align="left">Martin <xref ref-type="bibr" rid="B32">Klaudiny and Hilton (2012)</xref>
</td>
<td align="center">2,689</td>
<td align="center">310</td>
<td align="char" char=".">479.1</td>
<td align="char" char=".">0.80</td>
<td align="char" char=".">104</td>
</tr>
<tr>
<td align="left">Roxanne <xref ref-type="bibr" rid="B46">Starck and Hilton (2007)</xref>
</td>
<td align="center">2,475</td>
<td align="center">414</td>
<td align="char" char=".">428.1</td>
<td align="char" char=".">1.1</td>
<td align="char" char=".">103.3</td>
</tr>
<tr>
<td align="left">Thomas <xref ref-type="bibr" rid="B5">Boukhayma and Boyer (2015)</xref>
</td>
<td align="center">5,002</td>
<td align="center">212</td>
<td align="char" char=".">1,186.3</td>
<td align="char" char=".">0.55</td>
<td align="char" char=".">112.4</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5-6">
<title>5.6 Performance</title>
<p>Presented results were generated using a desktop PC with an Intel Core i7-6700K CPU, 64&#xa0;GB of RAM and an Nvidia Geforce GTX 1080 GPU. Our training time is approximately 4&#xa0;days on a single GPU. The non-optimised animation framework performance achieves &#x2248;10 frames per second (fps) at full resolution. The performance bottleneck is in the Pose2Appearance network from <xref ref-type="sec" rid="s3-3">Section 3.3</xref> as a result of the high number of convolutional layer and training parameters. The generative networks from <xref ref-type="sec" rid="s3-2">Section 3.2</xref> is capable of achieving &#x2248;35 frames per second (fps). The characteristic of the Pose2Appearance network allows for multi-scale texture resolution improving rendering performance and memory usage, see supplementary material for illustration. The generator is capable of reconstructing multiple resolutions of the appearance, increasing rendering performance and decreasing memory usage, allowing the possibility to use on platforms with memory constraints.</p>
</sec>
<sec id="s5-7">
<title>5.7 Limitations</title>
<p>Primary limitation is quality of the 4D volumetric video sequences for training. The synthesis will reproduce artefacts present in the input data such as shape error or appearance misalignment. Currently this is limited by the publicly available 4D video sequences but will improve as 4D volumetric video improves. The current implementation is not optimised for texture rendering due copy operations between CPU and GPU memory, with optimisation this could achieve &#x3e;30 fps for shape and appearance synthesis. Motion capture data synthesis may create undesired artefacts on the appearance and shape if the skeletal motion is outside the space of observed 4D motions as this requires extrapolation in the latent space. Currently the network is only able to represent one character a time, an interesting extension for future work would be to encode multiple characters in a single space, or a single person wearing multiple types of clothing.</p>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion</title>
<p>The proposed Deep4D representation enables interactive animation through motion graphs to generate dynamic shape and high-quality appearance. The 4D generative network supports interpolation in the latent space to synthesise novel intermediate motions allowing smooth transitions between captured sequences. The Pose2Appearance network synthesises high resolution textures for the learnt motion space, whilst preserving details of motion and realistic details. The proposed network is capable of a compact representation of multiple 4D volumetric video sequences achieving up-to two orders of magnitude compression compared to the captured 4D volumetric video. The generative network allows mapping of skeletal motion capture data to generate novel 4D volumetric video sequences with detailed dynamic shape and appearance. The approach achieves efficient representation and real-time rendering of 4D volumetric video in a motion graph for interactive animation. This overcomes the limitations of previous approaches to animate 4D volumetric video which require high storage and computational costs. Generative network usually suffer from discontinuities in areas where there is insufficient training data. This limitation is overcome by enforcing transitions through the motion graph which does not allow for extrapolation outside the space of observed 4D volumetric video. The proposed method is able to preserve shape details, motion and appearance as shown in the evaluation. We demonstrated the integration of the proposed generative network with traditional animation frameworks, improving on interpolation between different motions, and adding more information to the similarity metrics to improve the quality of motion transitions. The animation framework is independent of the network architecture, allowing for future improvements in either of the frameworks. For instance, the training performance of the neural network can be improved by reducing the number of convolutional layers, which consequently improves the run-time appearance rendering. The animation framework can be extended to parameterised motion, allowing increased interactivity and motion control.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in the CVSSP3D (<ext-link ext-link-type="uri" xlink:href="https://cvssp.org/data/cvssp3d/">https://cvssp.org/data/cvssp3d/</ext-link>) and INRIA (<ext-link ext-link-type="uri" xlink:href="https://hal.inria.fr/hal-01348837/file/Data_EigenAppearance.zip">https://hal.inria.fr/hal-01348837/file/Data_EigenAppearance.zip</ext-link>) repositories.</p>
</sec>
<sec id="s8">
<title>Ethics Statement</title>
<p>Written informed consent was obtained from the individual(s) for the publication of any potentially identifiable images or data included in this article.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>JR is the first author and responsible for the implementation and also wrote the first draft of the manuscript. MV and AH supervised the implementation, manuscript generation and contributed to the final draft of the manuscript. All authors contributed to the manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This research was supported by the EPSRC &#x201c;Audio-Visual Media Research Platform Grant&#x201d; (EP/P022529/1), &#x201c;Polymersive: Immersive Video Production Tools for Studio and Live Events&#x2019;&#x201d; (InnovateUK 105168), and &#x201c;AI4ME: AI for Personalised Media Experiences&#x201d; UKRI EPSRC (EP/V038087/1).</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors would also like to thank Adnane Boukhayma for providing the &#x201c;Thomas&#x201d; dataset used for evaluation. The work presented was undertaken at CVSSP.</p>
</ack>
<sec id="s13">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frvir.2021.739010/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frvir.2021.739010/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Video1.MP4" id="SM2" mimetype="application/MP4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arikan</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Forsyth</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>O&#x27;Brien</surname>
<given-names>J.&#x20;F.</given-names>
</name>
<name>
<surname>O&#x2019;Brien</surname>
<given-names>J.&#x20;F.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Motion Synthesis from Annotations</article-title>. <source>ACM Trans. Graph.</source> <volume>22</volume>, <fpage>402</fpage>&#x2013;<lpage>408</lpage>. <pub-id pub-id-type="doi">10.1145/882262.882284</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bordino</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Donato</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gionis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Leonardi</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Mining Large Networks with Subgraph Counting</article-title>,&#x201d; in <conf-name>2008 Eighth IEEE International Conference on Data Mining</conf-name>, <conf-loc>Pisa, Italy</conf-loc>, <conf-date>15&#x2013;19 December 2008</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>737</fpage>&#x2013;<lpage>742</lpage>. <pub-id pub-id-type="doi">10.1109/icdm.2008.109</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Boukhayma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Boyer</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Controllable Variation Synthesis for Surface Motion Capture</article-title>,&#x201d; in <conf-name>2017 International Conference on 3D Vision (3DV)</conf-name>, <conf-loc>Qingdao, China</conf-loc>, <conf-date>10&#x2013;12 October 2017</conf-date>, (<publisher-loc>IEEE</publisher-loc>), <fpage>309</fpage>&#x2013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.1109/3DV.2017.00043</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boukhayma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Boyer</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Surface Motion Capture Animation Synthesis</article-title>. <source>IEEE Trans. Vis. Comput. Graphics</source> <volume>25</volume>, <fpage>2270</fpage>&#x2013;<lpage>2283</lpage>. <pub-id pub-id-type="doi">10.1109/tvcg.2018.2831233</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Boukhayma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Boyer</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Video Based Animation Synthesis with the Essential Graph</article-title>,&#x201d; in <conf-name>2015 International Conference on 3D Vision</conf-name>, <conf-loc>Lyon, France</conf-loc>, <conf-date>19&#x2013;22 October 2015</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>478</fpage>&#x2013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1109/3dv.2015.60</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brock</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Donahue</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Simonyan</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Large Scale GAN Training for High Fidelity Natural Image Synthesis</source>. <publisher-loc>New Orleans, LA, USA</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Budd</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Klaudiny</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Global Non-rigid Alignment of Surface Sequences</article-title>. <source>Int. J.&#x20;Comput. Vis.</source> <volume>102</volume>, <fpage>256</fpage>&#x2013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-012-0553-4</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cagniart</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Boyer</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ilic</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Free-form Mesh Tracking: A Patch-Based Approach</article-title>,&#x201d; in <conf-name>2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>San Francisco, CA</conf-loc>, <conf-date>13&#x2013;18 June 2010</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>1339</fpage>&#x2013;<lpage>1346</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2010.5539814</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hidalgo</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Simon</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>S.-E.</given-names>
</name>
<name>
<surname>Sheikh</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Openpose: Realtime Multi-Person 2d Pose Estimation Using Part Affinity fields</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>43</volume>, <fpage>172</fpage>&#x2013;<lpage>186</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2019.2929257</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carranza</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Theobalt</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Magnor</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Seidel</surname>
<given-names>H.-P.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Free-viewpoint Video of Human Actors</article-title>. <source>ACM Trans. Graph.</source> <volume>22</volume>, <fpage>569</fpage>&#x2013;<lpage>577</lpage>. <pub-id pub-id-type="doi">10.1145/882262.882309</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Casas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tejera</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Guillemaut</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>4d Parametric Motion Graphs for Interactive Animation</article-title>,&#x201d; in <conf-name>Proceedings of the ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (ACM), I3D &#x2019;12</conf-name>, <conf-loc>Costa Mesa, CA</conf-loc>, <conf-date>9&#x2013;11 March 2012</conf-date>, (<publisher-name>Association for Computing Machinery</publisher-name>), <fpage>103</fpage>&#x2013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1145/2159616.2159633</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Casas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tejera</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Guillemaut</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Interactive Animation of 4d Performance Capture</article-title>. <source>IEEE Trans. Vis. Comput. Graphics</source> <volume>19</volume>, <fpage>762</fpage>&#x2013;<lpage>773</lpage>. <pub-id pub-id-type="doi">10.1109/TVCG.2012.314</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Casas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tejera</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Guillemaut</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Parametric Control of Captured Mesh Sequences for Real-Time Animation</article-title>,&#x201d; in <conf-name>Proceedings of the 4th international conference on Motion in Games</conf-name>, <conf-loc>Edinburgh, UK</conf-loc>, <conf-date>13&#x2013;15/11/2011</conf-date> (<publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>242</fpage>&#x2013;<lpage>253</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-25090-3_21</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Casas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Volino</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Collomosse</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>4d Video Textures for Interactive Character Appearance</article-title>. <source>Comput. Graphics Forum</source> <volume>33</volume>, <fpage>371</fpage>&#x2013;<lpage>380</lpage>. <pub-id pub-id-type="doi">10.1111/cgf.12296</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<comment>[Dataset]</comment> <collab>CMU Graphics Lab</collab> (<year>2001</year>). <source>Cmu Graphics Lab Motion Capture Database</source>. <publisher-loc>Pittsburgh, PA</publisher-loc>: <publisher-name>Carnegie Mellon University</publisher-name>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Collet</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chuang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sweeney</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Gillett</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Evseev</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Calabrese</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>High-quality Streamable Free-Viewpoint Video</article-title>. <source>ACM Trans. Graph.</source> <volume>34</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1145/2766945</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Aguiar</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Stoll</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Theobalt</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ahmed</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Seidel</surname>
<given-names>H.-P.</given-names>
</name>
<name>
<surname>Thrun</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Performance Capture from Sparse Multi-View Video</article-title>. <source>ACM Trans. Graph.</source> <volume>27</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1145/1360612.1360697</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dosovitskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Springenberg</surname>
<given-names>J.&#x20;T.</given-names>
</name>
<name>
<surname>Brox</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Learning to Generate Chairs with Convolutional Neural Networks</article-title>,&#x201d; in <conf-name>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name> (<publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1538</fpage>&#x2013;<lpage>1546</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2015.7298761</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ekman</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1980</year>). <source>The Face of Man: Expressions of Universal Emotions in a New guinea Village</source>. <publisher-loc>Incorporated</publisher-loc>: <publisher-name>Garland Publishing</publisher-name>. </citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Esser</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Haux</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Milbich</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ommer</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Towards Learning a Realistic Rendering of Human Behavior</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>409</fpage>&#x2013;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-11012-3_32</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I. J.</given-names>
</name>
<name>
<surname>Pouget-Abadie</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mirza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Warde-Farley</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ozair</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <source>Generative Adversarial Networks</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>. </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hilsmann</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fechteler</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Morgenstern</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Paier</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Feldmann</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Schreer</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Going beyond Free Viewpoint: Creating Animatable Volumetric Video of Human Performances</article-title>. <source>IET Comput. Vis.</source> <volume>14</volume>, <fpage>350</fpage>&#x2013;<lpage>358</lpage>. <pub-id pub-id-type="doi">10.1049/iet-cvi.2019.0786</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Holden</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Komura</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Saito</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Phase-functioned Neural Networks for Character Control</article-title>. <source>ACM Trans. Graph.</source> <volume>36</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1145/3072959.3073663</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Budd</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Global Temporal Registration of Multiple Non-rigid Surface Sequences</article-title>,&#x201d; in <conf-name>CVPR 2011</conf-name>, <conf-loc>Colorado Springs, CO</conf-loc>, <conf-date>20&#x2013;25 June 2011</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>3473</fpage>&#x2013;<lpage>3480</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2011.5995438</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Starck</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Human Motion Synthesis from 3D Video</article-title>,&#x201d; in <conf-name>2009 IEEE Conference on Computer Vision and Pattern Recognition</conf-name> (<publisher-loc>Miami, FL, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1478</fpage>&#x2013;<lpage>1485</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206626</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tejera</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Collomosse</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Hybrid Skeletal-Surface Motion Graphs for Character Animation from 4d Performance Capture</article-title>. <source>ACM Trans. Graph.</source> <volume>34</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/2699643</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Isola</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Efros</surname>
<given-names>A. A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Image-to-Image Translation with Conditional Adversarial Networks</source>. <publisher-loc>Honolulu, Hawaii</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Alahi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fei-Fei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Perceptual Losses for Real-Time Style Transfer and Super-resolution</source>. <publisher-loc>Amsterdam, Netherlands</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>. </citation>
</ref>
<ref id="B29">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Karras</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Aila</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Laine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lehtinen</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Progressive Growing of GANs for Improved Quality, Stability, and Variation</source>. <publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Karras</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Laine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Aila</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <source>A Style-Based Generator Architecture for Generative Adversarial Networks</source>. <publisher-loc>Salt lake city, Utah</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Welling</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). <source>Auto-Encoding Variational Bayes</source>. <publisher-loc>Banff, AB, Canada</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B32">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Klaudiny</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>High-detail 3d Capture and Non-sequential Alignment of Facial Performance</article-title>,&#x201d; in <conf-name>2012 Second International Conference on 3D Imaging, Modeling, Processing, Visualization Transmission</conf-name>, <conf-loc>Zurich, Switzerland</conf-loc>, <conf-date>13&#x2013;15 October 2012</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>17</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1109/3dimpvt.2012.67</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kovar</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gleicher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pighin</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Motion Graphs</article-title>. <source>ACM Trans. Graph.</source> <volume>21</volume>, <fpage>473</fpage>&#x2013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1145/566654.566605</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="book">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Laine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Feature-based Metrics for Exploring the Latent Space of Generative Models</source>. <publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lombardi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Saragih</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Simon</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sheikh</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep Appearance Models for Face Rendering</article-title>. <source>ACM Trans. Graph.</source> <volume>37</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1145/3197517.3201401</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Georgoulis</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Van Gool</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Schiele</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Fritz</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Disentangled Person Image Generation</source>. <publisher-loc>Salt Lake City, Utah</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Miyato</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kataoka</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Koyama</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yoshida</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Spectral Normalization for Generative Adversarial Networks</source>. <publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B38">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Muller</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2007</year>). <source>Information Retrieval for Music and Motion</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>. </citation>
</ref>
<ref id="B39">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Paier</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hilsmann</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Eisert</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Neural Face Models for Example-Based Visual Speech Synthesis</article-title>,&#x201d; in <conf-name>CVMP &#x2019;20: European Conference on Visual Media Production</conf-name> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>). <pub-id pub-id-type="doi">10.1145/3429341.3429356</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Prada</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kazhdan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chuang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Collet</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hoppe</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Motion Graphs for Unstructured Textured Meshes</article-title>. <source>ACM Trans. Graph.</source> <volume>35</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/2897824.2925967</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Regateiro</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Volino</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Dynamic Surface Animation Using Generative Networks</article-title>,&#x201d; in <conf-name>International Conference on 3D Vision (3DV)</conf-name>, <conf-loc>Quebec, Canada</conf-loc>, <conf-date>16&#x2013;19 September 2019</conf-date>, (<publisher-name>IEEE</publisher-name>). <pub-id pub-id-type="doi">10.1109/3dv.2019.00049</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Regateiro</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Volino</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Hybrid Skeleton Driven Surface Registration for Temporally Consistent Volumetric Video</article-title>,&#x201d; in <conf-name>2018 International Conference on 3D Vision (3DV)</conf-name> (<publisher-loc>Verona, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>514</fpage>&#x2013;<lpage>522</lpage>. <pub-id pub-id-type="doi">10.1109/3DV.2018.00065</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sainburg</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Thielk</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Theilman</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Migliori</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gentner</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Generative Adversarial Interpolative Autoencoding: Adversarial Training on Latent Space Interpolations Encourage Convex Latent Distributions</source>. <publisher-loc>New Orleans, Louisiana</publisher-loc>: <publisher-name>International Conference on Learning Representations (ICLR)</publisher-name>. </citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sakoe</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Chiba</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1990</year>). <source>Dynamic Programming Algorithm Optimization for Spoken Word Recognition</source>. <publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>, <fpage>159</fpage>&#x2013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/b978-0-08-051584-7.50016-4</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Siarohin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sangineto</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Lathuiliere</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sebe</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Deformable GANs for Pose-Based Human Image Generation</source>. <publisher-loc>Salt Lake City, Utah</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Starck</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Surface Capture for Performance-Based Animation</article-title>. <source>IEEE Comput. Grap. Appl.</source> <volume>27</volume>, <fpage>21</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1109/MCG.2007.68</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Starck</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Video-based Character Animation</article-title>,&#x201d; in <conf-name>SCA &#x2019;05: Proceedings of the 2005 ACM SIGGRAPH/Eurographics Symposium on Computer Animation</conf-name> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>49</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1145/1073368.1073375</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tan</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lai</surname>
<given-names>Y.-K.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Variational Autoencoders for Deforming 3d Mesh Models</article-title>,&#x201d; in <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Salt Lake City, Utah</conf-loc>, <conf-date>18&#x2013;22 June 2018</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>5841</fpage>&#x2013;<lpage>5850</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00612</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tanco</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2000</year>). &#x201c;<article-title>Realistic Synthesis of Novel Human Movements from a Database of Motion Capture Examples</article-title>,&#x201d; in <conf-name>Proceedings Workshop on Human Motion</conf-name>, <conf-loc>Austin, Texas</conf-loc>, <conf-date>7&#x2013;8 Decembre 2000</conf-date>, (<publisher-name>IEEE</publisher-name>), <fpage>137</fpage>&#x2013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1109/HUMO.2000.897383</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tejera</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hilton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Learning Part-Based Models for Animation from Surface Motion Capture</article-title>,&#x201d; in <conf-name>2013 International Conference on 3D Vision</conf-name> (<publisher-loc>Seattle, WA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>159</fpage>&#x2013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1109/3DV.2013.29</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tulyakov</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M.-Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Kautz</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <source>MoCoGAN: Decomposing Motion and Content for Video Generation</source>. <publisher-loc>Salt Lake City, Utah</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
<ref id="B52">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ulyanov</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lebedev</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Vedaldi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lempitsky</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Texture Networks: Feed-Forward Synthesis of Textures and Stylized Images</source>. <publisher-loc>New York, USA</publisher-loc>: <publisher-name>ACM</publisher-name>. </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vlasic</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Baran</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Matusik</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Popovic</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Articulated Mesh Animation from Multi-View Silhouettes</article-title>. <source>ACM Trans. Graph.</source> <volume>27</volume>, <fpage>1</fpage>&#x2013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1145/1360612.1360696</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vondrick</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pirsiavash</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Torralba</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Generating Videos with Scene Dynamics</source>. <publisher-loc>Barcelona, Spain</publisher-loc>: <publisher-name>ACM</publisher-name>. </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bodenheimer</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Synthesis and Evaluation of Linear Motion Transitions</article-title>. <source>ACM Trans. Graph.</source> <volume>27</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1145/1330511.1330512</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Witkin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Popovic</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1995</year>). &#x201c;<article-title>Motion Warping</article-title>,&#x201d; in <conf-name>SIGGRAPH &#x2019;95: Proceedings of the 22Nd Annual Conference on Computer Graphics and Interactive Techniques (ACM)</conf-name>, <conf-loc>Los Angeles, CA</conf-loc>, <conf-date>6&#x2013;11 August 1995</conf-date>, (<publisher-name>Association for Computing Machinery</publisher-name>), <fpage>105</fpage>&#x2013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1145/218380.218422</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Krahenbuhl</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Shechtman</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Efros</surname>
<given-names>A. A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Generative Visual Manipulation on the Natural Image Manifold</source>. <publisher-loc>Amsterdam, The Netherlands</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B58">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Isola</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Efros</surname>
<given-names>A. A.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial Networks</source>. <publisher-loc>Venice, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>. </citation>
</ref>
</ref-list>
</back>
</article>