<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2022.890695</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Dynamic Cloth Manipulation Considering Variable Stiffness and Material Change Using Deep Predictive Model With Parametric Bias</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Kawaharazuka</surname> <given-names>Kento</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1697455/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Miki</surname> <given-names>Akihiro</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Bando</surname> <given-names>Masahiro</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Okada</surname> <given-names>Kei</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1342108/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Inaba</surname> <given-names>Masayuki</given-names></name>
</contrib>
</contrib-group>
<aff><institution>JSK Robotics Laboratory, Department of Mechano-Informatics, Graduate School of Information Science and Technology, The University of Tokyo</institution>, <addr-line>Tokyo</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Solvi Arnold, Shinshu University, Japan</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Katsu Yamane, Robert Bosch, United States; Jihong Zhu, Delft University of Technology, Netherlands; Yulei Qiu, Delft University of Technology, Netherlands, in collaboration with reviewer JZ</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Kento Kawaharazuka <email>kawaharazuka&#x00040;jsk.imi.i.u-tokyo.ac.jp</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>890695</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>03</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Kawaharazuka, Miki, Bando, Okada and Inaba.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Kawaharazuka, Miki, Bando, Okada and Inaba</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Dynamic manipulation of flexible objects such as fabric, which is difficult to modelize, is one of the major challenges in robotics. With the development of deep learning, we are beginning to see results in simulations and some actual robots, but there are still many problems that have not yet been tackled. Humans can move their arms at high speed using their flexible bodies skillfully, and even when the material to be manipulated changes, they can manipulate the material after moving it several times and understanding its characteristics. Therefore, in this research, we focus on the following two points: (1) body control using a variable stiffness mechanism for more dynamic manipulation, and (2) response to changes in the material of the manipulated object using parametric bias. By incorporating these two approaches into a deep predictive model, we show through simulation and actual robot experiments that Musashi-W, a musculoskeletal humanoid with a variable stiffness mechanism, can dynamically manipulate cloth while detecting changes in the physical properties of the manipulated object.</p></abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>predictive model</kwd>
<kwd>cloth manipulation</kwd>
<kwd>variable stiffness</kwd>
<kwd>parametric bias</kwd>
<kwd>musculoskeletal humanoid</kwd>
</kwd-group>
<contract-sponsor id="cn001">Japan Science and Technology Agency<named-content content-type="fundref-id">10.13039/501100002241</named-content></contract-sponsor>
<contract-sponsor id="cn002">Japan Society for the Promotion of Science<named-content content-type="fundref-id">10.13039/501100001691</named-content></contract-sponsor>
<counts>
<fig-count count="15"/>
<table-count count="0"/>
<equation-count count="8"/>
<ref-count count="27"/>
<page-count count="16"/>
<word-count count="8031"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Manipulation of flexible objects such as fabric, which is difficult to modelize, is one of the major challenges in robotics. There are two main types of manipulation of flexible objects: static manipulation and dynamic manipulation. For each type, various model-based and learning-based methods have been developed. Static manipulation has been tackled for a long time, and various model-based methods exist (Inaba and Inoue, <xref ref-type="bibr" rid="B9">1987</xref>; Saha and Isto, <xref ref-type="bibr" rid="B21">2007</xref>; Elbrechter et al., <xref ref-type="bibr" rid="B4">2012</xref>). Learning-based methods have also been actively pursued in recent years and have been successfully applied to actual robots (Lee et al., <xref ref-type="bibr" rid="B17">2015</xref>; Tanaka et al., <xref ref-type="bibr" rid="B23">2018</xref>). On the other hand, there are not so many examples of dynamic manipulation, which is more difficult compared to static manipulation. As for model-based methods, the study of Yamakawa et al. on cloth folding and knotting is well known (Yamakawa et al., <xref ref-type="bibr" rid="B26">2010</xref>, <xref ref-type="bibr" rid="B27">2011</xref>). Learning-based methods include Kawaharazuka et al. (<xref ref-type="bibr" rid="B15">2019b</xref>), which uses deep predictive models, and Jangir et al. (<xref ref-type="bibr" rid="B11">2020</xref>), which uses reinforcement learning. Also, there are many examples where only static manipulation is performed even though dynamic models are used (Ebert et al., <xref ref-type="bibr" rid="B3">2018</xref>; Hoque et al., <xref ref-type="bibr" rid="B8">2020</xref>). There is an example where the primitive of the dynamic motion is generated manually and only the grasping point is trained (Ha and Song, <xref ref-type="bibr" rid="B5">2021</xref>).</p>
<p>In this study, we handle dynamic manipulation such as the spreading of a bed sheet or a picnic sheet. For the dynamic manipulation, there are several points lacking in human-like adaptive dynamic cloth manipulation that have not been addressed in the introduced previous studies. The previous studies have not been able to utilize the flexible bodies of robots to move their arms at high speed as humans do, and to perform manipulation based on an understanding of the characteristics of the material from a few trials, even if the material to be manipulated changes. Therefore, we focus on two points: (1) body control for more dynamic manipulation, and (2) adaptation to changes in the material of the manipulated object (<xref ref-type="fig" rid="F1">Figure 1</xref>). In (1), we aim at cloth manipulation by a robot with a variable stiffness mechanism, which is flexible like a human and can manipulate its flexibility at will. In this study, we perform manipulation by Musashi-W (MusashiDarm Kawaharazuka et al., <xref ref-type="bibr" rid="B14">2019a</xref> with wheeled base), a musculoskeletal humanoid with variable stiffness mechanism using redundant muscles and nonlinear elastic elements. We consider how the dynamic cloth manipulation is changed by using the additional stiffness value as the control command of the body. There have been examples of dynamic pitching behavior by changing hardware stiffness in the framework of model-based optimal control (Braun et al., <xref ref-type="bibr" rid="B2">2012</xref>), and several static tasks by changing software impedance using reinforcement learning (Mart&#x000ED;n-Mart&#x000ED;n et al., <xref ref-type="bibr" rid="B18">2019</xref>). On the other hand, there is no example focusing on changes in the hardware stiffness in a dynamic handling task of difficult-to-modelize objects that require learning-based control, such as in this study. In (2), we aim at adapting to changes in the material of the manipulated object using parametric bias (Tani, <xref ref-type="bibr" rid="B24">2002</xref>). The parametric bias is an additional bias term in neural networks, which has been mainly used for imitation learning to extract multiple attractor dynamics from various motion data (Ogata et al., <xref ref-type="bibr" rid="B20">2005</xref>; Kawaharazuka et al., <xref ref-type="bibr" rid="B12">2021</xref>). In this study, we use it to embed information about the material and physical properties of the cloth. When the robot holds a new cloth, the physical properties can be identified by manipulating the cloth several times, and dynamic cloth manipulation can be accurately performed. Parametric bias can be applied not only to simulations but also to actual robots where it is difficult to identify the parameters of cloth materials since it implicitly self-organizes from differences in dynamics of each data. In this study, we construct a deep predictive model that incorporates (1) and (2) and demonstrate the effectiveness by performing a more human-like dynamic cloth manipulation.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Dynamic cloth manipulation by the musculoskeletal humanoid Musashi-W considering body control with variable stiffness and adaptation to material change.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0001.tif"/>
</fig>
<p>The contribution of this study is as follows.</p>
<list list-type="bullet">
<list-item><p>Examination of the effect of adding the body stiffness value to the control command.</p></list-item>
<list-item><p>Adaptation to changes in cloth material using parametric bias.</p></list-item>
<list-item><p>Human-like adaptive dynamic cloth manipulation by learning a deep predictive model considering variable stiffness and material change.</p></list-item>
</list>
<p>In Section 2, we describe the structure of the deep predictive model, the details of variable stiffness control, the estimation of material parameters, and the body control for dynamic cloth manipulation. In Section 3, we discuss online learning of material parameters and changes in dynamic cloth manipulation due to variable stiffness and cloth material changes in simulation and the actual robot. In addition, a table setting experiment including dynamic cloth manipulation is conducted. The results of the experiments are discussed in Section 4, and the conclusion is given in Section 5.</p>
</sec>
<sec id="s2">
<title>2. Dynamic Cloth Manipulation Considering Variable Stiffness and Material Change</title>
<p>We call the network used in this study Dynamic Predictive Model with Parametric Bias (DPMPB). The entire system is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The overview of our system: deep predictive model with parametric bias (DPMPB), controller using DPMPB for dynamic cloth manipulation, variable stiffness controller for musculoskeletal humanoids, data collector for DPMPB, and online updater of parametric bias (PB).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0002.tif"/>
</fig>
<sec>
<title>2.1. Network Structure of DPMPB</title>
<p>Dynamic Predictive Model with Parametric Bias can be expressed by the following equation.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>h</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>p</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>p</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>t</italic> is the current time step, <bold><italic>s</italic></bold> is the state of the manipulated object and robot, <bold><italic>u</italic></bold> is the control command to the robot body, <bold><italic>p</italic></bold> is the parametric bias (PB), and <bold><italic>h</italic></bold><sub><italic>dpmpb</italic></sub> is the function representing the time series change in the state of the manipulated object and robot due to the control command. We use the state of the cloth and the state of the robot for <bold><italic>s</italic></bold> in this study of cloth manipulation. For the cloth state, we use <bold><italic>z</italic></bold><sub><italic>t</italic></sub>, which is the compressed value of the current image <italic>I</italic><sub><italic>t</italic></sub> by using AutoEncoder (Hinton and Salakhutdinov, <xref ref-type="bibr" rid="B6">2006</xref>). Although the robot state should be different for each robot, we use <bold><italic>f</italic></bold><sub><italic>t</italic></sub> and <bold><italic>l</italic></bold><sub><italic>t</italic></sub> for the musculoskeletal humanoid used in this study (<bold><italic>f</italic></bold><sub><italic>t</italic></sub> and <bold><italic>l</italic></bold><sub><italic>t</italic></sub> represent the current muscle tension and muscle length). Thus, <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>l</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Also, we set <inline-formula><mml:math id="M3"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>k</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> (<bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> represents the target joint angle and <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> represents the target body stiffness value). The details of the target joint angle and body stiffness are described in Section 2.2. The parametric bias is a value that can embed implicit differences in dynamics, which are common for the same material and different for different materials. By collecting data using various cloth materials, information on the dynamics of the material of the manipulated object is embedded in <bold><italic>p</italic></bold>. Note that each sensor value is used as input to the network after normalization using all the obtained data.</p>
<p>In this study, the DPMPB consists of 10 layers: 4 fully-connected layers, 2 LSTM layers (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B7">1997</xref>), and 4 fully-connected layers, in order. The number of units is set to {<italic>N</italic><sub><italic>u</italic></sub> &#x0002B; <italic>N</italic><sub><italic>s</italic></sub> &#x0002B; <italic>N</italic><sub><italic>p</italic></sub>, 300, 100, 30, 30 (number of units in LSTM), 30 (number of units in LSTM), 30, 100, 300, <italic>N</italic><sub><italic>s</italic></sub>} (where <italic>N</italic><sub>{<italic>u, s, p</italic>}</sub> is the number of dimensions of {<bold><italic>u</italic></bold>, <bold><italic>s</italic></bold>, <bold><italic>p</italic></bold>}). The activation function is hyperbolic tangent and the update rule is Adam (Kingma and Ba, <xref ref-type="bibr" rid="B16">2015</xref>). Regarding the AutoEncoder, the image is compressed by applying convolutional layers with kernel size 3 and stride 2 five times to a 128 &#x000D7; 96 binary image (the cloth part is extracted in color), then reducing the dimensionality to 256 and 3 units, in order, using the fully-connected layers, and finally restoring the image by the fully-connected layers and the deconvolutional layers. For all layers exempting the last layer, batch normalization (Ioffe and Szegedy, <xref ref-type="bibr" rid="B10">2015</xref>) is applied, and the activation function is ReLU (Nair and Hinton, <xref ref-type="bibr" rid="B19">2010</xref>). The dimension of <bold><italic>p</italic></bold> is set to 2 in this study, which should be sufficiently smaller than the number of cloth materials used for training. The execution period of Equation (1) is set to 5 Hz.</p>
</sec>
<sec>
<title>2.2. Variable Stiffness Control for Musculoskeletal Humanoids</title>
<p>Regarding the musculoskeletal humanoid with variable stiffness mechanism used in this study, we will describe the details of <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> and <bold><italic>k</italic></bold><sup><italic>ref</italic></sup>. In this study, we use the following relationship of joint angle <bold><italic>&#x003B8;</italic></bold>, muscle tension <bold><italic>f</italic></bold>, and muscle length <bold><italic>l</italic></bold> (Kawaharazuka et al., <xref ref-type="bibr" rid="B13">2018</xref>).</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold-italic"><mml:mi>l</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>h</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><bold><italic>h</italic></bold><sub><italic>bodyimage</italic></sub> is trained using the actual robot sensor information. When using this trained network for control, the target joint angle <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> and the target muscle tension <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> are determined and the corresponding target muscle length <inline-formula><mml:math id="M5"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>l</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>h</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is calculated. However, since this value is the muscle length to be measured, not the target muscle length, <bold><italic>l</italic></bold><sup><italic>send</italic></sup>, which takes into account the muscle elongation in muscle stiffness control (Shirai et al., <xref ref-type="bibr" rid="B22">2011</xref>), is sent to the actual robot (Kawaharazuka et al., <xref ref-type="bibr" rid="B13">2018</xref>) in practice.</p>
<p>Here, we consider how <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> is given. Normally, <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> should be the value required to realize <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup>. On the other hand, in Kawaharazuka et al. (<xref ref-type="bibr" rid="B13">2018</xref>), a constant value <italic>f</italic><sup><italic>const</italic></sup> is first given to <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> for all muscles to achieve <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> to a certain degree, and then the current muscle tension <bold><italic>f</italic></bold> is given as <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> to achieve <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> more accurately. The body stiffness can be changed to some extent depending on the value of <italic>f</italic><sup><italic>const</italic></sup> sent at the beginning. In addition, since the relationship between <bold><italic>s</italic></bold> and <bold><italic>u</italic></bold> is eventually acquired by learning in this study, it is not necessary to realize <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> precisely. Therefore, we use <italic>f</italic><sup><italic>const</italic></sup> as the body stiffness value <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> included in <bold><italic>u</italic></bold> and operate the robot with <inline-formula><mml:math id="M6"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>l</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>h</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. By acquiring training data while changing <italic>f</italic><sup><italic>const</italic></sup>, we can perform dynamic cloth manipulation considering the body stiffness value.</p>
<p>In the following, we show the results of experiments to see how much the operational stiffness of the arm changes with <italic>f</italic><sup><italic>const</italic></sup>. <xref ref-type="fig" rid="F3">Figure 3</xref> shows the operational stiffness ellipsoid of the left arm in the sagittal plane when <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> is set as the state with the elbow of the left arm of Musashi-W bent at 90 degrees and <italic>f</italic><sup><italic>const</italic></sup> is changed to {10, 30, 50, 70} [N]. The graph shows the displacement of the hand when a force of 1 N is applied from all directions. It can be seen that the size of the stiffness ellipsoid changes greatly with the change in <italic>f</italic><sup><italic>const</italic></sup>. Therefore, we use <italic>f</italic><sup><italic>const</italic></sup> as the body stiffness value <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> in this study.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Operational stiffness ellipsoid when adding force of 1 N while changing <italic>f</italic><sup><italic>const</italic></sup> &#x0003D; {10, 30, 50, 70} [N].</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0003.tif"/>
</fig>
</sec>
<sec>
<title>2.3. Training of DPMPB</title>
<p>By setting <bold><italic>u</italic></bold> randomly, or operating the robot with GUI or VR device, we collect the data of <bold><italic>s</italic></bold> and <bold><italic>u</italic></bold>. For a single, coherent manipulation trial <italic>k</italic>, performed with the same cloth, the data <italic>D</italic><sub><italic>k</italic></sub> &#x0003D; {(<bold><italic>s</italic></bold><sub>1</sub>, <bold><italic>u</italic></bold><sub>1</sub>), (<bold><italic>s</italic></bold><sub>2</sub>, <bold><italic>u</italic></bold><sub>2</sub>), &#x022EF;&#x02009;, (<bold><italic>s</italic></bold><sub><italic>T</italic><sub><italic>k</italic></sub></sub>, <bold><italic>u</italic></bold><sub><italic>T</italic><sub><italic>k</italic></sub></sub>)} (1 &#x02264; <italic>k</italic> &#x02264; <italic>K</italic>, where <italic>K</italic> is the total number of trials and <italic>T</italic><sub><italic>k</italic></sub> is the number of time steps for the trial <italic>k</italic>). Then, we obtain the data <italic>D</italic><sub><italic>train</italic></sub> &#x0003D; {(<italic>D</italic><sub>1</sub>, <bold><italic>p</italic></bold><sub>1</sub>), (<italic>D</italic><sub>2</sub>, <bold><italic>p</italic></bold><sub>2</sub>), &#x022EF;&#x02009;, (<italic>D</italic><sub><italic>K</italic></sub>, <bold><italic>p</italic></bold><sub><italic>K</italic></sub>)} for training. <bold><italic>p</italic></bold><sub><italic>k</italic></sub> is the parametric bias for the trial <italic>k</italic>, a variable that has a common value during that one trial and a different value for different trials. We use this data <italic>D</italic><sub><italic>train</italic></sub> to train the DPMPB. In the usual training, only the network weight <italic>W</italic> is updated, but in this study, <italic>W</italic> and <bold><italic>p</italic></bold><sub><italic>k</italic></sub> are updated simultaneously as below,</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>W</mml:mi><mml:mo>&#x02190;</mml:mo><mml:mi>W</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>p</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>p</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>p</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>L</italic> is the loss function (mean squared error, in this study) and &#x003B2; is the learning rate. <bold><italic>p</italic></bold><sub><italic>k</italic></sub> will be embedded with the difference in dynamics in each trial, i.e., the dynamics of the manipulated cloth in this study. <bold><italic>p</italic></bold><sub><italic>k</italic></sub> is trained with an initial value of <bold>0</bold>, and it is not necessary to directly provide parameters related to the dynamics of the cloth. Therefore, parametric bias can be applied to cases where the dynamics parameters are not known, as in the handling of actual cloth.</p>
</sec>
<sec>
<title>2.4. Online Estimation of Cloth Material</title>
<p>When manipulating a new cloth, if we do not know its correct dynamics, we will not be able to perform the manipulation correctly. Therefore, we need to obtain data on how the shape of the cloth changes when the cloth is manipulated, and estimate the dynamics of the cloth, i.e., <bold><italic>p</italic></bold> in this study, based on this data. First, we obtain the data <italic>D</italic><sub><italic>new</italic></sub> (<italic>D</italic><sub><italic>train</italic></sub> for one trial with <italic>K</italic> &#x0003D; 1) as in Section 2.3 when the cloth is manipulated by random commands, GUI, or the controller described later (Section 2.5). Using this data, we update only <bold><italic>p</italic></bold> while <italic>W</italic> is fixed, unlike the training process of Section 2.3. We do not execute Equation (3) but update only <bold><italic>p</italic></bold> by Equation (4) using the current <bold><italic>p</italic></bold> as the initial value. Since <bold><italic>p</italic></bold> is of a very low dimension compared to <italic>W</italic>, the overfitting problem can be prevented. In other words, by updating only the terms related to the cloth dynamics, we make the dynamics of DPMPB consistent with <italic>D</italic><sub><italic>new</italic></sub>. In this case, the update rule is MomentumSGD.</p>
</sec>
<sec>
<title>2.5. Dynamic Cloth Manipulation Using DPMPB</title>
<p>We describe the dynamic cloth manipulation control using DPMPB. First, we obtain the target image <italic>I</italic><sup><italic>ref</italic></sup>, which is the target state of the cloth, and compress it to <bold><italic>z</italic></bold><sup><italic>ref</italic></sup> by AutoEncoder. <bold><italic>f</italic></bold><sup><italic>ref</italic></sup> and <bold><italic>l</italic></bold><sup><italic>ref</italic></sup> are set to <bold>0</bold>, and <bold><italic>s</italic></bold><sup><italic>ref</italic></sup> is generated by combining them. Next, we determine the number of expansions of DPMPB <inline-formula><mml:math id="M9"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and set the initial value <inline-formula><mml:math id="M10"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> as <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which is <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> steps of <bold><italic>u</italic></bold> to be optimized (the abbreviation of <inline-formula><mml:math id="M13"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>). We calculate the loss function as follows, and update <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from this value while <italic>W</italic> and <bold><italic>p</italic></bold> are fixed.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(6)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold-italic"><mml:mi>g</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02190;</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:mfrac><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>g</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mo>||</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>g</mml:mi></mml:mstyle><mml:msub><mml:mrow><mml:mo>||</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>h</italic><sub><italic>loss</italic></sub> is the loss function (to be explained later), <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is a time series of the target cloth states with <bold><italic>s</italic></bold><sup><italic>ref</italic></sup> arranged in <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> steps, <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which is the time series of <bold><italic>s</italic></bold> predicted when <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is applied to the current state <bold><italic>s</italic></bold><sub><italic>t</italic></sub>, ||&#x000B7;||<sub>2</sub> is L2 norm, and &#x003B3; is the learning rate. Equation (6) can be computed by the backpropagation method of neural networks, and Equation (7) represents the gradient descent method. In other words, the time series control command is updated to make the predicted time series of <bold><italic>s</italic></bold> closer to the target value. Note that <italic>W</italic> is obtained in Section 2.3, and <bold><italic>p</italic></bold> is obtained in Section 2.4. Here, &#x003B3; can be constant, but in this study, we update <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> using multiple values of &#x003B3; for faster convergence. We use the value obtained by dividing [0, &#x003B3;<sub><italic>max</italic></sub>] into <inline-formula><mml:math id="M24"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> equal logarithmic intervals (&#x003B3;<sub><italic>max</italic></sub> is the maximum value of &#x003B3;) as &#x003B3;, update <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> using each &#x003B3;, and adopt <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> with the smallest <italic>L</italic> among them, repeating the process <inline-formula><mml:math id="M27"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> times. An appropriate &#x003B3; is selected for each iteration and Equation (7) is performed. Also, for the initial value <inline-formula><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, we use the value optimized in the previous step <inline-formula><mml:math id="M30"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (the abbreviation of <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>) by shifting it one step to the left and duplicating the last term, <inline-formula><mml:math id="M32"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. This allows us to achieve faster convergence, taking into account the previous optimization results. The final obtained target value <inline-formula><mml:math id="M33"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> at the current time step of <inline-formula><mml:math id="M34"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is sent to the actual robot.</p>
<p>Here, we consider dynamic cloth manipulation tasks such as spreading out a bed sheet or a picnic sheet in the air. Let <bold><italic>s</italic></bold><sup><italic>ref</italic></sup> be the target value that contains the image <bold><italic>z</italic></bold><sup><italic>ref</italic></sup> of the unfolded state of the cloth, and we consider trying to achieve the target value. In this case, if the cloth does not spread well at the first attempt, it would be not possible to make minor adjustments, and it is necessary to generate a series of motions to try to spread the cloth from the beginning. In other words, even if the loss is lowered to some extent by the first attempt, it is necessary to go through a state where the loss is increased by large motions in order to lower the loss further. Therefore, if we just define <inline-formula><mml:math id="M35"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> as a vector of the same <bold><italic>s</italic></bold><sup><italic>ref</italic></sup>, we end up with <inline-formula><mml:math id="M36"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>u</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> that hardly moves after the first attempt. The same problem occurs when <bold><italic>s</italic></bold><sup><italic>ref</italic></sup> is defined as the state of spreading a cloth in the air since that state is not always realized. Thus, it is necessary to realize the target state by periodically spreading the cloth over and over again. Based on this feature, in this study, we determine the number of periodic steps of the periodic motion, <inline-formula><mml:math id="M37"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and change <italic>h</italic><sub><italic>loss</italic></sub> for each period. In this study, <italic>h</italic><sub><italic>loss</italic></sub> is expressed as follows.</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M38"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>||</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>m</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02297;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo>||</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>||</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mo>||</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M39"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>f</mml:mi></mml:mstyle></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the value of {<bold><italic>z</italic></bold>, <bold><italic>f</italic></bold>} extracted from <inline-formula><mml:math id="M40"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>s</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <italic>w</italic><sub><italic>loss</italic></sub> is the weight constant. <bold><italic>m</italic></bold><sub><italic>t</italic></sub> (<inline-formula><mml:math id="M41"><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula>) is a vector with 1 occurring at every <inline-formula><mml:math id="M42"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> step and 0, otherwise. It is shifted to the left at each time step, and 0 or 1 is inserted from the right depending on <inline-formula><mml:math id="M43"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (e.g., if <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula>, then (0, 1, 0, 0) &#x02192; (1, 0, 0, 0) &#x02192; (0, 0, 0, 1) in this order). This makes it possible to bring the cloth state <bold><italic>z</italic></bold> closer to the target value every <inline-formula><mml:math id="M45"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> step and enables dynamic cloth manipulation that handles periodic states which can only be realized momentarily. The second term on the right-hand side of Equation (8) is a term to minimize muscle tension as much as possible.</p>
<p>In this study, we set <inline-formula><mml:math id="M46"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M47"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>30</mml:mn></mml:math></inline-formula>, &#x003B3;<sub><italic>max</italic></sub> &#x0003D; 1.0, <inline-formula><mml:math id="M48"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>, <italic>w</italic><sub><italic>loss</italic></sub> &#x0003D; 0.001, and <inline-formula><mml:math id="M49"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Experiments</title>
<sec>
<title>3.1. Experimental Setup</title>
<p>The robot and the types of cloth used in this study are shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. First, we conduct a cloth manipulation experiment using a simple 2-DOF robot in simulation with Mujoco (Todorov et al., <xref ref-type="bibr" rid="B25">2012</xref>). Here, the cloth is modeled as a collection of 3 &#x000D7;6 point masses, and the damping <italic>C</italic><sub><italic>damp</italic></sub> between those point masses and the weight <italic>C</italic><sub><italic>mass</italic></sub> of the entire cloth can be changed. Next, we use the actual robot Musashi-W, which is a musculoskeletal dual arm robot MusashiDarm (Kawaharazuka et al., <xref ref-type="bibr" rid="B14">2019a</xref>) with a mechanum wheeled base and a z-axis slider. The head is equipped with Astra S camera (Orbbec 3D Technology International, Inc.). In this study, we use two kinds of cloth: soft type (5 mm thick) and hard type (2 mm thick) of polyethylene foam (TAKASHIMA color sheet, WAKI Factory, Inc.). Polyethylene foam was used instead of actual cloth in this study due to the joint speed limitation of Musashi-W. For the soft type, we prepare one or two sheets, and for the hard type, we prepare one, two, or three sheets (denoted as soft-1, soft-2, hard-1, hard-2, hard-3).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Experimental setup: the simulated simple robot and the musculoskeletal humanoid Musashi-W used in this study and cloth materials of soft and hard type polyethylene foam.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0004.tif"/>
</fig>
<p>We move the arms of the simulated robot and Musashi-W only in the sagittal plane, with two degrees of freedom in the shoulder and elbow pitch axes. Although the control command of the simulated robot is two dimensional without hardware body stiffness, the control command of Musashi-W is three dimensional including the stiffness (hardware stiffness was difficult to reproduce in simulation). The image is compressed by AutoEncoder after color extraction, binarization, closing opening, and resizing (the front part of the cloth is red and the back part is a different color such as blue or black).</p>
</sec>
<sec>
<title>3.2. Simulation Experiment</title>
<p>The cloth characteristics are changed to <italic>C</italic><sub><italic>damp</italic></sub> &#x0003D; {0.03, 0.05, 0.07} and <italic>C</italic><sub><italic>mass</italic></sub> &#x0003D; {0.05, 0.10, 0.15} while randomly commanding <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> (Random) for about 50 s, and while human operation by GUI (joint angle is given by GUI) for about 50 s. Using a total of 4,500 data points, we trained DPMPB with the number of LSTM expansions set to 30, the number of batches being 300, and the number of epochs being 300 (these parameters are empirically set). Principle Component Analysis (PCA) is applied to the parametric bias for each cloth obtained in the training, and its arrangement in the two-dimensional plane is shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. Note that even if <bold><italic>p</italic></bold> is two-dimensional, the axes of the principal components, PC1 and PC2, become more defined by applying PCA. It can be seen that the PBs are neatly arranged mainly on the axis of PC1 according to the magnitude of <italic>C</italic><sub><italic>damp</italic></sub>. On the other hand, for the axis of PC2, some PBs are not arranged according to the magnitude of <italic>C</italic><sub><italic>mass</italic></sub>, and we can see that the dynamics change due to <italic>C</italic><sub><italic>mass</italic></sub> is complex. Also, the contribution ratio in PCA is 0.65 for PC1 and 0.35 for PC2, indicating that the influence of <italic>C</italic><sub><italic>mass</italic></sub> on the dynamics of the cloth is small.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Simulation experiment: the trained parametric bias when setting <italic>C</italic><sub><italic>damp</italic></sub> &#x0003D; {0.03, 0.05, 0.07} and <italic>C</italic><sub><italic>mass</italic></sub> &#x0003D; {0.05, 0.10, 0.15}, the trajectory of online updated parametric bias when setting (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; {(0.05, 0.05), (0.07, 0.10), (0.03, 0.15)}, and the trajectory of online updated parametric bias when setting (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; (0.03, 0.10) for the integrated experiment.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0005.tif"/>
</fig>
<p>Here, with each cloth of (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; {(0.05, 0.05), (0.07, 0.10), (0.03, 0.15)}, the Random motion and the online material estimation in Section 2.4 are executed for about 40 s. The trajectory of the parametric bias (traj) is shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. Note that the initial value of PB is <bold>0</bold> (the origin of PB is not necessarily at the origin of the figure, since PCA is applied). It can be seen that the current PB gradually approaches the respective PB values obtained during training in the dynamics of the current cloth. In particular, the accuracy of the online estimation is high for the axis of <italic>C</italic><sub><italic>damp</italic></sub>. On the other hand, for the axis of <italic>C</italic><sub><italic>mass</italic></sub>, though the material can be correctly recognized to some extent, the accuracy is less than for the axis of <italic>C</italic><sub><italic>damp</italic></sub>.</p>
<p>Next, we performed the control in Section 2.5 by setting the cloth characteristics to (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; {(0.03, 0.05), (0.07, 0.15)} and by setting the current PB value to the PB obtained during the training of the object to be manipulated. Two images, Target-1 and Target-2 of <xref ref-type="fig" rid="F6">Figure 6</xref>, are given as the target images of the cloth. We denote the case of control in this study as Control and the random trial as Random. For each of the two types of cloth, the Control and Random trials are performed for 50 s each, using the two target images. The percentage (y-axis) of <inline-formula><mml:math id="M52"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> lower than a certain threshold (x-axis) is shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. <xref ref-type="fig" rid="F6">Figure 6</xref> indicates that the more the graph expands in the upper-left direction (the larger the y-axis value becomes when the x-axis value is low), the more the target image is realized. Here, in order to verify the reproducibility, all experiments including learning and control are conducted five times, and the mean and variance of these trials are shown in the graphs. Note that the mean of <inline-formula><mml:math id="M53"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> in the initial condition shown in the left figure of <xref ref-type="fig" rid="F4">Figure 4</xref> is 1.55 for Target-1 and 1.36 for Target-2. In all cases, Control outperforms Random and succeeds in realizing the target image accurately. Since Target-1 is more difficult than Target-2, especially when (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; (0.03, 0.05), the difference is clearly shown. Since the larger <italic>C</italic><sub><italic>damp</italic></sub> causes slower deformation of the cloth and makes it easier to lift up, the target image is relatively correctly realized when is set (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) to (0.07, 0.15), compared to when it is set to (0.03, 0.05). The reproducibility of performance with respect to learning and control is also high, especially in areas with a low threshold (x-axis), where the variance is small.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Simulation experiment: the rate (y-axis) of <inline-formula><mml:math id="M50"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo></mml:math></inline-formula>threshold (x-axis) when conducting experiments of dynamic manipulation of two cloths Target-1 and Target-2 when setting (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; {(0.03, 0.05), (0.07, 0.15)} regarding Control and Random.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0006.tif"/>
</fig>
<p>Next, assuming that the current cloth properties are (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; (0.03, 0.10) and the current PB values are (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; (0.07, 0.15) obtained during training, we simultaneously perform the online material estimation in Section 2.4 and control in Section 2.5. Note that the target image is set as Target-1 of <xref ref-type="fig" rid="F6">Figure 6</xref>. The transition of PB (integrated-traj) is shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, and the transition of <inline-formula><mml:math id="M55"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. It can be seen that the current PB gradually approaches the value obtained during training with (<italic>C</italic><sub><italic>damp</italic></sub>, <italic>C</italic><sub><italic>mass</italic></sub>) &#x0003D; (0.03, 0.10). Also, as PB becomes more accurate, <inline-formula><mml:math id="M56"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> periodically shows smaller values. In other words, the online material estimation and the learning control can be performed together, and the control becomes more accurate as PB becomes more accurate.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Simulation experiment: the transition of <inline-formula><mml:math id="M51"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> for the integrated experiment of online estimation and dynamic cloth manipulation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0007.tif"/>
</fig>
<p>Finally, we show that <inline-formula><mml:math id="M57"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is close to the actual distance between the raw images since the control and its evaluation are basically performed with the error of the latent variable <bold><italic>z</italic></bold> in this study. We use Symmetric Chamfer Distance (Borgefors, <xref ref-type="bibr" rid="B1">1988</xref>), which represents the similarity between binary images, to calculate the distance between the current raw image <italic>I</italic> and its target value <italic>I</italic><sup><italic>ref</italic></sup>. The relationship between the logarithm of Chamfer Distance and <inline-formula><mml:math id="M58"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> when the target image is set to Target-2 and the cloth is randomly moved for 90 s is shown in <xref ref-type="fig" rid="F8">Figure 8</xref>. The two values are well correlated, with a correlation coefficient of 0.811.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>The correlation between the logarithm of chamfer distance and <inline-formula><mml:math id="M54"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0008.tif"/>
</fig>
</sec>
<sec>
<title>3.3. Actual Robot Experiment of Musashi-W</title>
<p>Motion data was collected by randomly commanding <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup> and <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> (Random) for about 80 s, and while human operation with GUI (joint angle is given by GUI, joint stiffness is randomly given) for about 80 s, for the cloths of soft-1, soft-2, hard-1, and hard-3, respectively. The random command <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> is set between [10, 70] [N]. Using a total of 3,000 data points, we trained DPMPB with the number of LSTM expansions being 20, the number of batches being 1,000, and the number of epochs being 600 (these parameters are empirically set). Principle Component Analysis (PCA) is applied to the parametric bias for each cloth obtained in the training, and its arrangement in the two-dimensional plane is shown in <xref ref-type="fig" rid="F9">Figure 9</xref>. It can be seen that soft-1, soft-2, hard-1, and hard-3 are placed at the four corners, along with the thickness and stiffness of the cloth.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Actual robot experiment: the trained parametric bias for soft-1, soft-2, hard-1, and hard-3, and the trajectory of online updated parametric bias for hard-1, hard-2, and hard-3.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0009.tif"/>
</fig>
<p>We execute the control in Section 2.5 and the online material estimation in Section 2.4 for about 30 s with each of the hard-1, hard-2, and hard-3 cloths. The trajectory of parametric bias is also shown in <xref ref-type="fig" rid="F9">Figure 9</xref>. Note that the initial value of PB is <bold>0</bold> (the origin of PB is not necessarily at the origin of the figure, since PCA is applied). The updated PBs aligned in a straight line according to the thickness of the cloth. The updated PB of hard-3 is close to the PB of hard-3 obtained during training, while the updated PB of hard-1 deviated slightly from the PB of hard-1 obtained during training and was close to the PB of hard-3. The updated PB of hard-2, which is the data not used in the training, is located midway between hard-1 and hard-3. Therefore, we consider that the space of PB is self-organized according to the cloth thickness and cloth stiffness.</p>
<p>First, we performed the control in Section 2.5 with the value of PB set to the PB obtained during the training of the object to be manipulated. From the initial state of the cloth, as shown in the left figure of <xref ref-type="fig" rid="F10">Figure 10</xref>, we use Section 2.5 to bring the cloth closer to the state of being spread out in the air as shown in the right figure of <xref ref-type="fig" rid="F10">Figure 10</xref>. For hard-1 and soft-1, examples of transitions of the target joint angle <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup>, the measured joint angle <bold><italic>&#x003B8;</italic></bold>, and the muscle tension <italic>f</italic><sup><italic>const</italic></sup> representing the target body stiffness are shown in <xref ref-type="fig" rid="F11">Figure 11</xref>. Note that &#x003B8;<sub><italic>s</italic>&#x02212;<italic>p</italic></sub> and &#x003B8;<sub><italic>e</italic>&#x02212;<italic>p</italic></sub> represent the pitch joint angles of the shoulder and elbow, respectively. The area surrounded by the red frame is the main part of the spreading motion, where the upraised shoulder and elbow are lowered down significantly (the larger the joint angle is, the lower the arm is). Here, if we look at <italic>f</italic><sup><italic>const</italic></sup>, we can see that in many cases, the stiffness is increased in the first half and decreased in the second half. This increases the hand speed and causes the cloth to spread out and float in midair, realizing the state shown in the right figure of <xref ref-type="fig" rid="F10">Figure 10</xref>. The absolute values of the hand velocities when Equation (7) is not performed for <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> and <italic>f</italic><sup><italic>const</italic></sup> is fixed (&#x0003D; 10 N) (w/o stiffness), and when the control in Section 2.5 is performed (w/ stiffness), are shown in <xref ref-type="fig" rid="F12">Figure 12</xref>. The figure shows the mean and variance of the maximum hand velocities during five 20-s experiments from the state shown in the left figure of <xref ref-type="fig" rid="F10">Figure 10</xref>. It can be seen that the speed is about 12% faster when the change in the body stiffness is used. Also, looking at the difference between soft-1 and hard-1 in <xref ref-type="fig" rid="F11">Figure 11</xref>, in hard-1, &#x003B8;<sub><italic>e</italic>&#x02212;<italic>p</italic></sub> and &#x003B8;<sub><italic>s</italic>&#x02212;<italic>p</italic></sub> move almost simultaneously and the process from raising to lowering the hand is quick. On the other hand, in soft-1, &#x003B8;<sub><italic>s</italic>&#x02212;<italic>p</italic></sub> starts to move slightly later than &#x003B8;<sub><italic>e</italic>&#x02212;<italic>p</italic></sub>, and the process from raising to lowering the hand is often slow. The reason is that the soft cloth can maintain flight time even if the hand is lowered slowly at the end, while the hard cloth cannot maintain flight time unless the hand is lowered immediately. In this way, we can see that the way of manipulating the object varies depending on the difference in PB.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Experimental setup for the actual robot experiment: the initial state of cloth and target image of spread-out cloth.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0010.tif"/>
</fig>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>Actual robot experiment: examples of transitions of <bold><italic>&#x003B8;</italic></bold><sup><italic>ref</italic></sup>, <bold><italic>&#x003B8;</italic></bold>, <italic>f</italic><sup><italic>const</italic></sup>, and <inline-formula><mml:math id="M59"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> when conducting an experiment of dynamic manipulation of hard-1 <bold>(upper graphs)</bold> and soft-1 <bold>(lower graphs)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0011.tif"/>
</fig>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p>Actual robot experiment: the average and SD of the maximum velocity of the end effector when conducting experiments of dynamic manipulation of hard-1 and soft-1.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0012.tif"/>
</fig>
<p>Next, we performed the cloth manipulation under several conditions and compared them. The trials with correct PBs (the PB of soft-1 is used for the manipulation of soft-1 and the PB of hard-1 is used for the manipulation of hard-1) are called Correct, the trials with wrong PBs (the PB of hard-1 is used for the manipulation of soft-1 and the PB of soft-1 is used for the manipulation of hard-1) are called Wrong, and random trials such as in Section 3.3 are called Random. The case where the PB is correct but Equation (7) is not performed for <bold><italic>k</italic></bold><sup><italic>ref</italic></sup> and <italic>f</italic><sup><italic>const</italic></sup> is fixed (10N) is called Correct w/o stiffness. Under these four conditions, the hard-1 and soft-1 cloths are manipulated for 25 s five times from the state shown in the left graph of <xref ref-type="fig" rid="F10">Figure 10</xref>. The average of the percentage (y-axis) with <inline-formula><mml:math id="M64"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> lower than a certain threshold (x-axis) is shown in <xref ref-type="fig" rid="F13">Figure 13</xref>. The means and variances of the proportions of <inline-formula><mml:math id="M65"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> taken from <xref ref-type="fig" rid="F13">Figure 13</xref> are shown in the upper graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, the means and variances of the minimum value of <inline-formula><mml:math id="M66"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are shown in the middle graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, and the means and variances of the maximum value of <inline-formula><mml:math id="M67"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are shown in the lower graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>. <xref ref-type="fig" rid="F13">Figure 13</xref> indicates that the more the graph expands in the upper-left direction (the larger the y-axis value becomes when the x-axis value is low), the more the target image is realized. Note that the mean of <inline-formula><mml:math id="M68"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> for the initial condition shown in the left figure of <xref ref-type="fig" rid="F10">Figure 10</xref> is 0.86. For both hard-1 and soft-1, Correct is the best, and Random is the worst. Additionally, looking at the upper graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, we can see the characteristics quantitatively. Although hard-1 is more difficult and, thus, has a lower realization rate in general, the results are consistent with the trend that Wrong and Correct w/o stiffness have lower realization rates than Correct. Looking at the middle graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, the tendency of the minimum value of <inline-formula><mml:math id="M69"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is also consistent with the upper graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, and the lower the minimum value is, the higher the percentage of realization is. On the other hand, for the lower graphs of <xref ref-type="fig" rid="F14">Figure 14</xref>, Correct is the highest, and Random is the lowest. This may indicate that the more accurate the method can realize the target image, the more likely it is to produce a state that is different from the target image since the state must be changed significantly in order to realize the target image.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>Actual robot experiment: the rate (y-axis) of <inline-formula><mml:math id="M60"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo></mml:math></inline-formula>threshold (x-axis) when conducting experiments of dynamic manipulation of hard-1 and soft-1 regarding Correct, Wrong, Correct without stiffness, and Random.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0013.tif"/>
</fig>
<fig id="F14" position="float">
<label>Figure 14</label>
<caption><p>Actual robot experiment: the rate of <inline-formula><mml:math id="M61"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> <bold>(upper graphs)</bold>, the minimum value of <inline-formula><mml:math id="M62"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> <bold>(middle graphs)</bold>, and the maximum value of <inline-formula><mml:math id="M63"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> <bold>(lower graphs)</bold> when conducting experiments of dynamic manipulation of hard-1 and soft-1.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0014.tif"/>
</fig>
</sec>
<sec>
<title>3.4. Integrated Table Setting Experiment</title>
<p>We performed a series of motion experiments incorporating dynamic cloth manipulation. Musashi-W picked up hard-1 and laid it on the table like a table cloth by the proposed method, and placed a basket of sweets on it. The scene is shown in <xref ref-type="fig" rid="F15">Figure 15</xref>. First, the robot recognizes the point cloud of the cloth and grasps the cloth between its thumb and index by visual feedback. Next, the robot spreads the cloth on the desk by dynamic cloth manipulation, which has been previously trained on a similar setup. Finally, the robot recognizes the point cloud of the basket and grasps the basket between both hands by visual feedback, and successfully places it on the table.</p>
<fig id="F15" position="float">
<label>Figure 15</label>
<caption><p>The integrated table setting experiment of grasping a cloth, manipulating it dynamically, and putting a basket on it.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-890695-g0015.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>We discuss the results obtained from the experiments in this study. First, simulation experiments show that the value of PB is self-organized according to the parameters of the cloth. The dynamics of the current cloth can be estimated online from the cloth manipulation data. On the other hand, we found that it is difficult to accurately estimate the parameters that do not contribute much to the change in dynamics. Through dynamic cloth manipulation experiments while changing the cloth parameters and the target image, our method can accurately realize the target image. By changing PB, the method can handle various cloth properties, and the target image can be given arbitrarily. The reproducibility of the performance on learning and control is also high. From the integrated experiment of online material estimation and dynamic cloth manipulation, we show that they can be performed simultaneously and that the closer the current PB is to the current cloth properties, the more accurately the target image can be realized.</p>
<p>From the actual robot experiments, we found that the value of PB is self-organized depending on the thickness and stiffness of clothes. As in the simulation experiments, the dynamics of the current cloth can be estimated online from the cloth manipulation data. We also found that the dynamics can be estimated at the correct point, which is the internally dividing point of the trained PBs, even for the data that is not available at the time of training. Through dynamic cloth manipulation experiments, we found that the behavior of the cloth manipulation varies depending on the PB which expresses the dynamics of the cloth. The joint angles and the speed were appropriately changed according to the material of the cloth to be manipulated. In addition, the stiffness value affects the speed of the hand movements, and when the stiffness value is optimized, the hand speed is increased by about 12% compared to the case without its optimization. When PB does not match the current cloth or the stiffness value is not optimized, the performance decreased compared to the case where PB is correct and the variable stiffness is used. In addition, we found that in order to realize the target image correctly, it is necessary to go through a state that is far from the target image. From these results, we found that the dynamics of the cloth material are embedded and estimated online through parametric bias, that variable stiffness control can be used explicitly to improve the hand speed, and that the target cloth image can be realized more accurately by using deep predictive model learning and backpropagation technique to the control command.</p>
<p>While our method can greatly improve the ability of dynamic cloth manipulation of robots, there are still some issues to be solved. First, in this study, the body stiffness is substituted by a certain single value, but in humans, the body stiffness can be set more flexibly. It is necessary to discuss the degree of freedom of the body stiffness in the future. Next, we consider that it is difficult for our method to deal with the case of large nonlinear changes such as the cloth leaving the hand or changing the grasping position. There is no method that can handle such highly nonlinear, discontinuous, and high-dimensional states in trials using only actual robots, and this is an issue for the future. Finally, we are still far from reaching human-like adaptive dynamic cloth manipulation. Humans are capable of stretching and spreading the cloth by focusing on small wrinkles, instead of just looking at the entire cloth universally. In addition, since the current hardware is not as agile and soft as humans, it will be necessary to develop both hardware and software.</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>In this study, we developed a deep predictive model learning method that incorporates quick manipulation using variable stiffness mechanisms, and response to changes in cloth material for more human-like dynamic cloth manipulation. For variable stiffness, the command value is calculated by the backpropagation technique using the joint angle and the body stiffness value as control commands. For the adaptation to change in cloth material, we embed information about the cloth dynamics into parametric bias. The dynamics of the cloth material can be estimated online even when it is not included in the training data, and it is shown that the characteristics of the dynamic cloth manipulation vary greatly depending on the material. In addition, when the stiffness value is appropriately set according to the motion phase, the speed can be increased by about 12% compared to the case without variable stiffness control, and the target cloth state can be realized more accurately. In the future, we would like to develop a method to handle more complicated cloth manipulation tasks in a unified manner.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>KK contributed to conception and design of this study, implemented the software, conducted the experiments, and wrote this manuscript. AM and MB conducted the experiments with KK. KO and MI supervised this project and contributed to conception of this study. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This research was partially supported by JST ACT-X grant no. JPMJAX20A5 and JSPS KAKENHI grant no. JP19J21672.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fnbot.2022.890695/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fnbot.2022.890695/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Video_1.MP4" id="SM1" mimetype="video/mp4" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Video 1</label>
<caption><p>Attached is a video of experiments.</p></caption>
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borgefors</surname> <given-names>G..</given-names></name></person-group> (<year>1988</year>). <article-title>Hierarchical chamfer matching: a parametric edge matching algorithm</article-title>. <source>IEEE Trans. Pattern. Anal. Mach. Intell</source>. <volume>10</volume>, <fpage>849</fpage>&#x02013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1109/34.9107</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Braun</surname> <given-names>D.</given-names></name> <name><surname>Howard</surname> <given-names>M.</given-names></name> <name><surname>Vijayakumar</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Optimal variable stiffness control: formulation and application to explosive movement tasks</article-title>. <source>Auton. Robots</source> <volume>33</volume>, <fpage>237</fpage>&#x02013;<lpage>253</lpage>. <pub-id pub-id-type="doi">10.1007/s10514-012-9302-3</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ebert</surname> <given-names>F.</given-names></name> <name><surname>Finn</surname> <given-names>C.</given-names></name> <name><surname>Dasari</surname> <given-names>S.</given-names></name> <name><surname>Xie</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>A.</given-names></name> <name><surname>Levine</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Visual foresight: model-based deep reinforcement learning for vision-based robotic control</article-title>. <source>arXiv preprint</source> arXiv:1812.00568. <pub-id pub-id-type="doi">10.48550/arXiv.1812.00568</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Elbrechter</surname> <given-names>C.</given-names></name> <name><surname>Haschke</surname> <given-names>R.</given-names></name> <name><surname>Ritter</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Folding paper with anthropomorphic robot hands using real-time physics-based modeling,&#x0201D;</article-title> in <source>Proceedings of the 2012 IEEE-RAS International Conference on Humanoid Robots</source> (<publisher-loc>Osaka</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>210</fpage>&#x02013;<lpage>215</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ha</surname> <given-names>H.</given-names></name> <name><surname>Song</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;FlingBot: the unreasonable effectiveness of dynamic manipulation for cloth unfolding,&#x0201D;</article-title> in <source>Proceedings of the 2021 Conference on Robot Learning</source> (<publisher-loc>London, UK</publisher-loc>).</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name></person-group> (<year>2006</year>). <article-title>Reducing the dimensionality of data with neural networks</article-title>. <source>Science</source> <volume>313</volume>, <fpage>504</fpage>&#x02013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id><pub-id pub-id-type="pmid">16873662</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput</source>. <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hoque</surname> <given-names>R.</given-names></name> <name><surname>Seita</surname> <given-names>D.</given-names></name> <name><surname>Balakrishna</surname> <given-names>A.</given-names></name> <name><surname>Ganapathi</surname> <given-names>A.</given-names></name> <name><surname>Tanwani</surname> <given-names>A.</given-names></name> <name><surname>Jamali</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;VisuoSpatial foresight for multi-step, multi-task fabric manipulation,&#x0201D;</article-title> in <source>Proceedings of the 2020 Robotics: Science and Systems</source>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Inaba</surname> <given-names>M.</given-names></name> <name><surname>Inoue</surname> <given-names>H.</given-names></name></person-group> (<year>1987</year>). <article-title>Rope handling by a robot with visual feedback</article-title>. <source>Adv. Rob</source>. <volume>2</volume>, <fpage>39</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1163/156855387X00057</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Szegedy</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Batch normalization: accelerating deep network training by reducing internal covariate shift,&#x0201D;</article-title> in <source>Proceedings of the 32nd International Conference on Machine Learning</source> (<publisher-loc>Lille</publisher-loc>), <fpage>448</fpage>&#x02013;<lpage>456</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jangir</surname> <given-names>R.</given-names></name> <name><surname>Aleny&#x000E4;</surname> <given-names>G.</given-names></name> <name><surname>Torras</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Dynamic cloth manipulation with deep reinforcement learning,&#x0201D;</article-title> in <source>Proceedings of the 2020 IEEE International Conference on Robotics and Automation</source> (<publisher-loc>Paris</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4630</fpage>&#x02013;<lpage>4636</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kawaharazuka</surname> <given-names>K.</given-names></name> <name><surname>Kawamura</surname> <given-names>Y.</given-names></name> <name><surname>Okada</surname> <given-names>K.</given-names></name> <name><surname>Inaba</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Imitation learning with additional constraints on motion style using parametric bias</article-title>. <source>IEEE Rob. Autom. Lett</source>. <volume>6</volume>, <fpage>5897</fpage>&#x02013;<lpage>5904</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2021.3087423</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kawaharazuka</surname> <given-names>K.</given-names></name> <name><surname>Makino</surname> <given-names>S.</given-names></name> <name><surname>Kawamura</surname> <given-names>M.</given-names></name> <name><surname>Fujii</surname> <given-names>A.</given-names></name> <name><surname>Asano</surname> <given-names>Y.</given-names></name> <name><surname>Okada</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Online self-body image acquisition considering changes in muscle routes caused by softness of body tissue for tendon-driven musculoskeletal humanoids,&#x0201D;</article-title> in <source>Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Madrid</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1711</fpage>&#x02013;<lpage>1717</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kawaharazuka</surname> <given-names>K.</given-names></name> <name><surname>Makino</surname> <given-names>S.</given-names></name> <name><surname>Tsuzuki</surname> <given-names>K.</given-names></name> <name><surname>Onitsuka</surname> <given-names>M.</given-names></name> <name><surname>Nagamatsu</surname> <given-names>Y.</given-names></name> <name><surname>Shinjo</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2019a</year>). <article-title>&#x0201C;Component modularized design of musculoskeletal humanoid platform musashi to investigate learning control systems,&#x0201D;</article-title> in <source>Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Macau</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>7294</fpage>&#x02013;<lpage>7301</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kawaharazuka</surname> <given-names>K.</given-names></name> <name><surname>Ogawa</surname> <given-names>T.</given-names></name> <name><surname>Tamura</surname> <given-names>J.</given-names></name> <name><surname>Nabeshima</surname> <given-names>C.</given-names></name></person-group> (<year>2019b</year>). <article-title>&#x0201C;Dynamic manipulation of flexible objects with torque sequence using a deep neural network,&#x0201D;</article-title> in <source>Proceedings of the 2019 IEEE International Conference on Robotics and Automation</source> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2139</fpage>&#x02013;<lpage>2145</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Adam: a method for stochastic optimization,&#x0201D;</article-title> in <source>Proceedings of the 3rd International Conference on Learning Representations</source> (<publisher-loc>San Diego, CA</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>15</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>A. X.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Gupta</surname> <given-names>A.</given-names></name> <name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Learning force-based manipulation of deformable objects from multiple demonstrations,&#x0201D;</article-title> in <source>Proceedings of the 2015 IEEE International Conference on Robotics and Automation</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>177</fpage>&#x02013;<lpage>184</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mart&#x000ED;n-Mart&#x000ED;n</surname> <given-names>R.</given-names></name> <name><surname>Lee</surname> <given-names>M. A.</given-names></name> <name><surname>Gardner</surname> <given-names>R.</given-names></name> <name><surname>Savarese</surname> <given-names>S.</given-names></name> <name><surname>Bohg</surname> <given-names>J.</given-names></name> <name><surname>Garg</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Variable impedance control in end-effector space: an action space for reinforcement learning in contact-rich tasks,&#x0201D;</article-title> in <source>Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Macau</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1010</fpage>&#x02013;<lpage>1017</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nair</surname> <given-names>V.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Rectified linear units improve restricted boltzmann machines,&#x0201D;</article-title> in <source>Proceedings of the 27th International Conference on Machine Learning</source> (<publisher-loc>Haifa</publisher-loc>), <fpage>807</fpage>&#x02013;<lpage>814</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ogata</surname> <given-names>T.</given-names></name> <name><surname>Ohba</surname> <given-names>H.</given-names></name> <name><surname>Tani</surname> <given-names>J.</given-names></name> <name><surname>Komatani</surname> <given-names>K.</given-names></name> <name><surname>Okuno</surname> <given-names>H. G.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;Extracting multi-modal dynamics of objects using RNNPB,&#x0201D;</article-title> in <source>Proceedings of the 2005 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Edmonton, AB</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>966</fpage>&#x02013;<lpage>971</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saha</surname> <given-names>M.</given-names></name> <name><surname>Isto</surname> <given-names>P.</given-names></name></person-group> (<year>2007</year>). <article-title>Manipulation planning for deformable linear objects</article-title>. <source>IEEE Trans. Rob</source>. <volume>23</volume>, <fpage>1141</fpage>&#x02013;<lpage>1150</lpage>. <pub-id pub-id-type="doi">10.1109/TRO.2007.907486</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shirai</surname> <given-names>T.</given-names></name> <name><surname>Urata</surname> <given-names>J.</given-names></name> <name><surname>Nakanishi</surname> <given-names>Y.</given-names></name> <name><surname>Okada</surname> <given-names>K.</given-names></name> <name><surname>Inaba</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Whole body adapting behavior with muscle level stiffness control of tendon-driven multijoint robot,&#x0201D;</article-title> in <source>Proceedings of the 2011 IEEE International Conference on Robotics and Biomimetics</source> (<publisher-loc>Karon Beach</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2229</fpage>&#x02013;<lpage>2234</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>D.</given-names></name> <name><surname>Arnold</surname> <given-names>S.</given-names></name> <name><surname>Yamazaki</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>EMD Net: an encode-manipulate-decode network for cloth manipulation</article-title>. <source>IEEE Rob. Autom. Lett</source>. <volume>3</volume>, <fpage>1771</fpage>&#x02013;<lpage>1778</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2018.2800122</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tani</surname> <given-names>J..</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Self-organization of behavioral primitives as multiple attractor dynamics: a robotexperiment,&#x0201D;</article-title> in <source>Proceedings of the 2002 International Joint Conference on Neural Networks</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>489</fpage>&#x02013;<lpage>494</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name> <name><surname>Erez</surname> <given-names>T.</given-names></name> <name><surname>Tassa</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;MuJoCo: a physics engine for model-based control,&#x0201D;</article-title> in <source>Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Vilamoura-Algarve</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5026</fpage>&#x02013;<lpage>5033</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yamakawa</surname> <given-names>Y.</given-names></name> <name><surname>Namiki</surname> <given-names>A.</given-names></name> <name><surname>Ishikawa</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Motion planning for dynamic knotting of a flexible rope with a high-speed robot arm,&#x0201D;</article-title> in <source>Proceedings of the 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Taipei</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>49</fpage>&#x02013;<lpage>54</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yamakawa</surname> <given-names>Y.</given-names></name> <name><surname>Namiki</surname> <given-names>A.</given-names></name> <name><surname>Ishikawa</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Motion planning for dynamic folding of a cloth with two high-speed robot hands and two high-speed sliders,&#x0201D;</article-title> in <source>Proceedings of the 2011 IEEE International Conference on Robotics and Automation</source> (<publisher-loc>Shanghai</publisher-loc>,: <publisher-name>IEEE</publisher-name>), <fpage>5486</fpage>&#x02013;<lpage>5491</lpage>.</citation>
</ref>
</ref-list> 
</back>
</article>