<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">854212</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2022.854212</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Model-free reinforcement learning for robust locomotion using demonstrations from trajectory optimization</article-title>
<alt-title alt-title-type="left-running-head">Bogdanovic et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2022.854212">10.3389/frobt.2022.854212</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Bogdanovic</surname>
<given-names>Miroslav</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1634929/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Khadiv&#x2009;</surname>
<given-names>Majid</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/857386/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Righetti&#x2009;</surname>
<given-names>Ludovic</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/193664/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Movement Generation and Control Group</institution>, <institution>Max Planck Institute for Intelligent Systems</institution>, <addr-line>T&#xfc;bingen</addr-line>, <country>Germany</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Machines in Motion Laboratory</institution>, <institution>Tandon School of Engineering</institution>, <institution>New York University</institution>, <addr-line>New York</addr-line>, <addr-line>NY</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/143509/overview">Alan Frank Thomas Winfield</ext-link>, University of the West of England, United Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/754447/overview">L&#xe9;ni Kenneth Le Goff</ext-link>, Edinburgh Napier University, United Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1762253/overview">Mark Gluzman</ext-link>, Cornell University, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Miroslav Bogdanovic, <email>mbogdanovic@tue.mpg.de</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Robot Learning and Evolution, a section of the journal Frontiers in Robotics and AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>31</day>
<month>08</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>9</volume>
<elocation-id>854212</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>07</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Bogdanovic, Khadiv&#x2009; and Righetti&#x2009;.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Bogdanovic, Khadiv&#x2009; and Righetti&#x2009;</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>We present a general, two-stage reinforcement learning approach to create robust policies that can be deployed on real robots without any additional training using a single demonstration generated by trajectory optimization. The demonstration is used in the first stage as a starting point to facilitate initial exploration. In the second stage, the relevant task reward is optimized directly and a policy robust to environment uncertainties is computed. We demonstrate and examine in detail the performance and robustness of our approach on highly dynamic hopping and bounding tasks on a quadruped robot.</p>
</abstract>
<kwd-group>
<kwd>legged locomotion</kwd>
<kwd>deep reinforcement learning</kwd>
<kwd>trajectory optimization</kwd>
<kwd>robust control policies</kwd>
<kwd>contact uncertainty</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Deep reinforcement learning (DRL) has recently shown great promises to control complex robotic tasks, e.g., object manipulation <xref ref-type="bibr" rid="B13">Kalashnikov et al. (2018)</xref>, quadrupedal <xref ref-type="bibr" rid="B11">Hwangbo et al. (2019)</xref> and bipedal <xref ref-type="bibr" rid="B31">Xie et al. (2020)</xref> locomotion. However, exploration remains a serious challenge in RL, especially for legged locomotion control, mainly due to the sparse rewards in problems with contact as well as the inherent under-actuation and instability of legged robots. Furthermore, to successfully transfer learned control policies to real robots, there is still no consensus among researchers about the choice of the action space <xref ref-type="bibr" rid="B21">Peng and van de Panne, (2017)</xref> and what (and how) to randomize in the training procedure to generate robust policies <xref ref-type="bibr" rid="B33">Xie et al. (2021b)</xref>.</p>
<p>Trajectory optimization (TO) is a powerful tool for generating stable motions for complex and highly constrained systems such as legged robot (<xref ref-type="bibr" rid="B30">Winkler et al., 2018</xref>; <xref ref-type="bibr" rid="B2">Carpentier and Mansard, 2018</xref>; <xref ref-type="bibr" rid="B24">Ponton et al., 2021</xref>). However, re-planning trajectories through a model predictive control (MPC) scheme is still a challenge, because the computation time for solving a high-dimensional non-linear program in real-time remains too high. Furthermore, apart from recent works explicitly taking into account contact uncertainty to design robust control policies (<xref ref-type="bibr" rid="B4">Drnach and Zhao, 2021</xref>; <xref ref-type="bibr" rid="B9">Hammoud et al., 2021</xref>), the inclusion of robustness objectives in trajectory optimization can quickly end up in problems that cannot be solved in real time for high-dimensional systems in multi-contact scenarios.</p>
<p>In this work, we propose a general approach allowing for the generation of robust policies starting from a single demonstration. The main idea is to use TO to generate trajectories for different tasks that are then used for exploration for DRL.</p>
<p>
<bold>Reinforcement learning for locomotion:</bold> The use of reinforcement learning for generating stable locomotion patterns is a relatively old problem (<xref ref-type="bibr" rid="B10">Hornby et al., 2000</xref>; <xref ref-type="bibr" rid="B14">Kohl and Stone, 2004</xref>). However, these works mostly use a hand-crafted policy with very few policy parameters that are tuned on the hardware. Later works use a notion of Poincare map to ensure the cyclic stability of the gaits (<xref ref-type="bibr" rid="B18">Morimoto et al., 2005</xref>; <xref ref-type="bibr" rid="B28">Tedrake et al., 2005</xref>). These approaches have enabled later a humanoid robot to walk <xref ref-type="bibr" rid="B17">Morimoto and Atkeson, (2009)</xref>, but their underlying function approximator cannot handle large number of policy parameters which limits their application. Furthermore, the robot hardware at that time were not capable of dynamic movements which made the researchers focus mostly on the walking problem. Later, (<xref ref-type="bibr" rid="B5">Fankhauser et al., 2013</xref>), used more scalable approaches (PI&#x2c6;2 algorithm by <xref ref-type="bibr" rid="B29">Theodorou et al. (2010)</xref>) for generating a hopping motion on a single planar leg. Recently, deep reinforcement learning has become the main approach for learning both locomotion and manipulation policies.</p>
<p>
<bold>Combining Model-based Control with DRL.</bold> One approach to benefit from the efficiency of the model-based control methods and the robustness of DRL policies is to use a hybrid method. Works within this setting can be split into two categories; 1) Desired trajectories are generated by DRL using a reduced order model of the robot, e.g., centroidal momentum dynamics <xref ref-type="bibr" rid="B32">Xie et al. (2021a)</xref>, 2) A residual policy adapts the trajectories from TO (<xref ref-type="bibr" rid="B7">Gangapurwala et al., 2020</xref>; <xref ref-type="bibr" rid="B6">Gangapurwala et al., 2021</xref>). Both approaches pass the generated trajectories to a whole-body controller to track the trajectories while satisfying constraints. The first category resolves the problem of exploration in DRL by using a reduced model which neglects the whole-body dynamics which is limiting for most highly dynamic locomotion tasks. The second category works well as long as the real robot behaviour remains close to the pre-generated trajectories. In cases that there exists a significant change in the environment or large external disturbances, those trajectories are not useful and the RL policy needs to learn to ignore them and find a whole new policy to learn the new behaviour. In such cases, it seems this approach is very limited.</p>
<p>
<bold>Learning from demonstrations.</bold> An interesting approach to address this issue is to utilize demonstrations for the given task (<xref ref-type="bibr" rid="B25">Schaal, 1997</xref>; <xref ref-type="bibr" rid="B12">Ijspeert et al., 2002</xref>; <xref ref-type="bibr" rid="B22">Peters and Schaal, 2008</xref>). By providing to the reinforcement learning algorithm basic motions required for completing the task, we remove the need for the algorithm to find it on its own using random exploration. Here, one can combine trajectory optimization with deep reinforcement learning, by utilizing motions generated by trajectory optimization as demonstrations used to further generate robust policies using deep reinforcement learning. Demonstrations can be utilized by reinforcement learning approaches in several different ways. In order to train a reinforcement learning policy to reproduce the demonstration behavior, the majority of approaches have some notion of time in the input of the policy. Some approaches explicitly give time-indexed states from the demonstration trajectory directly as input (<xref ref-type="bibr" rid="B20">Peng et al., 2020</xref>; <xref ref-type="bibr" rid="B16">Li et al., 2021</xref>). Alternatively, some works train the policy to reproduce the demonstration behavior using only a phase variable in the policy input (<xref ref-type="bibr" rid="B31">Xie et al., 2020</xref>; <xref ref-type="bibr" rid="B27">Siekmann et al., 2021</xref>). Removing any notion of time from the input makes it difficult to train robust policies for real systems (<xref ref-type="bibr" rid="B31">Xie et al., 2020</xref>).</p>
<p>
<bold>Robustness and time-dependence.</bold> There is however a crucial issue in learning control policies in such a way, in particular in the presence of environmental uncertainties. As an example, imagine we want to produce a hopping policy that can account for large uncertainties in the ground height. When contact is made at a different time than in the demonstration, the phase given as input to the policy will be different than the actual phase in the task. Instead of just going directly into a baseline hopping motion after making contact with the ground, the policy would need to force itself to get back into phase with the demonstration to have any luck to complete the task. Even removing any notion of time from the input does not on its own solve this issue. The policy would still need to get back in phase with the demonstration trajectory, but without the time-based inputs lack the information needed to be able to do so. Hence, fully removing time dependence from the demonstration trajectories in the final feedback policy is key in our approach to provide robustness with respect to contact timing uncertainties.</p>
<p>There are additional benefits in eliminating time dependence when training robust control policies. It allows us to more broadly randomize initial configurations of the system, something key in deploying learned policies on real robots. It also separates behavior given in the demonstration from the task goals, allowing us to improve beyond the demonstration for all the important aspects of the task at hand.</p>
<p>
<bold>Our approach.</bold> In this work we propose a general approach that combines TO and DRL in order to produce robust policies that can be deployed on real robots. We benefit from trajectories produced by TO to bootstrap DRL algorithms and avoid exploration issues. We then use DRL to, based on these demonstrations, produce policies robust to environmental uncertainties. In this way we get the best of both worlds. Starting from TO trajectories affords solving complex tasks DRL would otherwise struggle with. The two-stage DRL approach we propose then allows us to avoid the above issues that arise when learning based on demonstrations and learn entirely time-independent robust policies.</p>
<sec id="s1-1">
<title>1.1 Related work</title>
<p>Several recent works have used reinforcement learning to compute policies for locomotion tasks in simulation before deploying them successfully on a real quadruped (<xref ref-type="bibr" rid="B11">Hwangbo et al., 2019</xref>; <xref ref-type="bibr" rid="B15">Lee et al., 2020</xref>; <xref ref-type="bibr" rid="B20">Peng et al., 2020</xref>) or biped robot (<xref ref-type="bibr" rid="B31">Xie et al., 2020</xref>; <xref ref-type="bibr" rid="B16">Li et al., 2021</xref>; <xref ref-type="bibr" rid="B27">Siekmann et al., 2021</xref>).</p>
<p>In (<xref ref-type="bibr" rid="B31">Xie et al., 2020</xref>), similarly to our work, the authors start by learning a policy to track a reference motion and then further improve it to enable its transfer to the real system. Crucially, unlike our approach, some notion of time remains present in the policy input in all stages of training, either as a reference motion or a phase variable. The reward for tracking the original trajectory also remains present in further stages of training, preventing the policy to freely adapt away from it in order to optimize task performance. Finally, while the resulting policies show some robustness to external perturbations, there are no environment uncertainties present and no need to adapt the timing of the behavior to account for it.</p>
<p>In (<xref ref-type="bibr" rid="B16">Li et al., 2021</xref>), the authors aim to improve upon some aspects of (<xref ref-type="bibr" rid="B31">Xie et al., 2020</xref>). Instead of learning the policy output in the residual space, i.e. learning only the correction with respect to the demonstration trajectory, they learn the full control signal. While this makes it easier for the policy to adapt its actions away from the demonstration, the time-dependence is still present in the policy input and the issue of adaptation of timing of the demonstration remains.</p>
<p>In (<xref ref-type="bibr" rid="B20">Peng et al., 2020</xref>) the control policy is computed in a single stage of training, by learning to track the given demonstration behavior. For successful transfer to real robot they rely on domain adaptation, finding the latent encoding over the set of dynamics parameters that performs the best. There is however no randomization of the environment during training and all test are performed on flat ground. Similarly to the previous works, the notion of time remains present in the policies here as well, explicitly as a goal given to the policy containing robot states from the reference motion in several of the following time steps.</p>
<p>Unlike these approaches, the most successful recent learning based approach for locomotion in a challenging uncertain environment does not utilize demonstrations at all. The authors of (<xref ref-type="bibr" rid="B15">Lee et al., 2020</xref>) learn a robust locomotion policy for a quadruped robot that performs well on uneven and uncertain surfaces. The lack of any notion of pre-determined timing affords more room to the policy to adapt to environmental uncertainties (in this case even explicitly by outputting the frequency of the motion). While learning from scratch is possible in this case, it is true for one specific task (trotting without any flight phase) and with the structure imposed by the proposed controller. With the structure we mean that the policy only outputs stepping frequency for each foot and this is mapped to the joint space using inverse kinematics. For highly dynamic tasks with flight phases and impacts, with no specific structure imposed for joint behavior utilizing demonstrations is still necessary. Therefore, the issue we address in this work, i.e. finding ways to build fully adaptive policies starting from demonstration behaviors, remains a challenge.</p>
</sec>
<sec id="s1-2">
<title>1.2 Contributions</title>
<p>The main contributions of this paper are as follows:<list list-type="simple">
<list-item>
<p>&#x2022; We propose a framework to exploit the benefits of both TO and DRL to generate control policies that are robust to environmental uncertainties. A key aspect of our framework is to lose time-dependence from the initial trajectories and to build a policy that can adapt to large uncertainties in the environment.</p>
</list-item>
<list-item>
<p>&#x2022; We evaluate the method on two highly dynamic tasks on a quadruped and show that our framework can deal with random uneven terrains as well as external disturbances. To the best of our knowledge, these results are the first demonstrating the successful use of DRL to robustly realize such behaviors.</p>
</list-item>
</list>
</p>
</sec>
</sec>
<sec id="s2">
<title>2 Proposed algorithm</title>
<p>
<bold>Algorithm overview.</bold> Our proposed algorithm has three main parts as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. First we use TO to generate efficiently, based on a nominal model of the robot, initial trajectories for new tasks. The first stage of DRL training then proceeds to build a control policy around this trajectory, caching the solution to a light neural network. In the second stage of DRL training we replace the trajectory tracking optimization with one that directly optimizes task performance and introduce uncertainties in the training environment. This allows us to further adapt the policy from the first stage, making it robust and independent from the initial demonstration. As a final result, we get a policy that can be directly deployed on the real system without any additional training.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Schematic of our proposed framework. We start from a single demonstration trajectory generated by TO. In the first DRL stage we learn a policy that tracks that demonstration trajectory and successfully produces nominal behavior in simulation. To enable transfer to the real robot, the second DRL stage starts with the resulting policy from the first stage and tries to robustify the policy by randomizing contact and to optimize for performance by replacing demonstration tracking reward with task reward. We directly apply the output policy from the second stage to the robot without any domain randomization of robot parameters.</p>
</caption>
<graphic xlink:href="frobt-09-854212-g001.tif"/>
</fig>
<sec id="s2-1">
<title>2.1 Trajectory optimization</title>
<p>
<bold>Generating initial demonstration.</bold> We use trajectory optimization to generate a nominal motion for the desired task based on a nominal model of the robot and the environment. In this work, we use the trajectory optimization algorithm proposed in (<xref ref-type="bibr" rid="B24">Ponton et al., 2021</xref>) to compute such demonstrations. It is important to emphasize that any other trajectory optimization algorithm could also be used, as long as it provides a set of full-body trajectories. However, the more realistic the generated motion is, the easier it is for the first stage of DRL to find a policy that tracks it. In addition, we do not utilize control actions provided by the demonstration trajectory. Only the state trajectories are required. This allows us to use any control parametrization for the policy being learned, potentially different than the one used to generate the demonstration. For instance, it is well known that having the policy output the next desired state for a fixed PD controller (rather than torque) is beneficial in terms of transfer to the real world (<xref ref-type="bibr" rid="B11">Hwangbo et al., 2019</xref>; <xref ref-type="bibr" rid="B1">Bogdanovic et al., 2020</xref>; <xref ref-type="bibr" rid="B27">Siekmann et al., 2021</xref>). However, finding these actions in a trajectory optimization setting is not necessarily trivial.</p>
</sec>
<sec id="s2-2">
<title>2.2 DRL stage 1: Learning a policy to track a given trajectory</title>
<p>
<bold>Control policy.</bold> We use a neural network to parametrize the control policy. A relatively small network proved sufficient for the tasks we considered, with two layers of 64 units each. We keep the observation space of the policy simple, with the positions and velocities for each joint (<bold>q</bold>
<sup>
<italic>joint</italic>
</sup>, <inline-formula id="inf5">
<mml:math id="m5">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>) and only robot base variables relevant for the current task. For the hopping task, this consists of the base position and velocity along the <italic>Z</italic>-axis (<italic>z</italic>
<sup>
<italic>base</italic>
</sup>, <inline-formula id="inf6">
<mml:math id="m6">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>). For the bounding task we also add the angular position and velocity around <italic>Y</italic>-axis (<inline-formula id="inf7">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf8">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>). The policy outputs parameters for a PD controller in joint space. It gives desired joint positions at each step, while utilizing fixed P and D gains for control.</p>
<p>Throughout all the stages of training, we use an additional cost term incentivizing the policy to output values for desired joint positions that are actually tracked as well as possible. Specifically, we penalize the difference between the value given for the desired position by the policy at step <italic>t</italic> and the actual position achieved at the next step <italic>t</italic> &#x2b; 1:<disp-formula id="e1">
<mml:math id="m9">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(1)</label>
</disp-formula>We note that while this reward term incentivizes a policy to produce a trajectory that is well tracked, it does not prevent it to give values off the trajectory to create forces during contact when necessary. This has been a crucial aspect of our previous work (<xref ref-type="bibr" rid="B1">Bogdanovic et al., 2020</xref>) to enable direct policy transfer to real robots without domain randomization of robot parameters.</p>
<p>
<bold>Training procedure.</bold> We use Proximal Policy Optimization (PPO) (<xref ref-type="bibr" rid="B26">Schulman et al., 2017</xref>) to optimize the policies in both stages of our framework, but we do not have many requirements in the choice of the algorithm. PPO is an on-policy reinforcement learning method, working similarly to a trust-region method, but relying on a clipped objective function in order to simplify the algorithm and ensure better sample complexity. As we use the provided demonstration to resolve exploration issues, we do not need a, potentially off-policy, reinforcement learning algorithm with strong characteristics in this regard. We instead choose an on-policy algorithm with good convergence properties.</p>
<p>In this stage, the optimized reward consists of two parts: a part for tracking the time-based demonstration and the above-defined regularization term that is a part of the controller parametrization:<disp-formula id="e2">
<mml:math id="m10">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>6</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2296;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>7</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>8</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>9</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>10</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>11</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>12</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">q</mml:mi>
</mml:mrow>
<mml:mo>&#x307;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>&#x2009;as&#x2009;defined&#x2009;in&#x2009;</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mtext>1</mml:mtext>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>,</mml:mtext>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(2)</label>
</disp-formula>where <bold>x</bold>
<sup>
<italic>base</italic>
</sup> is the position of the robot base, <bold>q</bold>
<sup>
<italic>base</italic>
</sup> the base quaternion, <bold>
<italic>&#x3c9;</italic>
</bold>
<sup>
<italic>base</italic>
</sup> the base angular velocity and <bold>q</bold>
<sup>
<italic>joint</italic>
</sup> the joint positions. <italic>k</italic>
<sub>
<italic>ti</italic>1</sub>, &#x2026; , <italic>k</italic>
<sub>
<italic>ti</italic>12</sub> represent individual weight and scale constants for each term. We mark difference between two quaternions with &#x2296;.</p>
<p>Following <xref ref-type="bibr" rid="B19">Peng et al. (2018)</xref>, we make two further design choices that prove to be vital in making the training robustly work:<list list-type="simple">
<list-item>
<p>&#x2022; We initialize each episode at a randomly chosen point on the demonstration trajectory.</p>
</list-item>
<list-item>
<p>&#x2022; We terminate episodes early if the robot enters states that are not likely to be recoverable (based for example on tilt angle of the robot base) or just not conducive for learning (for example knees of the robot making contact with the ground).</p>
</list-item>
</list>
</p>
<p>
<bold>Output.</bold> In the first DRL stage, we aim to produce a policy that provides some nominal behavior on the task in simulation. However, in our experiments, these policies failed to transfer to the real robot. They remain static, cause shaky behavior on the robot, or result in motions with severe impacts. We give some examples in the accompanying video. As can be seen there, the behavior is not even close to the gaits in simulation which makes any quantitative analysis unfeasible. To solve these problems, we need an additional stage of training that generates policies that can transfer to the real robot.</p>
</sec>
<sec id="s2-3">
<title>2.3 DRL stage 2: Generating robust time-independent policy</title>
<p>To create policies that are successfully transferable to the real robot, we continue training starting from the policies outputted from the first stage. As a note, we preserve the entire policy, including the parameters controlling the variance of the action distribution that the policy outputs. While the lower variance from the resulting policies from stage 1 might lower exploration capabilities in stage 2, the initial policy already performs nominal behavior on the task, so there is no need for significant exploration away from it. Additionally, we found that any attempts to artificially increase the variance prior to stage 2 result in quick loss of the behavior from stage 1 that we are attempting to carry over. On the other hand, we found no issues with potentially low resulting variance from stage 1 preventing further adaptation of the policy in stage 2.</p>
<p>We further introduce the following changes in the training procedure:</p>
<p>
<bold>Initialization.</bold> We replace initialization on the demonstrated trajectories with initialization in a wider range of states. This allows us to better cover the range of states the policy might observe when deployed on the real system, allowing it to learn how to recover and continue the motion in those cases.</p>
<p>
<bold>Environment uncertainties.</bold> We introduce uncertainties in the environment in order to produce more robust motions when deployed on the real system. In this work, we are mainly concerned with randomization of the contact surface heights and friction.</p>
<p>
<bold>Time-independent task reward.</bold> We replace the time-based demonstration tracking reward with a time-independent, direct task reward. We have already noted the issues that can arise while trying to account for environment uncertainties while tracking a time-based demonstration trajectory. The policy is locked into trying to follow the specific time schedule regardless of the environment, whereas adapting it would produce much better recovery behavior. Switching to a time-independent reward, directly defining the task helps us deal with this.</p>
<p>Switching to this task reward has additional benefits. It allows us to directly optimize desirable aspects of the task, whereas the demonstration only needs to give us some nominal behavior on the task. The policy is free to change the behavior in a way that is needed to perform the task in the best possible way, without being penalized for not doing it in the same way as in the demonstration. We can also, as we will see later, produce varied behavior starting from a single demonstration by adapting this task reward.</p>
<p>
<bold>Regularization rewards.</bold> Finally, we add additional reward terms in this stage to further regularize the behavior of the learned policies. We aim to incentivize desirable aspects of policies in the tasks, like torque smoothness and smooth contact transitions (refer to <xref ref-type="sec" rid="s2-3">Section 2.3</xref> for details).</p>
</sec>
</sec>
<sec id="s3">
<title>3 Evaluation</title>
<sec id="s3-1">
<title>3.1 Tasks setup</title>
<p>We evaluate our approach on two different dynamic tasks on a quadruped robot: hopping and bounding. In the hopping task, the robot has to perform continuous hops on four legs, reaching a specific height and softly landing on the ground. Bounding consists of the robot performing oscillatory behavior around the pitch angle, with two full flight phases during a single period and each time making contact on exactly two legs (front or back). In both cases we start from a basic demonstration trajectory that provides state trajectories for the current task. Starting from that we apply our two-step training procedure in simulation to produce robust policies that we then test on a real robot.</p>
<p>We perform experiments on the open-source torque-controlled quadruped robot, Solo8 (<xref ref-type="bibr" rid="B8">Grimminger et al., 2020</xref>) (see <xref ref-type="fig" rid="F2">Figure 2</xref>), which is capable of very dynamic behaviors. For simulating the system we use PyBullet (<xref ref-type="bibr" rid="B3">Coumans and Bai, 2016</xref>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Examples of the robustness tests carried out on the quadruped robot Solo8 <bold>(A)</bold> Hopping on a surface comprised of a soft mattress and small blocks <bold>(B)</bold> Hopping with push recovery <bold>(C)</bold> Bounding on a surface comprised of a soft mattress and small blocks <bold>(D)</bold> Bounding with push recovery.</p>
</caption>
<graphic xlink:href="frobt-09-854212-g002.tif"/>
</fig>
<p>We use exactly the same training procedure in both cases, with the only differences arising from the need to allow for base rotation around one axis in the bounding task. This is a particular benefit of the approach we present here&#x2013;for a new task we only need a single new demonstration trajectory and a single simple reward term defining the task.</p>
<p>
<bold>DRL stage 1: Early termination.</bold> The only aspect in the first stage of training specific to the chosen tasks is how we perform early termination. We perform early stopping here based on the current tilt of the robot base (we increase the range appropriately for the bounding task), as well as when any part of the robot that is not the foot touches the ground.</p>
<p>
<bold>DRL stage 2: Initialization states.</bold> As noted in the method description, in the second stage of DRL training, we introduce a wider range of initialization states. In the two tasks investigated here, this consists of randomizing the initial height of the base of the robot, tilt of the base around <italic>x</italic>- and <italic>y</italic>-axis and randomness in the initial joint configuration. We preserve the early termination criteria from stage 1, only extending the range of allowed base tilt angles with the way it is increased in the initialization.</p>
<p>
<bold>DRL stage 2: Environment uncertainties.</bold> We also introduce uncertainties in the training environment. We randomize the ground position up and down in the range of [&#x2212;5&#xa0;cm, 5&#xa0;cm] (approximately 20% of the robot leg length). We also randomize the ground surface friction coefficient in the range [0.5, 1.0]. While we restrict ourselves only to this limited set of initial state and environment randomizations, as we will see in the later evaluations, this produces policies that are quite robust as they can also handle uneven ground or external perturbations.</p>
<p>
<bold>DRL stage 2: Reward structure (hopping).</bold> By using a demonstration trajectory to deal with exploration issues, we can define the individual task rewards to be very simple, without the need for any reward shaping.</p>
<p>For the hopping task we use the following reward<disp-formula id="e3">
<mml:math id="m11">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>h</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>We use the <italic>r</italic>
<sub>
<italic>hp</italic>
</sub> reward term to define the task<disp-formula id="e4">
<mml:math id="m12">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if&#x2009;</mml:mtext>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>base</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>otherwise</mml:mtext>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(4)</label>
</disp-formula>The reward at each timestep is proportional to the current height of the robot base (<italic>z</italic>
<sup>
<italic>base</italic>
</sup>), with constant weight <italic>k</italic>
<sub>
<italic>hp</italic>
</sub>. It is clipped to zero below a certain threshold <inline-formula id="inf9">
<mml:math id="m13">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, one that the robot can reach without leaving the ground. We additionally clip the value of this reward to be zero above a certain height threshold <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>base</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> to incentivize lower hops. We will additionally vary this threshold to produce policies with different hopping heights starting from the same demonstration trajectory.</p>
<p>We also introduce several reward terms to incentivize different desirable aspects in the resulting behavior. They reward the base to be close to its horizontal default posture (<italic>r</italic>
<sub>
<italic>ps</italic>
</sub>) and smooth contact transitions (<italic>r</italic>
<sub>
<italic>ct</italic>
</sub>) and torque smoothness (<italic>r</italic>
<sub>
<italic>ts</italic>
</sub>).</p>
<p>We reward the policy for being static in all the base dimensions (positions (<italic>x</italic>
<sup>
<italic>base</italic>
</sup>, <italic>y</italic>
<sup>
<italic>base</italic>
</sup>) and Euler angles (<inline-formula id="inf11">
<mml:math id="m15">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf12">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf13">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>)) except the one the motion is performed on (<italic>z</italic>-axis in this case). With <italic>k</italic>
<sub>
<italic>ps</italic>1</sub>, &#x2026; , <italic>k</italic>
<sub>
<italic>ps</italic>10</sub> being weight and scale constants.<disp-formula id="e5">
<mml:math id="m18">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>6</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>7</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>8</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>9</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>10</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</disp-formula>This term is crucial as it drives the policy to stay at the default posture as much as possible. Without it the policy could perform the task well in the simulation while always being close to falling over-which would likely happen when it was transferred to the real system.</p>
<p>The second key reward term asks for smooth contact transition (<italic>r</italic>
<sub>
<italic>ct</italic>
</sub>)<disp-formula id="e6">
<mml:math id="m19">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if&#x2009;</mml:mtext>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3e;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>otherwise</mml:mtext>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(6)</label>
</disp-formula>We do so by simply penalizing any contact force values <inline-formula id="inf14">
<mml:math id="m20">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> above a certain threshold <inline-formula id="inf15">
<mml:math id="m21">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, to penalize impact, with <italic>k</italic>
<sub>
<italic>ct</italic>
</sub> being a constant weight. Without this term we would have the feet hitting the surface hard on each landing-this is precisely what we observe in policies from DRL stage 1 where this reward term is not present. This is not the kind of behavior desired on the real system and these impacts can cause actual damage to the robot. These types of policies also transfer less well between simulation and real world. They can learn to perform the task well in simulation by generating hard impacts, but doing so exploits the weaknesses of the contact model in simulation, resulting in a poor performance when transferred to the real system. Smooth contact transitions enable a better transition between simulation and the real system. Thus, policies which incentivize those better transfer to the real system.</p>
<p>The third reward term (<italic>r</italic>
<sub>
<italic>ts</italic>
</sub>) prevents the policy to ask for a very quick change in the desired torque which is not realizable on the real robot with a limited control bandwidth<disp-formula id="e7">
<mml:math id="m22">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(7)</label>
</disp-formula>with <italic>&#x3c4;</italic> being the joint torque and <italic>k</italic>
<sub>
<italic>ts</italic>1</sub> and <italic>k</italic>
<sub>
<italic>ts</italic>2</sub> weight and scale constants respectively.</p>
<p>The final reward term <italic>r</italic>
<sub>
<italic>tt</italic>
</sub> is the same one we use in the first stage of training (as defined in (<xref ref-type="disp-formula" rid="e1">Eq. 1</xref>).</p>
<p>
<bold>DRL stage 2: Reward structure (bounding).</bold> For the bounding task, we only make changes to the parts of the reward defining the task:<disp-formula id="e8">
<mml:math id="m23">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>b</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>We define the task reward here in two parts, <italic>r</italic>
<sub>
<italic>bn</italic>
</sub> and <italic>r</italic>
<sub>
<italic>cc</italic>
</sub>. <italic>r</italic>
<sub>
<italic>bn</italic>
</sub> rewards the policy for being close to the path the demonstration takes in the <inline-formula id="inf16">
<mml:math id="m24">
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> space:<disp-formula id="e9">
<mml:math id="m25">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:munder>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf17">
<mml:math id="m26">
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> are points from the demonstration trajectory, <inline-formula id="inf18">
<mml:math id="m27">
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the current state of the base, with <italic>k</italic>
<sub>
<italic>bn</italic>
</sub> being a constant weight. We do this with no concept of time in this case, by just taking the distance to the closest point. This gives the policy freedom to perform the motion slower or faster, with different amplitude. We will later see that this results in a variety in bounding behaviors from repeated trainings, independent of the timing in the original demonstration.</p>
<p>The second part of the task reward, <italic>r</italic>
<sub>
<italic>cc</italic>
</sub>, is related to the contact state<disp-formula id="e10">
<mml:math id="m28">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>only&#x2009;front&#x2009;two&#x2009;legs&#x2009;in&#x2009;contact</mml:mtext>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>only&#x2009;back&#x2009;two&#x2009;legs&#x2009;in&#x2009;contact</mml:mtext>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>no&#x2009;legs&#x2009;in&#x2009;contact</mml:mtext>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>otherwise</mml:mtext>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(10)</label>
</disp-formula>with <italic>k</italic>
<sub>
<italic>cc</italic>
</sub> being a constant reward. It incentivizes the policy to, when making contact with the ground, only do so with front or back legs at the same time. Without anything to incentivize the policy to do this, we have observed DRL stage 1 policies reproducing the bounding motion while keeping all feet in contact with the ground. This reward part ensures appropriate contact states with flight phases in between.</p>
<p>We keep the other reward terms, ones used to incentivize desired aspects of the behavior, the same as in the hopping task (<italic>r</italic>
<sub>
<italic>tt</italic>
</sub> as defined in (<xref ref-type="disp-formula" rid="e1">Eq. 1</xref>), <italic>r</italic>
<sub>
<italic>ct</italic>
</sub>, <italic>r</italic>
<sub>
<italic>ts</italic>
</sub> as defined in (<xref ref-type="disp-formula" rid="e7">Eq. 7</xref>). The one change we make is to the reward incentivizing the robot base staying close to the default posture, <italic>r</italic>
<sub>
<italic>ps</italic>
</sub>
<disp-formula id="e11">
<mml:math id="m29">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>6</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>9</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>s</mml:mi>
<mml:mn>10</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
<p>We do not reward staying at default posture in the <inline-formula id="inf19">
<mml:math id="m30">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> direction in this case, as that is the angle the robot is moving around while bounding.</p>
<p>Details of the policy structure and values of all the reward function parameters can be found in <xref ref-type="table" rid="T1">Appendix Table A1</xref>.</p>
</sec>
<sec id="s3-2">
<title>3.2 Hopping task results</title>
<p>
<xref ref-type="fig" rid="F3">Figure 3</xref> shows results when the real robot is dropped from different heights to start the motion. We show the base height and estimated contact force for one of the legs as a function of time. As we do not have force sensors in the feet, we estimate the contact forces based on the torques the robot applies, using <inline-formula id="inf20">
<mml:math id="m31">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mspace width="0.17em"/>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.17em"/>
<mml:mi>&#x3c4;</mml:mi>
</mml:math>
</inline-formula>, where <italic>S</italic>
<sub>
<italic>i</italic>
</sub> and <italic>J</italic>
<sub>
<italic>i</italic>
</sub> are the joint selection matrix of the leg <italic>i</italic> and Jacobian of the foot <italic>i</italic>, respectively. Note that this estimation ignores the energy dissipated through damping of the robot structure and drive system. However, it provides an approximate measure of contact forces sufficient for the analyses of this paper. We further align the plots based on the later part of the motion&#x2013;the stable cycle the robot gets into.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Hopping experiments started from different initial heights. The top figure shows that after dropping the robot from a set of different heights ranging between 0.3 and 1&#xa0;m, the robot goes back to the nominal behavior of hopping at height of 0.5&#xa0;m in one or at most two cycles. The bottom figure shows that, dropping from different initial heights, the robot is able to adapt its landing such that the impact forces remain very low and almost invisible in the estimated force.</p>
</caption>
<graphic xlink:href="frobt-09-854212-g003.tif"/>
</fig>
<p>First, we can note that regardless of the drop height the robot goes into the same stable hopping cycle. What is more, it does so very quickly, as we can see all the individual rollouts matching after only two hops. We can also note here the benefits of time independence of the policy. It is what allows us to be able to start the motion from this large range of initial heights. It is also what enables this fast stabilization, as we can see that the two initial hops are on a different cycle&#x2013;one needed to stabilize the motion properly.</p>
<p>This test also highlights the general quality of contact interaction achieved with this approach. We can see that the impact forces, even on the highest drop (1&#xa0;m height), barely go over the force values for the stable hopping cycle (50&#xa0;cm height). This is purely learned behavior, as a result of impact penalties introduced in stage 2 of DRL training. It is not present in the demonstration and when we test DRL stage 1 policies on the real system high impact forces are generated and the policies are very fragile. Smooth contact transitions can also be observed in the accompanying video.</p>
<p>Further, the behavior is robust to uneven terrain and external pushes although this was never explicitly trained for. The robot is able to recover from significant tilt of the base arising from either external pushes or landing on an uneven surface (<xref ref-type="fig" rid="F2">Figures 2A,B</xref>). More extensive examples of recovery behavior can be seen in the accompanying video.</p>
<p>To present quantitatively the performance of the policy on a random uneven terrain, we scattered different objects with heights ranging from 1 to 5&#xa0;cm and executed a jumping policy on this surface. As it can be seen in <xref ref-type="fig" rid="F4">Figure 4</xref>, while the robot feet land on the ground in different heights, the policy manages to keep the robot base stable and perform jumps close to the desired height which is 50&#xa0;cm.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Three different execution of a jumping policy in the real world on a random uneven terrain with 5&#xa0;cm height uncertainty. As we can see, different feet land on different height in each jump, but the policy manages to keep the jumping height within a certain bound.</p>
</caption>
<graphic xlink:href="frobt-09-854212-g004.tif"/>
</fig>
<p>Finally, we demonstrate the variety of robust behaviors that can be optimized from the same demonstration by doing repeated trainings with different values of the <inline-formula id="inf21">
<mml:math id="m32">
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>base</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> threshold in the hopping task reward. With this simple change in the task reward, starting from one demonstration, we can produce hopping behaviors at different heights. Examples of this can be seen in the accompanying video.</p>
</sec>
<sec id="s3-3">
<title>3.3 Bounding task results</title>
<p>In <xref ref-type="fig" rid="F5">Figure 5</xref> we show results for a test where we drop the robot from different angles to start the motion. We perform the same test for two different final DRL stage 2 policies for this task. We can see that the policies can handle a wide range of initial base angles&#x2013;around 35&#xb0; in both directions. What is more, as was the case with the hopping task, we can see that here as well all the initializations end up in the same stable motion cycle.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Two different bounding behaviors on the robot with different initial conditions. Starting with a wide range of initial angles for the base in <italic>y</italic>-direction (roughly between &#x2212;35 and 35&#xa0;deg), the robot quickly converges back to the desired behavior.</p>
</caption>
<graphic xlink:href="frobt-09-854212-g005.tif"/>
</fig>
<p>The bounding motion exhibits similar robustness to uneven terrain and external perturbations as the hopping motion (<xref ref-type="fig" rid="F2">Figures 2C,D</xref>). Same as with the hopping task, the policies rely on their knowledge of how to handle a varied set of base states to recover from anything that arises from these conditions even through it was not explicitly trained for.</p>
<p>In this task, we would also like to demonstrate the variety of behaviors we can generate from one single demonstration. Unlike in the hopping task, where we made simple changes in the task reward to achieve different jumping heights, we instead give more freedom to the task reward and examine the variety of produced behaviors. As noted in the task reward definition, the policy has the freedom to produce slower or faster bounds with smaller or larger amplitudes. As seen in the accompanying video we arrive at a variety of bounding behaviors in this way, starting from the same initial demonstration trajectory.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>
<bold>Simplicity and generality of the approach.</bold> One of the main benefits of the approach is its generality and simplicity. The only elements needed for each new task, as can be seen from the two tasks studied here, are a single demonstration trajectory and a direct straightforward task reward.</p>
<p>The demonstration does not need to provide ideal performance on the task, as we optimize task performance in the later stage of training. It simply needs to provide a sequence of states, and not necessarily actions, that enable a decent performance. Providing only states is also simpler in cases where it is not trivial to calculate the exact forces to realize a particular motion.</p>
<p>As for the task reward, reward shaping is not needed, as the demonstration resolves any exploration issues that could occur as a result of sparse reward signal. We can directly reward aspects of the task of interest. We keep reward terms other than the task reward as general as possible, encoding characteristics of general good robot behavior. We expect those to remain constant across a varied range of tasks.</p>
<p>
<bold>Learning from scratch.</bold> In this paper, we proposed a framework to learn a policy for dynamic legged locomotion that is transferable to the real world. One might argue that for some locomotion tasks, it is possible to generate the policy without the need for the demonstration, e.g. <xref ref-type="bibr" rid="B15">Lee et al. (2020)</xref>. However, for the two tasks examined in this paper, with given task rewards, we failed to find successful policies when training from scratch, with policies being unable to learn any notion of the task, even in simplest conditions. It is particularly difficult to learn from scratch in highly dynamic tasks with flight phase which has been the main focus of this paper. In such case, the robot crashes into the ground repeatedly, ending the episode, without providing any information for the learning algorithm on how to fix it. Additionally, the majority of recent successful approaches for doing RL in locomotion tasks use demonstrations in some way, e.g., <xref ref-type="bibr" rid="B31">Xie et al. (2020)</xref>; <xref ref-type="bibr" rid="B20">Peng et al. (2020)</xref>. This suggests that learning from scratch is often not viable and using demonstrations is one of the predominant solutions being utilized.</p>
<p>
<bold>Reinforcement learning perspective.</bold> From the reinforcement learning perspective our approach presents a simple and effective way to deal with exploration issues in robotic tasks. We also remove the trajectory tracking reward in the second stage of our training, so, as seen in our experiments, we are able to change the policy away from the exact behavior defined in the demonstration.</p>
<p>
<bold>Trajectory optimization perspective.</bold> From the trajectory optimization perspective, our approach proposes a systematic way to consider different types of uncertainty and find a robust control policy for robotic tasks, especially those with contact. Furthermore, our approach caches the solution of a model-based approach for future use and eliminates the need for re-generating repetitive motions. We believe this is a practical way to combine the strength of trajectory optimization and reinforcement learning for continuous control problems; 1) Trajectory optimization is used to generate a desired behavior efficiently to achieve the task at hand 2) different types of realistic uncertainties are easily added to the simulation, e.g. contact timing uncertainty, and DRL is used to produce a robust feedback policy.</p>
</sec>
<sec id="s5">
<title>5 Conclusion</title>
<p>In this work, we presented a general approach for going from trajectories optimized using TO to robust learned policies on a real robot. We showed how we can start from a single trajectory and arrive at a robust policy that can be directly deployed on a real robot, without any need for additional training. Through extensive tests on a real quadruped robot, we demonstrated significant robustness in the behaviors produced by our approach. Importantly, we do so in setups, uneven ground and external pushes, for which the robot was not explicitly trained for. All this gives hope that such approaches could be used across varied robotic tasks to simply generate robust policies to be used on real hardware, bridging the gap between trajectory optimization and reinforcement learning in such tasks.</p>
<p>In future work, we would like to take more advantage of model-based approaches to make our framework more efficient. We intend to study how replacing the first stage of our algorithm with a form of Behavior Cloning (<xref ref-type="bibr" rid="B23">Pomerleau, (1988)</xref>) to imitate a whole-body MPC policy can improve efficiency.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s11">Supplementary Material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>MB, MK, and LR designed research; MB performed numerical simulations; MK prepared the demonstrations; MB and MK performed hardware experiments; MB, MK, and LR analyzed the results and wrote the paper.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the New York University, the European Union&#x2019;s Horizon 2020 research and innovation program (grant agreement 780684) and the National Science Foundation (grants 1825993, 1932187 and 1925079).</p>
</sec>
<ack>
<p>We would like to thank the contributors of the Open Dynamic Robot Initiative (ODRI) for the development of the hardware, electronics and the low-level software for controlling the robot Solo.</p>
</ack>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frobt.2022.854212/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frobt.2022.854212/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Video1.MP4" id="SM1" mimetype="application/MP4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bogdanovic</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Khadiv</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Righetti</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Learning variable impedance control for contact sensitive tasks</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>5</volume>, <fpage>6129</fpage>&#x2013;<lpage>6136</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2020.3011379</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carpentier</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mansard</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Multicontact locomotion of legged robots</article-title>. <source>IEEE Trans. Robot.</source> <volume>34</volume>, <fpage>1441</fpage>&#x2013;<lpage>1460</lpage>. <pub-id pub-id-type="doi">10.1109/tro.2018.2862902</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="web">
<comment>[Dataset]</comment> <person-group person-group-type="author">
<name>
<surname>Coumans</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2016&#x2013;2020</year>). <article-title>Pybullet, a python module for physics simulation for games, robotics and machine learning</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://pybullet.org">http://pybullet.org</ext-link>
</comment>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Drnach</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Robust trajectory optimization over uncertain terrain with stochastic complementarity</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>6</volume>, <fpage>1168</fpage>&#x2013;<lpage>1175</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2021.3056064</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fankhauser</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gehring</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bloesch</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hoepflinger</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Siegwart</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Reinforcement learning of single legged locomotion</article-title>,&#x201d; in <conf-name>2013 IEEE/RSJ International Conference on Intelligent Robots and Systems</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>188</fpage>&#x2013;<lpage>193</lpage>. </citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Gangapurwala</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Geisert</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Orsolino</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Fallon</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Havoutis</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Real-time trajectory adaptation for quadrupedal locomotion using deep reinforcement learning</article-title>,&#x201d; in <conf-name>2021 IEEE International Conference on Robotics and Automation (ICRA) (IEEE)</conf-name>, <fpage>5973</fpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gangapurwala</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Geisert</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Orsolino</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Fallon</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Havoutis</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Rloc: Terrain-aware legged locomotion using reinforcement learning and optimal control</article-title>. <comment>
<italic>arXiv preprint arXiv:2012.03094.</italic>
</comment> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grimminger</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Meduri</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Khadiv</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Viereck</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>W&#xfc;thrich</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Naveau</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>An open torque-controlled modular robot architecture for legged locomotion research</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>5</volume>, <fpage>3650</fpage>&#x2013;<lpage>3657</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2020.2976639</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hammoud</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Khadiv</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Righetti</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Impedance optimization for uncertain contact interactions through risk sensitive optimal control</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>6</volume>, <fpage>4766</fpage>&#x2013;<lpage>4773</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2021.3068951</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hornby</surname>
<given-names>G. S.</given-names>
</name>
<name>
<surname>Takamura</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yokono</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hanagata</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Yamamoto</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Fujita</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2000</year>). &#x201c;<article-title>Evolving robust gaits with aibo</article-title>,&#x201d; in <conf-name>Proceedings 2000 iCRA. millennium conference. iEEE international conference on robotics and automation. symposia proceedings (cat. no. 00CH37065) (IEEE)</conf-name>, <fpage>3040</fpage>&#x2013;<lpage>3045</lpage>. </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hwangbo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dosovitskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bellicoso</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tsounis</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Koltun</surname>
<given-names>V.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Learning agile and dynamic motor skills for legged robots</article-title>. <source>Sci. Robot.</source> <volume>4</volume>, <fpage>eaau5872</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.aau5872</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ijspeert</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Nakanishi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Schaal</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2002</year>). <source>Learning attractor landscapes for learning motor primitives</source>. (<publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>Tech. rep.</publisher-name>) </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kalashnikov</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Irpan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pastor</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ibarz</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Herzog</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Jang</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation</article-title>. <comment>
<italic>arXiv preprint arXiv:1806.10293.</italic>
</comment> </citation>
</ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kohl</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Stone</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2004</year>). &#x201c;<article-title>Policy gradient reinforcement learning for fast quadrupedal locomotion</article-title>,&#x201d; in <conf-name>IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA&#x2019;04. 2004 (IEEE)</conf-name>, <fpage>2619</fpage>&#x2013;<lpage>2624</lpage>. </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hwangbo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wellhausen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Koltun</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Learning quadrupedal locomotion over challenging terrain</article-title>. <source>Sci. Robot.</source> <volume>5</volume>, <fpage>eabc5986</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.abc5986</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>X. B.</given-names>
</name>
<name>
<surname>Abbeel</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Berseth</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Reinforcement learning for robust parameterized locomotion control of bipedal robots</article-title>. <comment>
<italic>arXiv preprint arXiv:2103.14295.</italic>
</comment> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morimoto</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Atkeson</surname>
<given-names>C. G.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Nonparametric representation of an approximated poincar&#xe9; map for learning biped locomotion</article-title>. <source>Auton. Robots</source> <volume>27</volume>, <fpage>131</fpage>&#x2013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1007/s10514-009-9133-z</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Morimoto</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Nakanishi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Endo</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Atkeson</surname>
<given-names>C. G.</given-names>
</name>
<name>
<surname>Zeglin</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Poincare-map-based reinforcement learning for biped walking</article-title>,&#x201d; in <conf-name>Proceedings of the 2005 IEEE International Conference on Robotics and Automation (IEEE)</conf-name>, <fpage>2381</fpage>. </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>X. B.</given-names>
</name>
<name>
<surname>Abbeel</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>van de Panne</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deepmimic: Example-guided deep reinforcement learning of physics-based character skills</article-title>. <source>ACM Trans. Graph.</source> <volume>37</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/3197517.3201311</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>X. B.</given-names>
</name>
<name>
<surname>Coumans</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>T.-W.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Learning agile robotic locomotion skills by imitating animals</article-title>. <comment>
<italic>arXiv preprint arXiv:2004.00784.</italic>
</comment> </citation>
</ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>X. B.</given-names>
</name>
<name>
<surname>van de Panne</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Learning locomotion skills using deeprl: Does the choice of action space matter?</article-title>&#x201d; in <conf-name>Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation (ACM)</conf-name>, <fpage>12</fpage>. </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peters</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Schaal</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Reinforcement learning of motor skills with policy gradients</article-title>. <source>Neural Netw.</source> <volume>21</volume>, <fpage>682</fpage>&#x2013;<lpage>697</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2008.02.003</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pomerleau</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>Alvinn: An autonomous land vehicle in a neural network</article-title>. <source>Adv. neural Inf. Process. Syst.</source> <volume>1</volume>, <fpage>305</fpage>&#x2013;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.5555/89851.89891</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ponton</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Khadiv</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Meduri</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Righetti</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Efficient multicontact pattern generation with sequential convex approximations of the centroidal dynamics</article-title>. <source>IEEE Trans. Robot.</source> <volume>1</volume>, <fpage>1661</fpage>&#x2013;<lpage>1679</lpage>. <pub-id pub-id-type="doi">10.1109/TRO.2020.3048125</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schaal</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Learning from demonstration</article-title>. <source>Adv. neural Inf. Process. Syst.</source>, <fpage>1040</fpage>&#x2013;<lpage>1046</lpage>. </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schulman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wolski</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Dhariwal</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Radford</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Klimov</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2017</year>).<article-title>Proximal policy optimization algorithms</article-title>. <comment>
<italic>arXiv preprint arXiv:1707.06347.</italic>
</comment> </citation>
</ref>
<ref id="B27">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Siekmann</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Warila</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fern</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hurst</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Blind bipedal stair traversal via sim-to-real reinforcement learning</article-title>,&#x201d; in <conf-name>Robotics: Science and Systems</conf-name>. </citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tedrake</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>T. W.</given-names>
</name>
<name>
<surname>Seung</surname>
<given-names>H. S.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Learning to walk in 20 minutes</article-title>,&#x201d; in <conf-name>Proceedings of the Fourteenth Yale Workshop on Adaptive and Learning Systems (Beijing)</conf-name>, <fpage>1939</fpage>&#x2013;<lpage>1412</lpage>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Theodorou</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Buchli</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Schaal</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>A generalized path integral control approach to reinforcement learning</article-title>. <source>J. Mach. Learn. Res.</source> <volume>11</volume>, <fpage>3137</fpage>&#x2013;<lpage>3181</lpage>. <pub-id pub-id-type="doi">10.5555/1756006.1953033</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Winkler</surname>
<given-names>A. W.</given-names>
</name>
<name>
<surname>Bellicoso</surname>
<given-names>C. D.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Buchli</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Gait and trajectory optimization for legged systems through phase-based end-effector parameterization</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>3</volume>, <fpage>1560</fpage>&#x2013;<lpage>1567</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2018.2798285</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Clary</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Dao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Morais</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hurst</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Panne</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Learning locomotion skills for cassie: Iterative design and sim-to-real</article-title>, in <conf-name>Conference on Robot Learning (PMLR)</conf-name>, <fpage>317</fpage>. </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Da</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Babich</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Garg</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>van de Panne</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>Glide: Generalizable quadrupedal locomotion in diverse environments with a centroidal model</article-title>. <comment>
<italic>arXiv preprint arXiv:2104.09771.</italic>
</comment> </citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Da</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>van de Panne</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Babich</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Garg</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021b</year>). &#x201c;<article-title>Dynamics randomization revisited: A case study for quadrupedal locomotion</article-title>,&#x201d; in <conf-name>2021 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>, <conf-loc>Xi'an, China</conf-loc>, (<publisher-name>IEEE</publisher-name>), <fpage>4955</fpage>&#x2013;<lpage>4961</lpage>. </citation>
</ref>
</ref-list>
<app-group>
<app>
<title>Appendix</title>
<table-wrap id="T1" position="float">
<label>TABLE A1</label>
<caption>
<p>Parameter values used in the training of the policies.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="2" align="left">Policy network (fully-connected) parameter</th>
<th colspan="2" align="left">Reward parameters (cont.)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Number of layers</td>
<td align="left">2</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>hp</italic>
</sub>
</td>
<td align="char" char=".">0.5</td>
</tr>
<tr>
<td align="left">Units per layer</td>
<td align="left">64</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>1</sub>
</td>
<td align="char" char=".">1.0</td>
</tr>
<tr>
<td align="left">Nonlinearity</td>
<td align="left">tanh</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>2</sub>
</td>
<td align="char" char=".">0.05</td>
</tr>
<tr>
<td align="left">Output dimension</td>
<td align="left">8</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>3</sub>
</td>
<td align="char" char=".">1.0</td>
</tr>
<tr>
<td align="left">Output nonlinearity</td>
<td align="left">tanh</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>4</sub>
</td>
<td align="char" char=".">0.05</td>
</tr>
<tr>
<td colspan="2" align="left">Reward parameters</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>5</sub>
</td>
<td align="char" char=".">1.0</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>tt</italic>
</sub>
</td>
<td align="left">2.25</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>6</sub>
</td>
<td align="char" char=".">0.05</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>1</sub> (hopping)</td>
<td align="left">0.3</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>7</sub>
</td>
<td align="char" char=".">1.0</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>1</sub> (bounding)</td>
<td align="left">0.4</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>8</sub>
</td>
<td align="char" char=".">0.05</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>2</sub>
</td>
<td align="left">5.0</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>9</sub>
</td>
<td align="char" char=".">1.0</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>3</sub> (hopping)</td>
<td align="left">0.1</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ps</italic>10</sub>
</td>
<td align="char" char=".">0.05</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>3</sub> (bounding)</td>
<td align="left">0.0</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ct</italic>
</sub>
</td>
<td align="char" char=".">0.2</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>4</sub>
</td>
<td align="left">0.2</td>
<td align="left">
<inline-formula id="inf1">
<mml:math id="m1">
<mml:msubsup>
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>
</td>
<td align="char" char=".">50.0</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>5</sub> (hopping)</td>
<td align="left">0.1</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>ts</italic>
</sub>
</td>
<td align="char" char=".">0.02</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>5</sub> (bounding)</td>
<td align="left">0.4</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>bn</italic>
</sub>
</td>
<td align="char" char=".">8.0</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>6</sub>
</td>
<td align="left">5.0</td>
<td align="left">
<italic>k</italic>
<sub>
<italic>cc</italic>
</sub>
</td>
<td align="char" char=".">0.5</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>7</sub> (hopping)</td>
<td align="left">0.1</td>
<td colspan="2" align="left">Termination angles</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>7</sub> (bounding)</td>
<td align="left">0.0</td>
<td align="left">
<inline-formula id="inf2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>
</td>
<td align="char" char=".">0.2</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>8</sub>
</td>
<td align="left">0.2</td>
<td align="left">
<inline-formula id="inf3">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> (hopping)</td>
<td align="char" char=".">0.2</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>9</sub> (hopping)</td>
<td align="left">0.3</td>
<td align="left">
<inline-formula id="inf4">
<mml:math id="m4">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> (bounding)</td>
<td align="char" char=".">0.6</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>9</sub> (bounding)</td>
<td align="left">0.2</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>10</sub>
</td>
<td align="left">0.5</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>11</sub> (hopping)</td>
<td align="left">0.1</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>11</sub> (bounding)</td>
<td align="left">0.2</td>
</tr>
<tr>
<td align="left">
<italic>k</italic>
<sub>
<italic>ti</italic>12</sub>
</td>
<td align="left">0.05</td>
</tr>
</tbody>
</table>
</table-wrap>
</app>
</app-group>
</back>
</article>