<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1102854</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2023.1102854</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Soft-body dynamics induces energy efficiency in undulatory swimming: A deep learning study</article-title>
<alt-title alt-title-type="left-running-head">Li et&#xa0;al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2023.1102854">10.3389/frobt.2023.1102854</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Guanda</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1970659/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shintake</surname>
<given-names>Jun</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/707609/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Hayashibe</surname>
<given-names>Mitsuhiro</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/84197/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Neuro-Robotics Lab</institution>, <institution>Department of Robotics</institution>, <institution>Graduate School of Engineering</institution>, <institution>Tohoku University</institution>, <addr-line>Sendai</addr-line>, <country>Japan</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Mechanical and Intelligent Systems Engineering</institution>, <institution>The University of Electro-Communications</institution>, <addr-line>Chofu</addr-line>, <country>Japan</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/660410/overview">Concepci&#xf3;n A. Monje</ext-link>, Universidad Carlos III de Madrid, Spain</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1152521/overview">Francesco Giorgio-Serchi</ext-link>, University of Edinburgh, United Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/801415/overview">Lisbeth Mena</ext-link>, Universidad Carlos III de Madrid, Spain</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Mitsuhiro Hayashibe, <email>mitsuhiro.hayashibe.e6@tohoku.ac.jp</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Soft Robotics, a section of the journal Frontiers in Robotics and AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1102854</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Li, Shintake and Hayashibe.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Li, Shintake and Hayashibe</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Recently, soft robotics has gained considerable attention as it promises numerous applications thanks to unique features originating from the physical compliance of the robots. Biomimetic underwater robots are a promising application in soft robotics and are expected to achieve efficient swimming comparable to the real aquatic life in nature. However, the energy efficiency of soft robots of this type has not gained much attention and has been fully investigated previously. This paper presents a comparative study to verify the effect of soft-body dynamics on energy efficiency in underwater locomotion by comparing the swimming of soft and rigid snake robots. These robots have the same motor capacity, mass, and body dimensions while maintaining the same actuation degrees of freedom. Different gait patterns are explored using a controller based on grid search and the deep reinforcement learning controller to cover the large solution space for the actuation space. The quantitative analysis of the energy consumption of these gaits indicates that the soft snake robot consumed less energy to reach the same velocity as the rigid snake robot. When the robots swim at the same average velocity of 0.024&#xa0;m/s, the required power for the soft-body robot is reduced by 80.4% compared to the rigid counterpart. The present study is expected to contribute to promoting a new research direction to emphasize the energy efficiency advantage of soft-body dynamics in robot design.</p>
</abstract>
<kwd-group>
<kwd>soft robot</kwd>
<kwd>energy efficiency</kwd>
<kwd>underwater robot</kwd>
<kwd>snake robot</kwd>
<kwd>deep reinforcement learning</kwd>
</kwd-group>
<contract-num rid="cn001">20H05458 21H00324</contract-num>
<contract-sponsor id="cn001">Japan Society for the Promotion of Science<named-content content-type="fundref-id">10.13039/501100001691</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Soft robotics, which is an emerging scientific field creating robots based on compliant materials, promises numerous applications such as object manipulation and human-robot interaction in industry, search and rescue activities in the natural environment, and rehabilitation in the medical <xref ref-type="bibr" rid="B29">Rus and Tolley, 2015</xref>; <xref ref-type="bibr" rid="B28">Rich&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B32">Shintake&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B9">Cianchetti&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B16">Jumet&#xa0;et&#xa0;al., 2022</xref>. Compared to traditional &#x201c;rigid&#x201d; robots, the compliant body of the soft robots is said to have better adaptability to the surrounding environment <xref ref-type="bibr" rid="B28">Rich&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B1">Aracri&#xa0;et&#xa0;al., 2021</xref>.</p>
<p>In this context, underwater soft robots, especially those based on biomimetics, are one of the promising applications in soft robotics <xref ref-type="bibr" rid="B1">Aracri&#xa0;et&#xa0;al., 2021</xref>. These biomimetic underwater robots exploit active deformations of their continuum body made of soft actuators. They have morphologies similar to those of their natural counterparts, such as fish <xref ref-type="bibr" rid="B14">Hubbard&#xa0;et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B17">Katzschmann&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B32">Shintake&#xa0;et&#xa0;al., 2018</xref>, snake <xref ref-type="bibr" rid="B8">Christianson&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B25">Nguyen and Ho, 2022</xref>, jellyfish <xref ref-type="bibr" rid="B34">Villanueva&#xa0;et&#xa0;al., 2011</xref>; <xref ref-type="bibr" rid="B12">Frame&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B7">Cheng&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B27">Ren&#xa0;et&#xa0;al., 2019</xref>, ray <xref ref-type="bibr" rid="B26">Park&#xa0;et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B21">Li&#xa0;et&#xa0;al., 2017</xref>, <xref ref-type="bibr" rid="B19">Li&#xa0;et&#xa0;al., 2021a</xref>, and flagellate <xref ref-type="bibr" rid="B2">Armanini&#xa0;et&#xa0;al., 2021</xref>.</p>
<p>Biomimetic soft underwater robots mimic the motion of aquatic animals with the expectation to achieve efficient swimming observed in nature. In this sense, compared to rigid underwater robots, the continuous deformation of the soft robots is expected to significantly improve the energy efficiency of swimming locomotion. Evidence suggests that the energy efficiency of soft-bodied robots improves with an augmented propulsive force due to fluid-inertial effects, as demonstrated in <xref ref-type="bibr" rid="B13">Giorgio-Serchi and Weymouth (2017)</xref>. However, there has been no quantitative comparative study on the energy efficiency between soft and rigid robots under the constraints of having the same mass, size, shape, and input.</p>
<p>In order to shed light on the problem, in this study we investigate the effect of compliance on the energy efficiency in a specific swimming mode through a control experiment performed in a simulation environment. Different aspects can be considered for the simulation. These include, for instance, size, shape, weight, type of actuator, and material and mechanical properties. As the first comparative study, we focus on the dynamics of underwater locomotion as the main aspect. For the swimming mode, we employ anguilliform, a swimming mode enabled by the undulation of the snake-like slender body <xref ref-type="bibr" rid="B31">Sfakiotakis&#xa0;et&#xa0;al., 1999</xref>. The thin structure of the slender body is expected to simplify the simulation model and subsequent analysis. We compare the energy efficiency between soft-body and rigid-body snake robots with the same mass, size, shape, and motor capacity.</p>
<p>To precisely model the actuation and subsequent deformation of the body, we employ dielectric elastomer actuators (DEAs), due to their thin feature applicable to anguilliform swimming <xref ref-type="bibr" rid="B8">Christianson&#xa0;et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B20">Li&#xa0;et&#xa0;al., 2021b</xref>. Accordingly, we designed a soft snake robot and a rigid snake robot in a physical simulation environment, as shown in <xref ref-type="fig" rid="F1">Figure&#xa0;1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Morphological structures of a soft snake robot (left) and a rigid snake robot (right) in the simulation environment.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g001.tif"/>
</fig>
<p>The frequency at which the elasticity of the actuator oscillates plays a significant role in determining its overall efficiency. There is a significant body of evidence that suggests that the maximum level of efficiency is achieved when the actuator is operating at its natural resonant frequency <xref ref-type="bibr" rid="B5">Bujard&#xa0;et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B37">Zhong&#xa0;et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B36">Zheng&#xa0;et&#xa0;al., 2022</xref>. This highlights the importance of the elasticity of the actuator in terms of energy efficiency. However, one of the challenges in utilizing this principle is the difficulty in accurately determining the natural frequency of soft actuators. To address this issue, our work employs a combination of deep reinforcement learning and grid search methods to identify the most energy-efficient gait for snake robots. Through the use of these advanced techniques, we are able to overcome the limitations of traditional methods and make sure we can find the optimized solution for both the soft and rigid snake robots.</p>
</sec>
<sec id="s2">
<title>2 Simulation method</title>
<p>Current simulation methods for soft robots are limited, particularly for soft snake underwater robots composed of multiple soft actuators. The coupling effect between actuators, segments, and the surrounding water makes it more difficult to simulate underwater snake robots. Researchers have provided a method for underwater soft-robot simulation in the <xref ref-type="bibr" rid="B10">Du&#xa0;et&#xa0;al., 2021</xref>. However, the proposed method has only been validated on one degree of freedom (DoF) robot and requires the use of a real robot to collect data, which does not match the requirements of our experiments.</p>
<p>In our previous study <xref ref-type="bibr" rid="B20">Li&#xa0;et&#xa0;al., 2021b</xref>, we proposed an approximate simulation method for soft robot underwater locomotion in MuJoCo. MuJoCo is a physics engine that simulates multi-joint dynamics with contact and is widely used in robotics and biomechanics <xref ref-type="bibr" rid="B33">Todorov&#xa0;et&#xa0;al., 2012</xref>. In the previous study, we confirmed the accuracy of this simulation method by comparing a simulated robot with a real DEA fish robot. The result showed that the simulated soft robot had almost the same dynamics as a real soft robot.</p>
<p>As such, in this study, we use MuJoCo to design the soft and rigid snake robots with the same size (length 27.0&#xa0;cm, width 1.5&#xa0;cm, and thickness 0.5&#xa0;cm), mass (7.93&#xa0;g), and output torque of actuators (12&#xa0;Nm). The models of these robots are shown in <xref ref-type="fig" rid="F2">Figure&#xa0;2</xref>. The soft snake robot has five segments, every of which actively deforms independently with the simulated actuation behavior of a DEA. Each segment is split into a connected series of thin-sliced elements to realize the deformable attribution of the DEA, which are alternately connected by active joints and passive joints in the rotation. The active joints were controlled by the output torque of the simulated motor set on it, which was used to simulate the output characteristics of the DEA. All active joints in one segment are simultaneously controlled by the same input signal. The characteristics of the DEAs can be adjusted by tuning the attributes of passive joints.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The simulation model of snake robots in MuJoCo. <bold>(A)</bold> The rigid snake robot. <bold>(B)</bold> The soft snake robot. One DoF per segment for both models.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g002.tif"/>
</fig>
<p>There are 12 active and 11 passive joints per segment for the soft-body case. However, the actuation is still one DoF per segment as the shared same control input is applied to the active joints. One segment has only one DoF, resulting in curved body deformation. Since the soft snake robot has five segments, there are five DoFs in total for the actuation space. The range of rotation of the active joints is set to [&#x2212;3&#xb0;, 3&#xb0;]. The output torque range of the active joint is set to [&#x2212;1, 1] Nm.</p>
<p>The rigid snake robot has six segments with five joints, as shown in <xref ref-type="fig" rid="F2">Figure&#xa0;2</xref>. A simulated motor is placed on each of the five joints. Each segment has one DoF by bending at the joint. In total, there are five DoFs for the actuation space, identical to the soft snake robot. The robot is controlled by changing the output torque of the motors. To ensure a fair comparison, the actuation range of the rigid snake robot&#x2019;s joints was set to [&#x2212;36&#xb0;, 36&#xb0;] to match the total actuation range of the soft snake robot, which has 12 active joints per segment, each with an actuation range of [&#x2212;3&#xb0;, 3&#xb0;] (&#x2212;36&#xb0;, 36&#xb0; in total). The range of the output torque of the active joint is set to [&#x2212;12, 12] Nm. Twelve times more torque capacity is set for the joint torque for the rigid robot to obtain the equivalent torque capacity of the soft robot, which has 12 distributed actuators with a 1&#xa0;Nm torque range.</p>
<p>The density and viscosity of the liquid in the simulation environment are set to 1,000&#xa0;kg/m<sup>3</sup> and 0.0009&#xa0;Pa&#x22c5; s, respectively. The simulated frequency is 100&#xa0;Hz, and the control frequency is 50&#xa0;Hz.</p>
</sec>
<sec id="s3">
<title>3 Controller design</title>
<p>We present two methods for controlling the underwater locomotion of snake robots. The methods are used to find the gaits of snake robots in water; then, the average velocity and the required average power are verified to evaluate the energy efficiency of the snake robot. The gait equation controller is model-based. It continuously generates joint torques for the snake robot according to a pre-defined equation. The DRL controller is data-based. During the training process of the DRL controller, the snake robot could explore different gaits beyond the limitations of the gait equation controller.</p>
<sec id="s3-1">
<title>3.1 Gait equation controller</title>
<p>The gait equation controller we used to drive the snake robots is modelled as<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>&#x3c4;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>sin</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mi>&#x3c6;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>&#x3c4;</italic>(<italic>n</italic>, <italic>t</italic>) represents the output torque of the motor on the <italic>n</italic>th joint of the snake robots at moment <italic>t</italic>. <italic>A</italic> denotes torque amplitude. <italic>&#x3c9;</italic> and <italic>&#x3d5;</italic> denote spatial and temporal frequencies, respectively.</p>
<p>We can make the snake robot swim in different gaits by changing the values of <italic>A</italic>, <italic>&#x3c9;</italic> and <italic>&#x3d5;</italic>. Therefore, we used the grid search method to determine the optimal parameters of the gait equation controller. The parameters and interval ranges of the grid search are listed in <xref ref-type="table" rid="T1">Table&#xa0;1</xref>. The selection of the parameters and interval ranges was done through a process of trial and error. The setting the amplitude or temporal frequency too high resulted in unstable simulations and too low interval range would increase the simulation time. We aimed to find the best balance between including a wide range of motion patterns and maintaining stability and efficiency in the simulation.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>The parameters for the grid search.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Paramaters</th>
<th align="left">Values</th>
<th align="left">Descriptions</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">
<italic>&#x3c9;</italic>
</td>
<td align="left">0.1, 0.6, 1.1, 1.6, 2.1, 2.6</td>
<td rowspan="3" align="left">Temporal frequency</td>
</tr>
<tr>
<td align="left">3.1, 3.6, 4.1, 4.6, 5.1, 5.6</td>
</tr>
<tr>
<td align="left">6.1, 6.6, 7.1, 7.6, 8.1, 8.6</td>
</tr>
<tr>
<td rowspan="4" align="left">
<italic>A</italic>
</td>
<td align="left">0.1, 0.2, 0.3, 0.4, 0.5, 0.6</td>
<td rowspan="4" align="left">Amplitude</td>
</tr>
<tr>
<td align="left">0.7, 0.8, 0.9, 1.0, 1.1, 1.2</td>
</tr>
<tr>
<td align="left">1.3, 1.4, 1.5, 1.6, 1.7, 1.8</td>
</tr>
<tr>
<td align="left">1.9, 2.0</td>
</tr>
<tr>
<td rowspan="2" align="left">
<italic>&#x3d5;</italic>
</td>
<td align="left">18, 36, 54, 72, 90, 108</td>
<td rowspan="2" align="left">Spatial&#xa0;frequency (in degrees)</td>
</tr>
<tr>
<td align="left">126, 144, 162, 180</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>3.2 DRL controller</title>
<p>The DRL controller is a neural network controller trained by the deep reinforcement learning algorithm. When using this method, robots learn the target skill through interaction with the environment <xref ref-type="bibr" rid="B24">Mnih&#xa0;et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B3">Arulkumaran&#xa0;et&#xa0;al., 2017</xref> In this process, the robot collects a large number of state-action pairs to evaluate the data according to the reward function. Through continuous iterative training, the robot can discover better state-action pairs until training converges.</p>
<p>The state is represented in the data collected by the robot from the environment, such as the robot&#x2019;s joint position and angular velocity, which correspond to the input layer of the neural network. Action is the command used to drive the robot actuators and corresponds to the output layer of the neural network.</p>
<p>Deep reinforcement learning is primarily used in robotics for locomotion control <xref ref-type="bibr" rid="B15">Jangir&#xa0;et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B18">Lee&#xa0;et&#xa0;al., 2020</xref> and navigation <xref ref-type="bibr" rid="B11">Fan&#xa0;et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B35">Zeng&#xa0;et&#xa0;al., 2021</xref>. When a robot is trained for locomotion, it can steadily find synergetic motor action patterns similar to those of animals <xref ref-type="bibr" rid="B6">Chai and Hayashibe, 2020</xref>. Deep reinforcement learning algorithms have successfully found better gaits for a rigid snake robot on land <xref ref-type="bibr" rid="B4">Bing&#xa0;et&#xa0;al., 2020</xref>.</p>
<p>Snake robots can explore more gaits that cannot be found by grid search. It is achieved by using the DRL controller as a model-free method to overcome the limitations of the established equations.</p>
<sec id="s3-2-1">
<title>3.2.1 Algorithm</title>
<p>The deep reinforcement learning algorithm we used to train the snake robots is proximal policy optimization (PPO) <xref ref-type="bibr" rid="B30">Schulman&#xa0;et&#xa0;al., 2017</xref>, which is an on-policy algorithm and is widely used to handle continuous action space tasks <xref ref-type="bibr" rid="B23">Mahmood&#xa0;et&#xa0;al., 2018</xref>.</p>
</sec>
<sec id="s3-2-2">
<title>3.2.2 Reward function</title>
<p>In our experiments, we designed two different reward functions <inline-formula id="inf1">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf2">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, for the robot to learn as many different gaits as possible.</p>
<p>In the<disp-formula id="e2">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="&#x7c;" close="&#x7c;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>Vel</italic>
<sub>
<italic>x</italic>
</sub> denotes the velocity of the center of mass of the snake robot along <italic>x</italic>-axis. The faster the snake robot moved in the positive direction of the <italic>x</italic>-axis, the larger the rewards the robot could receive.</p>
<p>The second term <inline-formula id="inf3">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="&#x7c;" close="&#x7c;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
</mml:math>
</inline-formula> is the energy consumed by the robot in one timestep &#x394;<italic>t</italic>. In our experiments, &#x394;<italic>t</italic> was 0.02&#xa0;s. The energy consumed reduced the reward received by the robot. <italic>&#x3b1;</italic> and <italic>&#x3b2;</italic> are scaling factors used to adjust the magnitude of the rewards.</p>
<p>In the<disp-formula id="e3">
<mml:math id="m6">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="&#x7c;" close="&#x7c;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:msqrt>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mi mathvariant="normal">exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(3)</label>
</disp-formula>we set the target velocity of the snake robots using a Gaussian distribution <italic>f</italic>&#xa0;(<italic>Vel</italic>
<sub>
<italic>x</italic>
</sub>;&#xa0;<italic>&#x3bc;</italic>, <italic>&#x3c3;</italic>), where <italic>Vel</italic>
<sub>
<italic>x</italic>
</sub> denotes the current velocity of the robot. <italic>&#x3bc;</italic> denotes the target velocity of the robot, as shown in <xref ref-type="fig" rid="F3">Figure&#xa0;3</xref>. The other terms are the same as those in <inline-formula id="inf4">
<mml:math id="m7">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The relationship between velocity and reward. <italic>&#x3bc;</italic> is the target velocity of snake robots. The position of the margin point is (0.003, 0.1).</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g003.tif"/>
</fig>
</sec>
<sec id="s3-2-3">
<title>3.2.3 Observation space</title>
<p>The observation space contains the information that the robot obtains from the environment to perform feedback control.</p>
<p>The observation space <inline-formula id="inf5">
<mml:math id="m8">
<mml:mi mathvariant="double-struck">O</mml:mi>
</mml:math>
</inline-formula> used to train the snake robots is<disp-formula id="e4">
<mml:math id="m9">
<mml:mtable class="align" columnalign="left">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi mathvariant="double-struck">O</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mfenced open="" close="]">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>It contains the displacements of the five segments on the <italic>y</italic>-axis, and the velocities on the x- and <italic>y</italic>-axis.</p>
</sec>
<sec id="s3-2-4">
<title>3.2.4 Action space</title>
<p>Action space <inline-formula id="inf6">
<mml:math id="m10">
<mml:mi mathvariant="double-struck">A</mml:mi>
</mml:math>
</inline-formula> has the dimensions of the number of actuators of the snake robot. Each element in the action space corresponds to an actuator in the robot. For both the soft and rigid snake robots, the size of the action space was five.</p>
</sec>
<sec id="s3-2-5">
<title>3.2.5 Training configuration</title>
<p>We deployed our training in RLlib <xref ref-type="bibr" rid="B22">Liang&#xa0;et&#xa0;al., 2018</xref>, which is a distributed reinforcement learning framework that could allocate computing resources conveniently.</p>
<p>Furthermore, we used a two-layer fully connected network with 256 units per layer as the hidden layer of the policy network. The input layer of the policy network had the same dimension as the observation space <inline-formula id="inf7">
<mml:math id="m11">
<mml:mi mathvariant="double-struck">O</mml:mi>
</mml:math>
</inline-formula> and the output layer had the same dimensions as the action space <inline-formula id="inf8">
<mml:math id="m12">
<mml:mi mathvariant="double-struck">A</mml:mi>
</mml:math>
</inline-formula>.</p>
<p>We obtained two groups of training results by adjusting the values of <italic>&#x3b1;</italic> and <italic>&#x3b2;</italic> in the reward functions <inline-formula id="inf9">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf10">
<mml:math id="m14">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Three random seeds are used for each training to reduce the uncertainty of the results.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Results and analysis</title>
<p>We compare and analyze the differences in energy consumption between the gait equation controllers and the deep reinforcement learning controllers. We determined the energy efficiency of the snake robots by comparing the relationship between the average velocity of the CoM and the average output power required to drive the system. A gait with a higher average velocity at the same output power has better energy efficiency. In other words, the gait could be managed with less output power at the same motion velocity, it also demonstrated better energy efficiency.</p>
<p>We calculated the average velocity <italic>V</italic> and average output power <italic>P</italic> of all the actuators on the snake robots using the following two equations:<disp-formula id="e5">
<mml:math id="m15">
<mml:mi>V</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msqrt>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msqrt>
<mml:mo>,</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m16">
<mml:mi>P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="&#x7c;" close="&#x7c;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>T</italic> denotes the number of time steps in the simulation. The experiments described in this section had 1,000 timesteps (20&#xa0;s) for each round of testing.</p>
<sec id="s4-1">
<title>4.1 Results of gait equation controller</title>
<p>Utilizing the grid search, the gait equation controller generated 3,600 different gaits for the soft snake robot and rigid snake robot. In <xref ref-type="fig" rid="F4">Figure&#xa0;4</xref>, we present the results of the average velocity of the CoM and the average power of the two snake robots using scatter plots for a maximum speed of less than 0.03&#xa0;m/s. The comparison of the lower contour lines of the two results shows that the soft snake robot requires significantly less energy than the rigid snake robot to reach the same velocity, and the gap in energy consumption is further magnified as the velocity increases.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The grid search results of the gait equation controller for the soft and rigid snake robot. Each point in the scatter plot corresponds to a set of parameters in the grid equation controller. The fold line is the lower contour line of the same color scatter plot.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g004.tif"/>
</fig>
</sec>
<sec id="s4-2">
<title>4.2 Results of deep reinforcement learning</title>
<p>In <xref ref-type="fig" rid="F5">Figure&#xa0;5</xref>, we illustrate the process of training DRL controllers for the rigid snake robot and soft snake robot using the two different reward functions <inline-formula id="inf11">
<mml:math id="m17">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf12">
<mml:math id="m18">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> mentioned in the previous section. The figures present the training results of the two reward functions with the seven groups of parameters. The solid line and shaded part represent the mean and variance of the training results with different random seeds. As the number of training iterations increases, the policy network can obtain higher rewards in each round of testing, indicating that the snake robot can accomplish the tasks expected from the reward functions. The training results converged after 500 iterations. We used the neural network at the endpoint as the DRL controller to evaluate the energy efficiency of the soft and rigid snake robots.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>
<bold>(A)</bold> Deep reinforcement learning training progress by two different reward functions for soft snake robot underwater locomotion. <bold>(B)</bold> Deep reinforcement learning training progress by two different reward functions for rigid snake robot underwater locomotion.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g005.tif"/>
</fig>
</sec>
<sec id="s4-3">
<title>4.3 Comparison</title>
<p>In <xref ref-type="fig" rid="F6">Figure&#xa0;6</xref>, we compare the test results of the gait equation controller and DRL controller of the snake robots using scatter plots. The test results demonstrate that almost all of the gaits generated by the DRL controller are within the range of the gait equation controller&#x2019;s results indicating that the gait equation controller using grid search with very narrow intervals can provide a sufficient exploration of potential action patterns. Moreover, the results also demonstrate that the deep learning solutions are well distributed around the lower bottom side of the average power for <xref ref-type="fig" rid="F6">Figure&#xa0;6</xref>. It is interesting to confirm that deep learning works well for exploring the swimming solution space. Noticeably, Reward Function <inline-formula id="inf13">
<mml:math id="m19">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is suitable for finding faster swimming patterns, and Reward Function <inline-formula id="inf14">
<mml:math id="m20">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is suited for lower energy consumption with slower swimming patterns.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>
<bold>(A)</bold> Comparision of the testing result of deep reinforcement learning controllers and gait equation controllers for rigid snake robot. <bold>(B)</bold> Comparision of the testing result of deep reinforcement learning controllers and gait equation controllers for soft snake robot. The plots with reward functions are corresponding to the deep learning solutions.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g006.tif"/>
</fig>
<p>Special attention should be given to the DRL controllers of the soft snake robot. The results were much closer to the boundary of the gait equation test results than in the rigid snake robot case. In other words, in the same training situation, it is more challenging for rigid snake robot to learn energy-efficient gaits using DRL. This is a further evidence suggests that the soft snake robot&#x2019;s physical attributes make it more energy-efficient than the rigid robot when moving underwater.</p>
<p>In Section II, we introduced the range of the joint rotation and output torque of the rigid snake robot which is calculated based on the joint rotation range, output torque, and the number of active joints of the soft snake robot. We changed the joint rotation range and output range of the rigid snake robot to exclude the potential effects of the joint range and output torque range on the experimental results and performed additional comparative experiments.</p>
<p>First, we change the joint rotation range of the rigid snake robot. The original joint angle rotation range was [&#x2212;36&#xb0;, 36&#xb0;]. Here we changed it to [&#x2212;12&#xb0;, 12&#xb0;], [&#x2212;24&#xb0;, 24&#xb0;], and [&#x2212;48&#xb0;, 48&#xb0;] and performed the same test separately. The experimental results in <xref ref-type="fig" rid="F7">Figure&#xa0;7A</xref> indicate that the original range of rotation exhibits the best performance among the four sets of values. Generally, an excessively large or small rotation angle range will cause more energy consumption. Then we changed the output torque range of the rigid snake robot&#x2019;s actuators from [&#x2212;12, 12] Nm to [&#x2212;1, 1] Nm, [&#x2212;3, 3] Nm, [&#x2212;6, 6] Nm, and [&#x2212;9, 9] Nm. In the results shown in <xref ref-type="fig" rid="F7">Figure&#xa0;7B</xref>, the lower contour lines of the scatter plots for different torque ranges exhibit a very similar growing trend.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>
<bold>(A)</bold> Comparision of different rotation ranges of rigid snake robot&#x2019;s motor. <bold>(B)</bold> Comparision of different output torque ranges of rigid snake robot&#x2019;s motor.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g007.tif"/>
</fig>
<p>The results of the smaller torque range are overlaid by those of the larger torque range. It is primarily because when we perform the grid search for the gait equation controller, the search range of the torque amplitude A is [0.1, 2] Nm with a 0.1 search interval, and the original search result already contains part of the result of a smaller torque range.</p>
<p>Since the other parameters of the grid search were unchanged when we performed these experiments, the fact that no better gaits were found implies we have already fully explored all the gaits that can be generated by the gait equation.</p>
</sec>
<sec id="s4-4">
<title>4.4 Analysis</title>
<p>Based on the above experimental results, the soft robot can use the advantages of its distinct deformed body to achieve better energy efficiency when moving underwater.</p>
<p>To analyze the reasons for this difference, four test results with similar average velocity and lowest output power were selected for the soft and rigid snake robots, respectively, for comparison. These results were individually generated using the DRL and gait equation controller, as shown in <xref ref-type="fig" rid="F8">Figure&#xa0;8</xref>. We calculated the results of the gait equation controller and found that when the snake robot had the same average velocity of 0.024&#xa0;m/s, the output power of the soft-body snake robot was only 19.60% of that of the rigid-body snake robot.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The test results for the four selected gaits, in order from left to right, are soft snake robot with gait equation controller, soft snake robot with DRL controller, rigid snake robot with gait equation controller and rigid snake robot with DRL controller.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g008.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F9">Figure&#xa0;9</xref>, we plot the variation in the CoM velocity of the snake robot in these four gaits. The results indicate that the rigid snake robot has greater variations in CoM velocity as it swims, which is an underlying reason behind its low energy efficiency. The larger velocity variations imply the situation where the robot body has the larger drag force against the water flow.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Comparison of the center-of-mass velocities of the four gaits.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g009.tif"/>
</fig>
<p>The reason for having different drag forces is the continuous shape of the soft body robot can provide a smoother gradient of the body shape, which can significantly reduce the drag from the water. As the soft robot swimming, it can efficiently transport the surrounding water backwards while minimizing the drag, as shown in <xref ref-type="fig" rid="F10">Figure&#xa0;10</xref>.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Test results for the four selected gaits. <bold>(A)</bold> The rigid snake robot is controlled by the gait equation controller. <bold>(B)</bold> The rigid snake robot is controlled by the DRL controller. <bold>(C)</bold> The soft snake robot is controlled by the gait equation controller. <bold>(D)</bold> The soft snake robot is controlled by the DRL controller.</p>
</caption>
<graphic xlink:href="frobt-10-1102854-g010.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>5 Conclusion</title>
<p>We demonstrated that soft-body induces swimming motions to have better energy efficiency when moving underwater by comparing the different gaits of soft and rigid snake robots. First, we designed a soft snake robot and a rigid snake robot with the same dimensions, mass, and actuation power in a simulation environment while keeping the same degrees of freedom for the actuation space. Subsequently, we determined the optimal energy utilization gaits of the two robots explored by grid search and deep reinforcement learning methods for covering the large solution space for the control inputs. By comparing these gaits, we discovered that the average output power required by the soft-body robot is much smaller than that required by the rigid-body robot when it reached the same speed. The main reason for this is that the soft robot has less velocity variation of the body when moving underwater and the continuous body shape of the soft robot reduces the drag from the water. Then the soft robot can effectively transport the surrounding water backwards while minimizing the drag. Our work confirms the advantages of soft-body dynamics against rigid-body dynamics in snake-like underwater robots. We believe that this study can contribute to promoting a new direction to emphasize the energy efficiency advantage of soft-body dynamics for robot design. In the future, we will further analyze the advantages of soft-body dynamics for energy-efficient motion generation by comparing different motor tasks, both underwater and on the ground.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The raw data supporting the conclusion of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the JSPS Grant-in-Aid for Scientific Research on Innovative Areas &#x201c;Hyper-Adaptability&#x201d; project (Grant Number 22H04764) and &#x201c;Science of Soft Robot&#x201d; project (Grant Number 21H00324).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frobt.2023.1102854/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frobt.2023.1102854/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Video1.MP4" id="SM1" mimetype="application/MP4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aracri</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Giorgio-Serchi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Suaria</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sayed</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Nemitz</surname>
<given-names>M. P.</given-names>
</name>
<name>
<surname>Mahon</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Soft robots for ocean exploration and offshore operations: A perspective</article-title>. <source>Soft Robot.</source> <volume>8</volume>, <fpage>625</fpage>&#x2013;<lpage>639</lpage>. <pub-id pub-id-type="doi">10.1089/soro.2020.0011</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Armanini</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Farman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Calisti</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Giorgio-Serchi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Stefanini</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Renda</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Flagellate underwater robotics at macroscale: Design, modeling, and characterization</article-title>. <source>IEEE Trans. Robotics</source> <volume>38</volume>, <fpage>731</fpage>&#x2013;<lpage>747</lpage>. <pub-id pub-id-type="doi">10.1109/tro.2021.3094051</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arulkumaran</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Deisenroth</surname>
<given-names>M. P.</given-names>
</name>
<name>
<surname>Brundage</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bharath</surname>
<given-names>A. A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deep reinforcement learning: A brief survey</article-title>. <source>IEEE Signal Process. Mag.</source> <volume>34</volume>, <fpage>26</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1109/msp.2017.2743240</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bing</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lemke</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Knoll</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Energy-efficient and damage-recovery slithering gait design for a snake-like robot based on reinforcement learning and inverse reinforcement learning</article-title>. <source>Neural Netw.</source> <volume>129</volume>, <fpage>323</fpage>&#x2013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2020.05.029</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bujard</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Giorgio-Serchi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Weymouth</surname>
<given-names>G. D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A resonant squid-inspired robot unlocks biological propulsive efficiency</article-title>. <source>Sci. Robotics</source> <volume>6</volume>, <fpage>eabd2971</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.abd2971</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hayashibe</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Motor synergy development in high-performing deep reinforcement learning algorithms</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>5</volume>, <fpage>1271</fpage>&#x2013;<lpage>1278</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2020.2968067</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>T.-W.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Untethered soft robotic jellyfish</article-title>. <source>Smart Mater. Struct.</source> <volume>28</volume>, <fpage>015019</fpage>. <pub-id pub-id-type="doi">10.1088/1361-665x/aaed4f</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Christianson</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Goldberg</surname>
<given-names>N. N.</given-names>
</name>
<name>
<surname>Deheyn</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tolley</surname>
<given-names>M. T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Translucent soft robots driven by frameless fluid electrode dielectric elastomer actuators</article-title>. <source>Sci. Robotics</source> <volume>3</volume>, <fpage>eaat1893</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.aat1893</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cianchetti</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Laschi</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Menciassi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dario</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Biomedical applications of soft robotics</article-title>. <source>Nat. Rev. Mater.</source> <volume>3</volume>, <fpage>143</fpage>&#x2013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1038/s41578-018-0022-y</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hughes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wah</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Matusik</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Rus</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Underwater soft robot modeling and control with differentiable simulation</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>6</volume>, <fpage>4994</fpage>&#x2013;<lpage>5001</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2021.3070305</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Distributed multi-robot collision avoidance via deep reinforcement learning for navigation in complex scenarios</article-title>. <source>Int. J. Robotics Res.</source> <volume>39</volume>, <fpage>856</fpage>&#x2013;<lpage>892</lpage>. <pub-id pub-id-type="doi">10.1177/0278364920916531</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Frame</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lopez</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Curet</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Engeberg</surname>
<given-names>E. D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Thrust force characterization of free-swimming soft robotic jellyfish</article-title>. <source>Bioinspiration biomimetics</source> <volume>13</volume>, <fpage>064001</fpage>. <pub-id pub-id-type="doi">10.1088/1748-3190/aadcb3</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Giorgio-Serchi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Weymouth</surname>
<given-names>G. D.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Underwater soft robotics, the benefit of body-shape variations in aquatic propulsion</article-title>,&#x201d; in <source>Soft robotics: Trends, applications and challenges</source> (<publisher-name>Springer</publisher-name>), <fpage>37</fpage>&#x2013;<lpage>46</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hubbard</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Fleming</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Palmre</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Pugal</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K. J.</given-names>
</name>
<name>
<surname>Leang</surname>
<given-names>K. K.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Monolithic ipmc fins for propulsion and maneuvering in bioinspired underwater robotics</article-title>. <source>IEEE J. Ocean. Eng.</source> <volume>39</volume>, <fpage>540</fpage>&#x2013;<lpage>551</lpage>. <pub-id pub-id-type="doi">10.1109/joe.2013.2259318</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jangir</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Aleny&#xe0;</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Torras</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Dynamic cloth manipulation with deep reinforcement learning</article-title>,&#x201d; in <source>2020 IEEE international conference on robotics and automation (ICRA)</source> (<publisher-name>IEEE</publisher-name>), <fpage>4630</fpage>&#x2013;<lpage>4636</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumet</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bell</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Sanchez</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Preston</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A data-driven review of soft robotics</article-title>. <source>Adv. Intell. Syst.</source> <volume>4</volume>, <fpage>2100163</fpage>. <pub-id pub-id-type="doi">10.1002/aisy.202100163</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Katzschmann</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>DelPreto</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>MacCurdy</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rus</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Exploration of underwater life with an acoustically controlled soft robotic fish</article-title>. <source>Sci. Robotics</source> <volume>3</volume>, <fpage>eaar3449</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.aar3449</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hwangbo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wellhausen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Koltun</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Learning quadrupedal locomotion over challenging terrain</article-title>. <source>Sci. robotics</source> <volume>5</volume>, <fpage>eabc5986</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.abc5986</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2021a</year>). <article-title>Self-powered soft robot in the mariana trench</article-title>. <source>Nature</source> <volume>591</volume>, <fpage>66</fpage>&#x2013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-020-03153-z</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Shintake</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hayashibe</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021b</year>). &#x201c;<article-title>Deep reinforcement learning framework for underwater locomotion of soft robot</article-title>,&#x201d; in <source>2021 IEEE international conference on robotics and automation (ICRA)</source> (<publisher-name>IEEE</publisher-name>), <fpage>12033</fpage>&#x2013;<lpage>12039</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Fast-moving soft electronic fish</article-title>. <source>Sci. Adv.</source> <volume>3</volume>, <fpage>e1602045</fpage>. <pub-id pub-id-type="doi">10.1126/sciadv.1602045</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Liaw</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nishihara</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Moritz</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fox</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Goldberg</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). &#x201c;<article-title>Rllib: Abstractions for distributed reinforcement learning</article-title>,&#x201d; in <source>International conference on machine learning</source> (<publisher-loc>Stockholm, Sweden</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>3053</fpage>&#x2013;<lpage>3062</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mahmood</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>Korenkevych</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Vasan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Bergstra</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Benchmarking reinforcement learning algorithms on real-world robots</article-title>,&#x201d; in <source>Conference on robot learning</source> (<publisher-loc>Z&#x00FC;rich, Switzerland</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>561</fpage>&#x2013;<lpage>591</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mnih</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Kavukcuoglu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Silver</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Graves</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Antonoglou</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Wierstra</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <source>Playing atari with deep reinforcement learning</source>. <comment>arXiv preprint arXiv:1312.5602</comment>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nguyen</surname>
<given-names>D. Q.</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>V. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Anguilliform swimming performance of an eel-inspired soft robot</article-title>. <source>Soft Robot.</source> <volume>9</volume>, <fpage>425</fpage>&#x2013;<lpage>439</lpage>. <pub-id pub-id-type="doi">10.1089/soro.2020.0093</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Park</surname>
<given-names>S.-J.</given-names>
</name>
<name>
<surname>Gazzola</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Di Santo</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Blevins</surname>
<given-names>E. L.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Phototactic guidance of a tissue-engineered soft-robotic ray</article-title>. <source>Science</source> <volume>353</volume>, <fpage>158</fpage>&#x2013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaf4292</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Sitti</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Multi-functional soft-bodied jellyfish-like swimming</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>2703</fpage>&#x2013;<lpage>2712</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-10549-7</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rich</surname>
<given-names>S. I.</given-names>
</name>
<name>
<surname>Wood</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Majidi</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Untethered soft robotics</article-title>. <source>Nat. Electron.</source> <volume>1</volume>, <fpage>102</fpage>&#x2013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1038/s41928-018-0024-1</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rus</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tolley</surname>
<given-names>M. T.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Design, fabrication and control of soft robots</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>467</fpage>&#x2013;<lpage>475</lpage>. <pub-id pub-id-type="doi">10.1038/nature14543</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Schulman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wolski</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Dhariwal</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Radford</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Klimov</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Proximal policy optimization algorithms</source>. <comment>arXiv preprint arXiv:1707.06347</comment>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sfakiotakis</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lane</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Davies</surname>
<given-names>J. B. C.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Review of fish swimming modes for aquatic locomotion</article-title>. <source>IEEE J. Ocean. Eng.</source> <volume>24</volume>, <fpage>237</fpage>&#x2013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1109/48.757275</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shintake</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cacucciolo</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Shea</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Floreano</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Soft biomimetic fish robot made of dielectric elastomer actuators</article-title>. <source>Soft Robot.</source> <volume>5</volume>, <fpage>466</fpage>&#x2013;<lpage>474</lpage>. <pub-id pub-id-type="doi">10.1089/soro.2017.0062</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Todorov</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Erez</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tassa</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Mujoco: A physics engine for model-based control</article-title>,&#x201d; in <source>2012 IEEE/RSJ international conference on intelligent robots and systems</source> (<publisher-name>IEEE</publisher-name>), <fpage>5026</fpage>&#x2013;<lpage>5033</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Villanueva</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Priya</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A biomimetic robotic jellyfish (robojelly) actuated by shape memory alloy composite actuators</article-title>. <source>Bioinspiration biomimetics</source> <volume>6</volume>, <fpage>036004</fpage>. <pub-id pub-id-type="doi">10.1088/1748-3182/6/3/036004</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Simultaneous navigation and radio mapping for cellular-connected uav with deep reinforcement learning</article-title>. <source>IEEE Trans. Wirel. Commun.</source> <volume>20</volume>, <fpage>4205</fpage>&#x2013;<lpage>4220</lpage>. <pub-id pub-id-type="doi">10.1109/twc.2021.3056573</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hayashibe</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Joint elasticity produces energy efficiency in underwater locomotion: Verification with deep reinforcement learning</article-title>. <source>Front. Robotics AI</source> <volume>9</volume>, <fpage>957931</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2022.957931</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fish</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Kerr</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Downs</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bart-Smith</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Tunable stiffness enables fast and efficient swimming in fish-like robots</article-title>. <source>Sci. Robotics</source> <volume>6</volume>, <fpage>eabe4088</fpage>. <pub-id pub-id-type="doi">10.1126/scirobotics.abe4088</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>