<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">748716</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2021.748716</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Tool-Use Model to Reproduce the Goal Situations Considering Relationship Among Tools, Objects, Actions and Effects Using Multimodal Deep Neural Networks</article-title>
<alt-title alt-title-type="left-running-head">Saito et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Tool-Use Model to Reproduce Goal</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Saito</surname>
<given-names>Namiko</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/526211/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Ogata</surname>
<given-names>Tetsuya</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/314077/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mori</surname>
<given-names>Hiroki</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/675714/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Murata</surname>
<given-names>Shingo</given-names>
</name>
<xref ref-type="aff" rid="aff5">
<sup>5</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/215526/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sugano</surname>
<given-names>Shigeki</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Department of Modern Mechanical Engineering, Waseda University, <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Department of Intermedia Art and Science, Waseda University, <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>National Institute of Advanced Industrial Science and Technology (AIST), <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff4">
<label>
<sup>4</sup>
</label>Future Robotics Organization, Waseda University, <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff5">
<label>
<sup>5</sup>
</label>Department of Electronics and Electrical Engineering, Keio University, <addr-line>Kanagawa</addr-line>, <country>Japan</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/537983/overview">Yan Wu</ext-link>, Institute for Infocomm Research (A&#x2217;STAR), Singapore</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1021830/overview">Wenyu Liang</ext-link>, Institute for Infocomm Research (A&#x2217;STAR), Singapore</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1436967/overview">Fen Fang</ext-link>, Institute for Infocomm Research (A&#x2217;STAR), Singapore</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Namiko Saito, <email>n_saito@sugano.mech.waseda.ac.jp</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Robot Learning and Evolution, a section of the journal Frontiers in Robotics and&#x20;AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>28</day>
<month>09</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>748716</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Saito, Ogata, Mori, Murata and Sugano.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Saito, Ogata, Mori, Murata and Sugano</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>We propose a tool-use model that enables a robot to act toward a provided goal. It is important to consider features of the four factors; tools, objects actions, and effects at the same time because they are related to each other and one factor can influence the others. The tool-use model is constructed with deep neural networks (DNNs) using multimodal sensorimotor data; image, force, and joint angle information. To allow the robot to learn tool-use, we collect training data by controlling the robot to perform various object operations using several tools with multiple actions that leads different effects. Then the tool-use model is thereby trained and learns sensorimotor coordination and acquires relationships among tools, objects, actions and effects in its latent space. We can give the robot a task goal by providing an image showing the target placement and orientation of the object. Using the goal image with the tool-use model, the robot detects the features of tools and objects, and determines how to act to reproduce the target effects automatically. Then the robot generates actions adjusting to the real time situations even though the tools and objects are unknown and more complicated than trained&#x20;ones.</p>
</abstract>
<kwd-group>
<kwd>tool-use</kwd>
<kwd>manipulation</kwd>
<kwd>multimodal learning</kwd>
<kwd>recurrent neural networks</kwd>
<kwd>deep neural networks</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<sec id="s1-1">
<title>1.1 Background</title>
<p>Tool-use is critical for realizing robots that can accomplish various tasks in complex surroundings. By using tools, humans are capable of compensating for missing bodily functions, greatly expanding the range of tasks they can perform <xref ref-type="bibr" rid="B12">Gibson and Ingold (1993)</xref> and <xref ref-type="bibr" rid="B30">Osiurak et&#x20;al. (2010)</xref>. By using tools, robots can overcome physical limitations and perform complex tasks. They could also adapt to environments without changing or adding actuators or other mechanisms, reducing their weight and size and making them safer, less expensive, and easier to be used in the real world. Robots capable of working alongside humans and performing daily tasks are increasingly becoming a key research topic in the robotics field <xref ref-type="bibr" rid="B47">Yamazaki et&#x20;al. (2012)</xref> and <xref ref-type="bibr" rid="B5">Cavallo et&#x20;al. (2014)</xref>. Tool-use by robots would result in introducing robots into everyday spaces to assist human&#x20;lives.</p>
<p>Much research has aimed at realizing robots that can perform tool-use tasks, most using preset environments and pre-made numerical models for the target tools and objects. Following those models, the robots calculate and move along optimal motion trajectories. This approach has realized highly accurate and fast movements, allowing robots to perform pan tosses <xref ref-type="bibr" rid="B31">Pan et&#x20;al. (2018)</xref>, make pancakes with cooking tools <xref ref-type="bibr" rid="B1">Beetz et&#x20;al. (2011)</xref>, cut bread with a knife <xref ref-type="bibr" rid="B33">Ramirez-Amaro et&#x20;al. (2015)</xref>, and serve food with a spatula <xref ref-type="bibr" rid="B27">Nagahama et&#x20;al. (2013)</xref> in past studies. This approach can be efficient and realize high performance if the robots are dedicated to a particular purpose in a specific environment, but as the number of tools and objects or task types increase, it becomes difficult to design numerical models of properties and environments for each condition. In addition, it is difficult to deal with the situations when somethings that did not supposed to happen, because the robots can only move as initially instructed by humans.</p>
<p>There are unlimited conceivable situations in which robots might perform various tasks, so preparing tailored tools and individually teaching robots how to use them is impossible. It is thus critical for robots to acquire &#x201c;tool-use ability,&#x201d; and to consider how to use tools to achieve a goal without human assistance, even when tools are seen for the first time. We therefore develop a robot to plan and make an action by itself only by showing the goal of the&#x20;tasks.</p>
</sec>
<sec id="s1-2">
<title>1.2 Relationships Among Tools, Objects, Actions and Effects in Tool-Use</title>
<p>To acquire tool-use ability, considering relationships among tools, objects, actions and effects is important. According to <xref ref-type="bibr" rid="B10">Gibson (1977)</xref> and J.<xref ref-type="bibr" rid="B11">Gibson (1979)</xref>, affordance is defined as the possibility of an act presented to an agent by an object or environment. The definition of affordance by Gibson is highly conceptual, so there have been attempts to clarify it <xref ref-type="bibr" rid="B44">Turvey (1992)</xref>, <xref ref-type="bibr" rid="B29">Norman (2013)</xref>, <xref ref-type="bibr" rid="B6">Chemero (2003)</xref>, and <xref ref-type="bibr" rid="B41">Stoffregen (2003)</xref>. Especialy, Chemero argued that affordance is not unique to the environment and it is a relationship between agent abilities and environmental features, providing agent actions. For example, consider a hammer. The hammer head is suitable for hitting objects with strong blows, and the long handle is suited to grabbing and swinging. Therefore, when dealing with nails, the hammer affords hitting. However, if you do not have a hammer, you could hit a nail with a block having similar features, such as hardness, ruggedness, and weight. The hammer can also be used for other purposes. If you want to pull and obtain something that is out of reach, you can use a hammer to pull it. So a long hammer can also propose agents to pull and provide an effect of pulling objects. In other words, tool-use arises not only from tool features, but also from the operated objects, actions to take and the expected effects (i.e.,&#x20;goal). Acquisition of affordance is one of the most essential topic in the cognitive field <xref ref-type="bibr" rid="B4">Bushnell and Boudreau (1993)</xref>, <xref ref-type="bibr" rid="B32">Piaget (1952)</xref>, <xref ref-type="bibr" rid="B34">Rat-Fischer et&#x20;al. (2012)</xref>, and <xref ref-type="bibr" rid="B21">Lockman (2000)</xref>, and has also been discussed in robotics field and has gained significant support <xref ref-type="bibr" rid="B15">Horton et&#x20;al. (2012)</xref>, <xref ref-type="bibr" rid="B24">Min et&#x20;al. (2016)</xref>, and <xref ref-type="bibr" rid="B16">Jamone et&#x20;al. (2018)</xref>. Inspired from this theory, we construct the tool-use model with DNNs simultaneously considering features of tools, objects, actions,and effects. By recognizing the relationships among the four factors, we expect the robot to deal with tool-use tasks including those never experienced before.</p>
</sec>
<sec id="s1-3">
<title>1.3 Research Objective</title>
<p>Our objective is to develop a robot capable of acquiring relationships of four factors: 1) tools the robot can use, 2) target objects that can be manipulated by those tools, 3) actions performed by the robot, and 4) effects by those actions. And to allow the robot to conduct tool-use actions according to the relationships to reproduce providing&#x20;goals.</p>
<p>We take the approach of developing a robot from own task conduction experiences and allowing it to acquire the relationships. The robot is developed in the following steps. First, multimodal sensorimotor data are recorded while the robot is controlled to conduct tool-use tasks, like it plays to repeat manipulating objects. Then deep neural networks (DNNs) are trained and used to construct a tool-use model, which uses the recorded dataset to learn sensorimotor coordination and relationships among tools, objects, actions, and effects. Then for tests, the robot generates motions for handling novel tools and objects by detecting their features with the tool-use model, and acts to achieve its goal. We verified that the robot can detect the features, recognize its goals, and act to achieve&#x20;them.</p>
<p>Some of the tools that infants firstly develop to use is stick- or rake-like tools such as forks or spoons, allowing them to extend their reach to grasp or move distant objects. Similarly, we start by allowing the robot to learn to move objects by such simple tools. Specifically, the robot uses I- or T-shaped tools to extend their reach and pull or push an object to roll, slide, or topple.</p>
<p>The research goal is to enable the robot to perform tool-use operations solely from provided goals and reproduce the situations. The experimenters only present a goal image to the robot. We use an image because it clearly and easily shows the&#x20;target position and orientation of the object. The robot needs to understand the relationships and detect features of the tool and the object, then generates an action toward the goal effect. To manipulate the object to the target position at the target orientation, the robot must adjust operating movements every time step according to the current situation. The goal cannot be completed just by replay actions as same as the trained&#x20;ones.</p>
<p>The following summarizes our contributions:<list list-type="simple">
<list-item>
<p>&#x2022; The tool-use model learns relationships of tools, objects, actions, and effects.</p>
</list-item>
<list-item>
<p>&#x2022; The robot generates tool-use actions based on the detected relationships and current situations even with unknown tools and objects.</p>
</list-item>
<list-item>
<p>&#x2022; The robot accomplishes tasks solely from a provided goal image, and reproduce the situations.</p>
</list-item>
</list>
</p>
</sec>
</sec>
<sec id="s2">
<title>2 Related Works</title>
<p>In this section, we show some related tool-use works except for which use initially prepared numerical models of actions for fixed target or environments.</p>
<sec id="s2-1">
<title>2.1 Understanding Tool Features</title>
<p>
<xref ref-type="bibr" rid="B25">Myers et&#x20;al. (2015)</xref> and <xref ref-type="bibr" rid="B49">Zhu et&#x20;al. (2015)</xref> investigated autonomous understanding of tool features, introducing frameworks for dividing and localizing tool parts from RGB-D camera images. These frameworks identify several tool parts that are critically involved in tool functions and recognize tool features relying on those parts. Robots could thereby understand how tools and their parts can be used. Similarly, we allow robots to recognize tool features from their appearance. Moreover, we not only allow the robot to recognize features but also generate tool-use motions.</p>
</sec>
<sec id="s2-2">
<title>2.2 Planning Tool-Use Motions With Analytic Models</title>
<p>
<xref ref-type="bibr" rid="B42">Stoytchev (2005)</xref>, <xref ref-type="bibr" rid="B3">Brown and Sammut (2012)</xref>, and <xref ref-type="bibr" rid="B26">Nabeshima et&#x20;al. (2007)</xref> realized robots capable of performing tool-use motions by constructing analytic models. In particular, <xref ref-type="bibr" rid="B42">Stoytchev (2005)</xref> controlled a robot to move a hockey puck with various shaped tools, making a table showing object movements corresponding to the shape and shift direction of tools used during observations. By following this table, the robot could carry the object to a target position. <xref ref-type="bibr" rid="B3">Brown and Sammut (2012)</xref> presented an algorithm for discovering tool-use, in which the system first identifies subgoals, then searches for motions matching next subgoals one-by-one. They realized a robot capable of grasping various shaped tools and carrying an object. <xref ref-type="bibr" rid="B26">Nabeshima et&#x20;al. (2007)</xref> constructed a computational model for calculating tool shapes and moments of inertia, developing a robot capable of manipulating tools to pull an object from an invisible shielded area, regardless of the shape of the grasped tool. However, these systems need prepared tables or computational models before using tools, making it difficult to deal with unknown tools. In addition, modeling errors may accumulate during execution, often resulting in fragile systems.</p>
</sec>
<sec id="s2-3">
<title>2.3 Generating Tool-Use Motions by Learning</title>
<p>Another line of research on tool-use motion generation is learning tool features from tool-use experience, which is the same as our approach. <xref ref-type="bibr" rid="B28">Nishide et&#x20;al. (2012)</xref> allowed a robot to experience sliding a cylindrical object with various shaped tools, training a DNN to estimate trajectories of object movements that change depending on tool shape and how it slid. <xref ref-type="bibr" rid="B43">Takahashi et&#x20;al. (2017)</xref> constructed a DNN model for learning differences in functions depending on tool grasping positions. This allowed a robot to grasp appropriate tool positions and to generate motions to move an object to a goal. <xref ref-type="bibr" rid="B22">Mar et&#x20;al. (2018)</xref> considered the orientation of tools as well. They controlled a robot to push a designated object in several directions with tools of several shapes and orientations, and recorded the shift length of object movements. They detected tool features by a self-organizing map corresponded to lengths of object shifts. They realized a robot capable of selecting directions to push the objects depending on the tool&#x2019;s shape and orientation. <xref ref-type="bibr" rid="B37">Saito et&#x20;al. (2018b)</xref> focused on tool selection, setting several initial and target positions for objects and allowing the robot to experience moving objects with several tools. They trained a DNN model that enabled the robot to select a tool of proper length and shape, depending on the designated direction and distance to the object. In every of these studies, robots with learning models learn and recognize tool features and generate suitable tool-use motions according to tool features. However, they dealt with one specified target object and thus did not consider relationships between tools and objects. These tool-use situations are thus limited, because if a different object is provided, it would be difficult to operate it. <xref ref-type="bibr" rid="B13">Goncalves et&#x20;al. (2014)</xref> and <xref ref-type="bibr" rid="B7">Dehban et&#x20;al. (2016)</xref> conducted research to enable robots to consider four factors: tools, objects, actions, and effects. <xref ref-type="bibr" rid="B13">Goncalves et&#x20;al. (2014)</xref> constructed Bayesian networks that express relationships among the four factors, predicting the effect when the other three factors are input. However, they aimed to determine features of tools and objects based on categories such as area, length, and circularity, which were set in advance by the experimenters. It was therefore difficult for the robot to autonomously self-acquire features without requiring predefined feature extraction routines, and also difficult to manipulate arbitrary objects with arbitrary tools without human assistance. <xref ref-type="bibr" rid="B7">Dehban et&#x20;al. (2016)</xref> also expressed relationships among the four factors using DNNs, realizing a robot that could predict or select one factor when the other three are given. In other words, the robot could predict an effect or select an action or tool to use. However, the action types were fixed in advance, so the robot could only select and follow the predesignated motions, making it difficult to deal with objects and tools in unknown positions or objects not moving as expected.</p>
<p>To address these problems, we construct a tool-use model that allows the robot to self-acquire the relationships among the four factors. The model is expected to detect the features of provided tools and objects and to generate actions depending on the situations. The present study is an extension of Ref. <xref ref-type="bibr" rid="B36">Saito et&#x20;al. (2018a)</xref>. to improve two points. The first improvement is related to goal images. In the previous study, provided goal images showed situations just after task executions, namely the final position of the robot arm, making it easy for the tool-use model to predict what kind of actions should be conducted. In the present study, we make the robot arm return to its initial joint position after task completion. The goal images thus show the arm at the initial position, and there are no hints regarding actions to take from the images except for object position and orientation. Second, we introduce a force sensor to realize task executions even when the object is occluded by the robot arm or tools. In the previous study, they used only image data as sensory input, making it difficult to operate small objects that can be occluded during movement. Many papers have shown that using both vision and force can improve the accuracy of object recognition in both cognitive field and robotics field <xref ref-type="bibr" rid="B9">Fukui and Shimojo (1994)</xref>, <xref ref-type="bibr" rid="B8">Ernst and Banks (2002)</xref>, <xref ref-type="bibr" rid="B20">Liu et&#x20;al. (2017)</xref>, and <xref ref-type="bibr" rid="B38">Saito et&#x20;al. (2021)</xref>. By constructing the tool-use model with multimodal DNNs, we realize the robot to operate much more complex tools and objects than in the past&#x20;study.</p>
</sec>
</sec>
<sec id="s3">
<title>3&#x20;Tool-Use Model</title>
<p>In this section, we describe the method for constructing the tool-use model with&#x20;DNNs.</p>
<sec id="s3-1">
<title>3.1 Overview of the Tool-Use Model</title>
<sec id="s3-1-1">
<title>3.1.1 Task Conduction With the Tool-Use Model</title>
<p>
<xref ref-type="fig" rid="F1">Figure&#x20;1</xref> shows the way to control the robot using the model, which is extended from Ref. <xref ref-type="bibr" rid="B36">Saito et&#x20;al. (2018a)</xref> to deal with both image and force sensor data efficiently. The tool-use model comprises two modules, a feature extraction module and a motion generation module. Since the number of dimensions in image data is considerably larger than in other data, the feature extraction module compresses image data, making the multimodal learning well-balanced at low computational cost. Then, the motion generation module simultaneously learns all time-series data, that is image feature data, joint angle data and force data. This module is expected to learn the relationship among the four factors by using the latent space values Cs(0), which detect features of them from given goal images. It is also expected to generate tool-use actions adjusting in real time by outputting the next joint angle data and move the robot according to the angle&#x20;data.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The tool-use model to control the robot to conduct tasks. It comprises a feature extraction module and a motion generation module. In tests, at first, the internal latent space value Cs(0) is explored using a goal image. The Cs(0) detects and expresses features of the tool, object, and actions needed to produce the target effects. For generating a motion, the feature extraction module compresses the number of dimensions in image data and extracts low-dimensional image feature data. Then the motion generation module simultaneously learns image feature data, joint angle data, and force data at the moment, and predicts next-step data. The robot is controlled according to the predicted joint angle data step-by-step.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g001.tif"/>
</fig>
</sec>
<sec id="s3-1-2">
<title>3.1.2 Training of the Tool-Use Model</title>
<p>To collect training data, we remotely control a robot to conduct object manipulation tasks with several tools and record the image data, joint angle data, and force data in advance. As for training, in the first step, the image data is used for training the feature extraction module. After that, the motion generation module is trained with time series of the image feature data obtained through the trained feature extraction module, joint angle data and force&#x20;data.</p>
</sec>
</sec>
<sec id="s3-2">
<title>3.2 Feature Extraction Module</title>
<p>We use a convolutional autoencoder (CAE) <xref ref-type="bibr" rid="B23">Masci et&#x20;al. (2011)</xref> to construct the feature extraction module. The CAE is a multilayered neural network with convolutional and fully connected layers, so it has advantages of both a convolutional neural network (CNN) <xref ref-type="bibr" rid="B19">Krizhevsky et&#x20;al. (2012)</xref>, which has high performance in image recognition, and an autoencoder (AE) <xref ref-type="bibr" rid="B14">Hinton and Salakhutdinov (2006)</xref>, which has a bottleneck structure and can reduce data dimensionality. After training, the trained module can represent appearance features such as the shape, size, position, and orientation of tools and objects even when unknown images are&#x20;input.</p>
<p>
<xref ref-type="table" rid="T1">Table&#x20;1</xref> shows the CAE structure, in which input data pass through the center layer with the fewest nodes, then outputs data with the original number of dimensions. The module is trained so that the output (<italic>y</italic>) restores the input image (<italic>x</italic>), by minimizing the mean squared error (MSE) <xref ref-type="bibr" rid="B2">Bishop (2006)</xref> as<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2211;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>Image feature data can be extracted from the center-layer nodes with low dimension. We use a sigmoid function as the activation function for only the center layer, whereas we use the ReLU function for all other layers. The CAE is trained using MSE with the optimizer for the Adaptive Moment Estimation (Adam) algorithm <xref ref-type="bibr" rid="B18">Kingma and Ba (2014)</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>The CAE structure. The module includes convolution layers and fully connected layers with linear processing. An input data pass through the center layer with the fewest nodes, then outputs data with the original number of dimensions.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Layer</th>
<th align="center">Input</th>
<th align="center">Output</th>
<th align="center">Processing</th>
<th align="center">Kernel size</th>
<th align="center">Stride</th>
<th align="center">Padding</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="center">(64, 48, 3)</td>
<td align="center">(32, 24, 32)</td>
<td align="center">Convolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">2</td>
<td align="center">(32, 24, 32)</td>
<td align="center">(16, 12, 64)</td>
<td align="center">Convolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">3</td>
<td align="center">(16, 12, 64)</td>
<td align="center">(8, 6, 128)</td>
<td align="center">Convolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">4</td>
<td align="center">(8, 6, 128)</td>
<td align="center">(4, 3, 256)</td>
<td align="center">Convolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">5</td>
<td align="center">3,072</td>
<td align="center">254</td>
<td align="center">Linear</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
</tr>
<tr>
<td align="left">6</td>
<td align="center">254</td>
<td align="center">20</td>
<td align="center">Linear</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
</tr>
<tr>
<td align="left">7</td>
<td align="center">20</td>
<td align="center">254</td>
<td align="center">Linear</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
</tr>
<tr>
<td align="left">8</td>
<td align="center">254</td>
<td align="center">3,072</td>
<td align="center">Linear</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
<td align="center">
<bold>-</bold>
</td>
</tr>
<tr>
<td align="left">9</td>
<td align="center">(4, 3, 256)</td>
<td align="center">(8, 6, 128)</td>
<td align="center">Deconvolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">10</td>
<td align="center">(8, 6, 128)</td>
<td align="center">(16, 12, 64)</td>
<td align="center">Deconvolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">11</td>
<td align="center">(16, 12, 64)</td>
<td align="center">(32, 24, 32)</td>
<td align="center">Deconvolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
<tr>
<td align="left">12</td>
<td align="center">(32, 24, 32)</td>
<td align="center">(64, 48, 3)</td>
<td align="center">Deconvolution</td>
<td align="center">(4,4)</td>
<td align="center">(2,2)</td>
<td align="center">(1,1)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We tested changing the numbers of the center-layer nodes by 10, 15, 20, 25, and finally set it 20 because it was the smallest number that could sufficiently express the different features of camera images. Therefore, the feature extraction module compresses high-dimensional (9,216 dimensions (64 width &#xd7; 48 height &#xd7; 3 channels)) raw image data to 20 dimensions. The module is trained for 5,000 epochs.</p>
</sec>
<sec id="s3-3">
<title>3.3 Motion Generation Module</title>
<p>For the motion generation module, we use a multiple timescale recurrent neural network (MTRNN) <xref ref-type="bibr" rid="B46">Yamashita and Tani (2008)</xref>, attaching additional fully connected layers. A MTRNN is a type of recurrent neural network that can predict the next state from a current and previous state. It contains three node types with different time constants: input&#x2013;output (IO) nodes, fast context (Cf) nodes, and slow context (Cs) nodes. Cf nodes with small time constants learn movement primitives in the data, whereas Cs nodes with large time constants learn sequences. By combining these three node types, long, complex time series data can be learned, the usefulness of which for manipulation has been confirmed in several studies <xref ref-type="bibr" rid="B48">Yang et&#x20;al. (2016)</xref>, <xref ref-type="bibr" rid="B43">Takahashi et&#x20;al. (2017)</xref>, and <xref ref-type="bibr" rid="B37">Saito et&#x20;al. (2018b</xref>,<xref ref-type="bibr" rid="B36">a</xref>, <xref ref-type="bibr" rid="B39">2020</xref>, <xref ref-type="bibr" rid="B38">2021)</xref>.</p>
<p>The motion generation module integrates time series of image feature data output from the feature extraction module (<italic>x</italic>
<sup>image</sup>), joint angle data (<italic>x</italic>
<sup>motor</sup>), and force data (<italic>x</italic>
<sup>force</sup>), and predicts next time-step data. Since image data contain more complex and varied information than do other data, we connect fully connected layers before and after IO nodes of only image feature&#x20;data.</p>
<p>
<xref ref-type="table" rid="T2">Table&#x20;2</xref> shows the structure of the motion generation module, which has settings for the time constants and numbers of each node. We tried to vary numbers of Cs nodes in the range of 8&#x2013;12, the time constant of Cs nodes in the range of 30&#x2013;60 in increments of 10, and numbers of Cf nodes in the range of 30&#x2013;60 in increments of 10. Finally the combination that minimized training error is adopted. If these numbers are too small, complex information cannot be learned, and if they are too large, the module is overtrained and cannot adapt to untrained data. The time constant of Cf had little effect, even when it was changed to around 5. The module is trained for 20,000 epochs.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The structure of the MTRNN attaching fully connected layers. MTRNN contains three node types with different time constants: input&#x2013;output (IO) nodes, fast context (Cf) nodes, and slow context (Cs) nodes. Since image data contain more complex and varied information than do other data, we connect fully connected layers before and after IO nodes of only image feature data (F1, F2, F3, B3, B2, and B1).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Node name</th>
<th align="center">Number of nodes</th>
<th align="center">Time constant</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<bold>F1, B3</bold>
</td>
<td align="center">20 (number of image features)</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">
<bold>F2, B2</bold>
</td>
<td align="center">30</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">
<bold>F3, B1</bold>
</td>
<td align="center">15</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">
<bold>IO nodes</bold>
</td>
<td align="center">24</td>
<td align="char" char=".">1</td>
</tr>
<tr>
<td align="left">
<bold>Cf nodes</bold>
</td>
<td align="center">50</td>
<td align="char" char=".">5</td>
</tr>
<tr>
<td align="left">
<bold>Cs nodes</bold>
</td>
<td align="center">10</td>
<td align="char" char=".">40</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3-3-1">
<title>3.3.1 Forward Calculation</title>
<p>In forward calculations of this module, the internal value is first calculated by fully connected layers (F1, F2, F3) as<disp-formula id="e2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">h</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mn mathvariant="normal">1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mn mathvariant="normal">2</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">h</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mn mathvariant="normal">2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mn mathvariant="normal">3</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>x</italic>
<sup>image</sup>(<italic>t</italic>) is the input image feature data, <italic>t</italic> is the time step, <italic>w</italic>
<sub>
<italic>ij</italic>
</sub> is the weight of the connection between the <italic>j</italic>th and <italic>i</italic>th neuron, and <italic>x</italic>
<sub>
<italic>j</italic>
</sub>(<italic>t</italic>) is the value input to the <italic>i</italic>th neuron by the <italic>j</italic>th neuron.</p>
<p>We concatenate the value, motor data, and force data, then input to the IO nodes of the MTRNN as<disp-formula id="e3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mn mathvariant="normal">3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>We then input the concatenated data to the MTRNN. First, the internal value of the <italic>i</italic>th neuron <italic>u</italic>
<sub>
<italic>i</italic>
</sub> is calculated as<disp-formula id="e4">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
<mml:msub>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>N</italic> is the index sets of neural units and <italic>&#x3c4;</italic>
<sub>
<italic>i</italic>
</sub> is the time constant of the <italic>i</italic>th neuron (<italic>i</italic>&#x20;&#x2208; IO,Cf, Cs). Then the output value is calculated as<disp-formula id="e5">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">h</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(5)</label>
</disp-formula>Then the output value <italic>y</italic>
<sub>IO</sub>(<italic>t</italic>) is then divided into three parts in charge of image feature data, motor data and force data by their dimensions as<disp-formula id="e6">
<mml:math id="m6">
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">v</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>Then the output value for an image feature data (<inline-formula id="inf1">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mi>O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>) is recalculated with fully connected layers (B1, B2, B3) as<disp-formula id="e7">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">O</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">h</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">2</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">h</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">3</mml:mn>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(7)</label>
</disp-formula>Finally, the predicted image feature data can be obtained as<disp-formula id="e8">
<mml:math id="m9">
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mn mathvariant="normal">3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
</sec>
<sec id="s3-3-2">
<title>3.3.2&#x20;Next-step Data Prediction</title>
<p>The next-step predicted data is calculated as<disp-formula id="e9">
<mml:math id="m10">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
<disp-formula id="equ1">
<mml:math id="m11">
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi>Y</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
</disp-formula>where 0 &#x2264; <italic>&#x3b1;</italic> &#x2264; 1 is the feedback rate and T(<italic>t</italic>) is an input datum, which means training data when we train the module, which in turn means actual data recorded when testing the tool-use model while moving the robot. The predicted value <inline-formula id="inf2">
<mml:math id="m12">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is calculated by multiplying the output of the preceding step <italic>Y</italic>(<italic>t</italic>) and the datum T(<italic>t</italic>) by the feedback rate <italic>&#x3b1;</italic>. The first term presents &#x201c;closed-loop prediction,&#x201d; in which the robot associates a data series with past data but not real-time information, which is like a robot simulating inside its head without moving its body. The second term presents &#x201c;open-loop prediction,&#x201d; by which the robot repeatedly predicts next-step data from the current situation one-by-one. We can use the feedback rate to adjust predictions. When we train the data, we set feedback rate <italic>&#x3b1;</italic> &#x3d; 0.1, meaning 90% of input data are previous closed-loop predictions and 10% are recorded training data. When testing a moving robot with actual data, we set the feedback rate <italic>&#x3b1;</italic> &#x3d; 0.2, meaning 80% of input data are closed-loop predictions and 20% are real-time raw&#x20;data.</p>
<p>We can control the robot according to predicted joint angle data, namely the value of <inline-formula id="inf3">
<mml:math id="m13">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. We then input <italic>X</italic> (<italic>t</italic>&#x20;&#x2b; 1) data to <xref ref-type="disp-formula" rid="e2">Eqs. 2</xref>, <xref ref-type="disp-formula" rid="e3">3</xref> as next input to the motion generation module. In other words, the predicted image feature data and force data (<inline-formula id="inf4">
<mml:math id="m14">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf5">
<mml:math id="m15">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>) are not used directly to control the robot but internally used to predict future data with closed-loop predictions. Repeating this process step-by-step, the robot can generate actions.</p>
</sec>
<sec id="s3-3-3">
<title>3.3.3 Backward Calculation</title>
<p>In backward calculation, we use the back propagation through time (BPTT) algorithm <xref ref-type="bibr" rid="B35">Rumelhart and McClelland (1987)</xref> to minimize the training error (E), calculated as<disp-formula id="e10">
<mml:math id="m16">
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">l</mml:mi>
<mml:mi mathvariant="normal">S</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>Y</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>We then update the weights as<disp-formula id="e11">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b7;</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
<mml:mi>E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(11)</label>
</disp-formula>where <italic>&#x3b7;</italic> is the learning rate, which we set as <italic>&#x3b7;</italic> &#x3d; 0.001, and <italic>n</italic> is the number of iterations. The initial value of the Cs layer (Cs(0)) is simultaneously updated to store features of the dynamics information as<disp-formula id="e12">
<mml:math id="m18">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b7;</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
<mml:mi>E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:math>
<label>(12)</label>
</disp-formula>At this time, we start training the module, setting all Cs(0) values to 0. We thus expect that features of tools, objects, actions, and effects will accumulate and self-organize in Cs(0) space, allowing the tool-use model to understand the relationship among the four factors. Therefore, we use the Cs(0) value as latent space. By inputting proper Cs(0) values to the trained network, it is possible to generate actions corresponding to the features of the four factors.</p>
</sec>
<sec id="s3-3-4">
<title>3.3.4 Exploring the Latent Space for Detecting Features From a Goal Image</title>
<p>When the robot deals with unknown tools or objects while testing this tool-use model, the Cs(0) value that best matches the task can be calculated from the trained network, setting error as<disp-formula id="e13">
<mml:math id="m19">
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>Y</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">T</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">l</mml:mi>
<mml:mi mathvariant="normal">S</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">a</mml:mi>
<mml:mi mathvariant="normal">l</mml:mi>
<mml:mi mathvariant="normal">S</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(13)</label>
</disp-formula>We need to provide initial image, joint angle and force data (T (1)) and final image data (T<sup>image</sup> (FinalStep)), which is the goal image. At this time, the output value <italic>y</italic>
<sup>image</sup> (FinalStep &#x2212;1) is calculated by setting the feedback rate <italic>&#x3b1;</italic> &#x3d; 0 in <xref ref-type="disp-formula" rid="e9">Eq. 9</xref>, which is fully closed-loop predictions. Then, by altering Cs(0) to minimize error as same as <xref ref-type="disp-formula" rid="e12">Eq. 12</xref>, the model can explore a proper Cs(0) value. This calculation is conducted for 20,000 epochs, by which time it has fully converged.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Experimental Setup</title>
<sec id="s4-1">
<title>4.1 System Design</title>
<p>We use a humanoid robot, NEXTAGE OPEN developed by Kawada Robotics <xref ref-type="bibr" rid="B17">KAWADA-Robotics (2020)</xref> and control it to conduct tasks with its right arm, which has 6 degrees of freedom. The robot has two cameras on its head, which have 9,216 dimensions (64 width &#xd7; 48 height &#xd7; 3 channels). We used only its right eye camera. A force sensor developed by WACOH-TECH <xref ref-type="bibr" rid="B45">WACOH-TECH (2020)</xref> attached to its wrist records 3-axis force data. A gripper developed by SAKE Robotics <xref ref-type="bibr" rid="B40">SAKE-Robotics (2020)</xref> is also attached.</p>
<p>We record joint angle data, an image, and force data every 0.1&#xa0;s, namely at a sampling frequency of 10&#xa0;Hz. Before inputting to the tool-use model, force data value are rescaled to [ &#x2212; 0.8, 0.8], and joint angle data are rescaled to [ &#x2212; 0.9, 0.9]. Image data are first simply scaled to [0, 255] for the feature extraction module. Then the output image feature data are rescaled to [ &#x2212; 0.8, 0.8] and input to the motion generation module.</p>
</sec>
<sec id="s4-2">
<title>4.2 Objects Used in the Experiments</title>
<p>
<xref ref-type="fig" rid="F2">Figure&#x20;2</xref> shows the tools and objects used in our experiments. For training, we prepare two kinds of tools (I- and T-shaped) and five kinds of objects (ball, small box, tall box, lying cylinder, and standing cylinder), all basic and simple shapes. We choose these to provide a variety of shapes and heights to cause different effects. As the two pictures on the left side of <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> show, tools or objects can change the effects. In the left picture, even though the object and the action are the same, the ball will not move if it is pulled with the I-shaped tool, but will roll if pulled with the T-shaped tool. In the middle picture, if we push left with the I-shaped tool, the box will slide but the ball will&#x20;roll.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Tools and objects used in the experiments. The upper images show the training setup. The middle images show untrained tools and objects similar to trained ones. These are used to evaluate accuracy of the tool-use model. Images at bottom show untrained tools and objects completely different from trained ones, used to test the model&#x2019;s generalizability.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Examples in which effects will differ by tools, objects, and actions. In the left picture, even though the object and the action are the same, the ball will not move if it is pulled with the I-shaped tool, but will roll if pulled with the T-shaped tool. In the middle picture, if we push left with the I-shaped tool, the box will slide but the ball will roll. In the right picture, even when using the same box and same I-shape tool, the effect of object behavior will differ depending on the height at which the robot pushes the box, toppling if pushed at a high point or sliding if pushed at a low&#x20;point.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g003.tif"/>
</fig>
<p>For evaluation experiments, some similar tools and objects, shown in the middle low in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, are prepared. We also prepared some totally different and more complex tools and objects, such as an umbrella, a wiper, a tree branch, a box much smaller than the trained one, a spray bottle, and a PET bottle. These are shown at the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>.</p>
</sec>
<sec id="s4-3">
<title>4.3 Task Design</title>
<p>The robot&#x2019;s task is to use a tool to move an object to the position and orientation shown in a goal image. The robot must generate actions with the right direction and the right positional height. For example, as on the right in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>, even when using the same box and same I-shape tool, the effect of object behavior will differ depending on the height at which the robot pushes the box, toppling if pushed at a high point or sliding if pushed at a low point. Moreover the robot needs to flexibly adjust the arm angles in real time not just fixing the direction and the height at once. Objects are easy to move with a little force and sometimes behave differently as expected.</p>
<p>We placed one object on a table in front of the robot so that its center position is set in the same position. In all tasks, the robot starts and ends movements at the same home position. We pass the robot a tool before it starts a task, so it initially grips it. This grip is maintained during all movements.</p>
</sec>
<sec id="s4-4">
<title>4.4 Training Dataset</title>
<p>To record training data, we remotely control the robot with a 3-dimensional mouse controller. For the training data, we designed four kinds of trajectories: sliding sideways or pulling toward the robot, with each action performed at either a high or low position, designed by the remote control. Then the robot is controlled to move according to the trajectories 5&#x20;times in each combination of tools and objects. Therefore, there are 200 training datasets (2 tools &#xd7; 5 objects &#xd7; 4 actions &#xd7; 5 trials for each task).</p>
<p>By keeping the robot stationary at its position after task completion until 10.7&#xa0;s from the start, we record the sensory motor data for 10.7&#xa0;s in every task, sampling each 0.1&#xa0;s. There thus are 107 steps for each&#x20;data.</p>
<p>
<xref ref-type="table" rid="T3">Table&#x20;3</xref> roughly categorizes the effects and summarizes each combination of tools, objects, and actions. There are several effects like &#x201c;shift to the left,&#x201d; &#x201c;shift to the front,&#x201d; &#x201c;roll to the left,&#x201d; &#x201c;roll to the front,&#x201d; &#x201c;topple to the left,&#x201d; &#x201c;topple to the front,&#x201d; and &#x201c;do not move.&#x201d; Although we categorized the effects to make them easy to understand, the actual effects differ one by&#x20;one.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Forty dataset combinations for training whose sensorimotor data is recorded in advance by remote controlling the robot. Effects are roughly categorized and colored differently. We allowed the robot to experience these combinations of tools, objects, actions, and effects for training the tool-use&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Tool Actions</th>
<th colspan="4" align="center">I-shape</th>
<th colspan="4" align="center">T-shape</th>
</tr>
<tr>
<th align="left">object</th>
<th align="center">Slide to the left low</th>
<th align="center">Slide to the left high</th>
<th align="center">Pull to the front low</th>
<th align="center">Pull to the front high</th>
<th align="center">Slide to the left low</th>
<th align="center">Slide to the left high</th>
<th align="center">Pull to the front low</th>
<th align="center">Pull to the front high</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Small box</td>
<td align="left">Shift tothe left</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Shift to the left</td>
<td align="left">Does not move</td>
<td align="left">Shift to the front</td>
<td align="left">Does not move</td>
</tr>
<tr>
<td align="left">Tall box</td>
<td align="left">Shift to the left</td>
<td align="left">Topple to the left</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Shift to the left</td>
<td align="left">Topple to the left</td>
<td align="left">Shift to the front</td>
<td align="left">Topple to the front</td>
</tr>
<tr>
<td align="left">Lying cylinder</td>
<td align="left">Roll to the left</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Roll to the left</td>
<td align="left">Does not move</td>
<td align="left">Shift to the front</td>
<td align="left">Does not move</td>
</tr>
<tr>
<td align="left">Standing cylinder</td>
<td align="left">Shift to the left</td>
<td align="left">Topple to the left</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Shift to the left</td>
<td align="left">Topple to the left</td>
<td align="left">Shift to the front</td>
<td align="left">Topple to he front</td>
</tr>
<tr>
<td align="left">Ball</td>
<td align="left">Roll to the left</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Does not move</td>
<td align="left">Roll to the left</td>
<td align="left">Does not move</td>
<td align="left">Roll to the front</td>
<td align="left">Does not move</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>There are two differences between &#x201c;shift&#x201d; and &#x201c;roll.&#x201d; &#x201c;Shift&#x201d; is a movement where an object moves beside a tool and stops movement just after the robot stretches its arm to its full extent. In contrast, &#x201c;roll&#x201d; is a movement where an object precedes the arm movement and keeps moving for a while after the robot fully extends its arm, so the object moves differently every time and there are a variety of goal positions.</p>
</sec>
<sec id="s4-5">
<title>4.5 Experimental Evaluation</title>
<p>When testing the tool-use model with actual data taken during moving robot, the setup for recording data and constructing the model is same as in training. However, we set the feedback rate <italic>&#x3b1;</italic>&#x20;&#x3d;&#x20;0.2 as described in <xref ref-type="sec" rid="s3-3-2">section 3.3.2</xref>. Therefore, 80% of input data are closed-loop predictions and 20% are real-time raw data. By increasing the feedback rate from that in training, the robot more easily adjusts to real-time situations.</p>
<p>Three evaluations are conducted:<list list-type="simple">
<list-item>
<p>&#x2022; analyze the training results,</p>
</list-item>
<list-item>
<p>&#x2022; evaluate task execution accuracy with similar objects or tools,&#x20;and</p>
</list-item>
<list-item>
<p>&#x2022; evaluate generalizability with totally different tools and objects.</p>
</list-item>
</list>
</p>
<sec id="s4-5-1">
<title>4.5.1 Analysis of Training Results</title>
<p>At first, we check if the training can be conducted well. We first confirm that the feature extraction module with the CAE can properly extract image features. We reconstruct the images input by the module and confirm whether the reconstructed images are similar to the input. If the module performed reconstruction well, that means it could accumulate the essential characteristics of images in low-dimensional image feature&#x20;data.</p>
<p>Second, we evaluate whether the motion generation module with the MTRNN can acquire relationships among tools, objects, actions and effects. We analyze the latent space, the initial step of Cs neuron value (Cs(0)) of each training data trained by <xref ref-type="disp-formula" rid="e12">Eq. 12</xref> by principal component analysis (PCA). Cs(0) values for similar training data are expected to be clustered and different values should be apart, meaning the Cs(0) space can well express features of the factors of each task. In other words, we can confirm whether the features are self-organized. We analyze three maps of Cs(0) values, each presenting features of tools, objects, and actions.</p>
</sec>
<sec id="s4-5-2">
<title>4.5.2 Accuracy Evaluation</title>
<p>We then conduct evaluation experiments moving the robot. Task execution accuracy is evaluated using unknown objects and tools which are similar to trained ones, shown in the middle of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. To carefully confirm the robot&#x2019;s ability to detect both objects and tools, the experiments are conducted with combinations of trained tools and unknown objects, and with unknown tools and trained objects.</p>
<p>We provide the tool-use model with goal images and explore proper Cs(0) value by <xref ref-type="disp-formula" rid="e13">Eq. 13</xref>. The model then generates actions using the explored Cs(0) values, and we confirm whether the robot can reproduce the expected effects. We measure success rates according to object behaviors during robot actions, and the final positions and orientations of the objects. Regarding object behavior, we confirm whether they clearly roll, shift, or topple. Regarding final positions, we regard shift trials as successful if the final object position is within one-thirds distance from the home position to the goal position. Regarding orientation, we regard trials as successful if the difference in inclination between final and target orientations is less than 30&#xb0;.</p>
<p>We also analyze explored Cs(0) values and check if the tool-use model can detect and express the features of tools, objects and expected actions by comparison with Cs(0) values in the training data. This is confirmed by superimposing PCA results for the explored Cs(0) on the three Cs(0) training data maps described&#x20;above.</p>
<p>This experiment is conducted by providing 12 goal images with an I-shaped tool, and 16 goal images with a T-shaped tool to show different effects. There are 20 combinations of objects and actions, but there are some same &#x201c;does not move&#x201d; effects, as shown in <xref ref-type="table" rid="T3">Table. 3</xref>. With the T-shaped tool there are also &#x201c;roll to the left&#x201d; and &#x201c;roll to the right&#x201d; effects that result in random goal positions. Therefore, there are 12 effects for the I-shaped tool and 16 for the T-shaped tool. All tasks are conducted 3&#x20;times in both experiments, with combinations of trained tools and unknown objects, and unknown tools and trained objects. Thus the experiment is performed 168 trials ((12 &#x2b; 16) &#xd7; 3&#x20;&#xd7;&#x20;2).</p>
</sec>
<sec id="s4-5-3">
<title>4.5.3 Generalization Evaluation</title>
<p>In the last evaluation experiment, generalizability of the tool-use model is confirmed. The procedure is same as in Accuracy Evaluation, except that both tools and objects are totally different, as shown at the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. We set goal images expecting the robot to slide a very small box in a low position to the left with an umbrella, slide a spray bottle in a high position to topple it with a wiper, and pull a PET bottle in high position to topple it with a tree branch. The experiment is conducted 3&#x20;times each, so there are nine trials (3&#x20;&#xd7;&#x20;3).</p>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5 Result</title>
<p>In this section, we show the result of the training and two evaluation experiments: Accuracy Evaluation and Generalization Evaluation.</p>
<sec id="s5-1">
<title>5.1 Training Analysis</title>
<p>The reconstructed images in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> suggest that the trained feature extraction module could reproduce the original input images. The input images shown in the figure are test data not used for training. All are well reconstructed, demonstrating that the feature extraction module could well express image features in output from the center&#x20;layer.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Images reconstruction by the feature extraction module with the CAE. Upper images are original test data with training tools and objects, which are input to the module. The bottom images are reconstructed by the module. All images show their shape and size well, which demonstrates that the feature extraction module can well express image features in output from the center&#x20;layer.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g004.tif"/>
</fig>
<p>We performed PCA on the internal latent space values, Cs(0) of the training data. The results are shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> as solid circles. The plots are colored according to each feature in the three maps: tools, objects, and actions. As a result, in every map plots of different features are separated and same features are clustered, demonstrating that positions in Cs(0) value maps can express the features. Focusing on the map axes, we can also say that PC1 and PC2 represent action features, with PC1 indicating action type (sliding or pulling) and PC2 representing heights. PC3 represents tool features, and PC4 and PC5 represent object features. Cs(0) values could simultaneously express tools, objects, and actions which can reproduce goal effects.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Results of PCA for the internal latent space values, Cs(0). We change the axis to show for expressing map of tools, objects and actions. Trained tasks (solid circles) are separated and clustered according to the features of tools, objects, and actions. Focusing on the map axes, we can also say that PC1 and PC2 represent action features, with PC1 indicating action type (sliding or pulling) and PC2 representing heights. PC3 represents tool features, and PC4 and PC5 represent object features. Cs(0) values could simultaneously express tools, objects, and actions which can reproduce goal effects. Values for explored Cs(0) in the tested tasks are superimposed on the maps as hollow circles and triangles, and black stars, inverted triangles, and diamonds. Many plots are plotted in the proper regions, suggesting good feature detection.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g005.tif"/>
</fig>
</sec>
<sec id="s5-2">
<title>5.2 Accuracy Evaluation</title>
<p>In the experiment of Accuracy Evaluation, the robot is expected to use similar unknown objects or tools. We confirmed whether the robot can reproduce situations in goal images. For example, as <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> shows, the goal image shows the small yellow box shifted to the left. The DNN model explored proper Cs(0) values using the goal images, allowing the robot to generate motions. <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> shows camera images while the robot is moving. In this example, the robot moved properly to make the final image similar to the goal image, matching the success definition rule. We also show the generated trajectory with dotted lines. The lines are similar to but sometimes shifted from the solid lines, that is training trajectory for shifting the trained small box to the left. Therefore, we can say that the robot could detect proper action type and adjust it depending on the real time situation.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Example result of Accuracy Evaluation using unknown objects and tools shown in the middle of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. These pictures show camera images while the robot was conducting the task. We provided a goal image showing the untrained small yellow box, shifted to the left from the initial position. The robot then detected features of the tool and the object and predicted a suitable action, generating an action that would properly slide the box in a low position to move it left, reproducing the goal situation. The bottom graph shows the generated angle trajectories with solid lines. We compare the lines to trained angle trajectories for shifting the small box shown in the upper of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> to the left with dotted lines. The trajectory lines are almost same but sometimes shifted. We can say that the robot could detect proper action type and adjust it depending on the real time situation.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g006.tif"/>
</fig>
<p>
<xref ref-type="table" rid="T4">Table&#x20;4</xref> summarizes success rates. The robot succeeded in performing 83% (70/84) of tasks with unknown objects and 79% (66/84) of tasks with unknown tools. Notably, the robot had more difficulty dealing with tall objects and T-shaped tools. We assume that the reason for this is that they contain topple effects. If we look only at tasks for that effect, the success rate is further reduced to 61% (22/36). The topple effect causes sudden changes in image and force sensor data, causing large sudden changes in DNN model input, making this much more difficult to learn than other effects.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Success rates for Accuracy Evaluation using unknown objects and tools shown in the middle of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. We confirm whether the robot can reproduce the expected effects. We measure success rates according to object behaviors during robot actions, and the final positions and orientations of the objects.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Untrained objects or tools</th>
<th align="center">Success rate</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<bold>Small box</bold>
</td>
<td align="center">15/15 (100%)</td>
</tr>
<tr>
<td align="left">
<bold>Tall box</bold>
</td>
<td align="center">16/21 (76%)</td>
</tr>
<tr>
<td align="left">
<bold>Toppled cylinder</bold>
</td>
<td align="center">13/15 (87%)</td>
</tr>
<tr>
<td align="left">
<bold>Standing cylinder</bold>
</td>
<td align="center">14/21 (67%)</td>
</tr>
<tr>
<td align="left">
<bold>Ball</bold>
</td>
<td align="center">12/12 (100%)</td>
</tr>
<tr>
<td align="left">
<bold>Total untrained objects</bold>
</td>
<td align="center">70/84 (83%)</td>
</tr>
<tr>
<td align="left">
<bold>I-shaped tool</bold>
</td>
<td align="center">31/36 (86%)</td>
</tr>
<tr>
<td align="left">
<bold>T-shaped tool</bold>
</td>
<td align="center">35/48 (73%)</td>
</tr>
<tr>
<td align="left">
<bold>Total untrained tools</bold>
</td>
<td align="center">66/84 (79%)</td>
</tr>
<tr>
<td align="left">
<bold>Combined total</bold>
</td>
<td align="center">136/168 (81%)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We also analyzed the explored Cs(0) values and plotted them on the space of trained ones, as shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. There are three trials of each task, but we plotted only one with successful results. All tasks succeeded at least once. Cs(0) values for tasks using unknown objects are shown as hollow circles, and tasks using unknown tools are shown as hollow triangles. Plots of tools and objects are colored according to their shape features. For actions, the plots are categorized and colored according to the most similar trained actions: pull low, pull high, slide low, or slide high. Almost all plots are properly represented in the same feature spaces, meaning the tool-use model could explore proper Cs(0) values and detect features of tools and objects, and expect suitable actions to conduct in untrained&#x20;tasks.</p>
<p>When plotting Cs(0) values for failed tasks as a test, there were three types of mistaken distributions. Many were plotted at totally different far spaces, indicating the tool-use model could not detect the features. Some were plotted on the middle of feature clusters especially in the action map. This often occurs for the &#x201c;does not move&#x201d; effect, where multiple actions are possible to achieve the target. For example, we allowed the robot to experience &#x201c;pull high,&#x201d; &#x201c;pull low,&#x201d; and &#x201c;slide high&#x201d; as training data for behaviors that can reproduce the &#x201c;does not move&#x201d; effect in a combination of the ball and the I-shaped tool. We therefore suspect that the robot could not select one action from among the candidates. When we forcibly moved the robot with the ambiguous Cs(0) value, it sometimes mixed some actions and other times just waved its arm near the initial position. Finally, regarding the third mistaken distribution, Cs(0) values are plotted in mistaken combinations of clusters. For example, when the real combination was &#x201c;T-shaped tool, standing cylinder, slide high, topple to the left,&#x201d; the model detected this relationship as the combination &#x201c;T-shaped tool, standing cylinder, pull low, topple forward.&#x201d; This happened because the robot correctly understood the relationships among the four factors, but misunderstood the effects.</p>
</sec>
<sec id="s5-3">
<title>5.3 Generalization Evaluation</title>
<p>In Generalization Evaluation, we used totally different tools and objects. These tools and objects are more complex than trained ones, and thus the objects behave differently from trained situations. Therefore the task cannot be completed just by replay the training actions and the robot needs to adjust its movement step-by-step to manipulate the object with the tool. <xref ref-type="fig" rid="F7">Figures 7</xref>&#x2013;<xref ref-type="fig" rid="F9">9</xref> show the results of the robot&#x2019;s action generation. In the first task, the robot slid the umbrella to the left from a low position and shifted the very small box to the left. In the second task, the robot slid the wiper to the left from a high position and toppled the spray bottle to the left. In the third task, the robot pulled the tree branch to the front from a high position and toppled the PET bottle forward. In all the case, the robot could properly move and reproduce the goal situations, adjusting to the features of tools and objects, and the real time situations. We confirmed tool-use ability of the robot with the tool-use model that can be generalized to use unknown tools and objects.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Result of Generalization Evaluation with a very small box and an umbrella shown in the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. The pictures are the robot&#x2019;s eye camera images and the whole view taken by an external camera. We provided a goal image showing the very small box shifted left. The robot could properly detect features and generate a motion to slide to the left, reproducing the situation in the goal&#x20;image.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g007.tif"/>
</fig>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Result of Generalization Evaluation with a spray bottle and a wiper shown in the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. We provided a goal image showing the spray bottle toppled to the left. The robot could properly detect features and generate motions to push to the left from a high point and reproduce the situation in the goal&#x20;image.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g008.tif"/>
</fig>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Result of Generalization Evaluation with a PET bottle and a tree branch shown in the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. We provided a goal image showing the PET bottle toppled forward. The robot could properly detect features and generate motions to pull forward from a high position, reproducing the situation in the goal&#x20;image.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g009.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F5">Figure&#x20;5</xref> shows the explored Cs(0) value plotted on the space of trained ones. Black stars show the combination of the very small box and the umbrella, black inverted triangles show the combination of the spray bottle and the wiper, and black diamonds show the combination of the PET bottle and the tree branch. The results suggest recognition by the robot, which had to detect each combination as &#x201c;tool close to I-shape, ball like object, slide to the left from a low position&#x201d; in the first task, &#x201c;tool close to I-shape, tall box like object, slide to the left from a high position&#x201d; in the second task, and &#x201c;tool close to T-shape, standing cylinder like object, pull forward from a high position&#x201d; in the third task. The tools and objects are recognized as ones similar in shape, and actions are also correctly matched. Most importantly, these detected combinations result in the expected effects. These results indicate that the tool-use model could detect the features of tools, objects, actions, and effects considering the four relationships, even if the provided tools and objects are unknown.</p>
<p>
<xref ref-type="fig" rid="F10">Figure&#x20;10</xref> shows the decoded images using predicted image feature data output by the motion generation module. The original camera images are taken during the Generalization Evaluation experiments. The decoded images represent the shape, size, position, and orientation of the tools and objects used. They also show the position of the robot arm. It can be confirmed that the motion was generated while capturing the real time state during the movement. Note that, the color and appearance of the tools or objects in the reconstructed images are close to the ones of training data because the feature extraction module recognizes the feature of the tool or object based on the training&#x20;data.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Decoded images using predicted image features output from the motion generation module. The original camera images are taken during the movement dealing with the unknown tools and the objects shown in the bottom of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. Although the decoded images shows similar color of the training tools and objects since the feature extraction module recognizes their features based on the training ones, they represent the shape, size, position, and orientation of the tools and objects in the robot camera view. They also show the position of the robot arm. It can be confirmed that the motion was generated while capturing the real time state during the movement.</p>
</caption>
<graphic xlink:href="frobt-08-748716-g010.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<title>6 Discussion</title>
<p>For summary, the tool-use model has two neural network modules: 1) a feature extraction module with a CAE trained to extract visual features from captured raw images, and 2) a motion generation module with a MTRNN and fully connected layers that integrates and predicts multimodal sensory-motor information. Through training of image reconstruction by the feature extraction module, the robot could extract image features from raw images captured by its camera. The motion generation module learns coordination of image feature data, joint angle data, and force sensor data, and performs next-step data predictions. In addition, the motion generation module can express the relationship of tools, objects, actions and effects with an internal latent space value, Cs(0). Using a provided goal image, the robot can generate actions by exploring Cs(0) values matched to the&#x20;task.</p>
<p>Analyzing after training the DNN using task experience data, the tool-use model is able to self-organize tools, objects, and actions, and to automatically create maps representing their features in the Cs(0) space. With Cs(0) values, the tool-use model expresses the combination of the tools, objects, and actions needed to produce the target effects. In other words, the tool-use model understands combinations of the tools, objects, actions, and effects, not individually, meaning it acquires the relationships among the four factors. Then, we performed two experiments that confirmed the learning model&#x2019;s accuracy and generalizability. The robot succeeded in 81% of tasks with unknown similar objects or tools. Moreover, the robot demonstrated task executions that reproduced target situations with unknown, complicated, totally different tools and objects. The results demonstrate that the features of tools and objects can be detected, and optimal actions can be generated based on the acquired relationships and the real time situations. In summary, the robot gained the relationships through its own experience, allowing it to consider combinations of tools, objects, and actions necessary to achieve its goal effects. This study is valuable as a novel robot control system. At first, the tool-use model does not require pre-calculation or pre-definition of tools and objects, nor motor command instructions from humans. We can assign tasks just by providing a goal image. Second, we can construct the tool-use model with relatively small training cost and make the model adaptable to unknown targets. Our work requires 200 training data for 10.7&#xa0;s each. It took us only half a day to record that data. On the other hand, like methods using reinforcement learning require a large amount of training data, and it is difficult to record the data with actual robots. Studies that use simulation environments can overcome the difficulty of collecting data, however the difficulty of having to make up for the gaps from the real environment remains. In addition, our tool-use model shows broader generalization capabilities compared to previous studies. By acquiring the relationships among the four factors, our model has achieved what was not realized in many other research: it can handle both tools and objects which are unknown, and it can flexibly adjust actions in real time. Finally, multimodal learning solved the occlusion problem, and increased the range of objects that can be handled and the actions that can be taken. In the previous study Ref. <xref ref-type="bibr" rid="B36">Saito et&#x20;al. (2018a)</xref>, the adopted sensor was only image without force, so they could not use small objects and had many restrictions on the types of actions, because if an object was occluded by the arm, it would be misrecognized. Thanks to these achievement, our model can contribute to the realization of robots that can handle various tasks, it can greatly impact practical applications for robots in everyday environments. It is also useful for production at small quantities and wide variety.</p>
<p>However, there are some limitations in this work. First, tools and objects that change features like color, shape, or size during movement cannot be used, because it is difficult to explore Cs(0) values and detect the features if the goal image significantly differs from the initial image. Second, if objects or tools significantly differ from the trained ones, it will be difficult to handle them. If the appearances of them or effects are completely different, the feature extraction module and the motion generation module need to be trained again. Moreover, if their weights are too light or friction is too small, making them too easy to move. The robot would then struggle to operate them, because the motion generation model cannot predict next joint angles that largely differ from current ones, so there is a limit to the speed range that can be output.</p>
<p>In future studies, we will improve the tool-use model so that robots select tools suited to the situations and objects. In this model, the robot is initially grasping a tool, so while the robot considers how to move to accomplish its task, it does not consider tool selection. Reference <xref ref-type="bibr" rid="B37">Saito et&#x20;al. (2018b)</xref> focused on only tool selection without considering relationships between objects, so we will combine these two studies so that robots must consider both how to move and which tool is best suited to accomplishing the task. Another area for future work is to have robots come up with novel ways to use tools. Realizing that ability would allow robots to use whatever tools happen to be available without instructions on their use, possibly generating unexpected behaviors.</p>
</sec>
<sec id="s7">
<title>7 Conclusion</title>
<p>We realized a robot that could acquire the relationships among tools, objects, actions and effects enabling tool-use, even for tools and objects being seen for the first time. The tool-use model learns sensorimotor coordination and the relationships among the four factors by training with data recorded during tool-use experiences. Unlike previous studies, our DNN model can simultaneously consider tools, objects, actions, and effects with no pre-definitions. Moreover, the robot can treat unknown tools and objects based on the relationships by generalizability of the tool-use model. Another advantage is that the robot can generate motions step-by-step, not just by choosing and replaying designed motor commands, so the robot can adjust its motions to cope with uncertain situations. Finally, the robot can find task goals and accomplish them just by providing goal images. This is one of the easiest ways to demonstrate to the robot the desired position and orientation. We confirmed the accuracy and generalization ability of the tool-use model with real robot experiments.</p>
</sec>
</body>
<back>
<sec id="s8">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>NS designed the research, performed the experiments, analyzed the data and wrote the paper. TO, HM, SM, and SS contributed to the supervision of the experiments and the writing of the article.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This research was partially supported by JST Moonshot R and D No. JPMJMS 2031, JSPS Grant-in-Aid for Scientific Research (A) No. 19H01130, and by the Research Institute for Science and Engineering of Waseda University.</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s13">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frobt.2021.748716/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frobt.2021.748716/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Video1.MP4" id="SM1" mimetype="application/MP4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Beetz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Klank</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Kresse</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Maldonado</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mosenlechner</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pangercic</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). &#x201c;<article-title>Robotic Roommates Making Pancakes</article-title>,&#x201d; in <conf-name>2011 IEEE-RAS 11th International Conference on Humanoid Robots (Humanoids)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>529</fpage>&#x2013;<lpage>536</lpage>. <pub-id pub-id-type="doi">10.1109/humanoids.2011.6100855</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bishop</surname>
<given-names>C. M.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Pattern Recognition and Machine Learning</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brown</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sammut</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>A Relational Approach to Tool-Use Learning in Robots</article-title>,&#x201d; in <conf-name>International Conference on Inductive Logic Programming</conf-name> (<publisher-name>Springer</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>15</lpage>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bushnell</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Boudreau</surname>
<given-names>J.&#x20;P.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>Motor Development and the Mind: The Potential Role of Motor Abilities as a Determinant of Aspects of Perceptual Development</article-title>. <source>Child. Dev.</source> <volume>64</volume>, <fpage>1005</fpage>&#x2013;<lpage>1021</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-8624.1993.tb04184.x</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cavallo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Limosani</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Manzi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bonaccorsi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Esposito</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Di Rocco</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Development of a Socially Believable Multi-Robot Solution from Town to home</article-title>. <source>Cogn. Comput.</source> <volume>6</volume>, <fpage>954</fpage>&#x2013;<lpage>967</lpage>. <pub-id pub-id-type="doi">10.1007/s12559-014-9290-z</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chemero</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>An Outline of a Theory of Affordances</article-title>. <source>Ecol. Psychol.</source> <volume>15</volume>, <fpage>181</fpage>&#x2013;<lpage>195</lpage>. <pub-id pub-id-type="doi">10.1207/s15326969eco1502_5</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dehban</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Jamone</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kampff</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>Santos-Victor</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Denoising Auto-Encoders for Learning of Objects and Tools Affordances in Continuous Space</article-title>,&#x201d; in <conf-name>2016 IEEE International Conference on Robotics and Automation (ICRA)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>4866</fpage>&#x2013;<lpage>4871</lpage>. <pub-id pub-id-type="doi">10.1109/icra.2016.7487691</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ernst</surname>
<given-names>M. O.</given-names>
</name>
<name>
<surname>Banks</surname>
<given-names>M. S.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Humans Integrate Visual and Haptic Information in a Statistically Optimal Fashion</article-title>. <source>Nature</source> <volume>415</volume>, <fpage>429</fpage>&#x2013;<lpage>433</lpage>. <pub-id pub-id-type="doi">10.1038/415429a</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Fukui</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shimojo</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1994</year>). &#x201c;<article-title>Recognition of Virtual Shape Using Visual and Tactual Sense under Optical Illusion</article-title>,&#x201d; in <conf-name>1994 3rd IEEE International Workshop on Robot and Human Communication</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>294</fpage>&#x2013;<lpage>298</lpage>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gibson</surname>
<given-names>G. D.</given-names>
</name>
</person-group> (<year>1977</year>). <article-title>Himba Epochs</article-title>. <source>Hist. Afr.</source> <volume>4</volume>, <fpage>67</fpage>&#x2013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.2307/3171580</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gibson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1979</year>). <source>The Ecological Approach to Visual Perception: Classic Edition</source>. <publisher-loc>Boston</publisher-loc>: <publisher-name>Houghton Mifflin</publisher-name>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gibson</surname>
<given-names>K. R.</given-names>
</name>
<name>
<surname>Ingold</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1993</year>). <source>Tools, Language and Cognition in Human Evolution</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>, <fpage>389</fpage>&#x2013;<lpage>447</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Goncalves</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Abrantes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Saponaro</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jamone</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bernardino</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Learning Intermediate Object Affordances: Towards the Development of a Tool Concept</article-title>,&#x201d; in <conf-name>4th International Conference on Development and Learning and on Epigenetic Robotics</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>482</fpage>&#x2013;<lpage>488</lpage>. <pub-id pub-id-type="doi">10.1109/devlrn.2014.6983027</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Reducing the Dimensionality of Data with Neural Networks</article-title>. <source>Science</source> <volume>313</volume>, <fpage>504</fpage>&#x2013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Horton</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Chakraborty</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Amant</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Affordances for Robots: A Brief Survey</article-title>,&#x201d; in <source>AVANT</source> (<publisher-loc>Bydgoszcz</publisher-loc>: <publisher-name>Pismo Awangardy Filozoficzno-Naukowej</publisher-name>), <fpage>70</fpage>&#x2013;<lpage>84</lpage>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jamone</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ugur</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Cangelosi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fadiga</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bernardino</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Piater</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Affordances in Psychology, Neuroscience, and Robotics: A Survey</article-title>. <source>IEEE Trans. Cogn. Dev. Syst.</source> <volume>10</volume>, <fpage>4</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1109/tcds.2016.2594134</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="web">
<collab>KAWADA-Robotics</collab> (<year>2020</year>). <article-title>Next Generation Industrial Robot Nextage</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://www.kawadarobot.co.jp/nextage/">http://www.kawadarobot.co.jp/nextage/</ext-link>
</comment>. </citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <source>Adam: A Method for Stochastic Optimization</source>. <publisher-name>CoRR abs/1412.6980</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Imagenet Classification with Deep Convolutional Neural Networks</article-title>,&#x201d; in <conf-name>25th International Conference on Neural Information Processing Systems - Volume 1</conf-name> (<publisher-name>Curran Associates Inc.</publisher-name>), <fpage>1097</fpage>&#x2013;<lpage>1105</lpage>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Visual-Tactile Fusion for Object Recognition</article-title>. <source>IEEE Trans. Automat. Sci. Eng.</source> <volume>14</volume>, <fpage>996</fpage>&#x2013;<lpage>1008</lpage>. <pub-id pub-id-type="doi">10.1109/tase.2016.2549552</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lockman</surname>
<given-names>J.&#x20;J.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>A Perception-Action Perspective on Tool Use Development</article-title>. <source>Child. Dev.</source> <volume>71</volume>, <fpage>137</fpage>&#x2013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1111/1467-8624.00127</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mar</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tikhanoff</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Natale</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>What Can I Do with This Tool? Self-Supervised Learning of Tool Affordances from Their 3-d Geometry</article-title>. <source>IEEE Trans. Cogn. Dev. Syst.</source> <volume>10</volume>, <fpage>595</fpage>&#x2013;<lpage>610</lpage>. <pub-id pub-id-type="doi">10.1109/tcds.2017.2717041</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Masci</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Meier</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Cire&#x15f;an</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Stacked Convolutional Auto-Encoders for Hierarchical Feature Extraction</article-title>,&#x201d; in <source>Artificial Neural Networks and Machine Learning &#x2013; ICANN</source> (<publisher-name>Springer</publisher-name>), <fpage>52</fpage>&#x2013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-21735-7_7</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Min</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yi</surname>
<given-names>C. a.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bi</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Affordance Research in Developmental Robotics: A Survey</article-title>. <source>IEEE Trans. Cogn. Dev. Syst.</source> <volume>8</volume>, <fpage>237</fpage>&#x2013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1109/tcds.2016.2614992</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Myers</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Teo</surname>
<given-names>C. L.</given-names>
</name>
<name>
<surname>Ferm&#xfc;ller</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Aloimonos</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Affordance Detection of Tool Parts from Geometric Features</article-title>,&#x201d; in <conf-name>2015 IEEE International Conference on Robotics and Automation</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1374</fpage>&#x2013;<lpage>1381</lpage>. <pub-id pub-id-type="doi">10.1109/icra.2015.7139369</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Nabeshima</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kuniyoshi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lungarella</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Towards a Model for Tool-Body Assimilation and Adaptive Tool-Use</article-title>,&#x201d; in <conf-name>IEEE 6th International Conference on Development and Learning</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>288</fpage>&#x2013;<lpage>293</lpage>. <pub-id pub-id-type="doi">10.1109/devlrn.2007.4354031</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Nagahama</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yamazaki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Okada</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Inaba</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Manipulation of Multiple Objects in Close Proximity Based on Visual Hierarchical Relationships</article-title>,&#x201d; in <conf-name>2013 IEEE International Conference on Robotics and Automation</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1303</fpage>&#x2013;<lpage>1310</lpage>. <pub-id pub-id-type="doi">10.1109/icra.2013.6630739</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nishide</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tani</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Takahashi</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Okuno</surname>
<given-names>H. G.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Tool-Body Assimilation of Humanoid Robot Using a Neurodynamical System</article-title>. <source>IEEE Trans. Auton. Ment. Dev.</source> <volume>4</volume>, <fpage>139</fpage>&#x2013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1109/tamd.2011.2177660</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Norman</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2013</year>). <source>The Design of Everyday Things</source>. <publisher-name>Basic Books</publisher-name>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Osiurak</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Jarry</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Le Gall</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Grasping the Affordances, Understanding the Reasoning: Toward a Dialectical Theory of Human Tool Use</article-title>. <source>Psychol. Rev.</source> <volume>117</volume>, <fpage>517</fpage>&#x2013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1037/a0019004</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Kinematic Control of a Dual-Arm Humanoid mobile Cooking Robot</article-title>,&#x201d; in <conf-name>I-CREATe 2018 Proceedings of the 12th International Convention on Rehabilitation Engineering and Assistive Technology</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>308</fpage>&#x2013;<lpage>311</lpage>. </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Piaget</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1952</year>). <source>The Origins of Intelligence in Children</source>. <publisher-loc>New York, US</publisher-loc>: <publisher-name>W. W. Norton</publisher-name>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ramirez-Amaro</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Dean-Leon</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Robust Semantic Representations for Inferring Human Co-manipulation Activities Even with Different Demonstration Styles</article-title>,&#x201d; in <conf-name>2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1141</fpage>&#x2013;<lpage>1146</lpage>. <pub-id pub-id-type="doi">10.1109/humanoids.2015.7363496</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rat-Fischer</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>O&#x2019;Regan</surname>
<given-names>J.&#x20;K.</given-names>
</name>
<name>
<surname>Fagard</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The Emergence of Tool Use during the Second Year of Life</article-title>. <source>J.&#x20;Exp. Child Psychol.</source> <volume>113</volume>, <fpage>440</fpage>&#x2013;<lpage>446</lpage>. <pub-id pub-id-type="doi">10.1016/j.jecp.2012.06.001</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rumelhart</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>McClelland</surname>
<given-names>J.&#x20;L.</given-names>
</name>
</person-group> (<year>1987</year>). <source>Learning Internal Representations by Error Propagation (MITP), Vol. 1. 1</source>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Saito</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Murata</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018a</year>). &#x201c;<article-title>Detecting Features of Tools, Objects, and Actions from Effects in a Robot Using Deep Learning</article-title>,&#x201d; in <conf-name>2018 Joint IEEE 8th International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/devlrn.2018.8761029</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Saito</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Murata</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018b</year>). &#x201c;<article-title>Tool-use Model Considering Tool Selection by a Robot Using Deep Learning</article-title>,&#x201d; in <conf-name>The IEEE-RAS 18th International Conference on Humanoid Robots (Humanoids)</conf-name> (<publisher-name>IEEE</publisher-name>). <pub-id pub-id-type="doi">10.1109/humanoids.2018.8625048</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saito</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Funabashi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mori</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>How to Select and Use Tools? : Active Perception of Target Objects Using Multimodal Deep Learning</article-title>. <source>IEEE Robot. Autom. Lett.</source> <volume>6</volume>, <fpage>2517</fpage>&#x2013;<lpage>2524</lpage>. <pub-id pub-id-type="doi">10.1109/lra.2021.3062004</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Saito</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Mori</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Wiping 3d-Objects Using Deep Learning Model Based on Image and Force and Joint Information</article-title>,&#x201d; in <conf-name>2020 IEEE/RSJ International Conference on Intelligent Robots and Systems</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>814</fpage>&#x2013;<lpage>819</lpage>. </citation>
</ref>
<ref id="B40">
<citation citation-type="web">
<collab>SAKE-Robotics</collab> (<year>2020</year>). <article-title>Sake Robotics: Robot Grippers, Ezgripper</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://sakerobotics.com/">https://sakerobotics.com/</ext-link>
</comment>. </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoffregen</surname>
<given-names>T. A.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Affordances as Properties of the Animal-Environment System</article-title>. <source>Ecol. Psychol.</source> <volume>15</volume>, <fpage>115</fpage>&#x2013;<lpage>134</lpage>. <pub-id pub-id-type="doi">10.1207/s15326969eco1502_2</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Stoytchev</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Behavior-grounded Representation of Tool Affordances</article-title>,&#x201d; in <conf-name>2005 IEEE International Conference on Robotics and Automation</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>3060</fpage>&#x2013;<lpage>3065</lpage>. </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Takahashi</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Tool-body Assimilation Model Considering Grasping Motion through Deep Learning</article-title>. <source>Robotics Autonomous Syst.</source> <volume>91</volume>, <fpage>115</fpage>&#x2013;<lpage>127</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2017.01.002</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Turvey</surname>
<given-names>M. T.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>Affordances and Prospective Control: An Outline of the Ontology</article-title>. <source>Ecol. Psychol.</source> <volume>4</volume>, <fpage>173</fpage>&#x2013;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1207/s15326969eco0403_3</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="web">
<collab>WACOH-TECH</collab> (<year>2020</year>). <article-title>Wacoh-tech Inc. Products, Dynpick</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://wacoh-tech.com/en/">https://wacoh-tech.com/en/</ext-link>
</comment>. </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yamashita</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tani</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Emergence of Functional Hierarchy in a Multiple Timescale Neural Network Model: A Humanoid Robot experiment</article-title>. <source>Plos Comput. Biol.</source> <volume>4</volume>, <fpage>e1000220</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000220</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yamazaki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ueda</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nozawa</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kojima</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Okada</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Matsumoto</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Home-assistant Robot for an Aging Society</article-title>. <source>Proc. IEEE</source> <volume>100</volume>, <fpage>2429</fpage>&#x2013;<lpage>2441</lpage>. <pub-id pub-id-type="doi">10.1109/jproc.2012.2200563</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>P.-C.</given-names>
</name>
<name>
<surname>Sasaki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Suzuki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kase</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Sugano</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ogata</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Repeatable Folding Task by Humanoid Robot Worker Using Deep Learning</article-title>. <source>IEEE Robotics Automation Lett. (Ra-l)</source> <volume>2</volume>, <fpage>397</fpage>&#x2013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2016.2633383</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.-C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Understanding Tools: Task-Oriented Object Modeling, Learning and Recognition.</article-title>,&#x201d; in <conf-name>2015 IEEE International Conference on Computer Vision and Pattern Recognition</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>2855</fpage>&#x2013;<lpage>2864</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2015.7298903</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>