<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">838059</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2022.838059</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Visual Rewards From Observation for Sequential Tasks: Autonomous Pile Loading</article-title>
<alt-title alt-title-type="left-running-head">Strokina et al.</alt-title>
<alt-title alt-title-type="right-running-head">Visual Rewards From Observation</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Strokina</surname>
<given-names>Nataliya</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1276559/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>Wenyan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1749890/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Pajarinen</surname>
<given-names>Joni</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1744540/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Serbenyuk</surname>
<given-names>Nikolay</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>K&#xe4;m&#xe4;r&#xe4;inen</surname>
<given-names>Joni</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Ghabcheloo</surname>
<given-names>Reza</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/224108/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Computing Sciences</institution>, <institution>Tampere University</institution>, <addr-line>Tampere</addr-line>, <country>Finland</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Electrical Engineering and Automation</institution>, <institution>Aalto University</institution>, <addr-line>Espoo</addr-line>, <country>Finland</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Automation Technology and Mechanical Engineering</institution>, <institution>Tampere University</institution>, <addr-line>Tampere</addr-line>, <country>Finland</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/137171/overview">Nadia Magnenat Thalmann</ext-link>, Universit&#xe9; de Gen&#xe8;ve, Switzerland</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/226655/overview">Baihan Lin</ext-link>, Columbia University, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1713004/overview">Mohammad Hossein Hamedani</ext-link>, CREATE, Italy</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Nataliya Strokina, <email>nataliya.strokina@tuni.fi</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Field Robotics, a section of the journal Frontiers in Robotics and AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>31</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>9</volume>
<elocation-id>838059</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Strokina, Yang, Pajarinen, Serbenyuk, K&#xe4;m&#xe4;r&#xe4;inen and Ghabcheloo.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Strokina, Yang, Pajarinen, Serbenyuk, K&#xe4;m&#xe4;r&#xe4;inen and Ghabcheloo</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>One of the key challenges in implementing reinforcement learning methods for real-world robotic applications is the design of a suitable reward function. In field robotics, the absence of abundant datasets, limited training time, and high variation of environmental conditions complicate the task further. In this paper, we review reward learning techniques together with visual representations commonly used in current state-of-the-art works in robotics. We investigate a practical approach proposed in prior work to associate the reward with the stage of the progress in task completion based on visual observation. This approach was demonstrated in controlled laboratory conditions. We study its potential for a real-scale field application, autonomous pile loading, tested outdoors in three seasons: summer, autumn, and winter. In our framework, the cumulative reward combines the predictions about the process stage and the task completion (terminal stage). We use supervised classification methods to train prediction models and investigate the most common state-of-the-art visual representations. We use task-specific contrastive features for terminal stage prediction.</p>
</abstract>
<kwd-group>
<kwd>visual rewards</kwd>
<kwd>learning from demonstration</kwd>
<kwd>reinforcement learning</kwd>
<kwd>field robotics</kwd>
<kwd>earth moving</kwd>
<kwd>visual representations</kwd>
</kwd-group>
<contract-sponsor id="cn001">Academy of Finland<named-content content-type="fundref-id">10.13039/501100002341</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>In classical Reinforcement Learning (RL) architecture (<xref ref-type="bibr" rid="B47">Sutton and Barto 2018</xref>), an agent acts upon an environment, receives feedback in form of reward, and observes the state of the environment (see <xref ref-type="fig" rid="F1">Figure 1A</xref>). The collected experience is used to update the agent&#x2019;s policy. The process repeats until the agent converges to the desired behavior. Reward encodes the task objective. For example, the closer the robot is to task completion the higher the reward is. RL has demonstrated impressive results in simulated environments (<xref ref-type="bibr" rid="B28">Kaiser et al., 2020</xref>; <xref ref-type="bibr" rid="B54">Yu and Rosendo 2021</xref>) and for robotic tasks in controlled laboratory conditions (<xref ref-type="bibr" rid="B55">Zhu et al., 2020</xref>); <xref ref-type="bibr" rid="B30">Koert et al., 2020</xref>; <xref ref-type="bibr" rid="B51">Veiga et al., 2020</xref>). In real-world large-scale applications, such as those found in field robotics, the full state of the environment cannot be received. Instead, the robot only obtains an observation of the environment state through the sensors (see <xref ref-type="fig" rid="F1">Figure 1B</xref>). Moreover, the tasks usually require long-horizon decision-making and the reward function is difficult to engineer. In literature, reward learning is referred to as an inverse RL problem <xref ref-type="bibr" rid="B35">Osa et al. (2018)</xref>. Learning the reward function online from interactions with the environment is challenging when the amount of training samples is limited and training time for the RL algorithm is restricted. The desirable solution would be to estimate the reward from environment observation using the prior collected experience, i.e., as learning from demonstration. In this paper, we address the problem of visual reward estimation for multi-stage robotic applications and study its reliability in significantly varying conditions. Our test case is the autonomous pile-loading task implemented on a real-scale robotic wheel-loader in changing outdoor weather conditions.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>RL architecture: <bold>(A)</bold> classical online RL where state and reward are returned by the environment; <bold>(B)</bold> a real-world scenario where instead of full state only its observation is obtained and reward is predicted from the observation. Our contribution lies in the reward prediction block marked in green.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g001.tif"/>
</fig>
<p>Pile loading is one of the most challenging tasks in earth moving automation for heavy-duty mobile machines. This is partly caused by the difficulty of modelling the interaction between the tool and the material (<xref ref-type="bibr" rid="B8">Dadhich et al., 2016</xref>) and partly because of high variation in worksites and weather conditions throughout the year. Weather conditions affect the material properties, the hydraulics properties of the machine, and the ground surface properties. The majority of the state-of-the-art works on pile loading or excavation automation are either model-based or use heuristics (<xref ref-type="bibr" rid="B45">Sotiropoulos and Asada 2020</xref>), and experimented in simulators or with toy setups. Therefore it is unclear how well these methods perform in real worksites. Recently several works implemented reinforcement learning based autonomous pile-loading (<xref ref-type="bibr" rid="B1">Azulay and Shapiro 2021</xref>; <xref ref-type="bibr" rid="B2">Backman et al., 2021</xref>), which were also demonstrated either on a toy or simulated set-up. Existing progress and our own experience (<xref ref-type="bibr" rid="B52">Yang et al., 2020</xref>; <xref ref-type="bibr" rid="B53">Yang et al., 2021</xref>) indicate complexity of this real-world problem. <xref ref-type="bibr" rid="B19">G. Dulac-Arnold and Mankowitz, (2019)</xref> reports nine challenges of real-world RL, among which sample efficiency, safety constraints, large or unknown delays in the system actuators, high-dimensional state and action spaces, etc. In this work, we focus on one of the challenges - reward learning. Specifically, we investigate vision-based reward estimation for a real-world set-up learned from demonstrations.</p>
<p>Our experimental set-up is illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref> where a robotic wheel-loader performs the task of loading a pile of material and lifting the boom up. The wheel loader is equipped with a stereo camera providing an egocentric view. This is a long-horizon task where a suitable reward function is hard to engineer even using expert knowledge. Several previous works suggest learning a reward together with the policy online (<xref ref-type="bibr" rid="B27">Ho and Ermon 2016</xref>; <xref ref-type="bibr" rid="B18">Fu et al., 2017</xref>; <xref ref-type="bibr" rid="B21">Ghasemipour et al., 2020</xref>). To the best of our knowledge, there is no demonstration of this method for the long-horizon task in highly varying real-world conditions. Moreover, having an initial approximation of the reward function is desirable in the long-horizon tasks. <xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref> proposed a stage-based visual reward estimation approach and demonstrated it on a door opening task in laboratory conditions. This approach is attractive for field robotics applications since it requires only minimum information from an expert about the stages of the task and initially can be learned offline. <xref ref-type="fig" rid="F3">Figure 3</xref> shows an example of such rewards for the pile-loading task. At each time step, a reward is associated with the stage of the task. Additionally, we study the sparse reward prediction based on the outcome of the task using task-specific visual features that previously demonstrated good performance in training a behavior cloning controller <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref>. In our work, unlike in <xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref>, we report results for several visual representations, including, time-contrastive representations, depth, and selected deep features. We propose that the intermediate stage and the sparse terminal stage rewards can be combined into cumulative reward.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Example of the typical stages in the pile loading task.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Demonstration examples and the predicted stages: visual observation (top), corresponding to them gas commands and joint angles (middle), predicted stages of progress (bottom).</p>
</caption>
<graphic xlink:href="frobt-09-838059-g003.tif"/>
</fig>
<p>To summarize, our work provides the following contributions:<list list-type="simple">
<list-item>
<p>&#x2022; we review the methods of reward estimation and visual representations used in learning-based approaches for robotics applications; additionally, we overview the progress of learning-based methods in autonomous earth moving;</p>
</list-item>
<list-item>
<p>&#x2022; we propose a framework where the cumulative reward combines two predictions from visual observation: the current stage of the progress and whether the task has been completed (terminal stage). We formulate the prediction as a supervised classification task and investigate the most common state-of-the-art visual representations. For the terminal stage prediction, we test task-specific visual features.</p>
</list-item>
<list-item>
<p>&#x2022; the framework has been implemented and tested on an actual scale autonomous wheel-loader during three seasons (summer, autumn, and winter).</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Related Work</title>
<p>
<italic>Reward learning&#x2013;</italic>Several methods propose learning the reward function by iteratively optimizing the reward and agent behavior while interacting with the environment, e.g., <xref ref-type="bibr" rid="B27">Ho and Ermon (2016)</xref>; <xref ref-type="bibr" rid="B18">Fu et al. (2017)</xref>; <xref ref-type="bibr" rid="B21">Ghasemipour et al. (2020)</xref>. They are based on an adversarial paradigm as in Generative Adversarial Networks (GANs) (<xref ref-type="bibr" rid="B23">Goodfellow et al., 2014</xref>). The generator learns the policy and the discriminator learns to differentiate expert transitions from a non-expert. The reward is associated with a confusion of a discriminator. These methods minimize f-divergence between the expert and the learning agent state-action distributions <xref ref-type="bibr" rid="B20">Ghasemipour et al. (2019)</xref>. With the generator trying to maximize the reward provided by the discriminator, the problem becomes a min-max optimization problem, which can lead to training instabilities and poor sample efficiency. While demonstrating state-of-the-art performance in simulated environments, this approach was demonstrated only for short-horizon small-scale tasks in real environments.</p>
<p>Another group of methods recovers a reward function based on the set of pre-recorded expert demonstrations. This approach is attractive for real-world robotics applications since it allows to learn the reward function offline without a need to interact with the environment. In literature, two approaches are investigated for offline reward prediction: 1) reward as a measure of discrepancy between the expert and agent behavior, and 2) reward associated with a stage of progress in task completion. In the first approach, expert demonstrations are considered to be finite distributions. The reward is based on a distance measure between the expert and agent distributions. <xref ref-type="bibr" rid="B7">Dadashi et al. (2021)</xref> use the Wasserstein distance as a measure between the state-action distributions of the expert and the agent. Unlike f-divergences, the Wasserstein distance <xref ref-type="bibr" rid="B37">Rubner et al. (2000)</xref> is based on the geometry of the metric space it operates on. To avoid excess computation the authors suggest minimizing the upper bound of Wasserstein distance. Upper bound means that greedy coupling of state-action pairs is used instead of optimal coupling.</p>
<p>
<xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref> associate the reward with a stage of progress in task completion. This could be implemented by explicit goal discovery as by <xref ref-type="bibr" rid="B38">Schmeckpeper et al. (2020)</xref>. <xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref> uses unsupervised clustering of image sequence based on the similarity measure between the frames, thus discovering the task stages. At deployment, image classification provides prediction about the stage of the process. The reward is assigned as a difference between the features of the test sample and the mean features of the stage cluster. The difference is then multiplied by two to the power of the stage number. This method was demonstrated on door opening and water pouring tasks in controlled laboratory conditions with almost no variation in visual conditions.</p>
<p>Many research works in RL for real-world applications use sparse reward indicating whether the task was accomplished successfully or not (<xref ref-type="bibr" rid="B50">Vecer&#xed;k et al., 2019</xref>; <xref ref-type="bibr" rid="B32">Lee et al., 2020</xref>). We refer to it as a terminal reward. Several works train a reward function as a classifier of the final observation, predicting whether it was successful or not. While the terminal reward is sufficient for short-horizon tasks, it does not help in long-horizon large-scale tasks. <xref ref-type="bibr" rid="B43">Singh et al. (2019)</xref> teach the robot manipulation skills while interacting with a human and asking for a manual reward label of the observed states. Baseline reward is provided as a classification of the current state to be the final successful state. During the training, the robot queries the human to provide a label for previously-unlabeled states with the highest probability of success according to the classifier. Although the robot succeeds in learning, it is unclear how this approach would scale to the applications with a much larger state-space.</p>
<p>We focus on offline reward estimation from visual observations and stage-based reward, similar to <xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref>. We made this choice since in our application we are dealing with the long-horizon task with a large state space.</p>
<p>
<italic>Visual representations&#x2013;</italic>in this work, we are interested in reward estimation from vision. Some robotics applications use pre-trained deep Convolutional Neural Network (CNN) features. For example, <xref ref-type="bibr" rid="B41">Sermanet et al. (2017b)</xref> uses the Inception network <xref ref-type="bibr" rid="B48">Szegedy et al. (2016)</xref> pre-trained for ImageNet classification <xref ref-type="bibr" rid="B11">Deng et al. (2009)</xref>. In a number of works, generative modeling is used where a latent variable model is trained to model a latent distribution (<xref ref-type="bibr" rid="B17">Finn et al., 2015</xref>; <xref ref-type="bibr" rid="B26">Higgins et al., 2017</xref>; <xref ref-type="bibr" rid="B32">Lee et al., 2020</xref>). The latent variables are utilized as representations. The latent models are usually trained together with policy or other goal optimization while interacting with the environment which is impractical in real-world applications. These representations try to capture the variations related to all the underlying factors in task learning.</p>
<p>The representations can be trained while optimizing a contrastive loss (<xref ref-type="bibr" rid="B39">Schroff et al., 2015</xref>; <xref ref-type="bibr" rid="B49">van den Oord et al., 2018</xref>; <xref ref-type="bibr" rid="B4">Belghazi et al., 2018</xref>) with user-defined information. This means that a developer has to identify the factors for which the variation is modeled. For example, <xref ref-type="bibr" rid="B40">Sermanet et al. (2017a)</xref> proposes an approach to robotic behaviors training from unlabeled videos recorded from multiple viewpoints using time-contrastive representations. The contrastive loss tries to minimize the distance between the frames belonging to the same time window and maximize the distance to the frames outside the time window. In our previous work <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref>, contrastive representations are used to train a vision-based imitation learning controller for autonomous pile loading. We train the representations in a Siamese neural network classifying the successful and unsuccessful pile-loading outcomes. We minimize the distance between the inner-class representations and maximize the distance between the outer-class representations. A similar approach was used in representation learning for the peg-in-hole task in <xref ref-type="bibr" rid="B50">Vecer&#xed;k et al. (2019)</xref>
<xref ref-type="fn" rid="FN1">
<sup>1</sup>
</xref>.</p>
<p>Recently, actionable representations have been proposed to capture the variations that are important for decision making (<xref ref-type="bibr" rid="B22">Ghosh et al., 2019</xref>). This is implemented by comparing the actions taken by a goal-conditioned policy for two different goal states. If two goal states require different actions, then they are functionally different and vice-versa. The representations are learned such that Euclidean distance between states in representation space corresponds to actionable distances between them. The actionable distances capture the differences between the actions required to reach the different states based on Kullback&#x2013;Leibler (KL) divergence. In <xref ref-type="bibr" rid="B14">Dwibedi et al. (2018)</xref> the actionable visual representations are learned based on the contrastive loss.</p>
<p>We will investigate the time-contrastive, pre-trained deep CNN features, and depth as well as Histogram of Oriented Gradients (HOG) representation, as a representative of classical edge-based features. These representations seem practical in real-world applications since they do not require training of policy together with the representations.</p>
<p>
<italic>Autonomous pile-loading state-of-the-art&#x2013;</italic>most of the autonomous pile loading works adopt heuristics (<xref ref-type="bibr" rid="B16">Fernando et al., 2018</xref>) or are model-based (<xref ref-type="bibr" rid="B44">Sotiropoulos and Asada 2019</xref>), and are experimented only in a simulator (<xref ref-type="bibr" rid="B16">Fernando et al., 2018</xref>) or toy-scale setups (<xref ref-type="bibr" rid="B12">D. Jud et al., 2017</xref>; <xref ref-type="bibr" rid="B45">Sotiropoulos and Asada 2020</xref>), which cannot capture the complicated phenomena of the real-world problem. Model-based approaches succeed in many robotics applications. However, in pile loading, the interaction between the bucket and the material is hard to model accurately. Several works attempt to learn this interaction using learning from demonstrations. <xref ref-type="bibr" rid="B8">Dadhich et al. (2016)</xref> fit linear regression models to the lift and tilt bucket commands recorded with a joystick. <xref ref-type="bibr" rid="B36">R. Fukui et al. (2015)</xref> use a neural network model that selects a pre-programmed excavation motion from a dataset of motions. <xref ref-type="bibr" rid="B9">Dadhich et al. (2019)</xref>; <xref ref-type="bibr" rid="B25">Halbach et al. (2019)</xref>; <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref> report real experiments of autonomous scooping with a real-scale Heavy Duty Machine (HDM). <xref ref-type="bibr" rid="B9">Dadhich et al. (2019)</xref> propose a shallow time-delay neural network controller. The controller uses the joint angles and velocities as inputs. After outdoor experiments, the authors conclude that for different conditions the network controller needs to be retrained. <xref ref-type="bibr" rid="B25">Halbach et al. (2019)</xref> train a shallow neural network controller (NNet) for bucket loading based on the joint angles and hydraulic drive transmission pressure. Two of our recent works by <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref> and <xref ref-type="bibr" rid="B53">Yang et al. (2021)</xref> present the state-of-the-art data-driven controller learning demonstrated on real-world pile-loader. Despite demonstrated successful performance in tested conditions and robustness against slight variations in weather conditions, the imitation learning-based, i.e., behavior cloning methods, by construction are not able to provide online adaptability to the variable conditions.</p>
<p>Several recent works started investigating the applicability and limitations of the reinforcement learning framework in autonomous excavation and pile-loading. Since the heavy-duty machine job mainly involves interaction with the material, RL seems a promising solution if only its limitations are addressed to make the system practically feasible. <xref ref-type="bibr" rid="B15">Egli and Hutter (2021)</xref> train in simulator an RL controller for the end-effector trajectory tracking of a real excavator. The training utilizes pre-recorded task demonstrations and was applied for motions generation in the air and with soil interaction in a grading task. <xref ref-type="bibr" rid="B2">Backman et al. (2021)</xref> present an approach to learn bucket-filling behavior for an underground loader in a simulated environment. As a reward, the authors use the bucket filling rate and energy consumption of the machine. <xref ref-type="bibr" rid="B31">Kurinov et al. (2020)</xref> train an agent for earthmoving in a simulated environment with sophisticated multi-body dynamics modeling. The reward function depends on the amount of simulated soil loaded and unloaded from the bucket. <xref ref-type="bibr" rid="B1">Azulay and Shapiro (2021)</xref> train an RL-based controller without pre-recorded demonstrations for bucket-filling first in a simulator and then test it on a toy set-up. The training in the simulator takes 3&#xa0;hours and the reward depends on whether the wheel-loader followed the necessary stages of the task and the amount of loaded soil.</p>
<p>Current state-of-the-art works on RL for learning excavation or pile-loading are mainly in simulated or toy environments. The real-world scenarios differ from simulation by a much larger state space with high stochasticity, the longer horizon of the tasks, difficulty in defining a reward function. To make a transfer to the real-world machines, progress should be made in 1) Rl algorithms to guarantee the sample efficiency of the methods, 2) in learning of appropriate state representations from multi-modal sensors to capture the most relevant environmental conditions, and 3) proper approaches to reward estimation that would be both sample-efficient and robust to varying conditions. In the discussed works, the reward is mainly defined using the progress of the machine through the stages of the task and the amount of the material in the bucket. In the real world, one way to follow the machine&#x2019;s progress is to identify the stages automatically, for example, by observing the environment. In this paper, we address the problem of reward estimation from vision by studying the available visual representations for stage-based reward estimation both for sparse terminal reward and intermediate stage reward.</p>
</sec>
<sec id="s3">
<title>3 Test-Study: Autonomous Pile Loading</title>
<sec id="s3-1">
<title>3.1 Problem Statement</title>
<p>In this paper, we adopt the finite Markov Decision Process (MDP) as an abstraction for the problem of goal-directed episodic learning (<xref ref-type="bibr" rid="B47">Sutton and Barto (2018)</xref>). This is a standard abstraction in reinforcement learning literature. An autonomous agent learns through interaction with an environment. At each time step <italic>t</italic> &#x3d; 0, 1, &#x2026; <italic>T</italic> the agent receives the state of the environment <italic>s</italic>
<sub>
<italic>t</italic>
</sub> &#x2208; <italic>S</italic>, selects an action to perform <italic>a</italic>
<sub>
<italic>t</italic>
</sub> &#x2208; <italic>A</italic>(<italic>s</italic>), and receives a feedback from the environment which is called a reward <inline-formula id="inf1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:math>
</inline-formula>. The reward received at each time step is called an immediate reward and the goal of the agent is to maximize the cumulative reward, which is defined as the sum of rewards obtained within the time horizon <italic>T</italic>. A discount factor can be applied to each immediate reward which regularizes how much the agent values immediate reward versus future rewards. The cumulative reward received by the agent can be expressed as<disp-formula id="e1">
<mml:math id="m2">
<mml:mi>r</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>In this work, a set of consecutive visual observations <inline-formula id="inf2">
<mml:math id="m3">
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> contains pairs of <italic>o</italic>
<sub>
<italic>n</italic>
</sub> visual observations and labels <italic>l</italic>
<sub>
<italic>n</italic>
</sub>. <italic>N</italic> is the total number of observation-label pairs available. The visual observation is an RGB left-frame image of a stereo camera providing an ego-centric view. To avoid introducing new notation, we use <italic>R</italic>
<sub>
<italic>t</italic>
</sub> (.) to define the mapping from state or observation to reward. We thus learn an approximator function <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> - a predictor learnt from visual input. In long-horizon tasks, especially in real-world problems, it is beneficial to link the reward to the progress in task completion. In our framework, the visual representations are learned from the intermediate and terminal stage classifiers. The intent is to study what visual representations are useful in highly varying conditions of field robotics applications. Additionally, a recent study by <xref ref-type="bibr" rid="B46">Stooke et al. (2021)</xref> showed that learning of visual representations and policy separately is more beneficial in an imitation learning task.</p>
<p>More specifically, beside the state observations <italic>o</italic>
<sub>
<italic>t</italic>
</sub> and transition dynamics that governs environment dynamics, we assume subgoals for our long horizon task together with a final goal-we call these stages. We also assume these sub-goals or intermediate stages are sequential, that is, one needs to be performed before the other for the system to succeed. We denote them by <italic>S</italic>
<sub>0</sub>, <italic>S</italic>
<sub>1</sub>, <italic>S</italic>
<sub>2</sub>, <italic>S</italic>
<sub>
<italic>T</italic>
</sub>, where <italic>S</italic>
<sub>
<italic>T</italic>
</sub> is the terminal stage. We use supervised methods, and therefore the stages define classes which are labeled manually. The reward <italic>r</italic> is associated with the stage of the work process the machine is performing. Each time step the system associates the visual observation <italic>o</italic>
<sub>
<italic>t</italic>
</sub> of the machine with the stage of the process using classification model. We have two tasks: prediciting the work process stage and whether the task has completed (terminal stage). Thus, we have two training sets: for stage prediction <italic>D</italic> &#x3d; {(<italic>f</italic> (<italic>o</italic>
<sub>
<italic>t</italic>
</sub>), <italic>c</italic>
<sub>
<italic>t</italic>
</sub>): <italic>c</italic>
<sub>
<italic>t</italic>
</sub> &#x2208; <italic>C</italic> &#x3d; {0, 1, 2}} and terminal stage prediction <italic>G</italic> &#x3d; {(<italic>f</italic> (<italic>o</italic>
<sub>
<italic>t</italic>
</sub>), <italic>c</italic>
<sub>
<italic>t</italic>
</sub>): <italic>c</italic>
<sub>
<italic>t</italic>
</sub> &#x2208; <italic>C</italic> &#x3d; { &#x2212; 1, 1}}. We refer to the prediction about the stage as stage-based reward <inline-formula id="inf4">
<mml:math id="m5">
<mml:msubsup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and the prediction about completion as terminal stage reward <inline-formula id="inf5">
<mml:math id="m6">
<mml:msubsup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mfenced open="" close=")">
</mml:mfenced>
</mml:math>
</inline-formula>. Following <xref ref-type="disp-formula" rid="e1">Eq. (1)</xref>, the cumulative reward is computed as follows:<disp-formula id="e2">
<mml:math id="m7">
<mml:mi>r</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
</sec>
<sec id="s3-2">
<title>3.2 Choice of Visual Representations</title>
<p>In this section, we describe the visual features that we have investigated for training a classification model for sub-stage and final state prediction. The first step of our classification pipeline is to compute an embedding <inline-formula id="inf6">
<mml:math id="m8">
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
</mml:math>
</inline-formula>, which will be passed to the classification stage. In this, <inline-formula id="inf7">
<mml:math id="m9">
<mml:mi mathvariant="script">X</mml:mi>
</mml:math>
</inline-formula> denotes the space of features. Next, we will present methods that we have investigated for learning of the embeddings for the intermediate stage reward, followed by those for the terminal stage reward. In this work, we will use the terms embedding and feature interchangeably since the embedding <italic>f</italic> (<italic>o</italic>
<sub>
<italic>t</italic>
</sub>) acts as a feature input in our system.</p>
<sec id="s3-2-1">
<title>3.2.1 Intermediate Stage-Based Reward</title>
<p>In this section, we will introduce several state-of-the-art embeddings used for intermediate stage reward. We investigate the pretrained deep CNN features, time-contrastive representations, depth, and Histogram of Oriented Gradients (HOG) descriptors. We use the same class labels as described above for feature learning. These labels contain manually specified stages.</p>
<p>
<italic>Selected deep visual features (VGG) -</italic> to extract deep feature embedding <inline-formula id="inf8">
<mml:math id="m10">
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
</mml:math>
</inline-formula> in our experiments, we used a pretrained VGG model (<xref ref-type="bibr" rid="B42">Simonyan and Zisserman (2015)</xref>). Following the same approach, we applied a separate selection procedure to identify the most discriminative features for the classification task. For this, we computed the mean and standard deviation for each feature <italic>i</italic> for each class on the training set<disp-formula id="e3">
<mml:math id="m11">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf9">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf10">
<mml:math id="m13">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are the mean and standard deviation of the features within the same class and <inline-formula id="inf11">
<mml:math id="m14">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf12">
<mml:math id="m15">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> those of the rest of the classes. The labels of the classes are known from the ground truth data. We selected 200 features with a top <italic>z</italic>
<sub>
<italic>i</italic>
</sub> score in each class. This led to 29 deep features with top scores in all classes. The value 200 was selected empirically such that there are enough features with high score in all classes. One drawback of this approach is that the selection procedure might suffer from the drift of feature space when the environment changes.</p>
<p>
<italic>Time-constrastive representations (TCR) -</italic> TCR has shown to provide robust embedding maps for imitation learning tasks, as well as object, face, action recognition and alignment <xref ref-type="bibr" rid="B40">Sermanet et al. (2017a)</xref>; <xref ref-type="bibr" rid="B39">Schroff et al. (2015)</xref>. Here we describe our experiments with TCR. To capture the dynamic nature of the task, an input to the classification system at time <italic>t</italic> is <italic>x</italic>
<sub>
<italic>t</italic>
</sub>, a stack of <italic>d</italic> image frames<disp-formula id="e4">
<mml:math id="m16">
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>o</italic>
<sub>
<italic>t</italic>
</sub> is an rgb image observations. In our experiments, the size of a single observation image <italic>o</italic>
<sub>
<italic>t</italic>
</sub> was reduced to 64 &#xd7; 64; the number of frames <italic>d</italic> was 5. Each step, we choose a random input sequence <italic>I</italic>
<sub>
<italic>t</italic>
</sub> &#x3d; (<italic>x</italic>
<sub>
<italic>t</italic>
</sub>, <italic>x</italic>
<sub>
<italic>t</italic>&#x2b;1</sub>, &#x2026; , <italic>x</italic>
<sub>
<italic>t</italic>&#x2b;<italic>n</italic>
</sub>) in the training set and sample a random anchor input stack <italic>x</italic>
<sub>
<italic>i</italic>
</sub>: <italic>t</italic> &#x2264; <italic>i</italic> &#x2264; <italic>t</italic> &#x2b; <italic>n</italic>. The positive input stack <inline-formula id="inf13">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is sampled from a window (<italic>x</italic>
<sub>
<italic>i</italic>&#x2212;<italic>L</italic>
</sub>.<italic>x</italic>
<sub>
<italic>i</italic>&#x2b;<italic>L</italic>
</sub>), where <italic>L</italic> is a fixed window size. The negative sample <inline-formula id="inf14">
<mml:math id="m18">
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is randomly chosen outside this window within sequence <italic>I</italic>
<sub>
<italic>t</italic>
</sub>. The anchor, the positive, and the negative inputs are added to the training set.</p>
<p>Let <italic>f</italic>
<sub>
<italic>&#x3b8;</italic>
</sub>(<italic>x</italic>
<sub>
<italic>i</italic>
</sub>) denote the feature extraction (embedding) network with parameters <italic>&#x3b8;</italic>. With an anchor <italic>x</italic>
<sub>
<italic>i</italic>
</sub>, a positive sample <inline-formula id="inf15">
<mml:math id="m19">
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and a negative sample <inline-formula id="inf16">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, we formulate the loss to draw positive samples closer to the anchor in the feature space and negative samples further away. Thus, the embedding <italic>f</italic> (<italic>x</italic>
<sub>
<italic>i</italic>
</sub>) shall satisfy <xref ref-type="bibr" rid="B40">Sermanet et al. (2017a)</xref>:<disp-formula id="e5">
<mml:math id="m21">
<mml:mo>&#x2225;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:msubsup>
<mml:mrow>
<mml:mo>&#x2225;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:msubsup>
<mml:mrow>
<mml:mo>&#x2225;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2200;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>where <italic>&#x3b1;</italic> is an empirical parameter and &#x393; is the set of all possible triplets in the training set. We used a state-of-the-art video frame interpolation method DAIN (Depth-Aware Video Frame Interpolation) <xref ref-type="bibr" rid="B3">Bao et al. (2019)</xref> to robustify the training since our training set is limited.</p>
<p>
<italic>Selected HOG features -</italic> We also experimented with HOG, as a representative of classical methods. HOG demonstrated good performance in visual recognition tasks <xref ref-type="bibr" rid="B34">Lowe (2004)</xref>; <xref ref-type="bibr" rid="B10">Dalal and Triggs (2005)</xref>. HOG stands for the Histogram of Oriented Gradient descriptors. The HOG descriptor (embedding) characterizes an image by the distribution of local intensity gradients or edge directions. This information was shown to be rather sufficient even without precise knowledge of the corresponding gradient or edge positions. The embedding <italic>f</italic> (<italic>o</italic>
<sub>
<italic>t</italic>
</sub>) is computed by dividing an image into small spatial regions, for each region accumulating a local histogram of gradient directions or edge orientations over the pixels of the region. The combined histogram entries form the embedding. To improve the invariance to illumination, and other conditions, the local responses within a fixed block are contrast-normalized using L1-regularization. Let <italic>v</italic> be the unnormalized descriptor vector, norm&#x2009;<italic>v</italic>
<sub>1</sub> be its 1-norm <italic>&#x3b7;</italic> be a small constant. Normalization is done by dividing each value in the vector by its norm: <inline-formula id="inf17">
<mml:math id="m22">
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>norm</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b7;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:math>
</inline-formula> We use <italic>sklearn</italic>1 implementation to compute the HOG representations.</p>
<p>
<italic>Depth features -</italic> to obtain the depth image we used MonoDepth, the state-of-the-art depth computation pretrained on the KITTI dataset for outdoor environments2. We then define the features to be an array of depth values in fixed image locations.</p>
</sec>
<sec id="s3-2-2">
<title>3.2.2 Sparse Terminal Reward</title>
<p>We define a sparse reward signal that indicates whether the task was completed in the current episode. Practically, at each time step, it predicts whether the terminal stage has been reached. The approach to predict the terminal stage is similar to the TCR. For this, we use task-specific visual features that we previously trained and demonstrated in behavior cloning of pile loading task <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref>. There we trained the features on summer data and were testing during the summer/autumn period. The training objective was that the network should learn to distinguish the successful samples (the bucket is full and task accomplished) from unsuccessful samples (the bucket is empty or task is unfinished). For the main target, the standard cross-entropy loss worked well. However, in addition to the cross-entropy loss for classification, a contrastive loss <xref ref-type="bibr" rid="B24">Hadsell et al. (2006)</xref> was added to constrain visual feature extraction. The features within the positive examples should be close to each other in the feature space and far from the negative examples. The samples were labeled manually as described in the experimental section. The contrastive loss has been used in one-shot learning tasks and metric learning tasks where feature distances become important <xref ref-type="bibr" rid="B29">Koch et al. (2015)</xref>
<xref ref-type="fn" rid="FN2">
<sup>2</sup>
</xref>.</p>
<p>Here, the learning process is similar to the time-contrastive representations (TCR). The difference is that in TCR we sample a triplet of instances and here we sample a pair. Let <italic>x</italic>
<sub>1</sub>, <italic>x</italic>
<sub>2</sub> &#x2208; <italic>Z</italic> be the input samples. As in TCR, we use a stack of RGB images as input. In TCR we used only the time-contrastive loss since the sampling of data was based on the time window approach and not based on identified classes. Here, positive samples belong to the terminal stage and the negative ones do not. We train the embeddings in the Siamese classification network <italic>f</italic>
<sub>
<italic>&#x3b8;</italic>
</sub>(<italic>x</italic>) where the standard cross-entropy loss <italic>L</italic>
<sub>
<italic>ce</italic>
</sub> (<xref ref-type="bibr" rid="B5">Bishop (2006)</xref>) is combined with the contrastive loss <italic>L</italic>
<sub>
<italic>c</italic>
</sub> in the following way:<disp-formula id="e6">
<mml:math id="m23">
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>&#x3bb;</italic>
<sub>1</sub> &#x3d; 0.6 and <italic>&#x3bb;</italic>
<sub>2</sub> &#x3d; 0.4 denote the weight of the losses.</p>
<p>The contrastive loss is defined by<disp-formula id="e7">
<mml:math id="m24">
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msup>
<mml:mrow>
<mml:mi>max</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.17em"/>
<mml:mspace width="0.17em"/>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.17em"/>
<mml:mspace width="0.17em"/>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m25">
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2225;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2225;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>x</italic>
<sub>1</sub> and <italic>x</italic>
<sub>2</sub> are input into a shared-weights siamese network and <italic>D</italic> (<italic>x</italic>
<sub>1</sub>, <italic>x</italic>
<sub>2</sub>) is the Euclidean distance between the features. By minimizing the above loss function, the network parameters are trained in a way that features of inputs within the positive samples come closer and those of the inputs belonging to different sets get farther. Parameter <italic>m</italic> (<italic>m</italic> &#x3e; 0) is a margin value based on the spring model analogy (<xref ref-type="bibr" rid="B24">Hadsell et al. (2006)</xref>). The margin defines a radius within which the negative samples contribute to the loss function. The value of m depends on the distribution of features in the contrastive classes. We selected empirically <italic>m</italic> &#x3d; 0.3. <italic>m</italic> corresponds to <italic>&#x3b1;</italic> parameter in TCR. We used also the video frame interpolation DAIN <xref ref-type="bibr" rid="B3">Bao et al. (2019)</xref> and image data augmentation to generate more training examples and robustify the training.</p>
</sec>
</sec>
<sec id="s3-3">
<title>3.3 Choice of Classification Methods</title>
<p>In this section, we describe the classification methods that we have used to learn the classification model.</p>
<p>
<italic>K-Nearest Neighbor (KNN) -</italic> KNN is a straightforward non-parametric method for classification where no assumption is made on the distribution of the data (<xref ref-type="bibr" rid="B13">Duda et al. (2012)</xref>). The decision about each new test sample is made based explicitly on the training samples. Its simplicity is an advantage of the method, whereas computational complexity at the test time is the downside. To classify a new test observation <italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>), we use Euclidean distance to identify a set of <italic>M</italic> nearest neighbors in the feature space. We assign the class label of the majority of the neighbors to the test sample.</p>
<p>
<italic>Support Vector Machine (SVM) -</italic> we use the multi-class SVM formulated as one-versus-one classification. The basic SVM is a two-class classification approach which we explain further. We will therefore formulate it first for the two-class problem of terminal reward classification. The SVM uses a linear discriminant function of the form<disp-formula id="equ1">
<mml:math id="m26">
<mml:mi>y</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
</disp-formula>where <italic>&#x3d5;</italic>(<italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>)) denotes a fixed feature-space transformation of our embedding (e.g. RBF or polynomial) and <italic>b</italic> is a bias term. The new sample <italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>) is classified according to the sign of <italic>y</italic> (<italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>)). Assumption is that there exists at least one choice of parameters <italic>w</italic> and <italic>b</italic> such that two classes are separable in high-dimensional space of <italic>&#x3d5;</italic>(<italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>)): <italic>y</italic> (<italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>)) &#x3e; 0 for samples having <italic>c</italic>
<sub>
<italic>i</italic>
</sub> &#x3d; 1 and <italic>y</italic> (<italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>)) &#x3c; 0 for samples having <italic>c</italic>
<sub>
<italic>i</italic>
</sub> &#x3d; &#x2212;1.</p>
<p>To train <italic>w</italic> and <italic>b</italic>, the following objective function is used<disp-formula id="equ2">
<mml:math id="m27">
<mml:munder>
<mml:mrow>
<mml:mi mathvariant="normal">argmax</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>
</p>
<p>It maximizes the margin between two classes, where the margin is defined by the supporting hyperplanes separating the classes.</p>
<p>This can be transformed into the quadratic programming problem and solved, for example, with the method of Lagrangian multipliers (please, see <xref ref-type="bibr" rid="B5">Bishop (2006)</xref> for further details). We used the sklearn3 multi-class implementation.</p>
<p>
<italic>Random Forest (RF) -</italic> RF classification is more robust to the data intrinsic ambiguities when different output values might be associated with the same input values. (<xref ref-type="bibr" rid="B6">Criminisi et al. (2012)</xref>). Each tree learns only from a partial set of features and therefore can learn important cues about the different stages. The random forest <inline-formula id="inf18">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is a collection of decision trees:<disp-formula id="equ3">
<mml:math id="m29">
<mml:msub>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>
</p>
<p>
<italic>&#x3b8;</italic>
<sub>
<italic>m</italic>
</sub> denotes the parameters of each tree <inline-formula id="inf19">
<mml:math id="m30">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>. A decision tree is a special graph structure consisting of a set of questions hierarchically organized. By answering the question a decision tree can evaluate a property of a sample or identify its class or category. Parameters of these questions or rules are the parameters of a decision tree. Each decision tree in the forest classifies a sample according to its own rules and based on a sub-set of training data. During the training, each tree is trained separately. In our work, given the input feature <italic>f</italic> (<italic>o</italic>
<sub>
<italic>i</italic>
</sub>), the classification result produced by the random forest <inline-formula id="inf20">
<mml:math id="m31">
<mml:msub>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is:<disp-formula id="e9">
<mml:math id="m32">
<mml:msub>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(9)</label>
</disp-formula>that is, the output of the random forest is the average of all class probabilistic predictions produced by the trees.</p>
<p>The training of <inline-formula id="inf21">
<mml:math id="m33">
<mml:msub>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is performed as following: 1) draw a bootstrap data <italic>D</italic>
<sub>
<italic>bs</italic>
</sub> from the training set; 2) grow a classification tree <inline-formula id="inf22">
<mml:math id="m34">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> to the bootstrapped data <italic>D</italic>
<sub>
<italic>bs</italic>
</sub>, fit each tree until the maximum depth is reached. Each tree <inline-formula id="inf23">
<mml:math id="m35">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> is trained as following: 1) randomly select <italic>n</italic> features from the <italic>k</italic> features (<italic>n</italic> &#x3c; <italic>k</italic>); 2) pick the best variable split-point among the <italic>n</italic>; 3) split the node into two child nodes. Each node of each tree <inline-formula id="inf24">
<mml:math id="m36">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> in our implementation is trained by minimizing the Gini impurity. For any node <italic>j</italic> and class <italic>C</italic>
<sub>
<italic>i</italic>
</sub>, <italic>p</italic>
<sub>
<italic>j</italic>
</sub> (<italic>C</italic>
<sub>
<italic>i</italic>
</sub>) is a fraction of samples at <italic>j</italic> that belongs to class <italic>C</italic>
<sub>
<italic>i</italic>
</sub>. With a total number of classes <italic>M</italic>, the Gini impurity in node <italic>j</italic> is the probability of incorrectly classifying a randomly chosen element in the dataset:<disp-formula id="equ4">
<mml:math id="m37">
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments</title>
<p>The goal of the performed experiments was to explore the performance of the currently used representations for stage-based reward evaluation in highly varying weather conditions. We demonstrate the results for the outdoor pile loading task. At the beginning of this section, we introduce the set-up details and task description, including data collection and resulting datasets. After that we 1) present the results of stage discovery using the studied visual representations and classification methods; 2) test the task-specific visual features for terminal stage prediction, and 3) demonstrate how cumulative reward is computed using the stage information and perform the qualitative analysis of the results.</p>
<sec id="s4-1">
<title>4.1 Set-Up</title>
<p>
<italic>Wheel-loader -</italic> The autonomous scooping was implemented on a robotic wheel-loader, a so-called GIM machine. It has the mechanics of a commercial wheel loader (Avant 635) and power transmission and controllers are custom-made at Tampere University. The bucket is positioned in the vertical plane by two joints, the boom joint and the bucket joint, and in the horizontal plane by drive (throttle/gas) and articulated by a frame steering mechanism4. The GIM machine is equipped with various sensors including, for example, GNSS (Global Navigation Satellite System), wheel odometry, IMU (Inertial Measurement Unit), and pressure sensors.</p>
<p>
<italic>Vision -</italic> We used a ZED stereo camera to get visual feedback. The ZED camera images of 2560, &#xd7;, 720 resolution were captured at the frame rate of 15<italic>fps</italic>. The left RGB image was used to generate the embeddings that are used as features in stage classification. The experiments were performed at the outdoor test site (see <xref ref-type="fig" rid="F4">Figure 4</xref>).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Robotic set-up at test site.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g004.tif"/>
</fig>
<p>
<italic>The control system -</italic> is composed of multiple layers. In the lowest level (digital and analog I/O and CAN), industrial micro-controllers implement power management and basic safety functions. In the PC control level, a target PC runs Simulink Real-time models, which run real-time tasks such as localization. Sub-systems communicate low-level sensor data and control commands via UDP protocol running on a Jetson AGX Xavier (8-Core ARM v8.2 64-bit NVIDIA Carmel CPU and 512-core NVIDIA Volta GPU with 64 Tensor Cores) on-board. All the data collection, learning, and closed-loop control are implemented on the Jetson PC.</p>
<p>
<italic>Workflow of experiments -</italic> during training, we learn the models to compute the TRC representations (<xref ref-type="sec" rid="s3-2-1">Sec. 3.2.1</xref>), the sparse terminal reward representations (<xref ref-type="sec" rid="s3-2-2">Sec. 3.2.2</xref>), and classification models (<xref ref-type="sec" rid="s3-3">Sec. 3.3</xref>). At test time, we follow the routine in <xref ref-type="statement" rid="Algorithm_1">Algorithm 1</xref> which is performed on each sample of the test data.</p>
<p>
<statement content-type="algorithm" id="Algorithm_1">
<label>Algorithm 1</label>
<p>Visual rewards for sequential tasks: workflow of experiments at test time.</p>
<p>
<inline-graphic xlink:href="frobt-09-838059-fx1.tif"/>
</p>
<p>
<italic>Performance metric -</italic> as a performance measure, we use classification accuracy in percent:<disp-formula id="e10">
<mml:math id="m38">
<mml:mtext>Accuracy</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext>Number&#x2009;of&#x2009;correct&#x2009;predictions</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>Total&#x2009;number&#x2009;of&#x2009;predictions</mml:mtext>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2217;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>In our work, the reward is assigned based on the classification prediction. If the predictions match the ground-truth label of the sample, they are correct.</p>
</statement>
</p>
</sec>
<sec id="s4-2">
<title>4.2 Data</title>
<p>
<italic>The bucket filling task -</italic> <xref ref-type="fig" rid="F2">Figure 2</xref> demonstrates three stages of the task: driving up to the pile, filling the bucket, and lifting the boom. This process takes about 30&#x2013;60&#xa0;s for a human operator depending on the conditions of the ground. We treat this task as an episodic, long-horizon, sequential task. Since it is a real-world task, there might be variations in operator performance resulting in noisy and non-consistent demonstrations. For example, while loading the pile (stage 2), the operator might still use gas in order to load a larger amount of load.</p>
<p>
<italic>Reset to initial conditions -</italic> in all our experiments, every episode starts from the same initial conditions. The bucket is unloaded; the wheel-loader is driven a certain distance away from the pile (1&#x2013;5&#xa0;m); the boom is placed in a lower position, and the bucket is leveled with the ground. The process of bringing the machine to its initial state takes several minutes. In all our experiments, we automated this procedure by pre-programming the machine to reach certain joint values using a state machine algorithm. The distance the machine drives away from the pile is defined manually.</p>
<p>
<italic>Datasets -</italic> we created three sets of data:<list list-type="simple">
<list-item>
<p>&#x2022; <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>: 70 human demonstrations collected during two summer days (See examples in <xref ref-type="fig" rid="F5">Figure 5</xref>);</p>
</list-item>
<list-item>
<p>&#x2022; <italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>: 25 roll-outs of an RF controller trained in <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref> (See examples in <xref ref-type="fig" rid="F6">Figure 6</xref>);</p>
</list-item>
<list-item>
<p>&#x2022; <italic>D</italic>
<sub>
<italic>winter</italic>
</sub>: 25 human demonstrations collected during one snowy winter day (See examples in <xref ref-type="fig" rid="F7">Figure 7</xref>).</p>
</list-item>
</list>
</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Examples of images from <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g005.tif"/>
</fig>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Examples of images from <italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g006.tif"/>
</fig>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Examples of images from <italic>D</italic>
<sub>
<italic>winter</italic>
</sub>.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g007.tif"/>
</fig>
<p>All the samples in the dataset are successfully performed bucket filling demonstrations. The amount of load in different demonstrations might vary. The collected data includes images from the ZED camera, joint positions, pressure signals, and commands. We used the commands produced by the operator or controller to manually annotate the stages of the task and whether the terminal condition was reached. In the stage identification algorithms as well as representation learning only visual data was used. Each image was labeled by 1) one of three classes corresponding to the stage of the process, and 2) by a binary label belonging to the terminal state or not (-1 or 1)<xref ref-type="fn" rid="fn3">
<sup>3</sup>
</xref>
<sup>,</sup>
<xref ref-type="fn" rid="fn4">
<sup>4</sup>
</xref>.</p>
</sec>
<sec id="s4-3">
<title>4.3 Preliminary Experiments With Manually Assigned Rewards</title>
<p>In prior work, we attempted the approach by <xref ref-type="bibr" rid="B43">Singh et al. (2019)</xref> where the robot can query a human about the reward for a given visual observation. A baseline reward was given by a classification prediction of whether the observed state is a final successful state. In the original paper, the reward was queried for states which were predicted to be close to successful. In our implementation, the user assigned a reward manually when the robot was proceeding through the correct stages. Our task had a much longer horizon and reward classification did not provide a sufficient accuracy rate due to higher variations in visual observation. We used DDPG RL algorithm (See <xref ref-type="bibr" rid="B33">Lillicrap et al., 2015</xref>). In our experiment, the robot failed to converge to any solution within several hours. This experiment motivated us to study the topic of reward discovery separately.</p>
</sec>
<sec id="s4-4">
<title>4.4 Stage-Based Intermediate Reward</title>
<p>This section presents quantitative results of stage discovery for the pile loading tasks. The TCR representations were trained on the <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>. The HOG and VGG features were selected based on the <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>. We experiment in the following scenarios: 1) classification methods trained on <italic>D</italic>
<sub>
<italic>summer</italic>
</sub> and tested on the rest of the seasons, and 2) classification methods trained on the mixture of seasons. Scenario 1 has practical importance. In real-world industrial applications, it is desirable to collect the data once, develop models based on it, and re-use them. Scenario one tests this opportunity for visual rewards. We used 2-fold cross-validation repeated five times in our reporting of testing and training results, i.e., 50% of data was selected for training and 50% for testing which was repeated five times.</p>
<p>
<italic>Scenario 1 -</italic> in <xref ref-type="table" rid="T1">Table 1</xref> the results of stage discovery is shown when we merged the labels for stage 2 and 3 into one class. The results for using three classes were quite unsatisfactory. By reducing the number of stages (basically discovering whether the machine is driving to the pile or performing the scooping action), a more reasonable performance was achieved. The results suggest that depth and selected deep VGG features provide the most reliable cues. The results show that the time-contrasting features overfit training data and performance degrades on test data.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Stage discovery accuracy <bold>trained on</bold> <bold>
<italic>D</italic>
</bold>
<sub>
<bold>
<italic>summer</italic>
</bold>
</sub>. Two stages. We used repeated 2-fold cross-validation, i.e., 50% of data was selected for training and 50% for testing which was repeated five times. Highest accuracy for each season marked in bold.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="5" align="center">Stage discovery accuracy [%]</th>
</tr>
<tr>
<th align="left"/>
<th align="left"/>
<th colspan="3" align="center">Tested on</th>
</tr>
<tr>
<th align="left">Feature</th>
<th align="center">Classifier</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>summer</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>winter</italic>
</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">HOG</td>
<td>KNN</td>
<td align="center">78</td>
<td align="center">50</td>
<td align="center">58</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="center">80</td>
<td align="center">55</td>
<td align="center">55</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="center">79</td>
<td align="center">53</td>
<td align="center">59</td>
</tr>
<tr>
<td align="left">VGG</td>
<td>KNN</td>
<td align="center">88</td>
<td align="center">
<bold>68</bold>
</td>
<td align="center">68</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="center">92</td>
<td align="center">62</td>
<td align="center">55</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="center">90</td>
<td align="center">63</td>
<td align="center">63</td>
</tr>
<tr>
<td align="left">TCR</td>
<td>KNN</td>
<td align="center">
<bold>99</bold>
</td>
<td align="center">59</td>
<td align="center">43</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="center">
<bold>99</bold>
</td>
<td align="center">67</td>
<td align="center">56</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="center">
<bold>99</bold>
</td>
<td align="center">67</td>
<td align="center">56</td>
</tr>
<tr>
<td align="left">Depth</td>
<td>KNN</td>
<td align="center">84</td>
<td align="center">61</td>
<td align="center">
<bold>73</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="center">82</td>
<td align="center">66</td>
<td align="center">67</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="center">89</td>
<td align="center">60</td>
<td align="center">60</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<italic>Scenario 2 -</italic> <xref ref-type="table" rid="T2">Table 2</xref> contains the results when the training and testing was done on a mix set of all seasons. We compare the discovery rate for two- and three-stage labeling. As expected, the performance improved significantly when the representatives of all seasons were present in the dataset. Another reason for the improvement is that the summer data share is highest in the whole set. However, both depth and selected VGG produced the highest accuracy of more than 95%. HOG and time-contrastive representations demonstrate slightly worse performance. However, the accuracy does not drop significantly between the training and test sets.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Stage discovery accuracy <bold>trained on all seasons</bold>. We used repeated 2-fold cross-validation, i.e., 50% of data was selected for training and 50% for testing which was repeated five times. Highest accuracy for each scenario marked in bold.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="4" align="center">
<italic>Stage discovery accuracy [%]</italic>
</th>
</tr>
<tr>
<th align="left">Feature</th>
<th align="center">Classifier</th>
<th align="center">Two stages</th>
<th align="center">Three starges</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">HOG</td>
<td>KNN</td>
<td align="char" char=".">88</td>
<td align="char" char=".">85</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="char" char=".">74</td>
<td align="char" char=".">64</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="char" char=".">82</td>
<td align="char" char=".">75</td>
</tr>
<tr>
<td align="left">VGG</td>
<td>KNN</td>
<td align="char" char=".">95</td>
<td align="char" char=".">93</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="char" char=".">96</td>
<td align="char" char=".">94</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="char" char=".">94</td>
<td align="char" char=".">91</td>
</tr>
<tr>
<td align="left">TCR</td>
<td>KNN</td>
<td align="char" char=".">85</td>
<td align="char" char=".">74</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="char" char=".">85</td>
<td align="char" char=".">77</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="char" char=".">77</td>
<td align="char" char=".">69</td>
</tr>
<tr>
<td align="left">Depth</td>
<td>KNN</td>
<td align="char" char=".">
<bold>98</bold>
</td>
<td align="char" char=".">
<bold>96</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td>SVM</td>
<td align="char" char=".">90</td>
<td align="char" char=".">68</td>
</tr>
<tr>
<td align="left"/>
<td>RF</td>
<td align="char" char=".">97</td>
<td align="char" char=".">
<bold>96</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As for the classification methods, there was no significant difference between the approaches. In most cases, the KNN and RF classifier demonstrated better performance. Time-contrastive features overfitted the training data in Scenario one and demonstrated worse performance in Scenario 2. The reason for this might be that, compared to prior work <xref ref-type="bibr" rid="B40">Sermanet et al. (2017a)</xref>, the view is ego-centric and the structure of the motion is not observed. A further study should address combining ego-centric and third-person views complementing each other.</p>
</sec>
<sec id="s4-5">
<title>4.5 Terminal Reward Estimation</title>
<p>In this experiment, we attempted to use the task-specific visual features previously learned and successfully demonstrated in our behavior cloning experiments (see <xref ref-type="bibr" rid="B52">Yang et al. (2020)</xref>). Each image received a label minus one or one, where one corresponds to task completion and minus one to the rest. Similarly to the previous section, we test two scenarios 1) trained only on summer data <italic>D</italic>
<sub>
<italic>summer</italic>
</sub> and tested on two other datasets <italic>D</italic>
<sub>
<italic>autumn</italic>
</sub> and <italic>D</italic>
<sub>
<italic>winter</italic>
</sub> and 2) trained and tested on a mixed set containing representatives of all weather conditions. A 2-fold cross-validation repeated five times was used. <xref ref-type="table" rid="T3">Table 3</xref> contains the results of this experiment. In Scenario one although the performance drops while transferring between seasons, the average best accuracy for the test seasons was about 83%. Training and testing on the mixed set of data (Scenario 2) allowed to improve the results significantly with top accuracy of 95%. When trained and tested on the same season (summer) the learned feature demonstrates 99% accuracy for the RF classifier.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Terminal state classification. We used repeated 2-fold cross-validation, i.e., 50% of data was selected for training and 50% for testing which was repeated five times. Highest accuracy for each scenario marked in bold.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="11" align="center">
<italic>Step-wise accuracy [%]</italic>
</th>
</tr>
<tr>
<th align="left"/>
<th colspan="3" align="center">Trained on summer, tested on</th>
<th colspan="3" align="center">Trained on autumn, tested on</th>
<th colspan="3" align="center">Trained on winter, tested on</th>
<th rowspan="2" align="center">Trained and tested on mix</th>
</tr>
<tr>
<th align="left">Classifier</th>
<th align="left">
<italic>D</italic>
<sub>
<italic>summer</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>winter</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>summer</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>winter</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>summer</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>autumn</italic>
</sub>
</th>
<th align="center">
<italic>D</italic>
<sub>
<italic>winter</italic>
</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">KNN</td>
<td align="center">98</td>
<td align="center">76</td>
<td align="center">80</td>
<td align="center">78</td>
<td align="center">97</td>
<td align="center">92</td>
<td align="center">74</td>
<td align="center">65</td>
<td align="center">98</td>
<td align="center">95</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="center">85</td>
<td align="center">66</td>
<td align="center">78</td>
<td align="center">82</td>
<td align="center">89</td>
<td align="center">94</td>
<td align="center">81</td>
<td align="center">82</td>
<td align="center">94</td>
<td align="center">84</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="center">99</td>
<td align="center">75</td>
<td align="center">90</td>
<td align="center">79</td>
<td align="center">95</td>
<td align="center">92</td>
<td align="center">79</td>
<td align="center">78</td>
<td align="center">97</td>
<td align="center">94</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-6">
<title>4.6 Combining the Stage-Based Rewards</title>
<p>For each visual observation <italic>o</italic>
<sub>
<italic>i</italic>
</sub>, we obtain a label <italic>c</italic>
<sub>
<italic>i</italic>
</sub> of the current stage the machine is performing and a prediction <inline-formula id="inf25">
<mml:math id="m39">
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> whether the task has been completed. At each time step, we can sum these predictions and compute cumulative rewards as a weighted sum (see <xref ref-type="disp-formula" rid="e2">Eq. (2)</xref>). <xref ref-type="fig" rid="F8">Figure 8</xref> presents examples of cumulative rewards for summer and winter scenarios. Despite the noisy stage prediction, the cumulative reward is not much affected. The analysis of stage classification is presented in the next section.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Examples of proposed cumulative reward based on the discovered stages: <bold>(A,B)</bold> summer, different light conditions, <bold>(C,D)</bold> winter, different locations.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g008.tif"/>
</fig>
</sec>
<sec id="s4-7">
<title>4.7 Qualitative Results and Discussion</title>
<p>In this section, we investigate the stage prediction in detail. Every subplot in <xref ref-type="fig" rid="F11">Figures 11</xref>&#x2013;<xref ref-type="fig" rid="F13">13</xref> contains the stage prediction for the specified combination of a visual representation and a classification method. Each subplot presents the results for each episode in the datasets. The end of each episode corresponds to the right-most end of the <italic>x</italic>-axis since the duration of episodes varies.</p>
<sec id="s4-7-1">
<title>4.7.1 Stage-Based Reward</title>
<p>In the following, we discuss the results according to the scenarios introduced in <xref ref-type="sec" rid="s4-4">Section 4.4</xref>.</p>
<p>
<italic>Scenario 1,</italic> &#x201c;<italic>training on the summer data</italic>&#x201d; <italic>-</italic> <xref ref-type="fig" rid="F9">Figures 9</xref>, <xref ref-type="fig" rid="F10">10</xref> contain the stage prediction results for the two- and the three-stage classification task. Qualitatively and quantitatively depth features combined with KNN provide the highest accuracy compared to other combinations. Three-stage classification results in too noisy outputs when training on summer only. HOG representations, which reflect the geometrical information in the images, visually perform better than TCR and VGG representations. TCR representations do not cope with the task in our settings. This might be because we train them from scratch on a small set of data. Although we use the data augmentation technique to increase sample variation, it seems not to help in the explored task. As for the classification method, RF and KNN provide similar results, however, RF is more computationally efficient than KNN.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Stage-based reward plotted for all seasons. Scenario 1: trained on the summer set of data <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>, 2 stages. The stages proceed as follows: <italic>S</italic>1 &#x2212; &#x3e; <italic>S</italic>2. <italic>x</italic>-axis stands for the time step in reversed order to match the samples of varying length. Color bar encodes the stages: <italic>S</italic>1 -gray, <italic>S</italic>2 - black.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g009.tif"/>
</fig>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Stage-based reward plotted for all seasons. Scenario 1: trained on the summer set of data <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>, 3 stages. The stages proceed as follows: <italic>S</italic>1 &#x2212; &#x3e; <italic>S</italic>2 &#x2212; &#x3e; <italic>S</italic>3. <italic>x</italic>-axis stands for the time step in reversed order to match the samples of varying length. Color bar encodes the stages: <italic>S</italic>1 - light gray, <italic>S</italic>2 - gray, <italic>S</italic>3 - black.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g010.tif"/>
</fig>
<p>
<italic>Scenario 2,</italic> &#x201c;<italic>training on the mixed data</italic>&#x201d; <italic>-</italic> when training is performed on the mixed set of data (see <xref ref-type="fig" rid="F11">Figures 11</xref>, <xref ref-type="fig" rid="F12">12</xref>), the results look significantly better, as expected. The depth and VGG combined with KNN demonstrate the best performance. HOG seems still better than TCR.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Stage-based reward plotted for all seasons. Scenario 2: trained on the mixed set of data, 2 stages. The stages proceed as follows: <italic>S</italic>1 &#x2212; &#x3e; <italic>S</italic>2. <italic>x</italic>-axis stands for the time step in reversed order to match the samples of varying length. Color bar encodes the stages: <italic>S</italic>1 -gray, <italic>S</italic>2 - black.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g011.tif"/>
</fig>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>Stage-based reward plotted for all seasons. Scenario 2: trained on the mixed set of data, 3 stages. The stages proceed as follows: <italic>S</italic>1 &#x2212; &#x3e; <italic>S</italic>2 &#x2212; &#x3e; <italic>S</italic>3. <italic>x</italic>-axis stands for the time step in reversed order to match the samples of varying length. Color bar encodes the stages: <italic>S</italic>1 - light gray, <italic>S</italic>2 - gray, <italic>S</italic>3 - black.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g012.tif"/>
</fig>
<p>
<italic>Summary -</italic> in our task settings we have only a limited set of data available. It seems that in such conditions training TCR even with augmentation does not seem to be feasible. Another consideration is that previous works demonstrated successful use of time-contrastive representations in the tasks where the whole structure of motion was visible in the collected images. And the representations were learned on the multi-view set of data. The applicability of such representations to more general settings had to be investigated. In our work, we attempted also to use HOG features. This representation demonstrated visually logical outputs (not just random performance but noisy sequential output). However, HOG features are rather heavy to compute at the run time. Depth and VGG demonstrated the best results which means that in robotics applications it is still better to rely on physical measurements. The pretrained deep VGG features seem to be able to capture the distance phenomenon and thus be closer to the depth performance.</p>
<p>However, three aspects have to be taken into account. 1) The ground truth labeling by a human might affect the classification results because in reality there is no clear border between the stages, and the expert that was marking the data had to make a decision where to draw a line between the stages. Therefore it makes sense to look at the qualitative results. 2) The results that we present do not contain any output filtering or smoothening which could have been implemented to avoid additional noise in the output. The noisy output can be also eliminated with a higher level logic based on, for example, the order of the stages. 3) the reward learned on the offline set of data shall be improved in learning online.</p>
</sec>
<sec id="s4-7-2">
<title>4.7.2 Terminal Reward</title>
<p>
<xref ref-type="fig" rid="F13">Figure 13</xref> demonstrates the results of the terminal stage prediction. With these experiments, we intended to verify whether the task-specific contrast features (trained for vehicle control) can be used to detect the terminal stage of the task. When all seasons are mixed the results look good and provide a high classification rate. However, when only trained on summer and tested on the rest of the seasons, the results are not satisfactory. In this case, the contrastive part of the loss seems to allow us to learn the invariant part of the visible aspects of the task. The results of terminal stage classification are more solid compared to stage classification also because it was easier for the expert to label the terminal stage of the task - when the task is completed.</p>
<fig id="F13" position="float">
<label>FIGURE 13</label>
<caption>
<p>Identification of the terminal stage plotted for all seasons. Upper row: scenario 1 - trained on the summer set of data <italic>D</italic>
<sub>
<italic>summer</italic>
</sub>). Lower row: Scenario 2 - trained on the mixed set of data. <italic>x</italic>-axis stands for the time step in reversed order to match the samples of varying length. Color bar encodes whether the machine reached terminal stage or not: non-terminal - gray, terminal - black.</p>
</caption>
<graphic xlink:href="frobt-09-838059-g013.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion</title>
<p>We explored the stage-based reward prediction from a visual observation implemented for a real-world long-horizon robotic task. In our set-up, a wheel-loader performs a pile loading task. The reward is predicted from visual observations. We experimented with several most common visual representations used for imitation and reinforcement learning. In prior work, both the stage-based reward and visual representations were tested in laboratory conditions. Here we report the results for the outdoor pile loading task performed during three seasons.</p>
<p>Our question was whether the visual features and reward prediction models can be transferred between seasons. The results suggest that neither of the commonly used visual representation allows transfer from summer to other seasons without a loss of performance. The smallest drop of accuracy was produced with depth features. The best performance was achieved by mixing the data from all seasons. In this case, the most reliable results are achieved with depth and deep pre-selected VGG features. Time-contrastive features seem not to be efficient when trained from scratch on a small set of data. They seem to be less effective when the visual data does not contain a third-person view reflecting the structure of motion. As for real-world implementation, one should consider how the training data is labeled and which combinations of feature/classifier are computationally feasible to use.</p>
<p>In the future, we plan to investigate the problem of automatic ground truth generation, for example, by retrospective sensory data. This problem seems to be the bottleneck for the majority of current industries except autonomous driving in urban environments where an abundance of data is available. We will also study the methods to associate the visual observation with a map of the observed location rather than just one reward number. We will explore what form of this cost map is most suitable for action planning.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by authors upon a request.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the Academy of Finland (project no. 336357, PROFI 6 - TAU Imaging Research Platform) and the Academy of Finland project no. 310620.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="FN1">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://scikit-image.org/docs/dev/auto_examples/features_detection/plot_hog.html">https://scikit-image.org/docs/dev/auto_examples/features_detection/plot_hog.html</ext-link>
</p>
</fn>
<fn id="FN2">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://github.com/nianticlabs/monodepth2">https://github.com/nianticlabs/monodepth2</ext-link>
</p>
</fn>
<fn id="fn3">
<label>3</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://scikit-learn.org/stable/modules/svm.html">https://scikit-learn.org/stable/modules/svm.html</ext-link>
</p>
</fn>
<fn id="fn4">
<label>4</label>
<p>Avant 635 is a multi-purpose loader and also used for pallet loading, thus the boom (manipulator) comes with an extra degree of freedom a telescopic (prismatic) boom, which is not used in this study since it is not common in earth moving.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Azulay</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Shapiro</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Wheel Loader Scooping Controller Using Deep Reinforcement Learning</article-title>. <source>IEEE Access</source> <volume>9</volume>. <pub-id pub-id-type="doi">10.1109/access.2021.3056625</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Backman</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lindmark</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bodin</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Servin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>M&#xf6;rk</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>L&#xf6;fgren</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Continuous Control of an Underground Loader Using Deep Reinforcement Learning</article-title>. <source>Machines</source> <volume>9</volume>, <fpage>216</fpage>. <pub-id pub-id-type="doi">10.3390/machines9100216</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lai</surname>
<given-names>W.-S.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M.-H.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Depth-aware Video Frame Interpolation</article-title>,&#x201d; in <conf-name>IEEE Conference on Computer Vision and Pattern Recognition, CVPR</conf-name>. <pub-id pub-id-type="doi">10.1109/cvpr.2019.00382</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Belghazi</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Baratin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rajeshwar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ozair</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Courville</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). &#x201c;<article-title>Mutual Information Neural Estimation</article-title>,&#x201d; in <conf-name>Proceedings of the 35th International Conference on Machine Learning</conf-name>. </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bishop</surname>
<given-names>C. M.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Pattern Recognition and Machine Learning</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Criminisi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shotton</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Konukoglu</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Decision Forests: A Unified Framework for Classification, Regression, Density Estimation, Manifold Learning and Semi-supervised Learning</article-title>,&#x201d; in <conf-name>Foundations and trends&#xae; in computer graphics and vision edn</conf-name>. </citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dadashi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Hussenot</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Geist</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pietquin</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Primal Wasserstein Imitation Learning</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations</conf-name>. </citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dadhich</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bodin</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Sandin</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Andersson</surname>
<given-names>U.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Machine Learning Approach to Automatic Bucket Loading</article-title>,&#x201d; in <conf-name>24th Mediterranean Conference on Control and Automation (MED)</conf-name>. <pub-id pub-id-type="doi">10.1109/med.2016.7535925</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dadhich</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sandin</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bodin</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Andersson</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Martinsson</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Field Test of Neural-Network Based Automatic Bucket-Filling Algorithm for Wheel-Loaders</article-title>. <source>Automation Constr.</source> <volume>97</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.10.013</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dalal</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Triggs</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Histograms of Oriented Gradients for Human Detection</article-title>,&#x201d; in <conf-name>IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR</conf-name>. </citation>
</ref>
<ref id="B11">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Socher</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.-J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Fei-Fei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Imagenet: A Large-Scale Hierarchical Image Database</article-title>,&#x201d; in <conf-name>2009 IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <pub-id pub-id-type="doi">10.1109/cvpr.2009.5206848</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>D. Jud</surname>
<given-names>P. L.</given-names>
</name>
<name>
<surname>Hottiger</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Planning and Control for Autonomous Excavation</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>2</volume>, <fpage>2151</fpage>. <pub-id pub-id-type="doi">10.1109/lra.2017.2721551</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Duda</surname>
<given-names>R. O.</given-names>
</name>
<name>
<surname>Hart</surname>
<given-names>P. E.</given-names>
</name>
<name>
<surname>Stork</surname>
<given-names>D. G.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Pattern Classification</source>. <edition>second edition</edition>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>Wiley</publisher-name>. </citation>
</ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dwibedi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tompson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lynch</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Sermanet</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Learning Actionable Representations from Visual Observations</article-title>,&#x201d; in <conf-name>IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS</conf-name>. <pub-id pub-id-type="doi">10.1109/iros.2018.8593951</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Egli</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A General Approach for the Automation of Hydraulic Excavator Arms Using Reinforcement Learning</article-title>. <source>IEEE Robotics Automation Lett.</source> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fernando</surname>
<given-names>H. A.</given-names>
</name>
<name>
<surname>Marshall</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Larsson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Towards Controlling Bucket Fill Factor in Robotic Excavation by Learning Admittance Control Setpoints</article-title>. <source>Springer Proc. Adv. Robotics</source> <volume>5</volume>. <pub-id pub-id-type="doi">10.1007/978-3-319-67361-5_3</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Finn</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>X. Y.</given-names>
</name>
<name>
<surname>Duan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Darrell</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Abbeel</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Learning Visual Feature Spaces for Robotic Manipulation with Deep Spatial Autoencoders</article-title>. <source>Corr. abs/1509.06113</source>. </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Learning Robust Rewards with Adversarial Inverse Reinforcement Learning</article-title>. <source>arXiv Prepr. arXiv:1710.11248</source>. </citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>G. Dulac-Arnold</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Mankowitz</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Challenges of Real-World Reinforcement Learning</article-title>,&#x201d; in <conf-name>International Joint Conference on Neural Networks (IJCNN)</conf-name>. </citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ghasemipour</surname>
<given-names>S. K. S.</given-names>
</name>
<name>
<surname>Zemel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A Divergence Minimization Perspective on Imitation Learning Methods</article-title>,&#x201d; in <conf-name>Conference on Robot Learning (CORL)</conf-name>. </citation>
</ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ghasemipour</surname>
<given-names>S. K. S.</given-names>
</name>
<name>
<surname>Zemel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>A Divergence Minimization Perspective on Imitation Learning Methods</article-title>,&#x201d; in <conf-name>Conference on Robot Learning (PMLR)</conf-name>. </citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ghosh</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Learning Actionable Representations with Goal Conditioned Policies</article-title>,&#x201d; in <conf-name>7th International Conference on Learning Representations, ICLR</conf-name>. </citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I. J.</given-names>
</name>
<name>
<surname>Pouget-Abadie</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mirza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Warde-Farley</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ozair</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). &#x201c;<article-title>Generative Adversarial Nets</article-title>,&#x201d; in <conf-name>Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2</conf-name>. </citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hadsell</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chopra</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>LeCun</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Dimensionality Reduction by Learning an Invariant Mapping</article-title>,&#x201d; in <conf-name>IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR</conf-name>. </citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Halbach</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>K&#xe4;m&#xe4;r&#xe4;inen</surname>
<given-names>J.-K.</given-names>
</name>
<name>
<surname>Ghabcheloo</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Neural Network Pile Loading Controller Trained by Demonstration</article-title>,&#x201d; in <conf-name>IEEE International Conference on Robotics and Automation, ICRA</conf-name>. <pub-id pub-id-type="doi">10.1109/icra.2019.8793468</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Higgins</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Pal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rusu</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Matthey</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Burgess</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Darla: Improving Zero-Shot Transfer in Reinforcement Learning</article-title>. <source>ArXiv abs/1707.08475</source>. </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ho</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ermon</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Generative Adversarial Imitation Learning</article-title>. <source>Adv. neural Inf. Process. Syst.</source> </citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kaiser</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Babaeizadeh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Milos</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Osinski</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Campbell</surname>
<given-names>R. H.</given-names>
</name>
<name>
<surname>Czechowski</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). &#x201c;<article-title>Model Based Reinforcement Learning for Atari</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations</conf-name>. </citation>
</ref>
<ref id="B29">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Koch</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zemel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Siamese Neural Networks for One-Shot Image Recognition</article-title>,&#x201d; in <conf-name>ICML deep learning workshop. vol. 2</conf-name>. </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Koert</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kircher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salikutluk</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>D&#x2019;Eramo</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Peters</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multi-channel Interactive Reinforcement Learning for Sequential Tasks</article-title>. <source>Front. Robotics AI</source> <volume>7</volume>, <fpage>97</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2020.00097</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kurinov</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Orzechowski</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>H&#xe4;m&#xe4;l&#xe4;inen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Mikkola</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Automated Excavator Based on Reinforcement Learning and Multibody System Dynamics</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>213998</fpage>. <pub-id pub-id-type="doi">10.1109/access.2020.3040246</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Florensa</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tremblay</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ratliff</surname>
<given-names>N. D.</given-names>
</name>
<name>
<surname>Garg</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ramos</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). &#x201c;<article-title>Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning</article-title>,&#x201d; in <conf-name>IEEE International Conference on Robotics and Automation, ICRA</conf-name>. <pub-id pub-id-type="doi">10.1109/icra40945.2020.9197125</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lillicrap</surname>
<given-names>T. P.</given-names>
</name>
<name>
<surname>Hunt</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Heess</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Erez</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tassa</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Continuous Control with Deep Reinforcement Learning</article-title>. <source>arXiv Prepr. arXiv:1509.02971</source>. </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lowe</surname>
<given-names>D. G.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Distinctive Image Features from Scale-Invariant Keypoints</article-title>. <source>Int. J. Comput. Vis.</source> <volume>60</volume>, <fpage>91</fpage>. <pub-id pub-id-type="doi">10.1023/b:visi.0000029664.99615.94</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Osa</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Pajarinen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Neumann</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bagnell</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Abbeel</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Peters</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>An Algorithmic Perspective on Imitation Learning</article-title>. <source>Found. Trends Robotics</source>. </citation>
</ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>R. Fukui</surname>
<given-names>M. N.</given-names>
</name>
<name>
<surname>Niho</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Uetake</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Imitation-based Control of Automated Ore Excavator to Utilize Human Operator Knowledge of Bedrock Condition Estimation and Excavating Motion Selection</article-title>,&#x201d; in <conf-name>IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</conf-name>. <pub-id pub-id-type="doi">10.1109/iros.2015.7354217</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rubner</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tomasi</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Guibas</surname>
<given-names>L. J.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>The Earth Mover&#x2019;s Distance as a Metric for Image Retrieval</article-title>. <source>Int. J. Comput. Vis.</source> <volume>40</volume>, <fpage>99</fpage>&#x2013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.1023/a:1026543900054</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Schmeckpeper</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Rybkin</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Daniilidis</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Finn</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Reinforcement Learning with Videos: Combining Offline Observations with Interaction</article-title>,&#x201d; in <conf-name>Conference on Robot Learning</conf-name>. </citation>
</ref>
<ref id="B39">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Schroff</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kalenichenko</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Philbin</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Facenet: A Unified Embedding for Face Recognition and Clustering</article-title>,&#x201d; in <conf-name>2015 IEEE Conference on Computer Vision and Pattern Recognition, CVPR</conf-name>. <pub-id pub-id-type="doi">10.1109/cvpr.2015.7298682</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sermanet</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lynch</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hsu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017a</year>). &#x201c;<article-title>Time-contrastive Networks: Self-Supervised Learning from Multi-View Observation</article-title>,&#x201d; in <conf-name>IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPRW</conf-name>. <pub-id pub-id-type="doi">10.1109/cvprw.2017.69</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sermanet</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017b</year>). &#x201c;<article-title>Unsupervised Perceptual Rewards for Imitation Learning</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations, ICLR</conf-name>. <pub-id pub-id-type="doi">10.15607/rss.2017.xiii.050</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Simonyan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zisserman</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Very Deep Convolutional Networks for Large-Scale Image Recognition</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations</conf-name>. </citation>
</ref>
<ref id="B43">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Finn</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>End-to-end Robotic Reinforcement Learning without Reward Engineering</article-title>,&#x201d; in <conf-name>Robotics: Science and Systems XV</conf-name>. <pub-id pub-id-type="doi">10.15607/rss.2019.xv.073</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sotiropoulos</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Asada</surname>
<given-names>H. H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A Model-free Extremum-Seeking Approach to Autonomous Excavator Control Based on Output Power Maximization</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>4</volume>, <fpage>1005</fpage>. <pub-id pub-id-type="doi">10.1109/lra.2019.2893690</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sotiropoulos</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Asada</surname>
<given-names>H. H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Autonomous Excavation of Rocks Using a Gaussian Process Model and Unscented Kalman Filter</article-title>. <source>IEEE Robotics Automation Lett.</source> <volume>5</volume>, <fpage>2491</fpage>. <pub-id pub-id-type="doi">10.1109/lra.2020.2972891</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Stooke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Abbeel</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Laskin</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Decoupling Representation Learning from Reinforcement Learning</article-title>,&#x201d; in <conf-name>Proceedings of the 38th International Conference on Machine Learning</conf-name>. </citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sutton</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Barto</surname>
<given-names>A. G.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Reinforcement Learning: An Introduction</source>. <edition>second edn</edition>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>. </citation>
</ref>
<ref id="B48">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Szegedy</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Vanhoucke</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ioffe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Shlens</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wojna</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Rethinking the Inception Architecture for Computer Vision</article-title>,&#x201d; in <conf-name>IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <pub-id pub-id-type="doi">10.1109/cvpr.2016.308</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>van den Oord</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vinyals</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Representation Learning with Contrastive Predictive Coding</article-title>. <source>CoRR</source>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1807.03748">http://arxiv.org/abs/1807.03748</ext-link>
</comment>. </citation>
</ref>
<ref id="B50">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vecer&#xed;k</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sushkov</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Barker</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Roth&#xf6;rl</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hester</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Scholz</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A Practical Approach to Insertion with Variable Socket Position Using Deep Reinforcement Learning</article-title>,&#x201d; in <conf-name>IEEE International Conference on Robotics and Automation, ICRA</conf-name>. </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Veiga</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Akrour</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Peters</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Hierarchical Tactile-Based Control Decomposition of Dexterous In-Hand Manipulation Tasks</article-title>. <source>Front. Robotics AI</source> <volume>7</volume>, <fpage>521448</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2020.521448</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Strokina</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Serbenyuk</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Ghabcheloo</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>K&#xe4;m&#xe4;r&#xe4;inen</surname>
<given-names>J.-K.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Learning a Pile Loading Controller from Demonstrations</article-title>,&#x201d; in <conf-name>IEEE Int. Conf. on Robotics and Automation, ICRA</conf-name>. <pub-id pub-id-type="doi">10.1109/icra40945.2020.9196907</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Strokina</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Serbenyuk</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Pajarinen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ghabcheloo</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vihonen</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). &#x201c;<article-title>Neural Network Controller for Autonomous Pile Loading Revised</article-title>,&#x201d; in <conf-name>IEEE International Conference on Robotics and Automation, ICRA</conf-name>. <pub-id pub-id-type="doi">10.1109/icra48506.2021.9561804</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rosendo</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Risk-aware Model-Based Control</article-title>. <source>Front. Robotics AI</source> <volume>8</volume>, <fpage>617839</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2021.617839</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shah</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hartikainen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). &#x201c;<article-title>The Ingredients of Real World Robotic Reinforcement Learning</article-title>,&#x201d; in <conf-name>International Conference on Learning Representations, ICLR</conf-name>. </citation>
</ref>
</ref-list>
</back>
</article>