<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1076780</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2023.1076780</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>AROS: Affordance Recognition with One-Shot Human Stances</article-title>
<alt-title alt-title-type="left-running-head">Pacheco-Ortega and Mayol-Cuevas</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2023.1076780">10.3389/frobt.2023.1076780</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Pacheco-Ortega</surname>
<given-names>Abel</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1974061/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mayol-Cuevas</surname>
<given-names>Walterio</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/871515/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Visual Information Lab</institution>, <institution>Department of Computer Science</institution>, <institution>University of Bristol</institution>, <addr-line>Bristol</addr-line>, <country>United Kingdom</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Amazon.com</institution>, <addr-line>Seattle</addr-line>, <addr-line>WA</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1794592/overview">Fuqiang Gu</ext-link>, Chongqing University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1834967/overview">Xiaogang Jin</ext-link>, Zhejiang University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1979162/overview">Vittorio Cuculo</ext-link>, University of Milan, Italy</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Abel Pacheco-Ortega, <email>abel.pachecoortega@bristol.ac.uk</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Robot Vision and Artificial Perception, a section of the journal Frontiers in Robotics and AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1076780</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>03</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Pacheco-Ortega and Mayol-Cuevas.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Pacheco-Ortega and Mayol-Cuevas</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>We present Affordance Recognition with One-Shot Human Stances (AROS), a one-shot learning approach that uses an explicit representation of interactions between highly articulated human poses and 3D scenes. The approach is one-shot since it does not require iterative training or retraining to add new affordance instances. Furthermore, only one or a small handful of examples of the target pose are needed to describe the interactions. Given a 3D mesh of a previously unseen scene, we can predict affordance locations that support the interactions and generate corresponding articulated 3D human bodies around them. We evaluate the performance of our approach on three public datasets of scanned real environments with varied degrees of noise. Through rigorous statistical analysis of crowdsourced evaluations, our results show that our one-shot approach is preferred up to 80% of the time over data-intensive baselines.</p>
</abstract>
<kwd-group>
<kwd>affordance detection</kwd>
<kwd>scene understanding</kwd>
<kwd>human interactions</kwd>
<kwd>visual perception</kwd>
<kwd>affordances</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Robot Vision and Artificial Perception</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Vision evolved to make inferences in a 3D world, and one of the most important assessments we can make is what can be done where. Detecting such environmental affordances allows the identification of locations that support actions, such as stand-able, walk-able, place-able, and sit-able. Human affordance detection is not only important in scene analysis and scene understanding but also potentially beneficial in object detection and labeling (<italic>via</italic> how objects can be used) and can eventually be useful for scene generation as well.</p>
<p>Recent approaches have worked toward providing such key competency to artificial systems <italic>via</italic> iterative methods, such as deep learning (<xref ref-type="bibr" rid="B35">Zhang&#xa0;et&#xa0;al., 2020a</xref>; <xref ref-type="bibr" rid="B3">Bochkovskiy&#xa0;et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B4">Carion&#xa0;et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B7">Du&#xa0;et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B20">Nekrasov&#xa0;et&#xa0;al., 2021</xref>). The effectiveness of these data-driven efforts is highly dependent on the number of classes, the number of examples per class, and their diversity. Usually, a dataset consists of thousands of examples, and the training process requires a significant amount of hand tuning and computing of resources. When a new category needs to be added, further sufficient samples need to be provided and training remade. The appeal for one-shot training methods is clear.</p>
<p>Often, human pose-in-scene detection is conflated with object detection or other semantic scene recognition, for example, training to detect sit-able locations through chair recognition, while this is a flawed approach for general action-scene understanding, first, since people can recognize numerous non-chair locations where they can sit, e.g., on tables or cabinets (<xref ref-type="fig" rid="F1">Figure&#xa0;1</xref>). Second, an object-driven approach may fail to consider that affordance detection depends on the object pose and its surroundings&#x2014;it should not detect a chair as sit-able if it is upside-down or if an object is over it. Finally, object detectors alone may struggle to perceive a potentially sit-able place if a particular object example was not covered during training.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>AROS is capable of detecting human&#x2013;scene interactions with one-shot learning. Given a scene, our approach can detect locations that support interactions and generate the interacting human body in a natural and plausible way. Images show examples of detected sit-able, reach-able, lie-able, and stand-able locations.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g001.tif"/>
</fig>
<p>To address these limitations, Affordance Recognition with One-shot Human Stances (AROS) uses a direct representation of human-scene affordances. It extracts an explainable geometrical description by analyzing proximity zones and clearance space between interacting entities. The approach allows training from one or very few data samples per affordance and is capable of handling noisy scene data as provided by real visual sensors, such as RGBD and stereo cameras.</p>
<p>In summary, our contributions are as follows: 1) we propose a one-shot learning geometric-driven affordance descriptor that captures both proximity zones and clearance space around human&#x2013;pose interactions. 2) We set a statistical framework that relies on both central tendency statistics and a statistical inference to evaluate the performance of the compared approaches. The tests show that our approach generates natural and physically plausible human&#x2013;scene interactions with better performance than intensively trained state-of-the-art methods. 3) Our approach demonstrates control on the kind of human&#x2013;scene interaction sought, which permits exploring scenes with a concatenation of affordances.</p>
</sec>
<sec id="s2">
<title>2 Related work</title>
<p>Following Gibson&#x2019;s suggestion that affordances are what we perceive when looking at scenes or objects (<xref ref-type="bibr" rid="B10">Gibson, 1977</xref>), the perception of human affordances with computational approaches has been extensively explored over the years. Before the popularity of data-intensive approaches, <xref ref-type="bibr" rid="B12">Gupta&#xa0;et&#xa0;al. (2011)</xref> employed an environment geometric estimation and a voxelized discretization of four human poses to measure the environment affordance capabilities. This human pose method was employed by <xref ref-type="bibr" rid="B8">Fouhey&#xa0;et&#xa0;al. (2015)</xref> to automatically generate thousands of labeled RGB frames from the NYUv2 dataset (<xref ref-type="bibr" rid="B30">Silberman&#xa0;et&#xa0;al., 2012</xref>) for training a neural network and a set of local discriminative templates that permits the detection of four human affordances. A related approach was explored by <xref ref-type="bibr" rid="B25">Roy and Todorovic (2016)</xref>, where detection was performed for five different human affordances through a pipeline of CNNs that includes the extraction of mid-level cues trained on the NYUv2 dataset (<xref ref-type="bibr" rid="B30">Silberman&#xa0;et&#xa0;al., 2012</xref>). <xref ref-type="bibr" rid="B19">Luddecke and Worgotter (2017)</xref> implemented a residual neural network for detecting 15 human affordances and trained using a look-up table that assigns affordances to object parts on the ADE20K dataset (<xref ref-type="bibr" rid="B41">Zhou&#xa0;et&#xa0;al., 2017</xref>).</p>
<p>Another research line has been the creation of action maps. <xref ref-type="bibr" rid="B27">Savva&#xa0;et&#xa0;al. (2014)</xref> generated affordance maps by learning relations between human poses and geometries in recorded human actions. <xref ref-type="bibr" rid="B23">Piyathilaka and Kodagoda (2015)</xref> used human skeleton models positioned in different locations in an environment to measure geometrical features and determine the support required. In <xref ref-type="bibr" rid="B24">Rhinehart and Kitani (2016)</xref>, egocentric videos as well as scenes, objects, and actions classifiers were used to build up the action maps.</p>
<p>There have been efforts to use functional reasoning for describing the purpose of elements in the environment that helped define them. <xref ref-type="bibr" rid="B11">Grabner&#xa0;et&#xa0;al. (2011)</xref> designed a geometric detector for sit-able objects, such as chairs, while further explorations performed by <xref ref-type="bibr" rid="B42">Zhu&#xa0;et&#xa0;al. (2016</xref>) and <xref ref-type="bibr" rid="B33">Wu&#xa0;et&#xa0;al. (2020</xref>) included physics engines to ponder constrains, such as collision, inertia friction, and gravity.</p>
<p>An important line of research is focused on generating human&#x2013;environment interactions, representative of affordances detected in the environment. <xref ref-type="bibr" rid="B32">Wang&#xa0;et&#xa0;al. (2017)</xref> proposed an affordance predictor and a 2D human interaction generator trained on more than 20K images extracted from sitcoms with and without humans interacting with the environment. <xref ref-type="bibr" rid="B18">Li&#xa0;et&#xa0;al. (2019)</xref> extended this work by developing a 3D human pose synthesizer that learns on the same dataset of images but generates human interactions into input scenes that are represented as RGB, RGBD, or depth images. <xref ref-type="bibr" rid="B17">Jiang&#xa0;et&#xa0;al. (2016)</xref> exploited the spatial correlation between elements and human interactions on RGBD images to generate human interactions and improve object labeling. These methods use human skeletons for representing body&#x2013;environment configurations, which reduces their representativeness since contacts, collisions, and naturalness of the interactions cannot be evaluated in a reliable manner.</p>
<p>In further studies, <xref ref-type="bibr" rid="B26">Ruiz and Mayol-Cuevas (2020)</xref> developed a geometric interaction descriptor for non-articulated, rigid object shapes. Given a 3D environment, the method demonstrated good generalization on detecting physically feasible object&#x2013;environment configurations. In the SMPL-X human body representation (<xref ref-type="bibr" rid="B21">Pavlakos&#xa0;et&#xa0;al., 2019</xref>), <xref ref-type="bibr" rid="B37">Zhang&#xa0;et&#xa0;al. (2020c)</xref> presented a context-aware human body generator that learned the distribution of 3D human poses conditioned to the scene depth and semantics via recordings from the PROX (<xref ref-type="bibr" rid="B13">Hassan&#xa0;et&#xa0;al., 2019</xref>) dataset. In a follow-up effort, <xref ref-type="bibr" rid="B36">Zhang&#xa0;et&#xa0;al. (2020b)</xref> developed a purely geometrical approach to model human&#x2013;scene interactions by explicitly encoding the proximity between the body and the environment, thus only using a mesh as input. Training CNNs and related data-driven methods require the use of most, if not all, of the labeled dataset; e.g., in PROX (<xref ref-type="bibr" rid="B13">Hassan&#xa0;et&#xa0;al., 2019</xref>), there are 100K image frames.</p>
</sec>
<sec id="s3">
<title>3 AROS</title>
<p>Detecting human affordances in an environment is to find locations capable of supporting a given interaction between a human body and the environment. For example, the study of finding &#x201c;suitable to sit&#x201d; locations identifies all those places where a human can sit, which can include a range of object &#x201c;classes&#x201d; (sofa, bed, chair, table, etc.). Our method is motivated to develop a descriptor that characterizes such general interactions without requiring object classes by using two key components and that is lightweight in terms of data requirements while outperforming alternative baselines.</p>
<p>These two components weigh the extraction of characteristics from areas with high (contact) and low (clearance) physical proximity between the entities in interaction (<xref ref-type="fig" rid="F2">Figure&#xa0;2</xref>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>2D and 3D illustrations of our one-shot training pipeline. (Left) Posed human body <italic>M</italic>
<sub>
<italic>h</italic>
</sub> interacting with an environment <italic>M</italic>
<sub>
<italic>e</italic>
</sub> on a reference point <italic>p</italic>
<sub>
<italic>train</italic>
</sub>. (Center) Only during training, we calculate the Voronoi diagram with sample points from both the environment and body surfaces to generate an IBS. (Right) We use the IBS to characterize the proximity zones and the surrounding space with provenance and clearance vectors. A weighted sample of these provenance and clearance vectors, <inline-formula id="inf1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, respectively, results in good generalization of the interaction.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g002.tif"/>
</fig>
<p>Importantly, the representation allows one-shot training per affordance, which is desirable to improve training scalability. Furthermore, our approach is capable of describing and detecting interactions between noisy data representations as obtained from visual depth sensors and highly articulated human poses.</p>
<sec id="s3-1">
<title>3.1&#xa0;A spatial descriptor for spatial interactions</title>
<p>We are inspired by recent methods that have revisited geometric features, such as the bisector surface for scene&#x2013;object indexing (<xref ref-type="bibr" rid="B40">Zhao&#xa0;et&#xa0;al., 2014</xref>) and affordance detection (<xref ref-type="bibr" rid="B26">Ruiz and Mayol-Cuevas, 2020</xref>). Initiating from a spatial representation makes sense if it helps reduce data training needs and simplify explanations&#x2014;as long as it can outperform data-intensive approaches. Our affordance descriptor expands on the Interaction Bisector Surface (IBS) (<xref ref-type="bibr" rid="B40">Zhao&#xa0;et&#xa0;al., 2014</xref>), an approximation of the well-known Bisector Surface (BS) (<xref ref-type="bibr" rid="B22">Peternell, 2000</xref>). Given two surfaces <inline-formula id="inf3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, the BS is the set of sphere centers that touch both surfaces at one point each. Due to its stability and geometrical characteristics, the IBS has been used in context retrieval, interaction classification, and functionality analysis (<xref ref-type="bibr" rid="B40">Zhao&#xa0;et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B16">Hu&#xa0;et&#xa0;al., 2015</xref>; <xref ref-type="bibr" rid="B15">Hu&#xa0;et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B39">Zhao&#xa0;et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B38">Zhao&#xa0;et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B26">Ruiz and Mayol-Cuevas, 2020</xref>). Our approach expands on these ideas and is geometrically intuitive and straightforward. It explicitly captures areas that are important to be in scene-contact and those that are not. Importantly, we show how this approach can be generalized from just one or a small number of samples to a large unseen number of scenes.</p>
<p>Our one-shot training process represents interactions by 3-tuples (<italic>M</italic>
<sub>
<italic>h</italic>
</sub>, <italic>M</italic>
<sub>
<italic>e</italic>
</sub>, and <italic>p</italic>
<sub>
<italic>train</italic>
</sub>), where <italic>M</italic>
<sub>
<italic>h</italic>
</sub> is a posed human-body mesh, <italic>M</italic>
<sub>
<italic>e</italic>
</sub> is an environment mesh, and <italic>p</italic>
<sub>
<italic>train</italic>
</sub> is the reference point on <italic>M</italic>
<sub>
<italic>e</italic>
</sub> that supports the interaction. Let <italic>P</italic>
<sub>
<italic>h</italic>
</sub> and <italic>P</italic>
<sub>
<italic>e</italic>
</sub> be the sets of samples on <italic>M</italic>
<sub>
<italic>h</italic>
</sub> and <italic>M</italic>
<sub>
<italic>e</italic>
</sub>, respectively, their IBS <inline-formula id="inf4">
<mml:math id="m4">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> is defined as<disp-formula id="e1">
<mml:math id="m5">
<mml:mi mathvariant="script">I</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi mathvariant="normal">min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi mathvariant="normal">min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>We use the Voronoi diagram <inline-formula id="inf5">
<mml:math id="m6">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula> generated with <italic>P</italic>
<sub>
<italic>h</italic>
</sub> and <italic>P</italic>
<sub>
<italic>e</italic>
</sub> to produce <inline-formula id="inf6">
<mml:math id="m7">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula>. By construction, every ridge in <inline-formula id="inf7">
<mml:math id="m8">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula> is equidistant to the couple of points that defined it. Then, <inline-formula id="inf8">
<mml:math id="m9">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> is composed of ridges in <inline-formula id="inf9">
<mml:math id="m10">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula> generated because of points from both <italic>P</italic>
<sub>
<italic>h</italic>
</sub> and <italic>P</italic>
<sub>
<italic>e</italic>
</sub>. An IBS can reach infinity, but we limit <inline-formula id="inf10">
<mml:math id="m11">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> by clipping it with the bounding sphere of <italic>M</italic>
<sub>
<italic>h</italic>
</sub> with tolerance <italic>ibs</italic>
<sub>
<italic>rf</italic>
</sub>.</p>
<p>The number and distribution of samples in <italic>P</italic>
<sub>
<italic>h</italic>
</sub> and <italic>P</italic>
<sub>
<italic>e</italic>
</sub> are crucial for a well-constructed discrete IBS. A low rate of sampled points degenerates on an IBS that pierces the boundaries of <italic>M</italic>
<sub>
<italic>h</italic>
</sub> or <italic>M</italic>
<sub>
<italic>e</italic>
</sub>. A higher density is critical in those zones where the proximity is high. To populate <italic>P</italic>
<sub>
<italic>h</italic>
</sub> and <italic>P</italic>
<sub>
<italic>e</italic>
</sub>, we first use a Poisson-disc sampling strategy (<xref ref-type="bibr" rid="B34">Yuksel, 2015</xref>) to generate <italic>ibs</italic>
<sub>
<italic>ini</italic>
</sub> evenly distributed samples on each mesh surface. Then, we perform a <italic>counter-part sampling</italic> that increases the number of samples in <italic>P</italic>
<sub>
<italic>e</italic>
</sub> by including the closest points on <italic>M</italic>
<sub>
<italic>e</italic>
</sub> to elements in <italic>P</italic>
<sub>
<italic>h</italic>
</sub>, and similarly, we incorporate in <italic>P</italic>
<sub>
<italic>h</italic>
</sub> the closest point on <italic>M</italic>
<sub>
<italic>h</italic>
</sub> to samples in <italic>P</italic>
<sub>
<italic>e</italic>
</sub>. We perform the <italic>counter-part sampling</italic> strategy <italic>ibs</italic>
<sub>
<italic>cs</italic>
</sub> times to generate a new <inline-formula id="inf11">
<mml:math id="m12">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula>. However, we observed that for intricate human&#x2013;scene poses, convergence to an IBS without mesh piercing is challenging. If the IBS is penetrating the scene, we perform a <italic>collision-point sampling</italic> strategy. This adds as sampling points, a sub-sample of points where collisions happen and their counter-part points (body or environment). We then simply recompute the IBS and repeat the <italic>counter-part sampling</italic> and <italic>collision-point sampling</italic> strategies until we find a candidate <inline-formula id="inf12">
<mml:math id="m13">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> that does not collide with <italic>M</italic>
<sub>
<italic>h</italic>
</sub> or <italic>M</italic>
<sub>
<italic>e</italic>
</sub>. This is a straightforward process that can be implemented efficiently.</p>
<p>To capture the regions of interaction proximity on our enhanced IBS as mentioned above, we use the notion of provenance vectors (<xref ref-type="bibr" rid="B26">Ruiz and Mayol-Cuevas, 2020</xref>). The <italic>provenance vectors</italic> of an interaction start from any point on <inline-formula id="inf13">
<mml:math id="m14">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> and finish on <italic>M</italic>
<sub>
<italic>e</italic>
</sub>. Formally,<disp-formula id="e2">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">I</mml:mi>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi mathvariant="normal">arg min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>a</italic> is the stating point of the delta vector <inline-formula id="inf14">
<mml:math id="m16">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> to the nearest point on <italic>M</italic>
<sub>
<italic>e</italic>
</sub>.</p>
<p>
<italic>Provenance vectors</italic> inform about the direction and distance of the interaction; the smaller the <inline-formula id="inf15">
<mml:math id="m17">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula>, the more important it is in the description. Let <inline-formula id="inf16">
<mml:math id="m18">
<mml:msubsup>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2282;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> be the subset of <italic>provenance vectors</italic> that finish on any point in <italic>P</italic>
<sub>
<italic>e</italic>
</sub>, and we perform a weighted randomized selection sampling of elements from <inline-formula id="inf17">
<mml:math id="m19">
<mml:msubsup>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> with the allocation of weights as follows:<disp-formula id="e3">
<mml:math id="m20">
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf18">
<mml:math id="m21">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> and <inline-formula id="inf19">
<mml:math id="m22">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> are the norms of the biggest and smallest vectors in <inline-formula id="inf20">
<mml:math id="m23">
<mml:msubsup>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, respectively. The selected <italic>provenance vectors</italic> <inline-formula id="inf21">
<mml:math id="m24">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> integrate to our affordance descriptor with an adjustment to normalize their positions, with the defined reference point <italic>p</italic>
<sub>
<italic>train</italic>
</sub> as follows:<disp-formula id="e4">
<mml:math id="m25">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi>n</mml:mi>
<mml:mi>u</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>num</italic>
<sub>
<italic>pv</italic>
</sub> is the number of samples from <inline-formula id="inf22">
<mml:math id="m26">
<mml:msubsup>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> to integrate. The <italic>provenance vectors</italic> alone, however, are insufficient to work successfully on highly articulated objects, such as human poses. They are unable to capture the whole nature of the interaction. We expand this concept by taking a more comprehensive description that considers both areas of the IBS, those that are proximal to surfaces and those that are not.</p>
<p>We include a set of vectors into our descriptor to define the clearance space necessary for performing the given interaction. Given <italic>S</italic>
<sub>
<italic>h</italic>
</sub>, an evenly sampled set of <italic>num</italic>
<sub>
<italic>cv</italic>
</sub> points on <italic>M</italic>
<sub>
<italic>h</italic>
</sub>, the <italic>clearance vectors</italic> that integrate to our descriptor <inline-formula id="inf23">
<mml:math id="m27">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> on the interaction are defined as follows:<disp-formula id="e5">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c8;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m29">
<mml:mi>&#x3c8;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.3333em"/>
<mml:mspace width="0.3333em"/>
<mml:mspace width="0.3333em"/>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if&#x2009;&#x2009;</mml:mtext>
<mml:mi>&#x3c6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3e;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mi>&#x3c6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.3333em"/>
<mml:mspace width="0.3333em"/>
<mml:mspace width="0.3333em"/>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>otherwise</mml:mtext>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>p</italic>
<sub>
<italic>train</italic>
</sub> is the defined reference point, <inline-formula id="inf24">
<mml:math id="m30">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the unit surface normal vector on sample <italic>s</italic>
<sub>
<italic>j</italic>
</sub>, <italic>d</italic>
<sub>
<italic>max</italic>
</sub> is the maximum norm&#xa0;of any <inline-formula id="inf25">
<mml:math id="m31">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, and <inline-formula id="inf26">
<mml:math id="m32">
<mml:mi>&#x3c6;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mi mathvariant="script">I</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the distance traveled by a ray with origin <italic>s</italic>
<sub>
<italic>j</italic>
</sub> and direction <inline-formula id="inf27">
<mml:math id="m33">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> until collision with <inline-formula id="inf28">
<mml:math id="m34">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula>.</p>
<p>Formally, our affordance descriptor, AROS, is defined as<disp-formula id="e7">
<mml:math id="m35">
<mml:mi>f</mml:mi>
<mml:mo>:</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2192;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf29">
<mml:math id="m36">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the unit normal vector on <italic>M</italic>
<sub>
<italic>e</italic>
</sub> at <italic>p</italic>
<sub>
<italic>train</italic>
</sub>. We calculate <inline-formula id="inf30">
<mml:math id="m37">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> for speeding up the detection process.</p>
</sec>
<sec id="s3-2">
<title>3.2 Human affordance detection</title>
<p>Let <inline-formula id="inf31">
<mml:math id="m38">
<mml:mi mathvariant="script">A</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> be an affordance descriptor; we define its rigid transformation with <inline-formula id="inf32">
<mml:math id="m39">
<mml:mi>&#x3c4;</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> being a translation vector and <italic>&#x3d5;</italic> being the rotation around <italic>z</italic> defined by <italic>R</italic>
<sub>
<italic>&#x3d5;</italic>
</sub>.</p>
<p>Given a point <italic>p</italic>
<sub>
<italic>test</italic>
</sub> on an environment mesh <italic>M</italic>
<sub>
<italic>test</italic>
</sub> and its unit surface normal vector <inline-formula id="inf33">
<mml:math id="m40">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">test</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, we determine that such a location supports a trained interaction <inline-formula id="inf34">
<mml:math id="m41">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> if we can find that (1) has a small angle difference between <inline-formula id="inf35">
<mml:math id="m42">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">test</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf36">
<mml:math id="m43">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, (2) once translated to <italic>p</italic>
<sub>
<italic>test</italic>
</sub> and oriented with <italic>&#x3d5;</italic>
<sub>
<italic>test</italic>
</sub>, there is a correct alignment of <inline-formula id="inf37">
<mml:math id="m44">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and (3) a gated number of the <inline-formula id="inf38">
<mml:math id="m45">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is in collision with <italic>M</italic>
<sub>
<italic>test</italic>
</sub>.</p>
<p>A significant angle difference between <inline-formula id="inf39">
<mml:math id="m46">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">test</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf40">
<mml:math id="m47">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> permits to short-cut the test and reject <italic>p</italic>
<sub>
<italic>test</italic>
</sub> with reference to <inline-formula id="inf41">
<mml:math id="m48">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula>. We establish <inline-formula id="inf42">
<mml:math id="m49">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> as the decision threshold for the angle difference. <inline-formula id="inf43">
<mml:math id="m50">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is adjustable based on the level of mesh noise.</p>
<p>If we observe a normal match between <inline-formula id="inf44">
<mml:math id="m51">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <italic>p</italic>
<sub>
<italic>test</italic>
</sub> vectors, we perform transformations over the interaction descriptor <inline-formula id="inf45">
<mml:math id="m52">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> with <italic>&#x3c4;</italic> &#x3d; <italic>p</italic>
<sub>
<italic>test</italic>
</sub> and <italic>n</italic>
<sub>
<italic>&#x3d5;</italic>
</sub> different <italic>&#x3d5;</italic> &#x3d; <italic>&#x3d5;</italic>
<sub>
<italic>test</italic>
</sub> values within [0, 2<italic>&#x3c0;</italic>]. Hence, per each 3-tuple <inline-formula id="inf46">
<mml:math id="m53">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> calculated, we generated a set of rays <italic>R</italic>
<sub>
<italic>pv</italic>
</sub> defined as follows:<disp-formula id="e8">
<mml:math id="m54">
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2033;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3bd;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="0.3333em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3bd;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2033;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(8)</label>
</disp-formula>where <inline-formula id="inf47">
<mml:math id="m55">
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2033;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the starting point and <inline-formula id="inf48">
<mml:math id="m56">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3bd;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> is the direction of each ray. We extend each ray in <italic>R</italic>
<sub>
<italic>pv</italic>
</sub> by <inline-formula id="inf49">
<mml:math id="m57">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> until collision with <italic>M</italic>
<sub>
<italic>test</italic>
</sub> as<disp-formula id="e9">
<mml:math id="m58">
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2033;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3bd;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">test</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em"/>
<mml:mspace width="0.3333em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mi>u</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(9)</label>
</disp-formula>and compare with the magnitude of each correspondent provenance vector in <inline-formula id="inf50">
<mml:math id="m59">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. When any element in <italic>R</italic>
<sub>
<italic>pv</italic>
</sub> extends further than a predetermined limit <italic>max</italic>
<sub>
<italic>long</italic>
</sub>, the collision with the environment is classified as non-colliding. We calculate the alignment score <italic>&#x3ba;</italic> as a sum difference between extended rays and <italic>provenance vectors</italic> with<disp-formula id="e10">
<mml:math id="m60">
<mml:mi>&#x3ba;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2200;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">long</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>The bigger the <italic>&#x3ba;</italic> value, the less the support for the interaction on the <italic>p</italic>
<sub>
<italic>test</italic>
</sub>. We experimentally determine interaction-wise thresholds for the sum of differences <italic>max</italic>
<sub>
<italic>&#x3ba;</italic>
</sub> and the number of missing ray collisions <italic>max</italic>
<sub>
<italic>missings</italic>
</sub> that permits us to score the affordance capabilities on <italic>p</italic>
<sub>
<italic>test</italic>
</sub>.</p>
<p>
<italic>Clearance vectors</italic> are meant to fast-detect collision configurations by ray&#x2013;mesh intersection calculation. Similar to <italic>provenance vectors</italic>, we generate a set of rays <italic>R</italic>
<sub>
<italic>cv</italic>
</sub>, whose origins and directions are determined by <inline-formula id="inf51">
<mml:math id="m61">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. We extend rays in <italic>R</italic>
<sub>
<italic>cv</italic>
</sub> until collision with the environment and calculate its extension <inline-formula id="inf52">
<mml:math id="m62">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. Extended rays with <inline-formula id="inf53">
<mml:math id="m63">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2264;</mml:mo>
<mml:mo stretchy="false">&#x2016;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2016;</mml:mo>
</mml:math>
</inline-formula> are considered as possible collisions. In practice, we also track an interaction-wise threshold to refuse affordance due to collisions <italic>max</italic>
<sub>
<italic>collisions</italic>
</sub>.</p>
<p>A sparse distribution of clearance vectors on bi-dimensional noisy meshes in a 3D space results in collisions that are not detected by <italic>clearance vectors</italic>. To improve, we enhance scenes with a set of <italic>spherical fillers</italic> that pad the scene (see <xref ref-type="fig" rid="F3">Figure&#xa0;3</xref>). More details are provided in <xref ref-type="sec" rid="s10">Supplementary&#xa0;Material</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Approach for detecting human affordances. To mitigate 3D scan noise, the scene is augmented with spherical fillers for detecting collisions and SDF values. Our method detects if a test point in the environment can support an interaction by translating the descriptor to the test position over different orientations and measuring its alignment and collision rate. Then, the best-scored configuration is optimized to generate a more natural and physically plausible interaction with the environment.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g003.tif"/>
</fig>
<sec id="s3-2-1">
<title>3.2.1 Pose optimization</title>
<p>After a positive detection, we generate the body mesh representation used in training at the testing location. This generally has low levels of contact with the unseen environment. These gaps are because our descriptor based its construction on the bisector surface between the interacting entities. We can eliminate the gap by translating the body until it touches the environment. However, this na&#xef;ve method generates configurations that visually lack naturalness, <xref ref-type="fig" rid="F3">Figure&#xa0;3</xref> (Pose with best score).</p>
<p>Every human&#x2013;environment configuration trained has an associated 3D human SMPL-X characterization that we keep and use to optimize the human pose as in the work of <xref ref-type="bibr" rid="B36">Zhang&#xa0;S.&#xa0;et&#xa0;al. (2020b</xref>) with the <italic>AdvOptim</italic> loss function, using the SDF values that have been pre-calculated in each scene with a grid of 256 &#xd7; 256 &#xd7; 256 positions.</p>
<p>Overall, we train a human interaction by generating its AROS descriptor from a single example, keeping the associated SMPL-X parameters of the body pose and defining the contact regions that the body has with the environment. After a positive detection with AROS, we use the associated SMPL-X body parameters and its contact regions to close the environment&#x2013;body gap and generate a more natural body pose, as shown in <xref ref-type="fig" rid="F3">Figure&#xa0;3</xref> (ouput). Our approach generalizes well on the description of interaction and generates natural and physically plausible body&#x2013;environment configurations over novel environments with just one example for training (see <xref ref-type="fig" rid="F4">Figure&#xa0;4</xref>).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Our one-shot learning approach generalizes well on affordance detection. Only one example of an interaction is used to generate an AROS descriptor that generalizes well for the detection of affordances over previously unseen environments.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g004.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments</title>
<p>We conduct experiments in various environment configurations to examine the effectiveness and usefulness of the affordance recognition performed by AROS. Our experiments include several perceptual studies, as well as a <italic>physical plausibility</italic> evaluation of the body&#x2013;environment configurations generated.</p>
<p>Datasets: The PROX dataset (<xref ref-type="bibr" rid="B13">Hassan&#xa0;et&#xa0;al., 2019</xref>) includes data from 20 recordings of subjects interacting within 12 scanned indoor environments. An SMPL-X body model (<xref ref-type="bibr" rid="B21">Pavlakos&#xa0;et&#xa0;al., 2019</xref>) is used to characterize the shape and pose of humans within each frame in recordings. Following the setup in the work of <xref ref-type="bibr" rid="B36">Zhang&#xa0;S.&#xa0;et&#xa0;al. (2020b)</xref>, we use the rooms MPH16, MPH1Library, N0SittingBooth, and N3OpenArea for testing purposes and training on data from other PROX scenes. We also perform evaluations on seven scanned scenes from the MP3D dataset (<xref ref-type="bibr" rid="B5">Chang&#xa0;et&#xa0;al., 2017</xref>) and five scenes from the Replica dataset (<xref ref-type="bibr" rid="B31">Straub&#xa0;et&#xa0;al., 2019</xref>). We calculate the <italic>spherical fillers</italic> and SDF values of all 3D scanned environments.</p>
<p>Training: We manually select 23 frames in which subjects interact in one of the following ways: sitting, standing, lying down, walking, or reaching. From these selected human&#x2013;scene interactions, we generate the AROS descriptors and retain the SMPL-X parameters associated with human poses.</p>
<p>To generate the IBS associated with each trained interaction, we use an initial sampling set of <italic>ibs</italic>
<sub>
<italic>ini</italic>
</sub> &#x3d; 400 on each surface, execute the <italic>counter-part sampling</italic> strategy <italic>ibs</italic>
<sub>
<italic>cs</italic>
</sub> &#x3d; 4 times, and crop the generated IBS <inline-formula id="inf54">
<mml:math id="m64">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula> with <italic>ibs</italic>
<sub>
<italic>rf</italic>
</sub> &#x3d; 1.2. The AROS descriptors are a compound of <italic>num</italic>
<sub>
<italic>pv</italic>
</sub> &#x3d; 512 <italic>provenance vectors</italic> and <italic>num</italic>
<sub>
<italic>cv</italic>
</sub> &#x3d; 256 <italic>clearance vectors</italic> that extend up to <italic>d</italic>
<sub>
<italic>max</italic>
</sub> &#x3d; 5&#xa0;[<italic>cm</italic>] each.</p>
<p>The interaction-wise thresholds <italic>max</italic>
<sub>
<italic>&#x3ba;</italic>
</sub>, <italic>max</italic>
<sub>
<italic>missings</italic>
</sub>, and <italic>max</italic>
<sub>
<italic>collisions</italic>
</sub> are established experimentally, and <italic>max</italic>
<sub>
<italic>long</italic>
</sub> is 1.2 times the radius of the sphere used to crop <inline-formula id="inf55">
<mml:math id="m65">
<mml:mi mathvariant="script">I</mml:mi>
</mml:math>
</inline-formula>. We use a moderate angle difference threshold of <inline-formula id="inf56">
<mml:math id="m66">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c0;</mml:mi>
<mml:mo>/</mml:mo>
<mml:mn>3</mml:mn>
</mml:math>
</inline-formula>, in <italic>n</italic>
<sub>
<italic>&#x3d5;</italic>
</sub> &#x3d; 8 different directions.</p>
<p>With 512 provenance vectors <inline-formula id="inf57">
<mml:math id="m67">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and 256 clearance vectors <inline-formula id="inf58">
<mml:math id="m68">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, the AROS descriptor characterizes an interaction with less than 40&#xa0;KB, including the SMPL-X parameters.</p>
<p>Baselines: We compare our approach with the state-of-the-art PLACE (<xref ref-type="bibr" rid="B36">Zhang&#xa0;et&#xa0;al., 2020b</xref>) and POSA (contact only) (<xref ref-type="bibr" rid="B14">Hassan&#xa0;et&#xa0;al., 2021</xref>). PLACE is a pure scene-centric method that only requires a reference point on a scanned environment to generate a human body performing around it. However, PLACE does not have control over the type of interaction detected/generated. We used naive and optimized versions of this approach in experiments (PLACE, PLACE SimOptim, and PLACE AdvOptim). POSA is a human-centric approach that, given a posed human body mesh, calculates the zones on the body where contact with the scene may occur and uses this feature map to place the body in the environment. We encourage a fair comparison by evaluating the naive and optimized POSA versions that consider only contact information and excludes semantic information (POSA and POSA optimized). In our studies, POSA was executed with the same human shapes and poses used to train AROS.</p>
<sec id="s4-1">
<title>4.1 Physical plausibility</title>
<p>We evaluate the physical plausibility of the compared approaches mainly by following the work of <xref ref-type="bibr" rid="B36">Zhang&#xa0;et&#xa0;al. (2020b)</xref> and <xref ref-type="bibr" rid="B37">Zhang&#xa0;et&#xa0;al. (2020c)</xref>. Given the SDF values of a scene and a body mesh generated, 1) the <italic>contact score</italic> is assigned to 1 if any mesh vertex has a negative SDF value and is evaluated as 0, otherwise, 2) the <italic>non-collision score</italic> is the ratio of vertices with a positive SDF value, and 3) in order to measure the severity of the body&#x2013;environment collision on positive contact, we include the <italic>collision-depth score</italic>, which averages the depth of the collisions between the scene and the generated body mesh.</p>
<sec id="s4-1-1">
<title>4.1.1 Ablation study</title>
<p>We evaluate the influence of <italic>clearance vectors</italic>, spherical fillers, and different optimizers on the PROX dataset. Three different optimization procedures are evaluated. The <italic>downward</italic> optimizer translates the generated body downward (-Z direction) until it comes in contact with the environment. The ICP optimizer uses the well-known Interactive Closest Point algorithm to align the body vertices with the environment mesh. The <italic>AdvOptim</italic> optimizer is described in <xref ref-type="sec" rid="s3-2-1">Section&#xa0;3.2.1</xref>.</p>
<p>
<xref ref-type="table" rid="T1">Table&#xa0;1</xref> shows that models without <italic>clearance vectors</italic> have the highest collision-depth scores on models with the same optimizer. AROS models present a reduction in contact and collision-depth scores in all cases that consider <italic>clearance vectors</italic> in their descriptors to avoid collision with the environment. Spherical fillers have a significant influence on avoiding collisions, producing the best scores in all metrics per optimizer. The ICP optimizer closes the body&#x2013;environment gaps but drastically reduces the performance on both collision scores, while the <italic>AdvOptim</italic> and <italic>downward</italic> optimizers keep a trade-off between collision and contact. The best performance is achieved with affordance descriptors composed of <italic>provenance</italic> and <italic>clearance vectors</italic>, tested in scanned environments enhanced with <italic>spherical fillers</italic>, and where interactions are optimized with the <italic>AdvOptim</italic> optimizer.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Ablation study evaluation scores (<sup>
<italic>&#x2191;</italic>
</sup>: benefit; <sup>
<italic>&#x2193;</italic>
</sup>: cost). The best trade-off between scores per optimizer are in boldface.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Descriptor integrated by</th>
<th align="center">Spherical filler</th>
<th align="center">Optimizer</th>
<th align="center">Non-collision<sup>
<italic>&#x2191;</italic>
</sup>
</th>
<th align="center">Contact<sup>
<italic>&#x2191;</italic>
</sup>
</th>
<th align="center">Collision-depth<sup>
<italic>&#x2193;</italic>
</sup>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<inline-formula id="inf59">
<mml:math id="m69">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="center">w/o</td>
<td align="center">0.9348</td>
<td align="center">0.7998</td>
<td align="center">1.4132</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf60">
<mml:math id="m70">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="left"/>
<td align="center">0.9504</td>
<td align="center">0.6901</td>
<td align="center">0.6757</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf61">
<mml:math id="m71">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">Yes</td>
<td align="left"/>
<td align="center">
<bold>0.9623</bold>
</td>
<td align="center">
<bold>0.5448</bold>
</td>
<td align="center">
<bold>0.1573</bold>
</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf62">
<mml:math id="m72">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="center">ICP<sup>a</sup>
</td>
<td align="center">0.5820</td>
<td align="center">1.0000</td>
<td align="center">7.3770</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf63">
<mml:math id="m73">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="left"/>
<td align="center">0.5775</td>
<td align="center">1.0000</td>
<td align="center">7.2180</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf64">
<mml:math id="m74">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">Yes</td>
<td align="left"/>
<td align="center">
<bold>0.6299</bold>
</td>
<td align="center">1.0000</td>
<td align="center">
<bold>6.2665</bold>
</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf65">
<mml:math id="m75">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="center">Downward</td>
<td align="center">0.9271</td>
<td align="center">0.9377</td>
<td align="center">1.4380</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf66">
<mml:math id="m76">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="left"/>
<td align="center">0.9496</td>
<td align="center">0.9036</td>
<td align="center">0.7089</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf67">
<mml:math id="m77">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">Yes</td>
<td align="left"/>
<td align="center">
<bold>0.9641</bold>
</td>
<td align="center">
<bold>0.8603</bold>
</td>
<td align="center">
<bold>0.1807</bold>
</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf68">
<mml:math id="m78">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="center">AdvOptim</td>
<td align="center">0.9552</td>
<td align="center">0.9638</td>
<td align="center">2.0249</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf69">
<mml:math id="m79">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">No</td>
<td align="left"/>
<td align="center">0.9717</td>
<td align="center">0.9508</td>
<td align="center">1.2325</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf70">
<mml:math id="m80">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">train</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</td>
<td align="center">Yes</td>
<td align="left"/>
<td align="center">
<bold>0.9818</bold>
</td>
<td align="center">
<bold>0.9403</bold>
</td>
<td align="center">
<bold>0.6341</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn1">
<label>a</label>
<p>ICP stands for the Iterative Closest Point.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-1-2">
<title>4.1.2 Comparison with the state of the art</title>
<p>We generated 1300 interacting bodies per model in each of the 16 scenes and reported the averages of calculated non-collision, contact, and collision-depth scores. The results are shown in <xref ref-type="table" rid="T2">Table&#xa0;2</xref>. In all datasets, interacting bodies generated using our approach provided a good trade-off with high non-collision but low contact and collision-depth scores.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Physical plausibility: Non-collision, contact, and collision-depth scores (<sup>
<italic>&#x2191;</italic>
</sup>: benefit; <sup>
<italic>&#x2193;</italic>
</sup>: cost) before and after optimization. The best results are in boldface.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="2" align="left"/>
<th colspan="3" align="center">Non-collision<sup>
<italic>&#x2191;</italic>
</sup>
</th>
<th colspan="3" align="center">Contact<sup>
<italic>&#x2191;</italic>
</sup>
</th>
<th colspan="3" align="center">Collision-depth<sup>
<italic>&#x2193;</italic>
</sup>
</th>
</tr>
<tr>
<th align="left">Model</th>
<th align="left">Optimizer</th>
<th align="center">PROX</th>
<th align="center">MP3D</th>
<th align="center">Replica</th>
<th align="center">PROX</th>
<th align="center">MP3D</th>
<th align="center">Replica</th>
<th align="center">PROX</th>
<th align="center">MP3D</th>
<th align="center">Replica</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PLACE</td>
<td align="left">w/o</td>
<td align="center">0.9207</td>
<td align="center">0.9625</td>
<td align="center">0.9554</td>
<td align="center">0.9125</td>
<td align="center">0.5116</td>
<td align="center">0.8115</td>
<td align="center">1.6285</td>
<td align="center">0.8958</td>
<td align="center">1.2031</td>
</tr>
<tr>
<td align="left">PLACE</td>
<td align="left">SimOptim</td>
<td align="center">0.9253</td>
<td align="center">0.9628</td>
<td align="center">0.9562</td>
<td align="center">0.9263</td>
<td align="center">0.5910</td>
<td align="center">0.8571</td>
<td align="center">1.8169</td>
<td align="center">1.0960</td>
<td align="center">1.5485</td>
</tr>
<tr>
<td align="left">PLACE</td>
<td align="left">AdvOptim</td>
<td align="center">0.9665</td>
<td align="center">0.9798</td>
<td align="center">0.9659</td>
<td align="center">0.9725</td>
<td align="center">0.5810</td>
<td align="center">0.9931</td>
<td align="center">1.6327</td>
<td align="center">1.1346</td>
<td align="center">1.6145</td>
</tr>
<tr>
<td align="left">POSA</td>
<td align="left">w/o</td>
<td align="center">
<bold>0.9820</bold>
</td>
<td align="center">0.9792</td>
<td align="center">0.9814</td>
<td align="center">0.9396</td>
<td align="center">0.9526</td>
<td align="center">0.9888</td>
<td align="center">1.1252</td>
<td align="center">1.5416</td>
<td align="center">2.0620</td>
</tr>
<tr>
<td align="left">POSA</td>
<td align="left">Optimized</td>
<td align="center">0.9753</td>
<td align="center">0.9725</td>
<td align="center">0.9765</td>
<td align="center">
<bold>0.9927</bold>
</td>
<td align="center">
<bold>0.9988</bold>
</td>
<td align="center">
<bold>0.9963</bold>
</td>
<td align="center">1.5343</td>
<td align="center">2.0063</td>
<td align="center">2.4518</td>
</tr>
<tr>
<td align="left">AROS</td>
<td align="left">w/o</td>
<td align="center">0.9615</td>
<td align="center">
<bold>0.9853</bold>
</td>
<td align="center">
<bold>0.9931</bold>
</td>
<td align="center">0.5654</td>
<td align="center">0.3287</td>
<td align="center">0.4860</td>
<td align="center">
<bold>0.1648</bold>
</td>
<td align="center">
<bold>0.1326</bold>
</td>
<td align="center">
<bold>0.2096</bold>
</td>
</tr>
<tr>
<td align="left">AROS</td>
<td align="left">AdvOptim</td>
<td align="center">0.9816</td>
<td align="center">
<bold>0.9853</bold>
</td>
<td align="center">0.9883</td>
<td align="center">0.9363</td>
<td align="center">0.6213</td>
<td align="center">0.8682</td>
<td align="center">0.6330</td>
<td align="center">0.8716</td>
<td align="center">0.8615</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4-2">
<title>4.2 Perception of naturalness</title>
<p>We use Amazon Mechanical Turk to compare and evaluate the naturalness of body&#x2013;environment configurations generated by our approach and baselines. We used only the best version of the compared methods (with optimizer). Each scene in our test set was used equally to select 162 locations around which the compared approaches generate human interactions. MTurk judges observed all human&#x2013;environment pairs generated through dynamic views, allowing us to showcase them from different perspectives. Each judge performed 11 randomly selected assessments, without repetition, that included two control questions to detect and exclude untrustworthy evaluators. Three different judges accomplished each of the evaluations. Our perceptual experiments include individual and comparison studies for each comparison carried out.</p>
<p>In our side-by-side comparison studies, interactions detected/generated from two approaches are exposed simultaneously. Then, MTurkers were asked to respond to the question &#x201c;Which example is more natural?&#x201d; by direct selection.</p>
<p>We used the same set of interactions for individual evaluation studies, where judges rated every individual human&#x2013;scene interaction by responding to &#x201c;The human is interacting very naturally with the scene. What is your opinion?&#x201d; with a 5-point Likert scale according to its agreement level: 1) strongly disagree, 2) disagree, 3) neither disagree nor agree, 4) agree, and 5) strongly agree.</p>
<sec id="s4-2-1">
<title>4.2.1 Randomly selected test locations</title>
<p>The first group of studies compares human&#x2013;scene configurations generated at randomly selected locations. On the side-by-side comparison study that contrasts AROS with PLACE, our approach was selected as more natural in 60.7% of all assessments. Compared to POSA, ours is selected in 72.6% of all tests performed. The results per dataset are shown in <xref ref-type="table" rid="T4">Table&#xa0;4</xref> (% preferences in random locations).</p>
<p>Individual evaluation studies also suggest that AROS produced more natural interactions (see <xref ref-type="table" rid="T3">Table&#xa0;3</xref>). The mean and standard deviations of these scores obtained by the judges to PLACE are 3.23 &#xb1; 1.35 in comparison with AROS, 3.39 &#xb1; 1.25, while in the second study, these statistics obtained by POSA were 2.79 &#xb1; 1.18 in contrast with AROS, 3.20 &#xb1; 1.18. Evaluation scores of AROS have a larger mean and a narrower standard deviation compared to baselines. However, these descriptive statistics must be cautiously used as evidence to determine a performance difference because it assumes that the distribution of scores approximately resembles a normal distribution and that the ordinal variable was perceived as numerically equidistant by judges. Regrettably, Shapiro&#x2013;Wilk tests (<xref ref-type="bibr" rid="B28">Shapiro and Wilk, 1965</xref>) performed on data show that the score distributions depart from normality in both evaluation studies, PLACE/AROS and POSA/AROS with <italic>p</italic> &#x3c; 0.01.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Cross-tabulation data of individual evaluation studies on randomly selected locations. The best are in boldface.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Individual evaluation study</th>
<th align="left">Model</th>
<th align="left"/>
<th align="left">1. Stronglydisagree</th>
<th align="center">2. Disagree</th>
<th align="center">3. Neither</th>
<th align="center">4. Agree</th>
<th align="center">5. Strongly agree</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="center">PLACE <italic>vs</italic>. AROS</td>
<td rowspan="2" align="left">PLACE</td>
<td align="left">Observed frequency</td>
<td align="center">68</td>
<td align="center">98</td>
<td align="center">70</td>
<td align="center">153</td>
<td align="center">97</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">14.0</td>
<td align="center">20.2</td>
<td align="center">14.4</td>
<td align="center">31.5</td>
<td align="center">19.9</td>
</tr>
<tr>
<td rowspan="2" align="left">AROS</td>
<td align="left">Observed frequency</td>
<td align="center">
<bold>43</bold>
</td>
<td align="center">
<bold>98</bold>
</td>
<td align="center">64</td>
<td align="center">
<bold>187</bold>
</td>
<td align="center">
<bold>94</bold>
</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">
<bold>8.8</bold>
</td>
<td align="center">
<bold>20.2</bold>
</td>
<td align="center">13.2</td>
<td align="center">
<bold>38.5</bold>
</td>
<td align="center">
<bold>19.3</bold>
</td>
</tr>
<tr>
<td rowspan="4" align="center">POSA <italic>vs</italic>. AROS</td>
<td rowspan="2" align="left">POSA</td>
<td align="left">Observed frequency</td>
<td align="center">64</td>
<td align="center">173</td>
<td align="center">89</td>
<td align="center">123</td>
<td align="center">37</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">13.2</td>
<td align="center">35.6</td>
<td align="center">18.3</td>
<td align="center">25.3</td>
<td align="center">7.6</td>
</tr>
<tr>
<td rowspan="2" align="left">AROS</td>
<td align="left">Observed frequency</td>
<td align="center">
<bold>29</bold>
</td>
<td align="center">
<bold>136</bold>
</td>
<td align="center">85</td>
<td align="center">
<bold>179</bold>
</td>
<td align="center">
<bold>57</bold>
</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">
<bold>6.0</bold>
</td>
<td align="center">
<bold>28.0</bold>
</td>
<td align="center">17.5</td>
<td align="center">
<bold>36.8</bold>
</td>
<td align="center">
<bold>11.7</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Based on this, we performed a chi-square test of homogeneity (<xref ref-type="bibr" rid="B9">Franke&#xa0;et&#xa0;al., 2012</xref>) with a significance level <italic>&#x3b1;</italic> &#x3d; 0.05, to determine if the distributions of evaluation scores are statistically similar. If we observe significance, the level of association between the approach and the distribution of the scores was determined by calculating Cramer&#x2019;s V value (<italic>V</italic>) (<xref ref-type="bibr" rid="B6">Cramer, 1946</xref>).</p>
<p>In this first set of randomly selected locations, data from the PLACE/AROS evaluation suggest that there is no statistically significant difference between score distributions (<inline-formula id="inf71">
<mml:math id="m81">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>9.34</mml:mn>
</mml:math>
</inline-formula>, <italic>p</italic> &#x3d; 0.053). A larger sample size may be necessary to observe statistical significance; however, this will be of negligible size effect. Nevertheless, data from the POSA/AROS evaluation study showed that our approach performs better than POSA (<inline-formula id="inf72">
<mml:math id="m82">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>32.33</mml:mn>
</mml:math>
</inline-formula>, <italic>p</italic> &#x3c; 0.001) with a medium level of association (<italic>V</italic> &#x3d; 0.1823).</p>
</sec>
<sec id="s4-2-2">
<title>4.2.2 Challenging test locations</title>
<p>A random sampling strategy is insufficient to fully evaluate the performance of pose affordances, since what matters for such methods is how they perform under realistic albeit challenging specific scene locations. For example, a test can be oversimplified and inadequate for evaluations if the sampled scene has relatively large empty spaces where only the floor or a big plane surface surrounds the test locations. Therefore, we crowdsource the evaluations in a new set of more realistic locations provided by a golden annotator (none of the authors) tasked with identifying areas of interest for human interactions (<xref ref-type="fig" rid="F5">Figure&#xa0;5</xref>). These locations are available for comparison as part of our dataset (<ext-link ext-link-type="uri" xlink:href="https://abelpaor.github.io/AROS/">https://abelpaor.github.io/AROS/</ext-link>).</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Selected by a golden annotator, green spots correspond to examples of meaningful, challenging locations for affordance detection.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g005.tif"/>
</fig>
<p>The results of the side-by-side comparison studies confirm that in 60.6% of the comparisons with PLACE, AROS was considered more natural overall. Compared to POSA, AROS was marked with better performance in 76.1% of all evaluations with a notorious difference in MP3D locations, where AROS was evaluated to be more natural in 80.2% of the assessments. The results per dataset are shown in <xref ref-type="table" rid="T4">Table&#xa0;4</xref> (% preferences in challenging locations).</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>MTurk side-by-side studies results in random and challenging locations. The best are in boldface.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="left"/>
<th colspan="3" align="center">% preferences in random locations</th>
<th colspan="3" align="center">% preferences in challenging locations</th>
</tr>
<tr>
<th align="left">Side-by-side comparison study</th>
<th align="left">Model</th>
<th align="center">MP3D</th>
<th align="center">PROX</th>
<th align="center">Replica</th>
<th align="center">MP3D</th>
<th align="center">PROX</th>
<th align="center">Replica</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="2" align="center">PLACE <italic>vs</italic>. AROS</td>
<td align="left">PLACE</td>
<td align="center">39.5</td>
<td align="center">32.7</td>
<td align="center">45.7</td>
<td align="center">38.9</td>
<td align="center">30.9</td>
<td align="center">36.4</td>
</tr>
<tr>
<td align="left">AROS</td>
<td align="center">
<bold>60.5</bold>
</td>
<td align="center">
<bold>67.3</bold>
</td>
<td align="center">
<bold>54.3</bold>
</td>
<td align="center">
<bold>61.1</bold>
</td>
<td align="center">
<bold>69.1</bold>
</td>
<td align="center">
<bold>63.6</bold>
</td>
</tr>
<tr>
<td rowspan="2" align="center">POSA <italic>vs</italic>. AROS</td>
<td align="left">POSA</td>
<td align="center">24.7</td>
<td align="center">29.6</td>
<td align="center">27.8</td>
<td align="center">19.8</td>
<td align="center">21.6</td>
<td align="center">30.2</td>
</tr>
<tr>
<td align="left">AROS</td>
<td align="center">
<bold>75.3</bold>
</td>
<td align="center">
<bold>70.4</bold>
</td>
<td align="center">
<bold>72.2</bold>
</td>
<td align="center">
<bold>80.2</bold>
</td>
<td align="center">
<bold>78.4</bold>
</td>
<td align="center">
<bold>69.8</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As in the randomly selected test locations, a descriptive analysis of the data from individual evaluation studies on these new locations suggests that AROS performs better than other approaches with larger mean values and narrower standard deviations. The mean and standard deviation of the scores obtained by the judges to PLACE are 2.97 &#xb1; 1.33&#xa0;in comparison with AROS, 3.44 &#xb1; 1.19, while in the second study, these statistics obtained by POSA were 2.79 &#xb1; 1.25&#xa0;in contrast with AROS, 3.5 &#xb1; 1.25. However, a Shapiro&#x2013;Wilk test performed on these data shows that the score distributions also depart from normality with <italic>p</italic> &#x3c; 0.01 in both studies, PLACE/AROS and POSA/AROS.</p>
<p>A chi-square test of homogeneity, with <italic>&#x3b1;</italic> &#x3d; 0.05, was used to determine whether both score distributions were statistically similar on the data from the PLACE/AROS evaluation study, providing evidence that there is a difference in score distributions (<inline-formula id="inf73">
<mml:math id="m83">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>35.92</mml:mn>
</mml:math>
</inline-formula>, <italic>p</italic> &#x3c; 0.001) with a medium level of association (<italic>V</italic> &#x3d; 0.192).</p>
<p>However, an omnibus <italic>&#x3c7;</italic>
<sup>2</sup> statistic does not provide information about the source of the difference between the score distributions. To this end, we performed a <italic>post hoc</italic> analysis following the standardized residuals method described in the work of <xref ref-type="bibr" rid="B1">Agresti (2018)</xref>. As suggested by <xref ref-type="bibr" rid="B2">Beasley and Schumacker (1995</xref>), we corrected our significance level (<italic>&#x3b1;</italic> &#x3d; 0.05) with the Sidak method (<xref ref-type="bibr" rid="B29">&#x160;id&#xe1;k, 1967</xref>) to its adjusted version <italic>&#x3b1;</italic>
<sub>
<italic>adj</italic>
</sub> &#x3d; 0.005, with critical value <italic>z</italic> &#x3d; 2.81. The study revealed a significant difference in the qualification of the interactions generated by PLACE and AROS, with ours being qualified as natural more frequently.</p>
<p>The residuals associated with AROS indicate, with significant difference, that the interactions generated by our approach were marked as &#x201c;not natural&#x201d; less frequently than expected: <italic>strongly disagree</italic> (<italic>z</italic> &#x3d; &#x2212;4.4, <italic>p</italic> &#x3c; 0.001) and <italic>disagree</italic> (<italic>z</italic> &#x3d; &#x2212;2.98, <italic>p</italic> &#x3d; 0.002). Data also show a significant difference in favorable evaluations, where PLACE has less frequently positive evaluations than predicted by the hypothesis of independence in <italic>agree</italic> (<italic>z</italic> &#x3d; &#x2212;3.04, <italic>p</italic> &#x3c; 0.001). We also observed a marginal significance, still in favor of AROS, in the frequency of <italic>strongly agree</italic> evaluations (<italic>z</italic> &#x3d; &#x2212;2.3, <italic>p</italic> &#x3d; 0.015).</p>
<p>Not surprisingly, the chi-square test of homogeneity (<italic>&#x3b1;</italic> &#x3d; 0.05) on the data from the POSA/AROS evaluation study revealed that there is strong evidence of a difference in score distributions (<inline-formula id="inf74">
<mml:math id="m84">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>75.13</mml:mn>
</mml:math>
</inline-formula>, <italic>p</italic> &#x3c; 0.001) with a larger level of association (<italic>V</italic> &#x3d; 0.278). The <italic>post hoc</italic> analysis with standardized residuals concludes that the naturalness of human&#x2013;scene interactions generated by AROS is, in the long term, better than that from POSA. <xref ref-type="table" rid="T5">Table&#xa0;5</xref> shows the cross-tabulated data of the scores observed by MTurkers and their standardized residual (critical value <italic>z</italic> &#x3d; 2.81 for <italic>&#x3b1;</italic>
<sub>
<italic>adj</italic>
</sub> &#x3d; 0.005).</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Cross-tabulation data of individual evaluation studies on challenging locations. A chi-square test of homogeneity on data provides evidence of difference in the distribution of scores with <italic>&#x3b1;</italic> &#x3d; 0.05. An analysis of residual indicates the source of such differences, an asterisk (&#x2a;) indicates conservative statistical significance at <italic>&#x3b1;</italic> &#x3d; 0.05, and a double asterisk (&#x2a;&#x2a;) denotes statistical significance with <italic>&#x3b1;</italic>
<sub>
<italic>adj</italic>
</sub> &#x3d; 0.005. The best are in boldface.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Individual <break/>evaluation study</th>
<th align="left">Model</th>
<th align="left"/>
<th align="center">1. Stronglydisagree</th>
<th align="center">2. Disagree</th>
<th align="center">3. Neither</th>
<th align="center">4. Agree</th>
<th align="center">5. Strongly agree</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="6" align="left">PLACE <italic>vs</italic>. AROS</td>
<td rowspan="3" align="left">PLACE</td>
<td align="left">Observed frequency</td>
<td align="center">81</td>
<td align="center">131</td>
<td align="center">54</td>
<td align="center">161</td>
<td align="center">59</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">16.7%</td>
<td align="center">27.0%</td>
<td align="center">11.1%</td>
<td align="center">33.1%</td>
<td align="center">12.1%</td>
</tr>
<tr>
<td align="left">Standardized residual</td>
<td align="center">4.44&#x2a;&#x2a;</td>
<td align="center">2.98&#x2a;&#x2a;</td>
<td align="center">&#x2212;1.08</td>
<td align="center">&#x2212;3.04&#x2a;&#x2a;</td>
<td align="center">&#x2212;2.43&#x2a;</td>
</tr>
<tr>
<td rowspan="3" align="left">AROS</td>
<td align="left">Observed frequency</td>
<td align="center">
<bold>36</bold>
</td>
<td align="center">
<bold>92</bold>
</td>
<td align="center">65</td>
<td align="center">207</td>
<td align="center">86</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">
<bold>7.4%</bold>
</td>
<td align="center">
<bold>18.9%</bold>
</td>
<td align="center">13.4%</td>
<td align="center">
<bold>42.6%</bold>
</td>
<td align="center">
<bold>17.7%</bold>
</td>
</tr>
<tr>
<td align="left">Standardized residual</td>
<td align="center">
<bold>&#x2212;4.44&#x2a;&#x2a;</bold>
</td>
<td align="center">
<bold>&#x2212;2.98&#x2a;&#x2a;</bold>
</td>
<td align="center">1.08</td>
<td align="center">
<bold>3.04&#x2a;&#x2a;</bold>
</td>
<td align="center">
<bold>2.43&#x2a;</bold>
</td>
</tr>
<tr>
<td rowspan="6" align="left">POSA <italic>vs</italic>. AROS</td>
<td rowspan="3" align="left">POSA</td>
<td align="left">Observed frequency</td>
<td align="center">86</td>
<td align="center">141</td>
<td align="center">93</td>
<td align="center">122</td>
<td align="center">44</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">17.7%</td>
<td align="center">29.0%</td>
<td align="center">19.1%</td>
<td align="center">25.1%</td>
<td align="center">9.1%</td>
</tr>
<tr>
<td align="left">Standardized residual</td>
<td align="center">4.95&#x2a;&#x2a;</td>
<td align="center">3.52&#x2a;&#x2a;</td>
<td align="center">1.70</td>
<td align="center">&#x2212;2.88</td>
<td align="center">&#x2212;6.57</td>
</tr>
<tr>
<td rowspan="3" align="left">AROS</td>
<td align="left">Observed frequency</td>
<td align="center">
<bold>35</bold>
</td>
<td align="center">
<bold>94</bold>
</td>
<td align="center">73</td>
<td align="center">
<bold>163</bold>
</td>
<td align="center">
<bold>121</bold>
</td>
</tr>
<tr>
<td align="left">% within model</td>
<td align="center">
<bold>7.2%</bold>
</td>
<td align="center">
<bold>19.3%</bold>
</td>
<td align="center">15.0%</td>
<td align="center">
<bold>33.5%</bold>
</td>
<td align="center">
<bold>24.9%</bold>
</td>
</tr>
<tr>
<td align="left">Standardized residual</td>
<td align="center">
<bold>&#x2212;4.95&#x2a;&#x2a;</bold>
</td>
<td align="center">
<bold>&#x2212;3.52&#x2a;&#x2a;</bold>
</td>
<td align="center">&#x2212;1.70</td>
<td align="center">
<bold>2.88&#x2a;&#x2a;</bold>
</td>
<td align="center">
<bold>6.57&#x2a;&#x2a;</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4-3">
<title>4.3 Qualitative results</title>
<p>Experiments verify that our approaches can realistically generate human bodies that interact within a given environment in a natural and physically plausible manner. AROS allows us to not only determine the location on the environment in which we want the interaction to happen (the where) but also select the specific type of interaction to be performed (the what).</p>
<p>The number and variety of interactions detected by AROS can easily be increased as a result of its one-shot training capacity. The more trained the interactions, the more the human&#x2013;scene configuration can detect/generate. <xref ref-type="fig" rid="F6">Figure&#xa0;6</xref> shows examples of different affordance detections around single locations.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>AROS shows good performance on a variety of novel scenes.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g006.tif"/>
</fig>
<p>AROS showed better performance in more realistic environment configurations where elements, such as chairs, sofas, tables, and walls, are presented and must be considered during the generation of body interactions. <xref ref-type="fig" rid="F7">Figure&#xa0;7</xref> shows some examples of interaction generated by AROS and baselines over challenging locations.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Qualitative challenging locations. PLACE (yellow), POSA (pink), and AROS (silver).</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g007.tif"/>
</fig>
<p>Alternatively, AROS can be used to concatenate affordances over several positions to generate useful affordance maps for action planners (see <xref ref-type="fig" rid="F8">Figure&#xa0;8</xref>). This can be used as a way to generate visualizations of action scripts or to plan the ergonomics and usability of spaces beyond individual objects.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>AROS can be used to create maps for action planning. Top: Many locations in an environment are evaluated for three different affordances (sit-able, walk-able, and reach-able). Bottom: AROS scores used to plan concatenated action milestones.</p>
</caption>
<graphic xlink:href="frobt-10-1076780-g008.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>5 Conclusion</title>
<p>In this work, we present AROS, a one-shot geometric-driven affordance descriptor that is built on the bisector surface and combines proximity zones and clearance space to improve the affordance characterization of human poses. We introduced a generative framework that poses 3D human bodies interacting within a 3D environment in a natural and physically plausible manner. AROS shows a good generalization in unseen novel scenes. Furthermore, adding a new interaction to AROS is straightforward, since it requires only one example. Via rigorous statistical analysis, results show that our one-shot approach outperforms data-intensive baselines, with human judges preferring AROS proposals 80% of the time over the baselines. AROS can be used to concatenate affordances over several positions. This can be used as a way to generate visualizations of action scripts in 3D scenes or to plan the ergonomics and usability of spaces beyond individual object affordances. We believe that explicit and interpretable description is valuable for complementing data-driven methods and opens avenues for further work, including combining the strengths of both approaches.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found at: <ext-link ext-link-type="uri" xlink:href="https://abelpaor.github.io/AROS/">https://abelpaor.github.io/AROS/</ext-link>.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>APO performed all experiments, data preparation, and coding. All authors contributed to the conception and design of the study. All authors contributed to writing, revising, reading, and reviewing the manuscript submitted.</p>
</sec>
<ack>
<p>APO thanks the Mexican Council for Science and Technology (CONACYT) for the scholarship provided for his postgraduate studies with the scholarship number 709908. WMC thanks the visual egocentric research activity partially funded by UK EPSRC EP/N013964/1. The authors thank Eduardo Ruiz-Libreros for sharing his efforts on the description of affordances. They also thank Angeliki Katsenou and Pilar Padilla Mendoza for their advice on the performed statistical analysis.</p>
</ack>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of interest</title>
<p>WMC was employed by Amazon.com.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10">
<title>Supplementary material</title>
<p>Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frobt.2023.1076780/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frobt.2023.1076780/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Video1.mp4" id="SM2" mimetype="application/mp4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Agresti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <source>An introduction to categorical data analysis</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x26; Sons</publisher-name>, <fpage>39</fpage>&#x2013;<lpage>41</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beasley</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Schumacker</surname>
<given-names>R. E.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Multiple regression approach to analyzing contingency tables: Post hoc and planned comparison procedures</article-title>. <source>J. Exp. Educ.</source>
<volume>64</volume>, <fpage>79</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1080/00220973.1995.9943797</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bochkovskiy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.-Y.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>H.-Y. M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>YOLOv4: Optimal speed and accuracy of object detection</article-title>. <comment>
<italic>arXiv</italic>
</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2004.10934">http://arxiv.org/abs/2004.10934</ext-link>
</comment>.</citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Carion</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Massa</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Synnaeve</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Usunier</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kirillov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zagoruyko</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>End-to-End object detection with transformers</article-title>,&#x201d; in <source>Computer Vision &#x2013; ECCV 2020</source> (<publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>213</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-58452-8</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chang</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Funkhouser</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Halber</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Niessner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Savva</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). &#x201c;<article-title>Matterport3D: Learning from RGB-D data in indoor environments</article-title>,&#x201d; in <source>International conference on 3D vision (3DV)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>667</fpage>&#x2013;<lpage>676</lpage>. <pub-id pub-id-type="doi">10.1109/3DV.2017.00081</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cramer</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1946</year>). &#x201c;<article-title>The two-dimensional case</article-title>,&#x201d; in <source>Mathematical methods of statistics</source> (<publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>Princeton university press</publisher-name>), <fpage>260</fpage>&#x2013;<lpage>290</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>T.-Y.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ghiasi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). &#x201c;<article-title>SpineNet: Learning scale-permuted backbone for recognition and localization</article-title>,&#x201d; in <source>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR)</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>11589</fpage>&#x2013;<lpage>11598</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01161</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fouhey</surname>
<given-names>D. F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>In Defense of the direct perception of affordances</article-title>. <comment>
<italic>arXiv preprint arXiv:1505.</italic>01085</comment>. <pub-id pub-id-type="doi">10.1002/eji.201445290</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Franke</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Christie</surname>
<given-names>C. A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The chi-square test: Often used and more often misinterpreted</article-title>. <source>Am. J. Eval.</source>
<volume>33</volume>, <fpage>448</fpage>&#x2013;<lpage>458</lpage>. <pub-id pub-id-type="doi">10.1177/1098214011426594</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gibson</surname>
<given-names>J. J.</given-names>
</name>
</person-group> (<year>1977</year>). &#x201c;<article-title>The theory of affordances</article-title>,&#x201d; in <source>Perceiving, acting and knowing. Toward and ecological psychology</source> (<publisher-loc>Mahwah, NJ</publisher-loc>: <publisher-name>Lawrence Eribaum Associates</publisher-name>).</citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Grabner</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Gall</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Van Gool</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>What makes a chair a chair?</article-title>,&#x201d; in <source>2011 IEEE conference on computer vision and pattern recognition (CVPR)</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1529</fpage>&#x2013;<lpage>1536</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2011.5995327</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Satkin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Efros</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Hebert</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>From 3D scene geometry to human workspace</article-title>,&#x201d; in <source>CVPR 2011</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1961</fpage>&#x2013;<lpage>1968</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2011.5995448</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hassan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Choutas</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Tzionas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Black</surname>
<given-names>M. J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Resolving 3D human pose ambiguities with 3D scene constraints</article-title>,&#x201d; in <source>Proceedings of the IEEE/CVF international conference on computer vision</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2282</fpage>&#x2013;<lpage>2292</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00237</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hassan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ghosh</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tesch</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tzionas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Black</surname>
<given-names>M. J.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Populating 3D scenes by learning human-scene interaction</article-title>,&#x201d; in <source>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>14708</fpage>&#x2013;<lpage>14718</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01447</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>van Kaick</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shamir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Learning how objects function via co-analysis of interactions</article-title>. <source>ACM Trans. Graph.</source>
<volume>35</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1145/2897824.2925870</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>van Kaick</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Shamir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Interaction context (ICON): Towards a geometric functionality descriptor</article-title>. <source>ACM Trans. Graph.</source>
<volume>34</volume>, <fpage>1</fpage>&#x2013;<lpage>83:12</lpage>. <pub-id pub-id-type="doi">10.1145/2766914</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Koppula</surname>
<given-names>H. S.</given-names>
</name>
<name>
<surname>Saxena</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Modeling 3d environments through hidden human context</article-title>. <source>IEEE Trans. Pattern Analysis Mach. Intell.</source>
<volume>38</volume>, <fpage>2040</fpage>&#x2013;<lpage>2053</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2501811</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M.-H.</given-names>
</name>
<name>
<surname>Kautz</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Putting humans in a scene: Learning affordance in 3d indoor environments</article-title>,&#x201d; in <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>12368</fpage>&#x2013;<lpage>12376</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01265</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Luddecke</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Worgotter</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Learning to segment affordances</article-title>,&#x201d; in <source>The IEEE international conference on computer vision (ICCV) workshops</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>769</fpage>&#x2013;<lpage>776</lpage>. <pub-id pub-id-type="doi">10.1109/ICCVW.2017.96</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Nekrasov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Schult</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Litany</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Leibe</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Engelmann</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Mix3D: Out-of-Context data augmentation for 3D scenes</article-title>,&#x201d; in <source>2021 international conference on 3D vision (3DV)</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>116</fpage>&#x2013;<lpage>125</lpage>. <pub-id pub-id-type="doi">10.1109/3DV53792.2021.00022</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pavlakos</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Choutas</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ghorbani</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bolkart</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Osman</surname>
<given-names>A. A. A.</given-names>
</name>
<name>
<surname>Tzionas</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Expressive body capture: 3d hands, face, and body from a single image</article-title>,&#x201d; in <source>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>10967</fpage>&#x2013;<lpage>10977</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01123</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peternell</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Geometric properties of bisector surfaces</article-title>. <source>Graph. Models</source> <volume>62</volume>, <fpage>202</fpage>&#x2013;<lpage>236</lpage>. <pub-id pub-id-type="doi">10.1006/gmod.1999.0521</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Piyathilaka</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kodagoda</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Affordance-map: Mapping human context in 3D scenes using cost-sensitive SVM and virtual human models</article-title>,&#x201d; in <source>2015 IEEE international conference on Robotics and biomimetics (ROBIO)</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2035</fpage>&#x2013;<lpage>2040</lpage>. <pub-id pub-id-type="doi">10.1109/ROBIO.2015.7419073</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rhinehart</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kitani</surname>
<given-names>K. M.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Learning action maps of large environments via first-person vision</article-title>,&#x201d; in <source>2016 IEEE conference on computer vision and pattern recognition (CVPR)</source> (<publisher-loc>New York, NJ</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>580</fpage>. <comment>&#x2013;588</comment>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.69</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Roy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Todorovic</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>A multi-scale CNN for affordance segmentation in RGB images</article-title>,&#x201d; in <source>European conference on computer vision</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>186</fpage>&#x2013;<lpage>201</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ruiz</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mayol-Cuevas</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Geometric affordance perception: Leveraging deep 3D saliency with the interaction tensor</article-title>. <source>Front. Neurorobotics</source>
<volume>14</volume>, <fpage>45</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2020.00045</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Savva</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>A. X.</given-names>
</name>
<name>
<surname>Hanrahan</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fisher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nie&#xdf;ner</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>SceneGrok: Inferring action maps in 3D environments</article-title>. <source>ACM Trans. Graph. (TOG)</source>
<volume>33</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1145/2661229.2661230</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shapiro</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>Wilk</surname>
<given-names>M. B.</given-names>
</name>
</person-group> (<year>1965</year>). <article-title>An analysis of variance test for normality (complete samples)</article-title>. <source>Biometrika</source>
<volume>52</volume>, <fpage>591</fpage>&#x2013;<lpage>611</lpage>. <pub-id pub-id-type="doi">10.2307/2333709</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>&#x160;id&#xe1;k</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1967</year>). <article-title>Rectangular confidence regions for the means of multivariate normal distributions</article-title>. <source>J. Am. Stat. Assoc.</source>
<volume>62</volume>, <fpage>626</fpage>&#x2013;<lpage>633</lpage>. <pub-id pub-id-type="doi">10.2307/2283989</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Silberman</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hoiem</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kohli</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fergus</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Indoor segmentation and support inference from RGBD images</article-title>,&#x201d; in <source>European conference on computer vision</source> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>746</fpage>&#x2013;<lpage>760</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-33715-4_54</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Straub</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Whelan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wijmans</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>The Replica dataset: A digital Replica of indoor spaces</article-title>. <comment>
<italic>arXiv preprint arXiv:1906.05797</italic>
</comment>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Girdhar</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Binge watching: Scaling affordance learning from sitcoms</article-title>,&#x201d; in <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2596</fpage>&#x2013;<lpage>2605</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.359</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Misra</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chirikjian</surname>
<given-names>G. S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Is that a chair? Imagining affordances using simulations of an articulated human body</article-title>,&#x201d; in <source>2020 IEEE international conference on Robotics and automation (ICRA)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>7240</fpage>&#x2013;<lpage>7246</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA40945.2020.9197384</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuksel</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Sample elimination for generating Poisson disk sample sets</article-title>. <source>Comput. Graph. Forum</source>
<volume>34</volume>, <fpage>25</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1111/cgf.12538</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2020a</year>). <article-title>ResNeSt: Split-Attention networks</article-title>. <comment>
<italic>arXiv</italic>
</comment>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Black</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020b</year>). &#x201c;<article-title>Place: Proximity learning of articulation and contact in 3D environments</article-title>,&#x201d; in <source>8th international conference on 3D Vision (3DV 2020)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>642</fpage>&#x2013;<lpage>651</lpage>. <pub-id pub-id-type="doi">10.1109/3DV50981.2020.00074</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hassan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Neumann</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Black</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020c</year>). &#x201c;<article-title>Generating 3D people in scenes without people</article-title>,&#x201d; in <source>The IEEE/CVF conference on computer vision and pattern recognition (CVPR)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6193</fpage>&#x2013;<lpage>6203</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00623</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>M. G.</given-names>
</name>
<name>
<surname>Komura</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Character-object interaction retrieval using the interaction bisector surface</article-title>. <source>Eurogr. Symposium Geometry Process.</source> <volume>36</volume>, <fpage>119</fpage>&#x2013;<lpage>129</lpage>. <pub-id pub-id-type="doi">10.1111/cgf.13112</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Guerrero</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Mitra</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Komura</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Relationship templates for creating scene variations</article-title>. <source>ACM Trans. Graph.</source>
<volume>35</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1145/2980179.2982410</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Komura</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Indexing 3D scenes using the interaction bisector surface</article-title>. <source>ACM Trans. Graph.</source>
<volume>33</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/2574860</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Puig</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Fidler</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Barriuso</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Torralba</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Scene parsing through ADE20K dataset</article-title>,&#x201d; in <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5122</fpage>&#x2013;<lpage>5130</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.544</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Terzopoulos</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.-C.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Inferring forces and learning human utilities from videos</article-title>,&#x201d; in <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3823</fpage>&#x2013;<lpage>3833</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.415</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>