<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2022.895126</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Building Human Visual Attention Map for Construction Equipment Teleoperation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Fan</surname> <given-names>Jiamin</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Xiaomeng</given-names></name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Su</surname> <given-names>Xing</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1058429/overview"/>
</contrib>
</contrib-group>
<aff><institution>College of Civil Engineering and Architecture, Zhejiang University</institution>, <addr-line>Hangzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Hanliang Fu, Xi&#x2019;an University of Architecture and Technology, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Zeyu Wang, Guangzhou University, China; Yilong Han, Tongji University, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Xing Su, <email>xsu@zju.edu.cn</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Decision Neuroscience, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>895126</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>03</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Fan, Li and Su.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Fan, Li and Su</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Construction equipment teleoperation is a promising solution when the site environment is hazardous to operators. However, limited situational awareness of the operator exists as one of the major bottlenecks for its implementation. Virtual annotations (VAs) can use symbols to convey information about operating clues, thus improving an operator&#x2019;s situational awareness without introducing an overwhelming cognitive load. It is of primary importance to understand how an operator&#x2019;s visual system responds to different VAs from a human-centered perspective. This study investigates the effect of VA on teleoperation performance in excavating tasks. A visual attention map is generated to describe how an operator&#x2019;s attention is allocated when VAs are presented during operation. The result of this study can improve the understanding of how human vision works in virtual or augmented reality. It also informs the strategies on the practical implication of designing a user-friendly teleoperation system.</p>
</abstract>
<kwd-group>
<kwd>construction equipment teleoperation</kwd>
<kwd>virtual annotation</kwd>
<kwd>situational awareness</kwd>
<kwd>visual attention</kwd>
<kwd>cognitive load</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="12"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="46"/>
<page-count count="13"/>
<word-count count="6190"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Construction equipment operators are inevitably exposed to danger when operating in an extreme environment (<xref ref-type="bibr" rid="B21">Kim et al., 2017</xref>). Teleoperation of construction equipment can effectively assist an operator in completing a task while avoiding dangerous situations (<xref ref-type="bibr" rid="B40">Wang and Dunston, 2006</xref>). Equipment teleoperation has been applied in many domains, such as space exploration, military defense, underwater operation, telerobotics in forestry and mining, telesurgery, and telepresence robots (<xref ref-type="bibr" rid="B25">Lichiardopol, 2007</xref>). For example, <xref ref-type="bibr" rid="B43">Woo-Keun et al. (2004)</xref> combined force with motion command into a fixed space robotic teleoperation system. <xref ref-type="bibr" rid="B22">Kot and Nov&#x00E1;k (2018)</xref> employed virtual reality and the HMD Oculus Rift in Tactical Robotic System. The examples illustrate the potential of teleoperation in construction to reduce operational risks and extend the ranges of construction activities.</p>
<p>Construction equipment teleoperation is still an open research area and is rarely applied in practical activities. Limited situational awareness encountered in a teleoperating environment is one of the main causes that hinder the application (<xref ref-type="bibr" rid="B18">Hong et al., 2020</xref>). Situational awareness is defined as &#x201C;the perception of the elements in the environment within a volume of time and space, the comprehension of their meaning, and the projection of their status in the near future&#x201D; (<xref ref-type="bibr" rid="B12">Endsley, 1988</xref>). During teleoperation, an operator has no direct perception of the environment but has to rely on visual information on one or multiple teleoperating screens. The operator&#x2019;s perceptual processing is decoupled from the physical environment, resulting in a low situational awareness that may lead to collisions and other accidents (<xref ref-type="bibr" rid="B42">Woods et al., 2004</xref>).</p>
<p>Existing studies have explored a variety of means to improve perceptual awareness, among which, the application of virtual annotation (VA) has demonstrated significant potential. A VA can present critical information from sensors as a visual cue to assist in teleoperation, compensating the operator&#x2019;s limited situational awareness. Research and practical examples have been reported in some tourist and navigation systems (<xref ref-type="bibr" rid="B29">Orlosky et al., 2014</xref>; <xref ref-type="bibr" rid="B41">Williams et al., 2017</xref>), surgery training systems (<xref ref-type="bibr" rid="B1">Andersen et al., 2016</xref>), and augmented reality (AR)-based entertainment applications (<xref ref-type="bibr" rid="B23">Larabi, 2018</xref>; <xref ref-type="bibr" rid="B37">Tylecek and Fisher, 2018</xref>).</p>
<p>However, the existing VA system may not be directly applied to construction equipment teleoperation. Challenges remain due to some unique features of operating a piece of construction equipment. An important one is related to human attention allocation. When using the VA-based tourist system, a user can place as much attention as necessary on visualizing and understanding the VA. In contrast, a construction equipment operator must place enough attention on the operating task under a usually stressful situation. A VA can be ignored if it fails to draw the operator&#x2019;s attention or can be very interruptive on the other hand. In addition, unlike the surgery system, construction equipment operation often involves a frequent change in locations and scenes, which may require more attention from an operator.</p>
<p>Many VA-related studies and applications in the construction field focus on function-oriented technologies and rarely contemplate the problem from a human-oriented perspective (<xref ref-type="bibr" rid="B17">Hong et al., 2021</xref>). Understanding how an operator&#x2019;s visual system responds to different VAs during construction equipment teleoperation remains a challenge. It has been found that many design features of a VA, such as shape, format, size, and appearing location, may affect the driver&#x2019;s understanding and therefore affect the effectiveness of VA use. Subjects wearing a head-mounted display suggested that text annotations be placed below the center of the screen (<xref ref-type="bibr" rid="B29">Orlosky et al., 2014</xref>). Highlighted lines around the edges of obstacles are easier to understand and react to than radar maps (<xref ref-type="bibr" rid="B18">Hong et al., 2020</xref>).</p>
<p>This work aims at building a visual attention map for the construction equipment teleoperation to depict how an operator allocates her/his visual attention during operation with VAs. The visual attention map can contribute to a scientific basis for understanding an operator&#x2019;s visual attention allocating mechanism under a stressful work situation. It also informs design strategies for practitioners to improve the user interface of next-generation teleoperating equipment.</p>
</sec>
<sec id="S2">
<title>Related Work</title>
<sec id="S2.SS1">
<title>Virtual Annotation Design</title>
<p>Virtual annotation system has been applied in many fields, such as aircraft operating, navigation, and surgery. A VA can supplement otherwise-inaccessible information to improve an operator&#x2019;s situational awareness in the teleoperation context. For instance, during the simulation of aircraft operation, the GPS usually uses text and graphics to annotate traffic conditions and route information (<xref ref-type="bibr" rid="B7">Christoph et al., 2007</xref>). The intuitive graphic annotations in the Surgical Wound Closure Training System show the exact grip point of the scalpel and the route of the scalpel cut, giving the trainee effective guidance on surgical operations and procedures (<xref ref-type="bibr" rid="B1">Andersen et al., 2016</xref>). The navigation system designed by <xref ref-type="bibr" rid="B4">Bolton et al. (2015)</xref> adopted anchored annotations to highlight landmarks and improved response times and success rates by 43.1 and 26.2%, respectively.</p>
<p>A well-designed VA can facilitate an operator&#x2019;s spatial understanding while requiring a manageable level of cognitive load. Meanwhile, it has been reported that the processing of VA during operation may distract operators and affect the operating performance. Text annotations on head-mounted displays can distract subjects and interfere with the reading task, potentially reducing the performance (<xref ref-type="bibr" rid="B29">Orlosky et al., 2014</xref>). Several subjects in the excavator teleoperation experiment reported that the virtual annotations were distracting during the operation and harmed performance (<xref ref-type="bibr" rid="B18">Hong et al., 2020</xref>).</p>
<p>Existing studies have identified several critical design features, such as format, size, and position, which may play a critical role in a user&#x2019;s mental process of understanding VAs. The representative formats of VA include image (<xref ref-type="bibr" rid="B35">Shapira et al., 2008</xref>; <xref ref-type="bibr" rid="B15">Fritsche et al., 2017</xref>), sign (<xref ref-type="bibr" rid="B45">Ziaei et al., 2011</xref>), and text or video with various properties (<xref ref-type="bibr" rid="B19">Hori and Shimizu, 1999</xref>). Both single and multi-formats are studied in the existing works. For instance, a single textual format is used to obtain hypertext information to create virtual reality concept maps (<xref ref-type="bibr" rid="B38">Verlinden et al., 1993</xref>). <xref ref-type="bibr" rid="B44">Yeh et al. (2013)</xref> used multi-formats, including color, text, and digits, to explore the effects of collaborative tasks. <xref ref-type="bibr" rid="B30">Pennington (2001)</xref> designed the cross-shaped VA and the ring-shaped VA to imply stopping the movement and whistling to warn the workers.</p>
<p>The main criterion for determining the size of VA is that they should be able to remind people to the greatest extent possible without interfering with the rest of the display (<xref ref-type="bibr" rid="B19">Hori and Shimizu, 1999</xref>). Some works fixed the size of VA, such as images of 640 &#x00D7; 480 pixels (<xref ref-type="bibr" rid="B16">Grasset et al., 2012</xref>), whereas some experiments adopted VAs with flexible sizes. Results have shown that larger VAs are more likely to be detected and responded to by subjects (<xref ref-type="bibr" rid="B29">Orlosky et al., 2014</xref>).</p>
<p>With different VA appearing or anchoring positions, users have experienced different distractions, affecting task performance. In the experiment by <xref ref-type="bibr" rid="B10">Driewer et al. (2005)</xref>, the anchoring position of the VA changed according to the screen, and the central position received the most attention from the subjects. The highlighting of the edges of an obstacle in the positive field of view is more visible to the operator than the radar map in the upper right corner (<xref ref-type="bibr" rid="B18">Hong et al., 2020</xref>). In an experiment where participants wore head-mounted displays to read newspapers while walking, participants often placed text annotation below the center of the screen, avoiding the top left and right corners (<xref ref-type="bibr" rid="B29">Orlosky et al., 2014</xref>).</p>
<p>Other aspects, such as color and contrast, are also the important factors when designing a VA. The association of traffic signal colors (red, yellow, and green) with meanings such as prohibitions or stops at intersections is globally recognized (<xref ref-type="bibr" rid="B30">Pennington, 2001</xref>), just as detected obstacles and danger zones turn red on maps (<xref ref-type="bibr" rid="B10">Driewer et al., 2005</xref>). On the other hand, it is found that humans may only focus on the areas of relatively high visual saliency and ignore other areas and views (<xref ref-type="bibr" rid="B34">Sato et al., 2020</xref>).</p>
<p>Within the context of teleoperating construction equipment, the VA system can help with object identification and target detection in a dynamic construction site. Operators can obtain spatial information about the surrounding environment with the help of VA. However, when VAs are presented to the operator, it raises another question: how does an operator&#x2019;s visual system allocate attention to the VA and the work scene?</p>
</sec>
<sec id="S2.SS2">
<title>Human Visual Attention</title>
<p>Researchers assumed an underlying relationship between attention allocation and teleoperation performances (<xref ref-type="bibr" rid="B33">Riley et al., 2004</xref>). It has been divided into four categories: preattention, inattention, divided attention, and focused attention (<xref ref-type="bibr" rid="B27">Matthews et al., 2003</xref>), and the different attention levels will lead to different information acceptance (<xref ref-type="bibr" rid="B20">Kahneman, 1973</xref>). At the preattention stage, people handle objects that are not inherently available for later processing and thus do not affect awareness. Inattention makes a person not conscious of a perceptual stimulus, but the information may affect behavior (<xref ref-type="bibr" rid="B13">Fernandez-Duque and Thornton, 2000</xref>). Divided attention distributes attention over several objects, and focused attention uses all attentional resources to focus on one stimulus (<xref ref-type="bibr" rid="B27">Matthews et al., 2003</xref>).</p>
<p>The information processing of VAs during operation is potentially related to an operator&#x2019;s visual attention allocating mechanism. In teleoperation, information is mainly obtained by the vision, and human attention determines what people concentrate on or ignore (<xref ref-type="bibr" rid="B2">Anderson, 1980</xref>). Attention may be especially critical when operators must focus on VAs to achieve an accurate assessment of the situation. Sometimes, they may be susceptible to the saliency effect. For example, salient information from one position may draw most of the operator&#x2019;s attention, and information from other locations is ignored (<xref ref-type="bibr" rid="B36">Thomas and Wickens, 2001</xref>).</p>
<p>The existing literature has proposed a bottom-up framework for visual attention study (<xref ref-type="bibr" rid="B3">Bergen and Julesz, 1983</xref>). It emphasizes exploring factors that attract attention, such as color and movement (<xref ref-type="bibr" rid="B11">El-Nasr and Yan, 2006</xref>). The related studies can be divided into two groups based on whether the research media is static abstract images or abstract videos with changing backgrounds (<xref ref-type="bibr" rid="B31">Rea et al., 2017</xref>). The static images used to be applied in natural conditions, and the videos are usually used in complex scenes with free movement (<xref ref-type="bibr" rid="B8">Chun, 2000</xref>; <xref ref-type="bibr" rid="B5">Burke et al., 2005</xref>).</p>
<p>Human visual attention requires a proper selection of measures. Researchers have adopted different metrics for evaluation, such as response rate, task accuracy with trajectory, work efficiency with time, operation time, collision number, and response time (<xref ref-type="bibr" rid="B6">Chen et al., 2007</xref>; <xref ref-type="bibr" rid="B28">Menchaca-Brandan et al., 2007</xref>; <xref ref-type="bibr" rid="B26">Long et al., 2011</xref>; <xref ref-type="bibr" rid="B46">Zornitza et al., 2014</xref>; <xref ref-type="bibr" rid="B39">Wallmyr et al., 2019</xref>). Among these studies, some have given different weights to the assessment indexes depending on their importance.</p>
<p>With computer vision techniques emerging in the past decade, some researchers have explored the human visual attention mechanism in 2D and 3D fields. Many experimental results are presented by visual attention maps or statistical charts. A visual attention map summarizes the most frequently visualized areas in an image by a group of subjects (<xref ref-type="bibr" rid="B9">Corredor et al., 2017</xref>). For example, <xref ref-type="bibr" rid="B11">El-Nasr and Yan (2006)</xref> took 2D and 3D games as experimental tasks to obtain two-dimensional and three-dimensional attention maps and then analyzed eye-movement patterns. A dynamic and sometimes hazardous construction sites often require a teleoperator to conduct information integration of the site scene and VA signals. An operation task has already placed a certain amount of cognitive load on an operator, and how much attention can the operator afford to spare on processing VAs? Investigating the visual attention allocating mechanism and building an attention map is of vital significance in such a context.</p>
</sec>
</sec>
<sec id="S3">
<title>Experiment</title>
<p>A virtual teleoperation platform was developed to carry out the experiment designed for this study. It allows the user to perform an excavating task repeatedly. Different VAs may appear during the experiment, and the user must conduct a certain action according to the appeared VA. The operating data were recorded throughout the whole time.</p>
<sec id="S3.SS1">
<title>Virtual Annotation Design</title>
<p>The design of VA in this study follows several principles. First, a VA shall convey straightforward information that any operator, at first sight, can understand. A total of two shapes, ring and cross, are tested in this experiment (refer to <xref ref-type="fig" rid="F1">Figure 1</xref>). The ring-shaped VA requires the operator to push the honk button while excavating, and the cross-shaped VA requires the operator to cease operation until the VA vanishes. Such a design guarantees that an operator can easily understand a VA as long as it is noticed. Accordingly, the generated map mainly presents information about allocating an operator&#x2019;s visual attention rather than a complex combination of visual attention, cognitive load, or other factors involved during &#x201C;thinking.&#x201D;</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Two types of VA: <bold>(A)</bold> ring-shaped <bold>(B)</bold> cross-shaped.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g001.tif"/>
</fig>
<p>Second, the VA should appear in the right location with proper size to be noticed with limited interference to the operator&#x2019;s view of the work scene. The VA in this study randomly exhibits different sizes of small, middle, and large (<xref ref-type="fig" rid="F2">Figure 2</xref>). The VA in the experiment will appear randomly at any location on the teleoperation screen to investigate the location&#x2019;s impact. In addition, we designed a colorless worksite with red VAs to avoid potential interference from different color contrasts on the site.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Different VA sizes: small, middle, and large.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g002.tif"/>
</fig>
</sec>
<sec id="S3.SS2">
<title>Experiment Design</title>
<p>The experiment consists of three sessions. Before the experiment started, subjects were required to fill out the pre-task questionnaire to provide information about gender, age, and previous 3D gaming experience. The first session presents all subjects with a short video introducing excavator operation and control (<xref ref-type="fig" rid="F3">Figure 3</xref>). Each subject is given 5 min to familiarize the operation.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Introduction video screenshots.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g003.tif"/>
</fig>
<p>The second session informed participants that the goal of the test is to move the balls from one trench to another as fast as possible while performing actions according to the VA that randomly appears on the screen. Then, 2 min is given for the subjects to practice operation with VAs.</p>
<p>The third session is the formal test of 10 min. <xref ref-type="fig" rid="F4">Figure 4</xref> demonstrates the interaction mechanism between the subject and the system. The system initiates the task and starts to display the cross and ring-shaped VAs in a random location with a random interval of 3&#x2013;9 s throughout the experiment. The subject operates the excavator through two joysticks. When a VA appears, the subject must respond within 6 s; otherwise, the VA will disappear, and it will be considered a failed case of VA response. The number of balls moved and correct VA responses are presented in the top left corner of the screen.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Interaction mechanism of the experiments.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g004.tif"/>
</fig>
</sec>
<sec id="S3.SS3">
<title>Experimental Platform</title>
<p>The teleoperation platform is deployed on a computer with 3.70 GHz Intel(R) Core(TM), 64G RAM, and NVIDIA GeForce RTX 2080 Ti with 11,048 MB VRAM. The excavator simulation software is developed in Unity. The UML class diagram in <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates the architecture of the software. The excavator model was downloaded from GitHub,<sup><xref ref-type="fn" rid="footnote1">1</xref></sup> including the excavators&#x2019; movement control. The experiment adopts a teleoperation view that resides in the cockpit.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>UML class diagram of the software platform.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g005.tif"/>
</fig>
<p>A pilot test with three participants was conducted before the formal experiment to ensure that the system functions properly. After the formal test, all screen video records were carefully reviewed to ensure that the collected data were accurate.</p>
</sec>
<sec id="S3.SS4">
<title>Subjects</title>
<p>The subjects were recruited from the pool of Zhejiang University students through invitations and flyers. A total of twenty subjects were recruited for the experiments, including 10 females and 10 males. The mean age of the subjects was 23.5 years. All participants have no construction equipment operation experience. The 3D gaming experience is divided into three types: &#x201C;never or rarely play,&#x201D; &#x201C;not very often but better than the first type,&#x201D; and &#x201C;regularly play and good at 3D games,&#x201D; as suggested by <xref ref-type="bibr" rid="B11">El-Nasr and Yan (2006)</xref>. Most subjects had previous 3D game experience (<xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>Previous 3D gaming experience of subjects.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g006.tif"/>
</fig>
</sec>
<sec id="S3.SS5">
<title>Human Visual Attention Assessment Indices</title>
<p>Response rate and response time are analyzed as the two major assessment indices. Response rate is the ratio of correct responses over failed responses. Response time refers to the duration between a VA appears and the subject responds to it. The response rate directly measures the subject&#x2019;s performance and the response time implies the difficulty of processing a VA. In addition, we also recorded how many balls were moved by each subject as an assessment of excavating productivity.</p>
</sec>
</sec>
<sec id="S4" sec-type="results">
<title>Results</title>
<sec id="S4.SS1">
<title>Descriptive Statistic Results</title>
<p>The descriptive statistics data of gender and 3D game experience are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. Since only two subjects regularly play 3D games, we combined the two groups of &#x201C;not very often&#x201D; and &#x201C;regularly play.&#x201D; No clear pattern was found.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Descriptive statistics of gender and 3D game experience.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g007.tif"/>
</fig>
<p>The response rate, response time, and excavation productivity of each subject were submitted to a <italic>t</italic>-test, as listed in <xref ref-type="table" rid="T1">Table 1</xref>. Gender demonstrates no significant effect in differentiating the performances of response rate, response time, and excavation productivity. Those who play more 3D games tended to respond quickly (<italic>p</italic> = 0.054), but the result was not statistically significant.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p><italic>T</italic>-test results.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="center">Indicator</td>
<td valign="top" align="center"><italic>p</italic>-value</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Gender</td>
<td valign="top" align="center">Response rate</td>
<td valign="top" align="center">0.758</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Response time</td>
<td valign="top" align="center">0.606</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Excavation productivity</td>
<td valign="top" align="center">0.205</td>
</tr>
<tr>
<td valign="top" align="left">3D Gaming experience</td>
<td valign="top" align="center">Response rate</td>
<td valign="top" align="center">0.862</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Response time</td>
<td valign="top" align="center">0.054</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Excavation productivity</td>
<td valign="top" align="center">0.555</td>
</tr>
</tbody>
</table></table-wrap>
</sec>
<sec id="S4.SS2">
<title>Response Rate</title>
<p><xref ref-type="table" rid="T2">Table 2</xref> lists the response results. It is noticed that the cross-shaped VA has a better response rate than the ring-shaped VA.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Descriptive statistics of response rates.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">VA</td>
<td valign="top" align="center">Size</td>
<td valign="top" align="center">Correct number</td>
<td valign="top" align="center">Failed number</td>
<td valign="top" align="center">Total</td>
<td valign="top" align="center">Response rate</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Cross</td>
<td valign="top" align="center">Small</td>
<td valign="top" align="center">226</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">251</td>
<td valign="top" align="center">0.900</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Middle</td>
<td valign="top" align="center">254</td>
<td valign="top" align="center">24</td>
<td valign="top" align="center">278</td>
<td valign="top" align="center">0.914</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Large</td>
<td valign="top" align="center">290</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">317</td>
<td valign="top" align="center">0.915</td>
</tr>
<tr>
<td valign="top" align="left">Ring</td>
<td valign="top" align="center">Small</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">224</td>
<td valign="top" align="center">0.571</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Middle</td>
<td valign="top" align="center">165</td>
<td valign="top" align="center">121</td>
<td valign="top" align="center">286</td>
<td valign="top" align="center">0.577</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Large</td>
<td valign="top" align="center">153</td>
<td valign="top" align="center">138</td>
<td valign="top" align="center">291</td>
<td valign="top" align="center">0.526</td>
</tr>
<tr>
<td valign="top" align="left">Total</td>
<td/>
<td valign="top" align="center">1,216</td>
<td valign="top" align="center">431</td>
<td valign="top" align="center">1,647</td>
<td valign="top" align="center">0.738</td>
</tr>
</tbody>
</table></table-wrap>
<p><xref ref-type="fig" rid="F8">Figure 8</xref> demonstrates the correct responses for different sizes of VA. The radius of the dots (40 mm) in the scatter chart is estimated based on the vision span theory (<xref ref-type="bibr" rid="B14">Frey and Bosse, 2018</xref>). The coordinate system in <xref ref-type="fig" rid="F8">Figure 8</xref> matches the resolution of the teleoperation screen, and the origin is the center position of the screen. The scattered points are the corresponding position where the VA appears on the screen. <xref ref-type="fig" rid="F9">Figure 9</xref> includes both correct and failed responses. The blue dots stand for the correct ones and the red dots for the failed responses.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>Visualization of correct response numbers: <bold>(A)</bold> cross VA, <bold>(B)</bold> ring VA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g008.tif"/>
</fig>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption><p>Visualization of all response numbers: <bold>(A)</bold> cross VA, <bold>(B)</bold> ring VA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g009.tif"/>
</fig>
<p>To better visualize the result, we divided the screen into 8 &#x00D7; 12 grids and calculated an adjusted correct response rate for each grid by subtracting the number of false responses from correct responses. <xref ref-type="fig" rid="F10">Figure 10</xref> forms the result into a contour map, using a spectrum of warm color to cold color to represent the adjusted correct response values from high to low.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption><p>Visual attention map and corresponding view.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g010.tif"/>
</fig>
<p>The map identifies four types of areas, as shown in <xref ref-type="fig" rid="F10">Figure 10</xref>. Areas 1 and 4 are close to the edge of the screen. Specifically, area 4 refers to the blind spot of excavator operation, where the excavator&#x2019;s boom blocks the view. An operator rarely needs to move the eyesight into these areas to perform an excavation task. They both have a low adjusted response rate, as expected. Area 2 is near and around the fovea vision field and has the highest response rate. The excavating action mostly happens within this area. An operator must pay enough attention to the area for proper interaction between the excavator and the environment. In addition, it is noticed that subareas A and B inside area 2 have high response rates. Subarea A corresponds to the score billboard, and subarea B corresponds to the location of the two trenches for digging and dumping, respectively. It makes sense that an operator pays more attention to the subareas. What remains to be explained is area 3, which is located in the fovea area but presents the lowest response rate.</p>
</sec>
<sec id="S4.SS3">
<title>Response Time</title>
<p><xref ref-type="table" rid="T3">Table 3</xref> demonstrated that most response times are less than 5 s. In general, the response time of the ring VA is longer than that of the cross VA, and the response time is shorter when the size is larger.</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>Descriptive statistics of response time.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">VA</td>
<td valign="top" align="center">Size</td>
<td valign="top" align="center">Num</td>
<td valign="top" align="center">Mean</td>
<td valign="top" align="center">Std.</td>
<td valign="top" align="center">Min</td>
<td valign="top" align="center">Max</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Cross</td>
<td valign="top" align="center">Small</td>
<td valign="top" align="center">226</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">4.42</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Middle</td>
<td valign="top" align="center">254</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">4.88</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Large</td>
<td valign="top" align="center">290</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">4.64</td>
</tr>
<tr>
<td valign="top" align="left">Ring</td>
<td valign="top" align="center">Small</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">2.37</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Middle</td>
<td valign="top" align="center">165</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">5.00</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Large</td>
<td valign="top" align="center">153</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.20</td>
<td valign="top" align="center">2.45</td>
</tr>
</tbody>
</table></table-wrap>
<p><xref ref-type="fig" rid="F11">Figure 11</xref> shows the scattered diagrams of the response time. The radius of the dot is calculated by dividing the 40 mm by each corresponding response time. A large radius stands for a short response time. As shown in <xref ref-type="fig" rid="F11">Figure 11</xref>, when a VA appears at the edge of the screen, the operator&#x2019;s response time will be prolonged accordingly. With the size increasing, the number of larger dots is also increasing. The cross VA, on average, needed a longer response time. It should be noted that the cross VA leads to a better response rate, according to <xref ref-type="table" rid="T2">Table 2</xref>. <xref ref-type="fig" rid="F12">Figure 12</xref> is the contour map for response time. No clear pattern can be found.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption><p>Visualization of response time: <bold>(A)</bold> cross VA, <bold>(B)</bold> ring VA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g011.tif"/>
</fig>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption><p>Map of the adjusted successful response time.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-895126-g012.tif"/>
</fig>
</sec>
</sec>
<sec id="S5">
<title>Data Interpretation and Discussion</title>
<p>This study investigated human visual attention with a VAs-aided teleoperation system. The results revealed that human attention allocation changed regularly with the different VA properties. This section analyzes the mechanism of human attention allocation in detail.</p>
<sec id="S5.SS1">
<title>Visual Attention During Excavator Operation</title>
<p><xref ref-type="fig" rid="F10">Figure 10</xref> demonstrates a clear pattern of an operator&#x2019;s visual attention during the excavating task. A primary finding is that the operating task significantly influences an operator&#x2019;s visual attention. In this experiment, an operator needs to move balls from the left to the right trench by performing actions of bucket digging, boom lifting, cabin rotation, and bucket dumping. The eyesight during the actions mainly fell into area 2, especially subarea B in <xref ref-type="fig" rid="F10">Figure 10</xref>. The high response rate in subarea A also supports this finding. In addition, it matches our existing knowledge about human visual attention that the best area for the human eye to recognize objects is &#x00B1; 10&#x00B0; horizontally and &#x2013;30&#x00B0; to + 10&#x00B0; around the standard line of sight in the vertical direction (<xref ref-type="bibr" rid="B32">Ren et al., 2012</xref>).</p>
<p>The influence on attention allocation by the operating task is likely to override the effect of color contrast. The site background is white in the experiment, and the excavator part is yellow. The red VA should be more conspicuous against the white background than the yellow background. Nevertheless, the experiment did not differentiate the performance based on the background color.</p>
<p>The reason causing a low response rate in area 3 remains unrevealed. After carefully reviewing the experiment video records several times, we still cannot identify a solid reason. We can only speculate that the saliency effect may contribute to this phenomenon. Although area 3 is in the center of the screen, an operator&#x2019;s visual attention is drawn to the trenches and the moving bucket most of the time. The trenches and the bucket trace form a ring around the center area, and the center area, just like areas 1 and 4, receives less attention from the operator. However, it requires further investigation to validate our speculation. In addition, sensing data can be collected during excavation tasks, such as eye-movement tracking, electroencephalograph (EEG), and electromyography (EMG), as suggested by <xref ref-type="bibr" rid="B24">Lee et al. (2022)</xref>. The sensing data could provide an opportunity for more straightforward observation.</p>
</sec>
<sec id="S5.SS2">
<title>Cross Virtual Annotations vs. Ring Virtual Annotations</title>
<p>According to <xref ref-type="table" rid="T2">Table 2</xref>, the cross VA shows a much better performance in response rate. Although the cross and ring are two popular VA shapes used in many existing studies, we observed remarkable differences in this experiment. The ring VA requires the operator to push the honk button and the cross VA to cease operation. Many subjects demonstrated a &#x201C;thinking&#x201D; process when they saw a ring VA, but very few needed to spend time on &#x201C;thinking&#x201D; for a cross VA. It is possible because the shape of the cross generally means &#x201C;stop&#x201D; in the cultural background and in many practical scenes, such as traffic lights and no trespassing signs. In addition, the VA color in this study is red, which may enhance the impression of &#x201C;stop.&#x201D; On the other hand, no matter how easy we imagine it can be to push the honk button to respond to a ring VA, the difficulty level raises dramatically when a subject is under a stressful condition during excavator operation.</p>
<p>A practical implication is that we need to carefully consider all human common sense and cultural backgrounds during the design of VA. The effect of any additional small cognitive load imposed on an operator in a stressful working condition may be escalated.</p>
</sec>
<sec id="S5.SS3">
<title>Visual Attention by Virtual Annotations Size</title>
<p>Intuitively, as the VA size increases, subjects are more likely to detect VAs. Some experiment data in <xref ref-type="table" rid="T2">Table 2</xref> and <xref ref-type="fig" rid="F8">Figure 8</xref> support this intuitive assumption; however, it seems that the marginal positive effect of increasing VA size is decreasing. With the three different sizes, the average response rates are 0.900, 0.914, and 0.915 for the cross-shaped VA and 0.571, 0.577, and 0.526 for the ring-shaped VA, respectively. The data present a trend of improvement from small to middle sizes but not from middle to large sizes.</p>
<p>Considering the vision span theory that the human field of view with sufficient reading resolution typically spans about 6 degrees of arc, the middle-sized VA in this study seems to be close to the best maximum size. It brings up a critical question what is the most appropriate VA size. We suggest a larger size in practice. In this experiment, the subjects expect VAs to appear during operation and are very likely to have allocated a certain amount of attention dedicated to VAs. When VAs may not show up with a regular pattern in a practical scene, it may require a more conspicuous way to present itself.</p>
</sec>
</sec>
<sec id="S6" sec-type="conclusion">
<title>Conclusion</title>
<p>The overarching goal of this study was to investigate the operator&#x2019;s visual attention when VAs are present during excavator teleoperation. A visual attention map is built based on the experiment results, considering the effect of VA size, shape, and appearing location. It is observed that the excavating task influences an operator&#x2019;s visual attention, and the shape of VA plays a critical role in allocating visual attention. It is also speculated that the benefit of increasing VA size may have an asymptotic level, and the optimum size is to be studied in the future.</p>
<p>A major question is why there is an attention vacuum area in the vision center. We suggest future investigations with more subjects, eye-movement tracking, and physiological measurement devices. Testing on different types of construction equipment will also be helpful.</p>
</sec>
<sec id="S7" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="S8">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by the Human Research Ethics Committee. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="S9">
<title>Author Contributions</title>
<p>JF: writing the draft, methodology, and data analysis. XL: data collection and editing. XS: conceptualization, editing, and supervision. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S10" sec-type="funding-information">
<title>Funding</title>
<p>This research was funded by the Center for Balance Architecture, Zhejiang University, China and the National Natural Science Foundation of China (grant no. 71971196).</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andersen</surname> <given-names>D.</given-names></name> <name><surname>Popescu</surname> <given-names>V.</given-names></name> <name><surname>Cabrera</surname> <given-names>M. E.</given-names></name> <name><surname>Shanghavi</surname> <given-names>A.</given-names></name> <name><surname>Gomez</surname> <given-names>G.</given-names></name> <name><surname>Marley</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Virtual annotations of the surgical field through an augmented reality transparent display.</article-title> <source><italic>Vis. Comput.</italic></source> <volume>32</volume> <fpage>1481</fpage>&#x2013;<lpage>1498</lpage>. <pub-id pub-id-type="doi">10.1007/s00371-015-1135-6</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name></person-group> (<year>1980</year>). <source><italic>Cognitive Psychology and its Implications.</italic></source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>W.H. Freeman</publisher-name>.</citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergen</surname> <given-names>J. R.</given-names></name> <name><surname>Julesz</surname> <given-names>B.</given-names></name></person-group> (<year>1983</year>). <article-title>Parallel versus serial processing in rapid pattern discrimination.</article-title> <source><italic>Nature</italic></source> <volume>303</volume> <fpage>696</fpage>&#x2013;<lpage>698</lpage>. <pub-id pub-id-type="doi">10.1038/303696a0</pub-id> <pub-id pub-id-type="pmid">6855915</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolton</surname> <given-names>A.</given-names></name> <name><surname>Burnett</surname> <given-names>G.</given-names></name> <name><surname>Large</surname> <given-names>D. R.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>An investigation of augmented reality presentations of landmark-based navigation using a head-up display</article-title>,&#x201D; in <source><italic>Proceedings of the 7th International Conference on Automotive User Interfaces and Interactive Vehicular Applications</italic></source>, (<publisher-loc>Nottingham</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>).</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burke</surname> <given-names>M.</given-names></name> <name><surname>Hornof</surname> <given-names>A.</given-names></name> <name><surname>Nilsen</surname> <given-names>E.</given-names></name> <name><surname>Gorman</surname> <given-names>N.</given-names></name></person-group> (<year>2005</year>). <article-title>High-cost banner blindness: ads increase perceived workload, hinder visual search, and are forgotten.</article-title> <source><italic>ACM Trans. Comput. Hum. Interact.</italic></source> <volume>12</volume> <fpage>423</fpage>&#x2013;<lpage>445</lpage>. <pub-id pub-id-type="doi">10.1145/1121112.1121116</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. Y. C.</given-names></name> <name><surname>Haas</surname> <given-names>E. C.</given-names></name> <name><surname>Barnes</surname> <given-names>M. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Human performance issues and user interface design for teleoperated robots.</article-title> <source><italic>IEEE Trans. Syst. Man Cybernet. Part C Appl. Rev.</italic></source> <volume>37</volume> <fpage>1231</fpage>&#x2013;<lpage>1245</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCC.2007.905819</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Christoph</surname> <given-names>Z.</given-names></name> <name><surname>Christopher</surname> <given-names>G.</given-names></name> <name><surname>Johannes</surname> <given-names>W.</given-names></name> <name><surname>Kerstin</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). &#x201C;<article-title>Navigation based on a sensorimotor representation: a virtual reality study</article-title>,&#x201D; in <source><italic>Proceedings of the SPIE 6492, Human Vision and Electronic Imaging XII, 64921G</italic></source>, (<publisher-loc>San Jose, CA</publisher-loc>).</citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chun</surname> <given-names>M. M.</given-names></name></person-group> (<year>2000</year>). <article-title>Contextual cueing of visual attention.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>4</volume> <fpage>170</fpage>&#x2013;<lpage>178</lpage>. <pub-id pub-id-type="doi">10.1016/S1364-6613(00)01476-5</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corredor</surname> <given-names>G.</given-names></name> <name><surname>Whitney</surname> <given-names>J.</given-names></name> <name><surname>Pedroza</surname> <given-names>V. L. A.</given-names></name> <name><surname>Madabhushi</surname> <given-names>A.</given-names></name> <name><surname>Castro</surname> <given-names>E. R.</given-names></name></person-group> (<year>2017</year>). <article-title>Training a cell-level classifier for detecting basal-cell carcinoma by combining human visual attention maps with low-level handcrafted features.</article-title> <source><italic>J. Med. Imag.</italic></source> <volume>4</volume>:<issue>021105</issue>. <pub-id pub-id-type="doi">10.1117/1.JMI.4.2.021105</pub-id> <pub-id pub-id-type="pmid">28382314</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Driewer</surname> <given-names>F.</given-names></name> <name><surname>Schilling</surname> <given-names>K.</given-names></name> <name><surname>Baier</surname> <given-names>H.</given-names></name></person-group> (<year>2005</year>). &#x201C;<article-title>Human-computer interaction in the PeLoTe rescue system</article-title>,&#x201D; in <source><italic>Proceedings of the IEEE International Safety, Security and Rescue Rototics, Workshop, 2005</italic></source>, <publisher-loc>Kobe</publisher-loc>, <fpage>224</fpage>&#x2013;<lpage>229</lpage>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>El-Nasr</surname> <given-names>M. S.</given-names></name> <name><surname>Yan</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). &#x201C;<article-title>Visual attention in 3D video games</article-title>,&#x201D; in <source><italic>Proceedings of the 2006 ACM SIGCHI International Conference on Advances in Computer Entertainment Technology</italic></source>, (<publisher-loc>Arlington, VA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>22</fpage>&#x2013;<lpage>es</lpage>.</citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Endsley</surname> <given-names>M. R.</given-names></name></person-group> (<year>1988</year>). <article-title>Design and evaluation for situation awareness enhancement.</article-title> <source><italic>Proc. Hum. Fact. Soc. Annu. Meet.</italic></source> <volume>32</volume> <fpage>97</fpage>&#x2013;<lpage>101</lpage>.</citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fernandez-Duque</surname> <given-names>D.</given-names></name> <name><surname>Thornton</surname> <given-names>I.</given-names></name></person-group> (<year>2000</year>). <article-title>Change detection without awareness: do explicit reports underestimate the representation of change in the visual system?</article-title> <source><italic>Vis. Cogn.</italic></source> <volume>7</volume> <fpage>323</fpage>&#x2013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1080/135062800394838</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frey</surname> <given-names>A.</given-names></name> <name><surname>Bosse</surname> <given-names>M.-L.</given-names></name></person-group> (<year>2018</year>). <article-title>Perceptual span, visual span, and visual attention span: three potential ways to quantify limits on visual processing during reading.</article-title> <source><italic>Vis. Cogn.</italic></source> <volume>26</volume> <fpage>1</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1080/13506285.2018.1472163</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fritsche</surname> <given-names>P.</given-names></name> <name><surname>Zeise</surname> <given-names>B.</given-names></name> <name><surname>Hemme</surname> <given-names>P.</given-names></name> <name><surname>Wagner</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Fusion of radar, LiDAR and thermal information for hazard detection in low visibility environments</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 IEEE International Symposium on Safety, Security and Rescue Robotics (SSRR)</italic></source>, <publisher-loc>Shanghai</publisher-loc>, <fpage>96</fpage>&#x2013;<lpage>101</lpage>.</citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grasset</surname> <given-names>R.</given-names></name> <name><surname>Langlotz</surname> <given-names>T.</given-names></name> <name><surname>Kalkofen</surname> <given-names>D.</given-names></name> <name><surname>Tatzgern</surname> <given-names>M.</given-names></name> <name><surname>Schmalstieg</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). &#x201C;<article-title>Image-driven view management for augmented reality browsers</article-title>,&#x201D; in <source><italic>Proceedings of the 2012 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)</italic></source>, <publisher-loc>Atlanta, GA</publisher-loc>, <fpage>177</fpage>&#x2013;<lpage>186</lpage>.</citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hong</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Su</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). &#x201C;<article-title>Virtual annotations as assistance for construction equipment teleoperation</article-title>,&#x201D; in <source><italic>Proceedings of the 24th International Symposium on Advancement of Construction Management and Real Estate</italic></source>, (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1177</fpage>&#x2013;<lpage>1188</lpage>.</citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hong</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Su</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Effect of virtual annotation on performance of construction equipment teleoperation under adverse visual conditions.</article-title> <source><italic>Autom. Constr.</italic></source> <volume>118</volume>:<issue>103296</issue>. <pub-id pub-id-type="doi">10.1016/j.autcon.2020.103296</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hori</surname> <given-names>S.</given-names></name> <name><surname>Shimizu</surname> <given-names>Y.</given-names></name></person-group> (<year>1999</year>). <article-title>Designing methods of human interface for supervisory control systems.</article-title> <source><italic>Control Eng. Pract.</italic></source> <volume>7</volume> <fpage>1413</fpage>&#x2013;<lpage>1419</lpage>. <pub-id pub-id-type="doi">10.1016/S0967-0661(99)00112-4</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kahneman</surname> <given-names>D.</given-names></name></person-group> (<year>1973</year>). <source><italic>Attention and Effort.</italic></source> <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>Prentice-Hall</publisher-name>.</citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Image-based construction hazard avoidance system using augmented reality in wearable device.</article-title> <source><italic>Autom. Constr.</italic></source> <volume>83</volume> <fpage>390</fpage>&#x2013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2017.06.014</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kot</surname> <given-names>T.</given-names></name> <name><surname>Nov&#x00E1;k</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Application of virtual reality in teleoperation of the military mobile robotic system TAROS.</article-title> <source><italic>Int. J. Adv. Robot. Syst.</italic></source> <volume>15</volume> <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1177/1729881417751545</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larabi</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Augmented reality for mobile devices: textual annotation of outdoor locations</article-title>,&#x201D; in <source><italic>Augmented Reality and Virtual Reality: Empowering Human, Place and Business</italic></source>, <role>eds</role> <person-group person-group-type="editor"><name><surname>Jung</surname> <given-names>T.</given-names></name> <name><surname>tom Dieck</surname> <given-names>M. C.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>353</fpage>&#x2013;<lpage>362</lpage>.</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>J. S.</given-names></name> <name><surname>Ham</surname> <given-names>Y.</given-names></name> <name><surname>Park</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Challenges, tasks, and opportunities in teleoperation of excavator toward human-in-the-loop construction automation.</article-title> <source><italic>Autom. Constr.</italic></source> <volume>135</volume>:<issue>104119</issue>. <pub-id pub-id-type="doi">10.1016/j.autcon.2021.104119</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lichiardopol</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <source><italic>A Survey on Teleoperation. DCT report.</italic></source> <publisher-loc>Eindhoven</publisher-loc>: <publisher-name>Technische Universiteit Eindhoven</publisher-name>.</citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>L. O.</given-names></name> <name><surname>Gomer</surname> <given-names>J. A.</given-names></name> <name><surname>Wong</surname> <given-names>J. T.</given-names></name> <name><surname>Pagano</surname> <given-names>C. C.</given-names></name></person-group> (<year>2011</year>). <article-title>Visual spatial abilities in uninhabited ground vehicle task performance during teleoperation and direct line of sight.</article-title> <source><italic>Presence Teleoper. Virtual Environ.</italic></source> <volume>20</volume> <fpage>466</fpage>&#x2013;<lpage>479</lpage>. <pub-id pub-id-type="doi">10.1162/PRES_a_00066</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matthews</surname> <given-names>T.</given-names></name> <name><surname>Dey</surname> <given-names>A. K.</given-names></name> <name><surname>Mankoff</surname> <given-names>J.</given-names></name> <name><surname>Rattenbury</surname> <given-names>T.</given-names></name></person-group> (<year>2003</year>). &#x201C;<article-title>A peripheral display toolkit</article-title>,&#x201D; in <source><italic>Proceedings of the 17th Annual ACM Symposium on User Interface Software and Technology</italic></source>, (<publisher-loc>Arlington, VA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>).</citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Menchaca-Brandan</surname> <given-names>M. A.</given-names></name> <name><surname>Liu</surname> <given-names>A. M.</given-names></name> <name><surname>Oman</surname> <given-names>C. M.</given-names></name> <name><surname>Natapoff</surname> <given-names>A.</given-names></name></person-group> (<year>2007</year>). &#x201C;<article-title>Influence of perspective-taking and mental rotation abilities in space teleoperation</article-title>,&#x201D; in <source><italic>Proceedings of the ACM/IEEE International Conference on Human-Robot Interaction</italic></source>, (<publisher-loc>Arlington, VA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>). <pub-id pub-id-type="doi">10.3357/AMHP.4557.2016</pub-id> <pub-id pub-id-type="pmid">27634696</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orlosky</surname> <given-names>J.</given-names></name> <name><surname>Kiyokawa</surname> <given-names>K.</given-names></name> <name><surname>Takemura</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Managing mobile text in head mounted displays: studies on visual preference and text placement.</article-title> <source><italic>SIGMOBILE Mob. Comput. Commun. Rev.</italic></source> <volume>18</volume> <fpage>20</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1145/2636242.2636246</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pennington</surname> <given-names>R.</given-names></name></person-group> (<year>2001</year>). <article-title>Signs of marketing in virtual reality.</article-title> <source><italic>J. Interact. Adv.</italic></source> <volume>2</volume> <fpage>33</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1080/15252019.2001.10722056</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rea</surname> <given-names>D. J.</given-names></name> <name><surname>Seo</surname> <given-names>S. H.</given-names></name> <name><surname>Bruce</surname> <given-names>N.</given-names></name> <name><surname>Young</surname> <given-names>J. E.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Movers, shakers, and those who stand still: visual attention-grabbing techniques in robot teleoperation</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 ACM/IEEE International Conference on Human-Robot Interaction</italic></source>, <publisher-loc>Vienna</publisher-loc>, <fpage>398</fpage>&#x2013;<lpage>407</lpage>.</citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>X.</given-names></name> <name><surname>Xue</surname> <given-names>Q.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Research of virtual driver&#x2019; s visual attention model.</article-title> <source><italic>Appl. Res. Comput.</italic></source> <volume>48</volume> <fpage>209</fpage>&#x2013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.3778/j.issn.1002-8331.2012.19.048</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riley</surname> <given-names>J. M.</given-names></name> <name><surname>Kaber</surname> <given-names>D. B.</given-names></name> <name><surname>Draper</surname> <given-names>J. V.</given-names></name></person-group> (<year>2004</year>). <article-title>Situation awareness and attention allocation measures for quantifying telepresence experiences in teleoperation.</article-title> <source><italic>Hum. Fact. Ergon. Manuf. Serv. Industr.</italic></source> <volume>14</volume> <fpage>51</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1002/hfm.10050</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sato</surname> <given-names>R.</given-names></name> <name><surname>Kamezaki</surname> <given-names>M.</given-names></name> <name><surname>Niuchi</surname> <given-names>S.</given-names></name> <name><surname>Sugano</surname> <given-names>S.</given-names></name> <name><surname>Iwata</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Cognitive untunneling multi-view system for teleoperators of heavy machines based on visual momentum and saliency.</article-title> <source><italic>Autom. Constr.</italic></source> <volume>110</volume>:<issue>103047</issue>. <pub-id pub-id-type="doi">10.1016/j.autcon.2019.103047</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shapira</surname> <given-names>A.</given-names></name> <name><surname>Rosenfeld</surname> <given-names>Y.</given-names></name> <name><surname>Mizrahi</surname> <given-names>I.</given-names></name></person-group> (<year>2008</year>). <article-title>Vision system for tower cranes.</article-title> <source><italic>J. Constr. Eng. Manag.</italic></source> <volume>134</volume> <fpage>320</fpage>&#x2013;<lpage>332</lpage>. <pub-id pub-id-type="doi">10.1061/(ASCE)0733-93642008134:5(320</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>L.</given-names></name> <name><surname>Wickens</surname> <given-names>C.</given-names></name></person-group> (<year>2001</year>). &#x201C;<article-title>Visual displays and cognitive tunneling: frames of reference effects on spatial judgments and change detection</article-title>,&#x201D; in <source><italic>Proceedings of the Human Factors and Ergonomics Society Annual Meeting</italic></source>, (<publisher-loc>Fort Belvoir, VA</publisher-loc>: <publisher-name>Defense Technical Information Center</publisher-name>), <volume>45</volume>. <pub-id pub-id-type="doi">10.1177/154193120104500415</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tylecek</surname> <given-names>R.</given-names></name> <name><surname>Fisher</surname> <given-names>R. B.</given-names></name></person-group> (<year>2018</year>). <article-title>Consistent semantic annotation of outdoor datasets via 2D/3D label transfer.</article-title> <source><italic>Sensors (Basel)</italic></source> <volume>18</volume>:<issue>2249</issue>. <pub-id pub-id-type="doi">10.3390/s18072249</pub-id> <pub-id pub-id-type="pmid">30002334</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verlinden</surname> <given-names>J.</given-names></name> <name><surname>Bolter</surname> <given-names>J. D.</given-names></name> <name><surname>van der Mast</surname> <given-names>C.</given-names></name></person-group> (<year>1993</year>). <source><italic>The World Processor an Interface for Textual Display and Manipulation in Virtual Reality. The Faculty of Technical Mathematics and Information.</italic></source> <publisher-loc>Atlanta, GA</publisher-loc>: <publisher-name>Georgia Tech Library</publisher-name>.</citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wallmyr</surname> <given-names>M.</given-names></name> <name><surname>Sitompul</surname> <given-names>T. A.</given-names></name> <name><surname>Holstein</surname> <given-names>T.</given-names></name> <name><surname>Lindell</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <source><italic>Evaluating Mixed Reality Notifications to Support Excavator Operator Awareness.</italic></source> <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>, <fpage>743</fpage>&#x2013;<lpage>762</lpage>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Dunston</surname> <given-names>P. S.</given-names></name></person-group> (<year>2006</year>). <article-title>Compatibility issues in augmented reality systems for AEC: an experimental prototype study.</article-title> <source><italic>Autom. Constr.</italic></source> <volume>15</volume> <fpage>314</fpage>&#x2013;<lpage>326</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2005.06.002</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>M.</given-names></name> <name><surname>Yao</surname> <given-names>K.</given-names></name> <name><surname>Nurse</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>ToARist: an augmented reality tourism app created through user-centred design.</article-title> <source><italic>arXiv</italic></source> [<comment>Preprint</comment>]. <volume>arXiv</volume>:<issue>1807.05759</issue>,</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woods</surname> <given-names>D. D.</given-names></name> <name><surname>Tittle</surname> <given-names>J.</given-names></name> <name><surname>Feil</surname> <given-names>M.</given-names></name> <name><surname>Roesler</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Envisioning human-robot coordination in future operations.</article-title> <source><italic>IEEE Trans. Syst. Man Cybernet. Part C Appl. Rev.</italic></source> <volume>34</volume> <fpage>210</fpage>&#x2013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCC.2004.826272</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woo-Keun</surname> <given-names>Y.</given-names></name> <name><surname>Goshozono</surname> <given-names>T.</given-names></name> <name><surname>Kawabe</surname> <given-names>H.</given-names></name> <name><surname>Kinami</surname> <given-names>M.</given-names></name> <name><surname>Tsumaki</surname> <given-names>Y.</given-names></name> <name><surname>Uchiyama</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2004</year>). <article-title>Model-based space robot teleoperation of ETS-VII manipulator.</article-title> <source><italic>IEEE Trans. Robot. Autom.</italic></source> <volume>20</volume> <fpage>602</fpage>&#x2013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1109/TRA.2004.824700</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeh</surname> <given-names>S.-C.</given-names></name> <name><surname>Hwang</surname> <given-names>W.-Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.-L.</given-names></name> <name><surname>Zhan</surname> <given-names>S.-Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Study of co-located and distant collaboration with symbolic support via a haptics-enhanced virtual reality task.</article-title> <source><italic>Interact. Learn. Environ.</italic></source> <volume>21</volume> <fpage>184</fpage>&#x2013;<lpage>198</lpage>. <pub-id pub-id-type="doi">10.1080/10494820.2012.705854</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ziaei</surname> <given-names>Z.</given-names></name> <name><surname>Hahto</surname> <given-names>A.</given-names></name> <name><surname>Mattila</surname> <given-names>J.</given-names></name> <name><surname>Siuko</surname> <given-names>M.</given-names></name> <name><surname>Semeraro</surname> <given-names>L.</given-names></name></person-group> (<year>2011</year>). <article-title>Real-time markerless augmented reality for remote handling system in bad viewing conditions.</article-title> <source><italic>Fusion Eng. Design</italic></source> <volume>86</volume> <fpage>2033</fpage>&#x2013;<lpage>2038</lpage>. <pub-id pub-id-type="doi">10.1016/j.fusengdes.2010.12.082</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zornitza</surname> <given-names>Y.</given-names></name> <name><surname>Dimitrios</surname> <given-names>B.</given-names></name> <name><surname>Christos</surname> <given-names>G.</given-names></name> <name><surname>Corn&#x00E9;</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>Empirical evaluation of smartphone augmented reality browsers in an urban tourism destination context.</article-title> <source><italic>Int. J. Mob. Hum. Comput. Interact.</italic></source> <volume>6</volume> <fpage>10</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.4018/ijmhci.2014040102</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="footnote1">
<label>1</label>
<p><ext-link ext-link-type="uri" xlink:href="https://github.com/mogoson/MGS-Machinery">https://github.com/mogoson/MGS-Machinery</ext-link></p></fn>
</fn-group>
</back>
</article>
