<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2023.1079998</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Cross-modal correspondence enhances elevation localization in visual-to-auditory sensory substitution</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Bordeau</surname> <given-names>Camille</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1898117/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Scalvini</surname> <given-names>Florian</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2171123/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Migniot</surname> <given-names>Cyrille</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2171790/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Dubois</surname> <given-names>Julien</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2171867/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ambard</surname> <given-names>Maxime</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/38668/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>LEAD-CNRS UMR5022, Universit&#x000E9; de Bourgogne</institution>, <addr-line>Dijon</addr-line>, <country>France</country></aff>
<aff id="aff2"><sup>2</sup><institution>ImViA EA 7535, Universit&#x000E9; de Bourgogne</institution>, <addr-line>Dijon</addr-line>, <country>France</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Benedikt Zoefel, UMR5549 Centre de Recherche Cerveau et Cognition (CerCo), France</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Takuya Koumura, NTT Communication Science Laboratories, Japan; Gabriel Arnold, Independent Researcher, Villebon-sur-Yvette, France</p></fn>

<corresp id="c001">&#x0002A;Correspondence: Camille Bordeau &#x02709; <email>bordeau.camille&#x00040;gmail.com</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Perception Science, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>26</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1079998</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Bordeau, Scalvini, Migniot, Dubois and Ambard.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Bordeau, Scalvini, Migniot, Dubois and Ambard</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Visual-to-auditory sensory substitution devices are assistive devices for the blind that convert visual images into auditory images (or soundscapes) by mapping visual features with acoustic cues. To convey spatial information with sounds, several sensory substitution devices use a Virtual Acoustic Space (VAS) using Head Related Transfer Functions (HRTFs) to synthesize natural acoustic cues used for sound localization. However, the perception of the elevation is known to be inaccurate with generic spatialization since it is based on notches in the audio spectrum that are specific to each individual. Another method used to convey elevation information is based on the audiovisual cross-modal correspondence between pitch and visual elevation. The main drawback of this second method is caused by the limitation of the ability to perceive elevation through HRTFs due to the spectral narrowband of the sounds.</p>
</sec>
<sec>
<title>Method</title>
<p>In this study we compared the early ability to localize objects with a visual-to-auditory sensory substitution device where elevation is either conveyed using a spatialization-based only method (Noise encoding) or using pitch-based methods with different spectral complexities (Monotonic and Harmonic encodings). Thirty eight blindfolded participants had to localize a virtual target using soundscapes before and after having been familiarized with the visual-to-auditory encodings.</p>
</sec>
<sec>
<title>Results</title>
<p>Participants were more accurate to localize elevation with pitch-based encodings than with the spatialization-based only method. Only slight differences in azimuth localization performance were found between the encodings.</p>
</sec>
<sec>
<title>Discussion</title>
<p>This study suggests the intuitiveness of a pitch-based encoding with a facilitation effect of the cross-modal correspondence when a non-individualized sound spatialization is used.</p>
</sec></abstract>
<kwd-group>
<kwd>Virtual Acoustic Space</kwd>
<kwd>spatial hearing</kwd>
<kwd>sound spatialization</kwd>
<kwd>image-to-sound conversion</kwd>
<kwd>cross-modal correspondence</kwd>
<kwd>assistive technology</kwd>
<kwd>visual impairment</kwd>
<kwd>sound source localization</kwd>
</kwd-group>
<contract-sponsor id="cn001">European Regional Development Fund<named-content content-type="fundref-id">10.13039/501100008530</named-content></contract-sponsor>
<contract-sponsor id="cn002">Conseil r&#x000E9;gional de Bourgogne-Franche-Comt&#x000E9;<named-content content-type="fundref-id">10.13039/501100011773</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="69"/>
<page-count count="18"/>
<word-count count="14489"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Visual-to-auditory Sensory substitution devices (SSDs) are assistive tools for blind people. They convert visual information into auditory information in order to convey spatial information about the surrounding environment when vision is impaired. The visual-to-auditory conversion relies on the mapping of selected visual features with specific auditory cues. Visual information is usually acquired using a camera capturing the visual scene in front of the person. Then the scene converted into auditory information is transmitted to the user through soundscapes (or auditory images) delivered with headphones.</p>
<p>Various visual-to-auditory encodings are used by the existing visual-to-auditory SSDs to convey spatial information. Some of them use encoding schemes based on a Virtual Acoustic Space (VAS). A VAS consists in the simulation of a binaural acoustic signature of a virtual sound source located in a 3D space. In the context of visual-to-auditory SSDs, this is mainly used to simulate sound sources at the location of the obstacles. This simulation is achieved by spatializing the sound through the incorporation of spatial auditory cues in the original monophonic sound. Then a synthesized stereophonic signal simulating the distortions occurring while receiving the audio signal by the two ears is obtained. Among the SSDs used in localization experiments, the Synaestheatre (Hamilton-Fletcher et al., <xref ref-type="bibr" rid="B22">2016a</xref>; Richardson et al., <xref ref-type="bibr" rid="B53">2019</xref>), the Vibe (Hanneton et al., <xref ref-type="bibr" rid="B24">2010</xref>) and the one presented by Mhaish et al. (<xref ref-type="bibr" rid="B43">2016</xref>) spatialize azimuth (lateral position) and elevation (vertical position). Other SSDs only spatialize the azimuth: the See differently device (Rouat et al., <xref ref-type="bibr" rid="B55">2014</xref>), the one studied in Ambard et al. (<xref ref-type="bibr" rid="B5">2015</xref>), and the recent one presented in Scalvini et al. (<xref ref-type="bibr" rid="B57">2022</xref>).</p>
<p>The generation of a VAS is based on the reproduction of binaural acoustic cues related to the relative sound source location such as timing, intensity and spectral features (for an in-depth explanation of the auditory localization mechanisms see Blauert, <xref ref-type="bibr" rid="B12">1996</xref>). Those features arise from audio signal distortions mainly caused by the reflection and absorption of the head, pinna and torso and are partly reproducible using Head-Related Transfer Functions (HRTFs). HRTFs are transfer functions characterizing these signal distortions as a function of the position of the sound source relatively to the two ears. They are usually obtained by conducting multiple binaural recordings with a sound source carefully placed in various positions while repeatedly producing the same sound.</p>
<p>Due to the technical difficulty in acquiring these recordings in good conditions, non-individualized HRTFs acquired in controlled conditions with another listener or a manikin are frequently used. However, these HRTFs failed to simulate the variability of individual-specific spectrum distortions that are related to individual morphologies (head, torso and pinna). Consequently, the localization of simulated sound sources using non-individualized HRTFs is often inaccurate with front-back and up-down confusions that are less resolvable (Wenzel et al., <xref ref-type="bibr" rid="B67">1993</xref>), and a less perceptible externalization (Best et al., <xref ref-type="bibr" rid="B11">2020</xref>). Nonetheless, due to the robustness of the binaural cues, azimuth localization accuracy is well preserved compared to the perception of elevation since azimuth perception relies less on the individual-specific spectrum distortions (Makous and Middlebrooks, <xref ref-type="bibr" rid="B40">1990</xref>; Wenzel et al., <xref ref-type="bibr" rid="B67">1993</xref>; Middlebrooks, <xref ref-type="bibr" rid="B44">1999</xref>). Therefore, visual-to-auditory encodings only based on the creation of a VAS have the advantage to rely on acoustic cues that mimic natural acoustic features for sound source localization, nevertheless in practice the elevation perception can be impaired.</p>
<p>To compensate for this difficulty some visual-to-auditory SSDs use additional acoustic cues to convey spatial information. For instance, pitch modulation is often used to convey elevation location (Meijer, <xref ref-type="bibr" rid="B41">1992</xref>; Abboud et al., <xref ref-type="bibr" rid="B1">2014</xref>; Ambard et al., <xref ref-type="bibr" rid="B5">2015</xref>). This mapping between elevation location and auditory pitch is based on the audiovisual cross-modal correspondence between pitch and elevation (see Spence, <xref ref-type="bibr" rid="B60">2011</xref> for a review on audiovisual cross-modal correspondences). Humans show a tendency to associate high pitch with high spatial locations and low pitch with low spatial locations. For example, they tend to exhibit faster response times in an audio-visual Go/No-Go task when the visual and auditory stimuli are congruent, i.e., higher pitch with higher visual location, and lower pitch with lower visual location (Miller, <xref ref-type="bibr" rid="B47">1991</xref>). They also tend to discriminate more accurately and quickly the location of a visual stimulus (high <italic>vs</italic>. low location) when the pitch of a presented sound is congruent with the visual elevation (Evans and Treisman, <xref ref-type="bibr" rid="B19">2011</xref>). Also, humans tend to respond to high pitch sounds with a high-located response button instead of a lower-located response button (Rusconi et al., <xref ref-type="bibr" rid="B56">2006</xref>). The pitch-based encoding used in the vOICe SSD (Meijer, <xref ref-type="bibr" rid="B41">1992</xref>) has been suggested somewhat intuitive in a recognition task (Stiles and Shimojo, <xref ref-type="bibr" rid="B64">2015</xref>). Nevertheless, the main drawback of a pitch-based encoding is caused by the limitation of the abilities to perceive elevation through HRTFs due to the audio spectral narrowband (Algazi et al., <xref ref-type="bibr" rid="B4">2001b</xref>). Although some acoustic cues for elevation perception are present in low frequencies below 3,500 Hz (Gardner, <xref ref-type="bibr" rid="B20">1973</xref>; Asano et al., <xref ref-type="bibr" rid="B6">1990</xref>), localization abilities are higher when the spectral content contains high frequencies above 4,000 Hz (Hebrank and Wright, <xref ref-type="bibr" rid="B25">1974</xref>; Middlebrooks and Green, <xref ref-type="bibr" rid="B45">1990</xref>). Since the ability to perceive the elevation through HRTFs is higher with broadband sounds containing high frequencies, the spectral content of the sound used in the visual-to-auditory encoding might modulate the perception of elevation through HRTFs. No study has directly compared encodings only based on HRTFs with encodings adding a pitch modulation and it remains unclear if the simulation of natural acoustic cues is less efficient for object localization than a more artificial sonification method using the cross-modal correspondence between pitch and elevation.</p>
<p>Many studies investigating static object localization abilities have already been conducted with blindfolded sighted persons using visual-to-auditory SSDs. Various types of tasks have already been used, for example discrimination tasks with forced choice (Proulx et al., <xref ref-type="bibr" rid="B51">2008</xref>; Levy-Tzedek et al., <xref ref-type="bibr" rid="B36">2012</xref>; Ambard et al., <xref ref-type="bibr" rid="B5">2015</xref>; Mhaish et al., <xref ref-type="bibr" rid="B43">2016</xref>; Richardson et al., <xref ref-type="bibr" rid="B53">2019</xref>), grasping tasks (Proulx et al., <xref ref-type="bibr" rid="B51">2008</xref>), index or tool pointing tasks (Auvray et al., <xref ref-type="bibr" rid="B8">2007</xref>; Hanneton et al., <xref ref-type="bibr" rid="B24">2010</xref>; Brown et al., <xref ref-type="bibr" rid="B13">2011</xref>; Pourghaemi et al., <xref ref-type="bibr" rid="B50">2018</xref>; Comm&#x000E8;re et al., <xref ref-type="bibr" rid="B16">2020</xref>), or head-pointing tasks (Scalvini et al., <xref ref-type="bibr" rid="B57">2022</xref>). Those studies showed the high potential of SSDs to localize an object and interact with it. However, long trainings were often conducted before the localization tasks to learn the visual-to-auditory encoding schemes: from 5 min in Pourghaemi et al. (<xref ref-type="bibr" rid="B50">2018</xref>) to 3 h in Auvray et al. (<xref ref-type="bibr" rid="B8">2007</xref>). On the contrary, in the study of Scalvini et al. (<xref ref-type="bibr" rid="B57">2022</xref>) the experimenter only explained verbally the encoding schemes to the participants.</p>
<p>Virtual environments are more and more used to investigate the abilities to perceive the environment with a visual-to-auditory SSD (Maidenbaum et al., <xref ref-type="bibr" rid="B37">2014</xref>; Kristj&#x000E1;nsson et al., <xref ref-type="bibr" rid="B31">2016</xref>) since they allow a complete control of the experimental environment (e.g., number of objects, object locations...) (Maidenbaum and Amedi, <xref ref-type="bibr" rid="B38">2019</xref>) and a more accurate assessment of localization abilities with precise pointing methods. They have been used in standardization tests to compare the abilities to interpret information provided by SSDs in navigation or localization tasks (Caraiman et al., <xref ref-type="bibr" rid="B15">2017</xref>; Richardson et al., <xref ref-type="bibr" rid="B53">2019</xref>; Jicol et al., <xref ref-type="bibr" rid="B29">2020</xref>; Real and Araujo, <xref ref-type="bibr" rid="B52">2021</xref>).</p>
<p>The current study aimed at investigating the intuitiveness of different types of visual-to-auditory encodings for the elevation in the context of object localization with a SSD. Therefore, we conducted a localization task in a virtual environment with blindfolded participants testing a spatialization-based encoding and a pitch-based encoding. This study also aimed at assessing whether a higher spectral complexity of the sound used in a pitch-based encoding could improve the localization performance. Therefore, 2 types of pitch-based encodings were investigated: one monotonic and one harmonic with 3 octaves. We measured the localization performance for the azimuth and for the elevation. For each of these measures, we studied the effect of the visual-to-auditory encoding before and after an audio-motor familiarization of short duration.</p>
<p>Since the audio spatialization method was not based on individualized HRTFs, and since the pitch-based encodings were not explained to the participants, localization performance for the elevation was expected to be impaired. However, a facilitation effect of the pitch-based encodings for the elevation localization accuracy was hypothesized. Among the two pitch-based encodings, a higher elevation localization accuracy was predicted with the harmonic encoding since the sound has a higher spectral complexity. Also, the intuitiveness of the azimuth perception for all the encodings was hypothesized since it is based on less individual-specific acoustic spatial cues than elevation perception.</p>
</sec>
<sec id="s2">
<title>2. Method</title>
<sec>
<title>2.1. Participants</title>
<p>Thirty eight participants were divided into two groups: the Monotonic group (19, age: <italic>M</italic> = 25.5, <italic>SD</italic> = 3.04, 6 female, 19 right-handed) and the Harmonic group (19, age: <italic>M</italic> = 24.4, <italic>SD</italic> = 3.27, 10 female, 18 right-handed). No participant reported impairments of hearing or any history of psychiatric illness or neurological disorder. The experimental protocol was approved by the local ethical committee Comit&#x000E9; d&#x00027;Ethique pour la Recherche de Universit&#x000E9; Bourgogne Franche-Comt&#x000E9; (CERUBFC-2021-12-21-050) and followed the ethical guidelines of the Declaration of Helsinki. Written informed consent was obtained from all the participants before the experiment. No monetary compensation was given to the participants.</p>
</sec>
<sec>
<title>2.2. Visual-to-auditory conversion in the virtual environment</title>
<p>The visual-to-auditory SSD used took place in a virtual environment created in UNITY3D and including the target to localize, a virtual camera, and a tracked pointing tool. Four HTC VIVE base stations were used to track the participants&#x00027; head and the pointing tool on which HTC VIVE Trackers 2.0 were attached. Participants did not carry a headset and therefore could not explore visually the virtual environment. The pointing task can be separated in several steps that are explained in detail below: the virtual target placement, the video acquisition from a virtual camera, the video processing, the visual-to-auditory conversion and the participants&#x00027; response collection using the pointing tool.</p>
<sec>
<title>2.2.1. Virtual target</title>
<p>The virtual target that participants had to localize was a 3D propeller shape of 25 cm in diameter composed of 4 bars with a length of 25 cm and a rectangular section of 5 &#x000D7; 5 cm that was self-rotating at a speed of 10&#x000B0; per video frame (see <xref ref-type="fig" rid="F1">Figure 1A</xref>). The use of an angular shaped target that is self-rotating generated a modification of successive video frames without changing the center position of the target. The orientation of the target was managed in order to continuously face the virtual camera while being displayed. Since participants could not see the virtual target, it was only perceivable through the soundscapes.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>(A)</bold> The virtual target was a self-rotating 3D propeller shape. The straight arrow shows the forward axis facing the virtual camera. The circular arrow shows the self-rotation direction. <bold>(B)</bold> In each localization test, 25 target positions (blue circles) were tested, including 5 elevation positions (horizontal dotted ellipses) tested at 5 azimuth positions (vertical dotted ellipses). <bold>(C)</bold> In the familiarization session, participants placed the virtual target during 60 seconds on the grid by moving the pointing tool. Participants&#x00027; head were tracked during the localization tests and familiarization sessions (blue square).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0001.tif"/>
</fig>
</sec>
<sec>
<title>2.2.2. Video acquisition</title>
<p>The virtual camera position was set at the beginning of each trial using the position of the head tracker attached on the participants&#x00027; forehead. Images were acquired with a virtual camera with a field of view of 90 &#x000D7; 74&#x000B0; (Horizontal &#x000D7; Vertical) and a frame rate of 60 Hz. The resulting image was using a grayscale encoding (0&#x02013;255 gray levels) of a depth map (0.2 m = 0, 5.0 m = 255) of the virtual scene although in this experiment we did not manipulate the depth parameter.</p>
</sec>
<sec>
<title>2.2.3. Video processing</title>
<p>Video processing principles are similar to those used by Ambard et al. (<xref ref-type="bibr" rid="B5">2015</xref>), aiming to convey only new visual information from one frame to another. Video frames are grayscale images with gray levels ranging from 0 to 255. Pixels of the current frame are only conserved if the gray level pixel-by-pixel absolute difference with the previous frame (frame differencing) is larger than a threshold of 10. The processed image is then rescaled to a 160 &#x000D7; 120 (Horizontal &#x000D7; Vertical) grayscale image where 0-gray-level pixels are called &#x0201C;inactive&#x0201D; (i.e., no new visual information contained) and the others are &#x0201C;active&#x0201D; graphical pixels (i.e., containing new visual information). Active graphical pixels are then converted into spatialized sounds following a visual-to-auditory encoding in order to generate a soundscape (<xref ref-type="fig" rid="F2">Figure 2</xref>), as explained in the following section.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Two examples of processed video frames and their corresponding soundscapes. The two processed video frames are depicted in the left side of the figure with a target located on the upper right <bold>(top image)</bold> and bottom left <bold>(bottom image)</bold>. Active and inactive graphical pixels are depicted in gray and black, respectively. Two graphical pixels are highlighted in the video frames (orange in the <bold>top image</bold>, blue in the <bold>bottom image</bold>) and the corresponding auditory pixel waveforms are depicted in the right part of the figure in orange and blue. The corresponding soundscape waveforms (in gray) and soundscape spectrograms of the video frames are also depicted in the right part of the figure. Auditory pixel waveforms, soundscapes waveforms and spectrograms are displayed separately for the Noise encoding <bold>(left column)</bold>, the Harmonic encoding <bold>(middle column)</bold> and the Monotonic encoding <bold>(right column)</bold> and for left (L) and right (R) ear channels separately.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0002.tif"/>
</fig>
</sec>
<sec>
<title>2.2.4. Visual-to-auditory conversion</title>
<p>The visual-to-auditory conversion consists in the transformation of the processed video stream into a synchronized audio stream that acoustically encodes the extracted graphical features. Each graphical pixel is associated with an &#x0201C;auditory pixel&#x0201D; which is a stereophonic sound with auditory cues specific to the position of the graphical pixel it is associated with. The conversion from a graphical pixel to an auditory pixel follows an encoding that is explained step-by-step in the following sections. Each graphical pixel of a video frame is first associated with a corresponding monophonic audio pixel (detailed in Section 2.2.4.1). The spatialization of the sound using HRTFs is then used to generate a stereophonic audio pixel that simulates a sound source with azimuth and elevation corresponding to the position of the graphical pixel in the virtual camera&#x00027;s field of view (detailed in Section 2.2.4.2). All the stereophonic pixels of a video frame are then compiled to obtain an audio frame (detailed in Section 2.2.4.3). Successive audio frames are then mixed together to generate a continuous audio stream (i.e., the soundscape). Two examples of stereophonic auditory pixels are provided in <xref ref-type="fig" rid="F2">Figure 2</xref> for each of the three encodings, as well as two examples of soundscapes depending on the location of an object in the field of view of the virtual camera.</p>
<sec>
<title>2.2.4.1. Monophonic pixel synthesizing</title>
<p>Three visual-to-auditory encodings were tested in this study: the Noise encoding and 2 Pitch encodings (the Monotonic encoding and the Harmonic encoding). These methods varied in the elevation encoding scheme and in the spectral complexity of the monophonic auditory pixels but all three methods used afterwards the same method for the sound spatialization.</p>
<p>For the Noise encoding, the simulated sound source (i.e., monophonic auditory pixel) in the VAS was a white noise signal generated by inverting a Fourier representation of the auditory pixel with a flat spectrum and random phases.</p>
<p>For the Monotonic encoding, each monophonic auditory pixel was a sinusoidal waveform audio signal (i.e., a pure tone) with a random phase and a frequency related to the elevation of the corresponding graphical pixel in the processed image. For this purpose, we used a linear Mel scale ranging from 344 mel (bottom) to 1,286 mel (top) corresponding to frequencies from 250 to 1,492 Hz.</p>
<p>For the Harmonic encoding, we used the same monophonic auditory pixels as in the Monotonic encoding but instead of a pure tone, we added to it two other frequencies at the 2 following octaves with the same intensity and random phases.</p>
<p>Since the loudness depends on the frequency components of the audio signal, we minimized the differences in loudness between auditory pixels using the <italic>pyloudnorm</italic> Python-package (Steinmetz and Reiss, <xref ref-type="bibr" rid="B62">2021</xref>). Auditory pixel spectrums were then adjusted to compensate for the frequency response of the headphones we used in this experiment (SONY MDR-7506).</p>
</sec>
<sec>
<title>2.2.4.2. Auditory pixel spatialization</title>
<p>The azimuth and elevation associated with each pixel were computed based on the coordinates of the corresponding graphical pixel in the camera&#x00027;s field of view. Monophonic auditory pixels were then spatialized by convolving them with the corresponding KEMAR HRTFs from the CIPIC database (Algazi et al., <xref ref-type="bibr" rid="B3">2001a</xref>). This database provides HRTFs recordings with a sound source located in various azimuths and elevations ranging in steps of 5 and 5.625&#x000B0;, respectively. For each pixel, the applied HRTFs were estimated from the database by computing a 4 points time-domain interpolation in which the Interaural Level Difference (ILD) and the convolution signals were separately interpolated using bilinear interpolations before being reassembled as in Sodnik et al. (<xref ref-type="bibr" rid="B59">2005</xref>) but using a 2D interpolation instead of a 1D interpolation.</p>
</sec>
<sec>
<title>2.2.4.3. Audio frame mixing</title>
<p>Each auditory pixel lasted 34.83 ms including a 5 ms cosine fade-in and a 5 ms cosine fade-out. All the auditory pixels corresponding to the active graphical pixels of the processed current video frame were compiled to form an audio frame. After their compilation, these fade-in and fade-out were still present at the beginning and at the end of the audio frame and they were used to overlap successive audio frames while limiting the artifacts of the auditory transition.</p>
</sec>
</sec>
<sec>
<title>2.2.5. Pointing tool and response collection</title>
<p>The pointing tool was a tracked gun pistol. Participants were instructed to indicate the perceived target position by pointing to it with the gun, with stretched arm. Participants logged their response by pressing a button with their index finger. They were instructed to hold the pointing tool with their dominant hand. The response position was defined as the intersection point of a virtual ray originating at the tip of the pointing tool and a virtual 1-m radius sphere with the origin at the location of the virtual camera. The response positions were declined in the elevation response and the azimuth response. The elevation and azimuth signed errors were also computed as the difference between the target position and the response position (in elevation and azimuth separately). A negative elevation signed error indicated a downward shift, and a negative azimuth signed error indicated a shift to the left in the response position. Unsigned errors were computed as the absolute value of the signed errors of each trial.</p>
</sec>
</sec>
<sec>
<title>2.3. Experimental procedure</title>
<p>The experiment consisted in a 45-min session during which participants were seated comfortably in a chair at the center of a room surrounded by the virtual reality tracking system. The participants were equipped with SONY MDR-7506 headphones used to deliver soundscapes. <xref ref-type="fig" rid="F3">Figure 3</xref> illustrates the timeline of the experimental session. Each participant had to test two visual-to-auditory encodings: the Noise encoding, and a Pitch encoding (Monotonic or Harmonic encoding depending on the group they belonged). Participants from the Monotonic group had to test the Noise encoding and the Monotonic encoding, and participants from the Harmonic group had to test the Noise encoding and the Harmonic encoding. The order of the two tested encodings was counterbalanced between participants so half participants of each group started with the Noise encoding and the other half started with the Pitch encoding.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Experimental timeline. Participants had to sequentially test the Noise encoding and one of the two Pitch encodings (Monotonic or Harmonic). Participants of the Monotonic and Harmonic groups tested the Monotonic encoding and the Harmonic encoding respectively. For each encoding, participants practiced the localization test two times, before and after a familiarization session.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0003.tif"/>
</fig>
<p>For each encoding, the participants practiced 2 times the localization test: one without any familiarization or explanation of the encoding and one after a familiarization session. At the beginning of the experiment, participants were instructed to localize a virtual target by pointing to it while being blindfolded. The experimenter explained that they will not be able to see the virtual target, but that they will only hear it and that the sound will depend on the position of the target. No indication was given about the way visual-to-auditory encodings worked. Participants were seated and blindfolded using an opaque blindfold fixed with a rubber band and could remove it during breaks. Participants were instructed to keep their head as still as possible during the localization tests. For control purposes, participants&#x00027; head position was recorded with the tracker every 200 ms to check that they kept their head still. We measured the maximum distance of the head from its mean position for each trial and we found an average maximum distance of approximately 1.5 cm showing that the instructions were rigorously followed.</p>
<sec>
<title>2.3.1. Localization test</title>
<p>The localization test consisted in 50 trials during which blindfolded participants had to localize the virtual target using soundscapes provided by the visual-to-auditory SSD. During each localization test, the target was located at 25 different positions distributed on a grid of 5 azimuths (&#x02212;40, &#x02212;20, 0, &#x0002B;20, and &#x0002B;40&#x000B0;) and 5 elevations (&#x02212;25, &#x02212;12.5, 0, &#x0002B;12.5, and &#x0002B;25&#x000B0;). <xref ref-type="fig" rid="F1">Figure 1B</xref> illustrates the grid with the 25 tested positions. As an example, the position [0&#x000B0;, 0&#x000B0;] corresponded to the central position, i.e., the virtual target was centered with the participant&#x00027;s head tracker. For the position [&#x02212;40&#x000B0;, &#x0002B;12.5&#x000B0;], the target was 40&#x000B0; leftward and 12.5&#x000B0; upward from the central position ([0&#x000B0;, 0&#x000B0;]). The order of the tested positions was randomized and each position was tested 2 times per localization test. The target was placed at 1-meter-distance from the participant&#x00027;s head tracker for all positions (on the virtual 1-meter radius sphere used to collect the response positions).</p>
<p>Each trial started with a 500 ms 440 Hz beep sound, indicating the beginning of the trial. After a 500 ms silent period, the virtual target was displayed at one of the 25 tested positions. Participants were instructed to point with the pointing tool to the perceived location of the target with stretched arm. No time limit was imposed for responding but participants were asked to respond as fast and accurately as possible. The virtual target was displayed until participants pressed the trigger of the pointing tool. The response position was recorded (see Section 2.2.5 for response position computing) and the target disappeared. After a 1,000 ms inter-trial break, the next trial began with the 500 ms beep sound. No feedback was provided regarding response accuracy.</p>
</sec>
<sec>
<title>2.3.2. Familiarization session</title>
<p>In between the 2 localization tests of each of the 2 tested encodings, participants practiced a familiarization session which consisted in a 60-s period during which participants freely moved the pointing tool in the front field. <xref ref-type="fig" rid="F1">Figure 1C</xref> illustrates the familiarization session. The virtual target was continuously placed (i.e., no need to press the trigger) on a 1-meter radius sphere centered with the camera position, on the axis of the pointing tool. Consequently, when participants moved their arm, the target was continuously placed at the corresponding position on the 1-meter radius sphere and they could hear the soundscape provided by the encoding corresponding to the processed target images within the camera&#x00027;s field of view. The virtual camera position was updated one time at the beginning of the 60-s timer.</p>
</sec>
</sec>
<sec>
<title>2.4. Data analysis</title>
<p>Statistical analysis were performed using R (version 3.6.1) (Team, <xref ref-type="bibr" rid="B65">2020</xref>). Localization performance during localization tests was assessed separately for azimuth and elevation dimensions, with error-based and regression-based metrics, both fitted with Linear mixed models (LMMs) in order to take into account participants as random factor. All trials of all participants were included in the models without averaging the response positions or the unsigned errors by participant. The LMMs were fitted using the <italic>lmerTest</italic> R-package (Kuznetsova et al., <xref ref-type="bibr" rid="B34">2017</xref>). We used an ANOVA with Satterthwaite approximation of degrees-of-freedom to estimate the effects. <italic>Post-hoc</italic> analysis were conducted using the <italic>emmeans</italic> R-package (version 1.7.4) (Lenth, <xref ref-type="bibr" rid="B35">2022</xref>) with Tukey HSD correction.</p>
<sec>
<title>2.4.1. Error-based metrics with unsigned and signed errors</title>
<p>Localization performance was assessed through unsigned and signed errors. The elevation signed errors and azimuth signed errors were computed as the difference between target position and response position in each trial. A negative elevation signed error indicated a downward shift, and a negative azimuth signed error indicated a shift to the left in the response. Only descriptive statistics were conducted on the signed errors. The unsigned errors were computed as the absolute value of the signed error for each trial. They were investigated using LMMs including Encoding (Noise or Pitch), Group (Monotonic or Harmonic) and Phase (Before or After the familiarization) as fixed factors. Therefore, the positions of the target were not included as a factor in the LMMs of the unsigned error. Participants were considered as random effect in both models.</p>
</sec>
<sec>
<title>2.4.2. Regression-based metrics with response positions</title>
<p>LMMs were also used for the analysis of the response positions. LMMs included Encoding (Noise or Pitch), Group (Monotonic or Harmonic), Phase (Before or After the familiarization), and Target position as fixed effects. The target elevation only, and the target azimuth only, were included in the elevation response LMM, and in the azimuth response LMM, respectively. Participants were considered as random effect in both models. We used the LMMs predictions to approximate the elevation and the azimuth gains and biases. The gains and biases were obtained by computing the trends (slopes) and intercepts of the models expressing the response position as a function of target position. Note that an optimal localization performance would be obtained with a gain value of 1.0 and a bias of 0.0&#x000B0;.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Performance in elevation localization</title>
<p>The elevation unsigned errors are depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>, left, all target positions combined. <xref ref-type="table" rid="T1">Table 1</xref> shows the elevation signed and unsigned errors for each Target elevation, Phase, Encoding and Group. The ANOVA on elevation unsigned errors showed a significant interaction effect of Phase &#x000D7; Encoding &#x000D7; Group [<italic>F</italic><sub>(1, 7556)</sub> = 6.23, <italic>p</italic> = 0.0126, <inline-formula><mml:math id="M1"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.0008]. <italic>Post-hoc</italic> analysis were conducted to investigate the interaction.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Unsigned error in elevation <bold>(left)</bold> and in azimuth <bold>(right)</bold> as a function of the encoding, all target positions combined. Mean unsigned errors (in degree) before (non-surrounded) and after (surrounded) are depicted separately for the Monotonic group (squares) and Harmonic group (circles) and for the three visual-to-auditory encodings: the Noise (blue), the Monotonic (orange) and the Harmonic (red) encodings. Error bars show standard error of the unsigned error.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0004.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Elevation signed error and unsigned error (in degree) for each encoding and target elevation, before, and after the familiarization session.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Encoding</bold></th>
<th valign="top" align="center"><bold>Target elevation</bold><break/> <bold>(degree)</bold></th>
<th valign="top" align="center" colspan="2"><bold>Elevation signed error (degree) Mean</bold> &#x000B1;<bold>standard deviation</bold></th>
<th valign="top" align="center" colspan="2"><bold>Elevation unsigned error (degree)</bold><break/> <bold>Mean</bold> &#x000B1;<bold>standard deviation</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td/>
<td valign="top" align="center"><bold>Before familiarization</bold></td>
<td valign="top" align="center"><bold>After familiarization</bold></td>
<td valign="top" align="center"><bold>Before familiarization</bold></td>
<td valign="top" align="center"><bold>After familiarization</bold></td>
</tr>
 <tr>
<td valign="middle" align="left" rowspan="5">Monotonic</td>
<td valign="top" align="center">&#x0002B;25</td>
<td valign="top" align="center">&#x02212;26.71 &#x000B1; 41.40</td>
<td valign="top" align="center">&#x02212;13.21 &#x000B1; 20.02</td>
<td valign="top" align="center">37.79 &#x000B1; 31.55</td>
<td valign="top" align="center">18.55 &#x000B1; 15.17</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;12.5</td>
<td valign="top" align="center">&#x02212;28.09 &#x000B1; 33.09</td>
<td valign="top" align="center">&#x02212;16.55 &#x000B1; 22.28</td>
<td valign="top" align="center">33.45 &#x000B1; 27.62</td>
<td valign="top" align="center">21.99 &#x000B1; 16.89</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;18.83 &#x000B1; 36.41</td>
<td valign="top" align="center">&#x02212;11.94 &#x000B1; 22.73</td>
<td valign="top" align="center">31.63 &#x000B1; 26.01</td>
<td valign="top" align="center">19.60 &#x000B1; 16.55</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;12.5</td>
<td valign="top" align="center">&#x02212;15.14 &#x000B1; 34.95</td>
<td valign="top" align="center">&#x02212;13.23 &#x000B1; 24.95</td>
<td valign="top" align="center">27.30 &#x000B1; 26.50</td>
<td valign="top" align="center">20.73 &#x000B1; 19.14</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;25</td>
<td valign="top" align="center">&#x02212;8.82 &#x000B1; 34.34</td>
<td valign="top" align="center">&#x02212;15.81 &#x000B1; 15.18</td>
<td valign="top" align="center">27.51 &#x000B1; 22.29</td>
<td valign="top" align="center">17.89 &#x000B1; 21.65</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Harmonic</td>
<td valign="top" align="center">&#x0002B;25</td>
<td valign="top" align="center">&#x02212;17.66 &#x000B1; 47.39</td>
<td valign="top" align="center">&#x02212;13.73 &#x000B1; 26.16</td>
<td valign="top" align="center">36.66 &#x000B1; 34.76</td>
<td valign="top" align="center">23.03 &#x000B1; 18.45</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;12.5</td>
<td valign="top" align="center">&#x02212;7.18 &#x000B1; 45.40</td>
<td valign="top" align="center">&#x02212;6.61 &#x000B1; 26.46</td>
<td valign="top" align="center">33.41 &#x000B1; 31.47</td>
<td valign="top" align="center">20.73 &#x000B1; 17.67</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;5.68 &#x000B1; 45.71</td>
<td valign="top" align="center">&#x02212;10.28 &#x000B1; 28.07</td>
<td valign="top" align="center">32.91 &#x000B1; 32.14</td>
<td valign="top" align="center">24.06 &#x000B1; 17.67</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;12.5</td>
<td valign="top" align="center">&#x02212;1.10 &#x000B1; 48.75</td>
<td valign="top" align="center">&#x02212;16.43 &#x000B1; 21.80</td>
<td valign="top" align="center">33.44 &#x000B1; 35.40</td>
<td valign="top" align="center">22.23 &#x000B1; 15.81</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;25</td>
<td valign="top" align="center">2.99 &#x000B1; 48.76</td>
<td valign="top" align="center">&#x02212;15.85 &#x000B1; 16.08</td>
<td valign="top" align="center">34.28 &#x000B1; 34.72</td>
<td valign="top" align="center">18.44 &#x000B1; 13.01</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Noise (Monotonic group)</td>
<td valign="top" align="center">&#x0002B;25</td>
<td valign="top" align="center">&#x02212;42.69 &#x000B1; 52.91</td>
<td valign="top" align="center">&#x02212;33.92 &#x000B1; 27.00</td>
<td valign="top" align="center">56.49 &#x000B1; 37.73</td>
<td valign="top" align="center">37.54 &#x000B1; 21.64</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;12.5</td>
<td valign="top" align="center">&#x02212;32.93 &#x000B1; 45.45</td>
<td valign="top" align="center">&#x02212;23.24 &#x000B1; 23.54</td>
<td valign="top" align="center">43.63 &#x000B1; 35.23</td>
<td valign="top" align="center">28.00 &#x000B1; 17.56</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;25.49 &#x000B1; 43.51</td>
<td valign="top" align="center">&#x02212;12.98 &#x000B1; 24.78</td>
<td valign="top" align="center">36.61 &#x000B1; 34.62</td>
<td valign="top" align="center">23.09 &#x000B1; 15.73</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;12.5</td>
<td valign="top" align="center">&#x02212;20.61 &#x000B1; 43.35</td>
<td valign="top" align="center">&#x02212;7.09 &#x000B1; 22.23</td>
<td valign="top" align="center">32.93 &#x000B1; 34.88</td>
<td valign="top" align="center">19.02 &#x000B1; 13.46</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;25</td>
<td valign="top" align="center">&#x02212;8.40 &#x000B1; 47.90</td>
<td valign="top" align="center">3.11 &#x000B1; 22.32</td>
<td valign="top" align="center">31.30 &#x000B1; 37.15</td>
<td valign="top" align="center">16.86 &#x000B1; 14.90</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Noise (Harmonic group)</td>
<td valign="top" align="center">&#x0002B;25</td>
<td valign="top" align="center">&#x02212;34.74 &#x000B1; 46.47</td>
<td valign="top" align="center">&#x02212;33.67 &#x000B1; 28.35</td>
<td valign="top" align="center">47.86 &#x000B1; 32.71</td>
<td valign="top" align="center">36.96 &#x000B1; 23.87</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;12.5</td>
<td valign="top" align="center">&#x02212;24.89 &#x000B1; 45.12</td>
<td valign="top" align="center">&#x02212;21.25 &#x000B1; 27.08</td>
<td valign="top" align="center">40.59 &#x000B1; 31.66</td>
<td valign="top" align="center">27.87 &#x000B1; 20.16</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;15.68 &#x000B1; 42.71</td>
<td valign="top" align="center">&#x02212;12.20 &#x000B1; 25.02</td>
<td valign="top" align="center">32.41 &#x000B1; 31.87</td>
<td valign="top" align="center">22.15 &#x000B1; 16.81</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;12.5</td>
<td valign="top" align="center">&#x02212;7.03 &#x000B1; 44.28</td>
<td valign="top" align="center">&#x02212;3.40 &#x000B1; 25.66</td>
<td valign="top" align="center">28.84 &#x000B1; 34.27</td>
<td valign="top" align="center">19.68 &#x000B1; 16.76</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;25</td>
<td valign="top" align="center">0.69 &#x000B1; 43.84</td>
<td valign="top" align="center">8.72 &#x000B1; 25.46</td>
<td valign="top" align="center">27.84 &#x000B1; 33.81</td>
<td valign="top" align="center">20.10 &#x000B1; 17.85</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The elevation response positions are depicted in <xref ref-type="fig" rid="F5">Figure 5</xref>. The ANOVA showed a significant interaction effect of Phase &#x000D7; Target Elevation &#x000D7; Encoding [<italic>F</italic><sub>(1, 7548)</sub> = 38.84, <italic>p</italic> &#x0003C; 0.0001, <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.005]. We conducted <italic>post-hoc</italic> analysis to investigate the elevation gain (the trend of the model) and bias (the intercept of the model) depending on the Phase and the Encoding. Although the interaction effect of Phase &#x000D7; Target Elevation &#x000D7; Encoding &#x000D7; Group was not significant [<italic>F</italic><sub>(1, 7548)</sub> = 0.50, <italic>p</italic> = 0.48, <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.00007], <italic>post-hoc</italic> analysis were also performed for a control purpose in order to check for differences between the Monotonic and Harmonic groups. The elevation response positions are provided separately for each participant in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Figures S1</xref>, <xref ref-type="supplementary-material" rid="SM2">S2</xref>.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Elevation response position as a function of target elevation in the Monotonic group <bold>(A)</bold> and the Harmonic group <bold>(B)</bold>. Mean elevation response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the three visual-to-auditory encodings: the Noise (blue squares), the Monotonic (orange circles) and the Harmonic (red circles) encodings. Error bars show standard error of elevation response position. Solid lines represent the elevation gains with the Noise (blue), the Monotonic (orange) and Harmonic (red) encodings. Black dashed lines indicate the optimal elevation gain 1.0.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0005.tif"/>
</fig>
<sec>
<title>3.1.1. Elevation localization performance before the familiarization</title>
<p>Before the practice of the familiarization session, and depending on the encoding, the elevation unsigned errors were comprised between 31.54 &#x000B1; 27.19&#x000B0; and 40.19 &#x000B1; 37.02&#x000B0;. For the Monotonic group, the elevation unsigned errors were significantly lower with the Monotonic encoding (<italic>M</italic> = 31.54, <italic>SD</italic> = 27.19) than with the Noise encoding (<italic>M</italic> = 40.19, <italic>SD</italic> = 37.03) [<italic>t</italic><sub>(7556)</sub> = 7.457, <italic>p</italic> &#x0003C; 0.0001], suggesting a lower accuracy with the Noise encoding. There was no significant difference in the Harmonic group regarding the elevation unsigned error between the Harmonic encoding (<italic>M</italic> = 34.14, <italic>SD</italic> = 33.69) and the Noise encoding (<italic>M</italic> = 35.51, <italic>SD</italic> = 33.69).</p>
<p>The elevation response positions before the familiarization are depicted in the left panels of the <xref ref-type="fig" rid="F5">Figures 5A</xref>, <xref ref-type="fig" rid="F5">B</xref> for the Monotonic group and the Harmonic group, respectively. The elevation gains were significantly different from 0.0 for all encodings: 0.62 [95% CI = [0.5, 0.74], <italic>t</italic><sub>(7548)</sub> = 10.118, <italic>p</italic> &#x0003C; 0.0001] with the Harmonic encoding, 0.61 [95% CI = [0.49, 0.73], <italic>t</italic><sub>(7548)</sub> = 9.94, <italic>p</italic> &#x0003C; 0.0001] with the Monotonic encoding, and 0.29 [95% CI = [0.17, 0.41], <italic>t</italic><sub>(7548)</sub> = 4.728, <italic>p</italic> &#x0003C; 0.0001] and 0.35 [95% CI = [0.23, 0.47], <italic>t</italic><sub>(7548)</sub> = 5.746, <italic>p</italic> &#x0003C; 0.0001] with the Noise encoding of the Harmonic and Monotonic groups, respectively. It suggests that participants could discriminate different elevation positions with the three encodings even before the familiarization.</p>
<p>However, elevation gains were significantly lower than the optimal gain 1.0 with all encodings: with the Harmonic encoding [<italic>t</italic><sub>(7548)</sub> = &#x02212;6.173, <italic>p</italic> &#x0003C; 0.0001], with the Monotonic encoding [<italic>t</italic><sub>(7548)</sub> = &#x02212;6.351, <italic>p</italic> &#x0003C; 0.0001], and with the Noise encoding of the Harmonic group [<italic>t</italic><sub>(7548)</sub> = &#x02212;11.562, <italic>p</italic> &#x0003C; 0.0001] and of the Monotonic group [<italic>t</italic><sub>(7548)</sub> = &#x02212;10.544, <italic>p</italic> &#x0003C; 0.0001]. It depicts a situation where although some variations in elevation seemed to be perceived with the three encodings, participants had difficulties to estimate it before the familiarization.</p>
<p>The participants tended to localize the elevation with a higher performance with the Harmonic or Monotonic encoding than with the Noise encoding. Indeed, the participants from the Harmonic group showed a higher elevation gain with the Harmonic encoding than with the Noise encoding with a significant difference of 0.33 [<italic>t</italic><sub>(7548)</sub> = &#x02212;3.811, <italic>p</italic>= 0.0008]. For the Monotonic group, the elevation gain was also significantly higher with the Monotonic encoding than with the Noise encoding with a difference of 0.26 [<italic>t</italic><sub>(7548)</sub> = &#x02212;2.97, <italic>p</italic>= 0.016]. There was no significant difference regarding the elevation gain between the Harmonic and the Monotonic encodings.</p>
<p>The participants tended to underestimate the elevation position of the targets with the three encodings, as indicated by downward bias and negative elevation errors. In the Monotonic group, the elevation bias were &#x02212;26.02&#x000B0; (95% CI = [&#x02212;31.9, &#x02212;20.16]) with the Noise encoding and &#x02212;19.52&#x000B0; (95% CI = [&#x02212;25.4, &#x02212;13.65]) with the Monotonic encoding. In the Harmonic group, the elevation bias with the Noise encoding and with the Harmonic encoding were &#x02212;16.33&#x000B0; (95% CI = [&#x02212;22.2, &#x02212;10.47]) and &#x02212;5.73&#x000B0; (95% CI = [&#x02212;11.6, 0.14]), respectively. With the exception of the Harmonic encoding for which there was just a trend [<italic>t</italic><sub>(44.9)</sub> = 1.97, <italic>p</italic>= 0.055], all the elevation bias mentioned above were significantly negative [all |<italic>t</italic><sub>(44.9)</sub>| &#x0003E; 5.61, all <italic>p</italic> &#x0003C; 0.0001].</p>
<p>To sum up, participants appeared partially able to perceive a variation of the elevation position of the target with the three encodings before the audio-motor familiarization. Interestingly, participants seemed better able to localize the elevation with the Harmonic and Monotonic encodings.</p>
</sec>
<sec>
<title>3.1.2. Elevation localization performance after the familiarization</title>
<p>After the familiarization, the elevation unsigned errors were significantly higher with the Noise encoding than with the 2 pitch-based encodings (Monotonic or Harmonic encodings). With the Noise encoding, the elevation unsigned errors were 24.90 &#x000B1; 18.40&#x000B0; in the Monotonic group and 25.35 &#x000B1; 20.31&#x000B0; in the Harmonic group. With the Harmonic and Monotonic encodings, the elevation unsigned errors were 21.70 &#x000B1; 16.72&#x000B0; and 19.75 &#x000B1; 16.25&#x000B0; respectively. In the Monotonic group, the elevation unsigned errors were significantly lower with the Monotonic encoding (<italic>M</italic> = 19.75, <italic>SD</italic> = 16.25) than with the Noise encoding (<italic>M</italic> = 24.90, <italic>SD</italic> = 18.40) [<italic>t</italic><sub>(7556)</sub> = 4.44, <italic>p</italic> &#x0003C; 0.0001]. Unlike before the familiarization, the difference was also significant in the Harmonic group. The elevation unsigned errors with the Harmonic encoding (<italic>M</italic> = 21.70, <italic>SD</italic> = 16.72) were lower than with the Noise encoding (<italic>M</italic> = 25.35, <italic>SD</italic> = 20.31), [<italic>t</italic><sub>(7556)</sub> &#x0003D; 3.15, <italic>p</italic>= 0.0016]. Interestingly, the elevation unsigned errors significantly decreased after the familiarization with all the encodings [all |<italic>t</italic><sub>(7556)</sub>| &#x0003E; 8.75, all <italic>p</italic> &#x0003C; 0.0001], suggesting that participants localized more accurately the elevation after the familiarization.</p>
<p>The elevation response positions after the familiarization are depicted in the <xref ref-type="fig" rid="F5">Figures 5A</xref>, <xref ref-type="fig" rid="F5">B</xref> for the Monotonic and Harmonic groups, respectively. After the familiarization, the elevation gains were still significantly higher than 0.0 [all |<italic>t</italic><sub>(7548)</sub>| &#x0003E; 2.9152, all <italic>p</italic> &#x0003C; 0.0036] with all encodings in the 2 groups. The elevation gains were 1.112 (95% CI = [0.99, 1.23]) with the Harmonic encoding and 1.015 (95% CI = [0.89, 1.14]) with the Monotonic encoding. For the participants of the Harmonic group and the Monotonic group, the elevation gains with the Noise encoding were 0.179 (95% CI = [0.06, 0.30]), and 0.278 (95% CI = [0.16, 0.40]), respectively.</p>
<p>The elevation gains were significantly higher with the Harmonic and Monotonic encodings than with the Noise encoding. We measured a difference of 0.74 [<italic>t</italic><sub>(7548)</sub> = &#x02212;8.49, <italic>p</italic> &#x0003C; 0.0001] in the Harmonic group and a difference of 0.93 [<italic>t</italic><sub>(7548)</sub> = &#x02212;10.75, <italic>p</italic> &#x0003C; 0.0001] in the Monotonic group. Inter-group analysis showed that the difference in elevation gain between the Monotonic and the Harmonic encodings did not significantly differ [<italic>t</italic><sub>(7548)</sub> = 1.12, <italic>p</italic>= 0.95].</p>
<p>The elevation gains with the Harmonic and the Monotonic encodings significantly improved after the familiarization to get closer than the optimal gain 1.0. With the Harmonic encoding, the elevation gain significantly increased from 0.62 to 1.112 [<italic>t</italic><sub>(7548)</sub> = 5.66, <italic>p</italic> &#x0003C; 0.0001] after which it was not significantly different from the optimal gain 1.0 [<italic>t</italic><sub>(7548)</sub> = 1.832, <italic>p</italic>= 0.067]. With the Monotonic encoding, the elevation gain significantly increased from 0.61 to 1.015 [<italic>t</italic><sub>(7548)</sub> = 4.665, <italic>p</italic> &#x0003C; 0.0001], and was also no more significantly different from the optimal gain 1.0 [<italic>t</italic><sub>(7548)</sub> = 0.246, <italic>p</italic>= 0.806]. However with the Noise encoding in both groups, the familiarization did not improve the elevation gains. In the Harmonic and Monotonic groups, the elevation gains decreased from 0.29 to 0.179 and from 0.35 to 0.278, respectively, but, as previously reported, the decreases were not significant.</p>
<p>Participants kept tending to underestimate the elevation position of the targets with all three encodings, as indicated by persistent negative bias. In the Monotonic group, the elevation bias with the Noise encoding and with the Monotonic encoding were &#x02212;14.82&#x000B0; (95% CI = [&#x02212;20.7, &#x02212;8.96]) and &#x02212;14.15&#x000B0; (95% CI = [&#x02212;20.0, &#x02212;8.28]), respectively. In the Harmonic group, the elevation bias with the Noise encoding and with the Harmonic encoding were &#x02212;12.36&#x000B0; (95% CI = [&#x02212;18.2, &#x02212;6.49]) and &#x02212;12.58&#x000B0; (95% CI = [&#x02212;18.4, &#x02212;6.72]), respectively. All the elevation bias were significantly negative [all |<italic>t</italic><sub>(44.9)</sub>| &#x0003E; 4.24, all <italic>p</italic> &#x0003C; 0.0001].</p>
<p>To sum up, after the familiarization, the perception of elevation with the Harmonic and Monotonic encodings improved with elevation gains getting closer to the optimal gain. However, the familiarization did not induce any significant improvement in the perception of elevation with the Noise encoding, with persistent low elevation gains in both groups. Additionally, the underestimation elevation bias decreased with the Monotonic and Noise encodings, but not with the Harmonic encoding for which it increased.</p>
</sec>
</sec>
<sec>
<title>3.2. Performance in azimuth localization</title>
<p>The azimuth unsigned errors are depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>, right, all target positions combined. <xref ref-type="table" rid="T2">Table 2</xref> shows the azimuth signed and unsigned errors for each Target azimuth, Phase, Encoding and Group. The ANOVA on azimuth unsigned errors showed a significant interaction effect of Phase &#x000D7; Encoding [<italic>F</italic><sub>(1, 7556)</sub> = 5.15, <italic>p</italic>= 0.023, <inline-formula><mml:math id="M4"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.00068], but the interaction including the group was not significant.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Azimuth signed error and unsigned error (in degree) for each encoding and target azimuth, before, and after the familiarization session.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Encoding</bold></th>
<th valign="top" align="center"><bold>Target azimuth (degree)</bold></th>
<th valign="top" align="center" colspan="2"><bold>Azimuth signed error (degree) Mean</bold> &#x000B1;<bold>standard deviation</bold></th>
<th valign="top" align="center" colspan="2"><bold>Azimuth unsigned error (degree)</bold><break/> <bold>Mean</bold> &#x000B1;<bold>standard deviation</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td/>
<td valign="top" align="center"><bold>Before familiarization</bold></td>
<td valign="top" align="center"><bold>After familiarization</bold></td>
<td valign="top" align="center"><bold>Before familiarization</bold></td>
<td valign="top" align="center"><bold>After familiarization</bold></td>
</tr>
 <tr>
<td valign="middle" align="left" rowspan="5">Monotonic</td>
<td valign="top" align="center">&#x0002B;40</td>
<td valign="top" align="center">17.79 &#x000B1; 23.73</td>
<td valign="top" align="center">4.39 &#x000B1; 18.25</td>
<td valign="top" align="center">23.16 &#x000B1; 18.5</td>
<td valign="top" align="center">13.97 &#x000B1; 12.5</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;20</td>
<td valign="top" align="center">25.34 &#x000B1; 25.50</td>
<td valign="top" align="center">12.06 &#x000B1; 18.35</td>
<td valign="top" align="center">28.71 &#x000B1; 21.61</td>
<td valign="top" align="center">17.14 &#x000B1; 13.69</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;7.75 &#x000B1; 25.39</td>
<td valign="top" align="center">&#x02212;5.02 &#x000B1; 19.33</td>
<td valign="top" align="center">17.95 &#x000B1; 19.52</td>
<td valign="top" align="center">15.21 &#x000B1; 12.90</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;20</td>
<td valign="top" align="center">&#x02212;25.25 &#x000B1; 21.81</td>
<td valign="top" align="center">&#x02212;15.02 &#x000B1; 18.10</td>
<td valign="top" align="center">26.93 &#x000B1; 19.69</td>
<td valign="top" align="center">18.68 &#x000B1; 14.27</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;40</td>
<td valign="top" align="center">&#x02212;16.27 &#x000B1; 20.33</td>
<td valign="top" align="center">&#x02212;5.34 &#x000B1; 19.40</td>
<td valign="top" align="center">20.68 &#x000B1; 15.79</td>
<td valign="top" align="center">15.23 &#x000B1; 13.11</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Harmonic</td>
<td valign="top" align="center">&#x0002B;40</td>
<td valign="top" align="center">23.26 &#x000B1; 19.59</td>
<td valign="top" align="center">5.43 &#x000B1; 21.49</td>
<td valign="top" align="center">24.9 &#x000B1; 17.45</td>
<td valign="top" align="center">15.83 &#x000B1; 15.48</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;20</td>
<td valign="top" align="center">27.17 &#x000B1; 20.26</td>
<td valign="top" align="center">12.98 &#x000B1; 18.80</td>
<td valign="top" align="center">28.17 &#x000B1; 19.61</td>
<td valign="top" align="center">17.23 &#x000B1; 14.98</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;6.89 &#x000B1; 23.03</td>
<td valign="top" align="center">&#x02212;6.02 &#x000B1; 18.49</td>
<td valign="top" align="center">16.89 &#x000B1; 17.36</td>
<td valign="top" align="center">14.15 &#x000B1; 12.97</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;20</td>
<td valign="top" align="center">&#x02212;32.76 &#x000B1; 23.38</td>
<td valign="top" align="center">&#x02212;16.61 &#x000B1; 19.93</td>
<td valign="top" align="center">33.16 &#x000B1; 22.81</td>
<td valign="top" align="center">21.6 &#x000B1; 14.35</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;40</td>
<td valign="top" align="center">&#x02212;27.45 &#x000B1; 23.93</td>
<td valign="top" align="center">&#x02212;7.03 &#x000B1; 22.08</td>
<td valign="top" align="center">29.25 &#x000B1; 21.68</td>
<td valign="top" align="center">16.68 &#x000B1; 16.05</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Noise (Monotonic group)</td>
<td valign="top" align="center">&#x0002B;40</td>
<td valign="top" align="center">23.01 &#x000B1; 25.34</td>
<td valign="top" align="center">7.27 &#x000B1; 16.58</td>
<td valign="top" align="center">27.91 &#x000B1; 19.79</td>
<td valign="top" align="center">14.35 &#x000B1; 11.01</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;20</td>
<td valign="top" align="center">28.35 &#x000B1; 29.17</td>
<td valign="top" align="center">11.96 &#x000B1; 19.59</td>
<td valign="top" align="center">30.95 &#x000B1; 26.38</td>
<td valign="top" align="center">18.11 &#x000B1; 14.06</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;8.89 &#x000B1; 24.82</td>
<td valign="top" align="center">&#x02212;8.47 &#x000B1; 14.79</td>
<td valign="top" align="center">16.78 &#x000B1; 20.31</td>
<td valign="top" align="center">12.04 &#x000B1; 12.05</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;20</td>
<td valign="top" align="center">&#x02212;31.27 &#x000B1; 25.61</td>
<td valign="top" align="center">&#x02212;21.37 &#x000B1; 17.24</td>
<td valign="top" align="center">33.42 &#x000B1; 22.71</td>
<td valign="top" align="center">22.32 &#x000B1; 15.98</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;40</td>
<td valign="top" align="center">&#x02212;21.22 &#x000B1; 24.10</td>
<td valign="top" align="center">&#x02212;11.55 &#x000B1; 16.88</td>
<td valign="top" align="center">25.53 &#x000B1; 19.45</td>
<td valign="top" align="center">16.82 &#x000B1; 11.59</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="5">Noise (Harmonic group)</td>
<td valign="top" align="center">&#x0002B;40</td>
<td valign="top" align="center">24.13 &#x000B1; 23.85</td>
<td valign="top" align="center">9.86 &#x000B1; 24.77</td>
<td valign="top" align="center">28.65 &#x000B1; 18.12</td>
<td valign="top" align="center">19.17 &#x000B1; 18.49</td>
</tr>
 <tr>

<td valign="top" align="center">&#x0002B;20</td>
<td valign="top" align="center">29.80 &#x000B1; 24.77</td>
<td valign="top" align="center">15.02 &#x000B1; 17.67</td>
<td valign="top" align="center">31.78 &#x000B1; 22.16</td>
<td valign="top" align="center">18.94 &#x000B1; 13.35</td>
</tr>
 <tr>

<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02212;11.84 &#x000B1; 27.02</td>
<td valign="top" align="center">&#x02212;7.04 &#x000B1; 20.76</td>
<td valign="top" align="center">21.10 &#x000B1; 20.57</td>
<td valign="top" align="center">16.34 &#x000B1; 14.57</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;20</td>
<td valign="top" align="center">&#x02212;36.95 &#x000B1; 20.92</td>
<td valign="top" align="center">&#x02212;22.36 &#x000B1; 18.41</td>
<td valign="top" align="center">37.28 &#x000B1; 20.31</td>
<td valign="top" align="center">25.06 &#x000B1; 14.48</td>
</tr>
 <tr>

<td valign="top" align="center">&#x02212;40</td>
<td valign="top" align="center">&#x02212;30.29 &#x000B1; 17.83</td>
<td valign="top" align="center">&#x02212;12.40 &#x000B1; 23.27</td>
<td valign="top" align="center">30.71 &#x000B1; 17.09</td>
<td valign="top" align="center">20.71 &#x000B1; 16.28</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The azimuth response positions are depicted in <xref ref-type="fig" rid="F6">Figure 6</xref>. The ANOVA yielded a significant interaction effect of Phase &#x000D7; Target Azimuth &#x000D7; Encoding [<italic>F</italic><sub>(1, 7548)</sub> = 12.69, <italic>p</italic>= 0.0004, <inline-formula><mml:math id="M5"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.00005]. <italic>Post-hoc</italic> analysis were conducted to investigate the azimuth gain (the trend of the model) and bias (the intercept of the model) depending on the Phase and the Encoding. Although the interaction effect of Phase &#x000D7; Target Elevation &#x000D7; Encoding &#x000D7; Group was not significant [<italic>F</italic><sub>(1, 7548)</sub> = 1.64, <italic>p</italic> = 0.20, <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> = 0.0002], we conducted <italic>post-hoc</italic> analysis to check for differences between the Monotonic and Harmonic groups for a control purpose. The azimuth response positions are provided separately for each participant in the <xref ref-type="supplementary-material" rid="SM3">Supplementary Figures S3</xref>, <xref ref-type="supplementary-material" rid="SM4">S4</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Azimuth response position as a function of target azimuth in the Monotonic group <bold>(A)</bold> and the Harmonic group <bold>(B)</bold>. Mean azimuth response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the three visual-to-auditory encodings: the Noise (blue squares), the Monotonic (orange circles) and the Harmonic (red circles) encodings. Error bars show standard error of azimuth response position. Solid lines represent the azimuth gains with the Noise (blue), the Monotonic (orange) and Harmonic (red) encodings. Black dashed lines indicate the optimal azimuth gain 1.0.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0006.tif"/>
</fig>
<sec>
<title>3.2.1. Azimuth localization performance before the familiarization</title>
<p>Before the practice of the familiarization session, and depending on the encoding, the azimuth unsigned errors were comprised between 23.48 &#x000B1; 19.48&#x000B0; and 29.91 &#x000B1; 20.38&#x000B0;. In the Monotonic group, the azimuth unsigned errors were significantly lower with the Monotonic encoding (<italic>M</italic> = 23.48, <italic>SD</italic> = 19.48) than with the Noise encoding (<italic>M</italic> = 26.92, <italic>SD</italic> = 22.58) [<italic>t</italic><sub>(7556)</sub> = 4.64, <italic>p</italic> &#x0003C; 0.0001]. The azimuth unsigned errors in the Harmonic group were also significantly lower [<italic>t</italic><sub>(7556)</sub> = 4.73, <italic>p</italic> &#x0003C; 0.0001] with the Harmonic encoding (<italic>M</italic> = 26.41, <italic>SD</italic> = 20.63) than with the Noise encoding (<italic>M</italic> = 29.91, <italic>SD</italic> = 33.69).</p>
<p>The azimuth response positions over all participants before the familiarization are depicted in the left panels of the <xref ref-type="fig" rid="F6">Figures 6A</xref>, <xref ref-type="fig" rid="F6">B</xref> for the Monotonic and Harmonic groups, respectively. Before the familiarization, the participants were able to interpret soundscapes to localize the target azimuth. First, the participants perceived different azimuth positions. Indeed, azimuth gains were significantly different from 0.0 with all encodings: 1.81 (95% CI = [1.75, 1.87], <italic>t</italic><sub>(7548)</sub> = 70.397, <italic>p</italic> &#x0003C; 0.0001) with the Harmonic encoding, 1.59 [95% CI = [1.54, 1.65], <italic>t</italic><sub>(7548)</sub> = 61.996, <italic>p</italic> &#x0003C; 0.0001] with the Monotonic encoding, and 1.88 [95% CI = [1.82, 1.93], <italic>t</italic><sub>(7548)</sub> = 73.0624, <italic>p</italic> &#x0003C; 0.0001] and 1.74 [95% CI = [1.68, 1.80], <italic>t</italic><sub>(7548)</sub> = 67.710, <italic>p</italic> &#x0003C; 0.0001] with the Noise encoding for the participants in the Harmonic and Monotonic groups, respectively.</p>
<p>Interestingly, the azimuth gains were significantly higher than the optimal gain (i.e., higher than 1.0) with the Harmonic encoding [<italic>t</italic><sub>(7548)</sub> = 31.492, <italic>p</italic> &#x0003C; 0.0001], the Monotonic encoding [<italic>t</italic><sub>(7548)</sub> = 23.091, <italic>p</italic> &#x0003C; 0.0001], and the Noise encoding in the Harmonic group [<italic>t</italic><sub>(7548)</sub> = 34.156, <italic>p</italic> &#x0003C; 0.0001] and the Monotonic group [<italic>t</italic><sub>(7548)</sub> = 28.804, <italic>p</italic> &#x0003C; 0.0001]. These gains higher than the optimal gain reflect a lateral overestimation (i.e., left targets localized too much on the left and right targets localized too much on the right) that can be seen with the three encodings.</p>
<p>In the Monotonic group, the overestimation observed with the Noise encoding was significantly higher than with the Monotonic encoding [<italic>t</italic><sub>(7548)</sub> = 4.04, <italic>p</italic>= 0.0003]. In the Harmonic group, the overestimation with the Noise encoding compared to the Harmonic encoding was also higher but not significantly. Inter-group comparison of the azimuth gains obtained with the Noise encoding shows a small but significant higher azimuth gain in the Harmonic group [difference of 0.14: <italic>t</italic><sub>(7548)</sub> = 3.784, <italic>p</italic>= 0.0009]. As inter-group comparison, we also observed a slight but significant higher overestimation pattern with the Harmonic encoding in comparison with the Monotonic encoding [difference of 0.22: <italic>t</italic><sub>(7548)</sub> = 5.94, <italic>p</italic> &#x0003C; 0.0001].</p>
<p>Another interesting result is the tendency to show a left shift as indicated by negative azimuth bias with the three encodings. With the Noise encoding in the Harmonic group, the leftward azimuth bias of &#x02212;5.03&#x000B0; was significant [<italic>t</italic><sub>(47.2)</sub> = 2.84, <italic>p</italic>= 0.0066]. However, leftward azimuth bias with the other encodings were not significantly different from 0.0&#x000B0; (Harmonic group, Noise encoding: &#x02212;3.23&#x000B0;; Monotonic group, Noise encoding: &#x02212;2.0&#x000B0;; Monotonic group, Monotonic encoding: &#x02212;1.233&#x000B0;).</p>
<p>In summary, before the familiarization and with the three encodings, participants were able to localize the azimuth of the target accurately with a tendency to overestimate the lateral eccentricity and a tendency to point too much on the left.</p>
</sec>
<sec>
<title>3.2.2. Azimuth localization performance after the familiarization</title>
<p>After the participants practiced the familiarization session, and depending on the encoding, the azimuth unsigned errors were comprised between 16.05 &#x000B1; 13.38&#x000B0; and 20.05 &#x000B1; 15.77&#x000B0;. In the Harmonic group, the azimuth unsigned errors were significantly lower with the Harmonic encoding (<italic>M</italic> = 17.16, <italic>SD</italic> = 14.97) than with the Noise encoding (<italic>M</italic> = 20.05, <italic>SD</italic> = 15.77) [<italic>t</italic><sub>(7556)</sub> = 3.91, <italic>p</italic>= 0.0001]. The azimuth unsigned errors were not significantly different anymore between the Monotonic encoding (<italic>M</italic> = 16.05, <italic>SD</italic> = 13.38) and the Noise encoding (<italic>M</italic> = 16.73, <italic>SD</italic> = 13.50). Importantly, the azimuth unsigned errors significantly decreased after the familiarization session for all three encodings [all |<italic>t</italic><sub>(7556)</sub>| &#x0003E; 10.06, all <italic>p</italic> &#x0003C; 0.0001], suggesting that participants localized more accurately the azimuth after the familiarization.</p>
<p>The azimuth response positions after the familiarization are depicted in the right panels of the <xref ref-type="fig" rid="F6">Figures 6A, B</xref> for the Monotonic and Harmonic groups, respectively. As expected, after the familiarization, participants were still able to localize different azimuth positions by interpreting soundscapes. Azimuth gains were still significantly different from 0.0 with the Harmonic encoding [1.27, 95% CI = [1.22, 1.32], <italic>t</italic><sub>(7548)</sub> = 49.51, <italic>p</italic> &#x0003C; 0.0001], with the Monotonic encoding [1.23, 95% CI = [1.18, 1.28], <italic>t</italic><sub>(7548)</sub> = 47.96, <italic>p</italic> &#x0003C; 0.0001], and with the Noise encoding for the participants in the Harmonic group [1.41, 95% CI = [1.36, 1.46], <italic>t</italic><sub>(7548)</sub> = 54.84, <italic>p</italic> &#x0003C; 0.0001] and in the Monotonic group [1.35, 95% CI = [1.30, 1.41], <italic>t</italic><sub>(7548)</sub> = 52.711, <italic>p</italic> &#x0003C; 0.0001], respectively.</p>
<p>The overestimation pattern was still present, as indicated by azimuth gains still significantly higher than the optimal gain 1.0 with all encodings: the Harmonic encoding [<italic>t</italic><sub>(7548)</sub> = 10.61, <italic>p</italic> &#x0003C; 0.0001], the Monotonic encoding [<italic>t</italic><sub>(7548)</sub> = 9.06, <italic>p</italic> &#x0003C; 0.0001], the Noise encoding in the Harmonic group [<italic>t</italic><sub>(7548)</sub> = 15.93, <italic>p</italic> &#x0003C; 0.0001] and in the Monotonic group [<italic>t</italic><sub>(7548)</sub> = 13.81, <italic>p</italic> &#x0003C; 0.0001].</p>
<p>Although the lateral overestimation was still significant, it significantly decreased compared to the same localization test before the familiarization. Indeed, the azimuth gains decreased and reached values closer than the optimal gain 1.0 with the 3 encodings. There were significant decreases in azimuth gains of a magnitude of 0.54 [<italic>t</italic><sub>(7548)</sub> = 14.77, <italic>p</italic> &#x0003C; 0.0001] and 0.36 [<italic>t</italic><sub>(7548)</sub> = 9.92, <italic>p</italic> &#x0003C; 0.0001] with the Harmonic and Monotonic encodings, respectively. The decreases in azimuth gains with the Noise encoding in the Harmonic and the Monotonic groups were also significant with a decrease of a magnitude of, respectively, 0.47 [<italic>t</italic><sub>(7548)</sub> = 12.897, <italic>p</italic> &#x0003C; 0.0001] and 0.39 [<italic>t</italic><sub>(7548)</sub> = 10.61, <italic>p</italic> &#x0003C; 0.0001].</p>
<p>Additionally, after the familiarization, participants tended to localize the azimuth with a higher performance with the Harmonic and Monotonic encodings than with the Noise encoding. This is suggested by a more pronounced lateral overestimation with the Noise encoding in both groups: the azimuth gains were 0.14 higher [<italic>t</italic><sub>(7548)</sub> = 3.76, <italic>p</italic>= 0.0042] and 0.12 higher [<italic>t</italic><sub>(7548)</sub> = 3.36, <italic>p</italic>= 0.018] with the Noise encoding in comparison with the Harmonic and Monotonic encodings, respectively.</p>
<p>The slight tendency to show a left shift bias in azimuth was still present with the three encodings. With the Noise encoding in the Monotonic group, the leftward azimuth bias of &#x02212;4.43&#x000B0; was significant [<italic>t</italic><sub>(47.2)</sub> = 2.502, <italic>p</italic>= 0.0159], but in the Harmonic group the bias of &#x02212;3.38&#x000B0; was just a tendency [<italic>t</italic><sub>(47.2)</sub> = 1.911, <italic>p</italic>= 0.0621]. The leftward azimuth bias with the Harmonic encoding (&#x02212;2.25&#x000B0;) and Monotonic encoding (&#x02212;1.79&#x000B0;) were also not significant.</p>
<p>To sum up the accuracy in azimuth localization, participants were able to localize target azimuths accurately even before the audio-motor familiarization. After the familiarization, the accuracy increased with a decrease in both the tendency to overestimate the lateral position of lateral targets and the tendency to point too much on the left.</p>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>In this study, we investigated the early stage of use of visual-to-auditory SSDs based on the creation of a VAS (Virtual Acoustic Space) for object localization in a virtual environment. Based on soundscapes created using non-individualized HRTFs, we investigated blindfolded participants&#x00027; abilities to localize a virtual target with three encoding schemes: one conveying elevation with spatialization only (Noise encoding), and two conveying elevation with spatialization and pitch modulation (Monotonic and Harmonic encodings). The two pitch-based encodings varied regarding the sound spectrum complexity: one narrowband with monotones (Monotonic encoding) and one more complex with 2 additional octaves (Harmonic encoding). In order to compare the localization abilities for the azimuth and the elevation with the different visual-to-auditory encodings, we collected the response positions and angular errors of the participants during a task consisting in the localization of a virtual target placed at different azimuths and elevations in their front-field.</p>
<sec>
<title>4.1. Elevation localization abilities using the visual-to-auditory encodings</title>
<sec>
<title>4.1.1. Elevation localization performance only based on non-individualized HRTFs is impaired</title>
<p>With the spatialization-based only encoding (Noise encoding), the target was localized before the familiarization with an elevation unsigned error between 27.84 &#x000B1; 33.81&#x000B0; and 56.49 &#x000B1; 37.73&#x000B0;. After the familiarization, the elevation unsigned errors decreased to reach values comprised between 16.86 &#x000B1; 14.90&#x000B0; and 37.54 &#x000B1; 21.64&#x000B0;. As a comparison, in Mendon&#x000E7;a et al. (<xref ref-type="bibr" rid="B42">2013</xref>) where the same HRTFs database was used with a white noise sound, the mean elevation unsigned error of participants was 29.3&#x000B0; before practicing a training. The elevation unsigned errors in Geronazzo et al. (<xref ref-type="bibr" rid="B21">2018</xref>) without any familiarization and with a white noise sound were comprised between 15.58 &#x000B1; 12.47&#x000B0; and 33.75 &#x000B1; 16.17&#x000B0; depending on participants, which is comparable to our results after the familiarization. However, as shown by elevation gains below 0.4 before or after familiarization, the participants had difficulties to discriminate different elevations with this encoding.</p>
<p>The abilities to localize the elevation of an artificially spatialized sound are known to be impaired in comparison with azimuth (Wenzel et al., <xref ref-type="bibr" rid="B67">1993</xref>). Those difficulties arise from the spectral distortions that are specific to individual body morphology (Blauert, <xref ref-type="bibr" rid="B12">1996</xref>; Xu et al., <xref ref-type="bibr" rid="B68">2007</xref>). When using non-individualized HRTFs, these spectral distortions are different from the participant&#x00027;s specific distortions, causing misinterpretation of elevation location. Additionally, the abilities to localize the elevation position of a sound source (virtual or real) are modulated by the spectral content of the sound (Middlebrooks and Green, <xref ref-type="bibr" rid="B46">1991</xref>; Blauert, <xref ref-type="bibr" rid="B12">1996</xref>).</p>
<p>In our study, the difficulty with the spatialization-based only encoding to localize the elevation of the target, even after the audio-motor familiarization, could be explained by a too brief training period to get used to the new auditory cues. Actually, some studies showed an improvement of localization abilities with non-individualized or modified HRTFs after 3 weeks of training in Majdak et al. (<xref ref-type="bibr" rid="B39">2013</xref>) or Romigh et al. (<xref ref-type="bibr" rid="B54">2017</xref>), or after 2 weeks in Shinn-Cunningham et al. (<xref ref-type="bibr" rid="B58">1998</xref>) or 1 week in Kumpik et al. (<xref ref-type="bibr" rid="B33">2010</xref>), and about 5 h in Bauer et al. (<xref ref-type="bibr" rid="B10">1966</xref>). Moreover, Mendon&#x000E7;a et al. (<xref ref-type="bibr" rid="B42">2013</xref>) showed the positive long term effect (1-month long) of training in azimuth and elevation localization abilities with a sound source spatialized using the same HRTFs database that was used in the current study. It suggests that the exclusive use of HRTFs to encode spatial information in SSDs might require a long training period or a long process to acquire individualized HRTFs.</p>
</sec>
<sec>
<title>4.1.2. Positive effects of cross-modal correspondence on elevation localization</title>
<p>The participants&#x00027; abilities to localize the elevation of the target using the 2 pitch-based encodings were significantly better than with a broadband sound spatialization encoding. Before the audio-motor familiarization, with the narrowband encoding (Monotonic) and the more complex encoding (Harmonic), the unsigned errors in elevation were comprised between 27.30 &#x000B1; 26.50&#x000B0; and 37.79 &#x000B1; 31.55&#x000B0; depending on the target elevation.</p>
<p>Before the familiarization, participants did not receive any information about the way the sound was modulated depending on the target location. In other words, they did not know that low pitch sounds were associated with low elevation locations, and conversely. However, the individual results of each participant for the elevation (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figures S1</xref>, <xref ref-type="supplementary-material" rid="SM2">S2</xref>) suggest that even before the familiarization, several participants interpreted the pitch to perceive the target elevation, using high pitch for high elevation and low pitch for low elevation. We suppose that participants were able to guess that the pitch of the sound varied with the target elevation because the experimenter explicitly told them that sound features were modulated as a function of the location of the target although no details regarding this modulation were provided. Two participants (S12 from the Harmonic group and S15 from the Monotonic group) reversed the pitch encoding by associating a low pitch to high elevations and a high pitch to low elevations, but they reversed this miss-representation after the familiarization. Our study showed that after the audio-motor familiarization, the elevation unsigned errors significantly decreased with both pitch-based encodings to reach values comprised between 17.67 &#x000B1; 22.23&#x000B0; and 24.06 &#x000B1; 17.67&#x000B0;, which are lower elevation unsigned errors than the mean elevation error of 25.2&#x000B0; immediately after the training in Mendon&#x000E7;a et al. (<xref ref-type="bibr" rid="B42">2013</xref>).</p>
<p>In the visual-to-auditory SSD domain, the artificial pitch mapping of elevation is used by several existing visual-to-auditory SSDs and relies on the audiovisual cross-modal correspondence between visual elevation and pitch (Spence, <xref ref-type="bibr" rid="B60">2011</xref>; Deroy et al., <xref ref-type="bibr" rid="B18">2018</xref>). In the current study, the frequency range was between 250 Hz and about 1,500 Hz with the Monotonic encoding and between 250 Hz and about 6,000 Hz with the Harmonic encoding (i.e., 1,500 Hz &#x000D7; 2 &#x000D7; 2). The floor value of 250 Hz was chosen to provide frequency steps of at least 3 Hz between each of the 120 auditory pixels in a column, to fit to the human frequency discrimination abilities (Howard and Angus, <xref ref-type="bibr" rid="B26">2009</xref>). We used the Mel scale (Stevens et al., <xref ref-type="bibr" rid="B63">1937</xref>) to take into account the perceived scaling in sound frequency discrimination. All the SSDs using a pitch mapping of elevation use different frequency ranges, resolutions (i.e., number of used frequencies) and frequency steps. The vOICe SSD (Meijer, <xref ref-type="bibr" rid="B41">1992</xref>) uses a larger frequency range than the current study (from 500 to 5,000 Hz) following an exponential scale with a 64-frequency resolution. The EyeMusic SSD (Abboud et al., <xref ref-type="bibr" rid="B1">2014</xref>) uses a pentatonic musical scale with 24 frequencies from 65.785 Hz to 1577.065 Hz. The SSD proposed in Ambard et al. (<xref ref-type="bibr" rid="B5">2015</xref>) also uses 120 frequency steps but following the Bark scale (Zwicker, <xref ref-type="bibr" rid="B69">1961</xref>) and with a larger frequency range (from 250 Hz to about 2,500 Hz). Technically, increasing the range of frequencies might increase discrimination abilities between target elevations and improve localization abilities. Although, as sound frequency increases the sound feels unpleasant (Kumar et al., <xref ref-type="bibr" rid="B32">2008</xref>). We can postulate that SSD users should be able to modulate some of the parameters in order to adapt the encoding scheme to their own auditory abilities and perceptual preferences.</p>
<p>Our results suggest that a pitch mapping of elevation can quickly be interpreted, even without any explicit explanation of the mapping rules. They also suggest that the pitch mapping provides acoustic cues that are easily interpretable at the early stage of use of a SSD to localize an object. In terms of spatial perception, our study shows that adding abstract acoustic cues to convey spatial information can be more efficient than an imperfect synthesizing of natural acoustic cues. It is difficult to assert that the differences in the results between the Noise encoding and the Pitch encodings are entirely due to the cross-modal correspondence between elevation and pitch since modifying the timbre of the sound by reducing its spectral content also modified how the HRTFs spatialize the sound. Therefore, it would be interesting to investigate the localization performance with monotonic or harmonic sounds in which the pitch is constant (i.e., not related to the elevation of the target) and by conducting an experiment where HRTFs convolution is computed to convey azimuth only, with for instance a constant elevation of 0&#x000B0;.</p>
</sec>
<sec>
<title>4.1.3. Insights about the pitch-elevation cross-modal correspondence</title>
<p>Although the aim of this study was not to directly investigate the multisensory perceptual process, the results might bring insights about the pitch-elevation cross-modal correspondence. In the SSD research, it has been suggested that the pitch-based elevation mapping is intuitive in an object recognition task (Stiles and Shimojo, <xref ref-type="bibr" rid="B64">2015</xref>). Based on the results of the current study, it also seems intuitive in a localization task. However, it remains to be further investigated with, for instance, a comparison of elevation localization abilities with a similar pitch-based elevation encoding and another encoding where the direction of the pitch mapping is reversed (i.e., low pitch for high elevation and high pitch for low elevation). The current study also raises the question regarding the automaticity of the cross-modal correspondences as disccussed in Spence and Deroy (<xref ref-type="bibr" rid="B61">2013</xref>). In the current study, the facilitation effect of the cross-modal correspondence probably relies on voluntary multisensory perceptual processes. The way the instructions were given to the participants intrinsically induced a goal-directed voluntary strategy in order to infer which modifications in the sound could convey information about the location of the object.</p>
<p>These insights about multisensory process should also be investigated in the blind. Since the pitch-elevation cross-modal correspondence has been suggested to be weak in this population (Deroy et al., <xref ref-type="bibr" rid="B17">2016</xref>), and since auditory spatial perception of the elevation can be impaired in this population (Voss, <xref ref-type="bibr" rid="B66">2016</xref>), it remains to investigate whether similar results would be obtained with blind participants. For this reason, the procedure of the current study was designed in a way to be reproducible with blind participants.</p>
</sec>
<sec>
<title>4.1.4. No positive effects of harmonics on elevation localization</title>
<p>The elevation-pitch encoding adds a salient auditory cue while reducing the frequency range where the HRTFs spectrum alterations can operate. To study the effect of the spectral complexity we used an encoding with harmonic sounds (monotonic and 2 following octaves) meant to be a trade-off in terms of spectral complexity between the broadband sound of the Noise encoding and the monotones of the Monotonic encoding. Although pure tones were used in the Monotonic encoding, it is important to keep in mind that soundscapes were not pure tones. Indeed, soundscapes were made of adjacent auditory pixels, resulting in narrowband but multi-frequency soundscapes (see <xref ref-type="fig" rid="F2">Figure 2</xref>).</p>
<p>The results did not show inter-group differences in the localization accuracy between the Monotonic and the Harmonic encodings. It suggests that adding 2 octaves to the original sound (i.e., the Monotonic encoding) did not modulate the ability to perceive the elevation of the target. Using more complex tones with several sub-octave intervals in the Harmonic encoding might sufficiently modify the sound spectrum to obtain a significant difference with the Monotonic encoding. It could also be interesting to investigate the ability to perceive the elevation of the target with an encoding using sounds containing frequencies higher than the current ceiling frequency (6,000 Hz). However, as mentioned in Section 4.1.1, it seems that the benefits that could arise from the application of the HRTFs on a sound with a broader spectrum could only be perceivable after a long training period.</p>
</sec>
</sec>
<sec>
<title>4.2. Azimuth localization using the visual-to-auditory encodings is accurate but overestimated</title>
<p>Depending on the encoding and the target eccentricity, the magnitude of the azimuth unsigned errors was comprised between 16.78 &#x000B1; 20.31&#x000B0; and 37.29 &#x000B1; 20.32&#x000B0;. As a comparison, Mendon&#x000E7;a et al. (<xref ref-type="bibr" rid="B42">2013</xref>) spatialized white noise sounds using the same HRTFs database and their participants localized the azimuth with a mean unsigned error of 21.3&#x000B0; before the training practice. In Geronazzo et al. (<xref ref-type="bibr" rid="B21">2018</xref>), the azimuth unsigned errors of participants varied between 3.67 &#x000B1; 2.97&#x000B0; and 35.98 &#x000B1; 45.32&#x000B0;. In the SSD domain, Scalvini et al. (<xref ref-type="bibr" rid="B57">2022</xref>) found a mean azimuth error of 6.72 &#x000B1; 5.82&#x000B0; in a task consisting in localizing a target with the head. In the current study, after the familiarization, the magnitude of azimuth unsigned errors decreased and was comprised between 12.04 &#x000B1; 12.05&#x000B0; and 25.06 &#x000B1; 14.48&#x000B0; depending on the azimuth eccentricity which is comparable to the azimuth unsigned errors found in Geronazzo et al. (<xref ref-type="bibr" rid="B21">2018</xref>), without training. In Mendon&#x000E7;a et al. (<xref ref-type="bibr" rid="B42">2013</xref>), immediately after the training, the mean azimuth unsigned errors also decreased and reached a magnitude of 15.3&#x000B0; which is also comparable to the current results.</p>
<p>In the current study, without any familiarization, and with the three visual-to-auditory encodings, participants were able to discriminate the different azimuths as suggested by gains higher than the optimal value of 1.0. After the familiarization, and with the three visual-to-auditory encodings, participants were able to localize the azimuth of the target with average azimuth gains comprised between 1.23 and 1.41 which were higher than the null value and than the optimal gain 1.0. It shows that the sound spatialization method used in the current study based on HRTFs from the CIPIC database (Algazi et al., <xref ref-type="bibr" rid="B3">2001a</xref>) partly reproduced the natural cues used in free-field sound azimuth localization. These results are not surprising since azimuth is mainly conveyed through binaural cues including the Interaural Level Difference (ILD) and the Interaural Time Difference (ITD) that reflect audio signal differences between the two ears. ITD is mainly used when the spectral content of the audio signal does not include frequencies higher than 1,500 Hz and ILD is mainly used for frequencies higher than 3,000 Hz (Blauert, <xref ref-type="bibr" rid="B12">1996</xref>). The used frequencies ranged from 250 Hz to about 1,500 Hz with the Monotonic encoding, which is in a frequency domain where ITDs are mainly used to perceive the azimuth. With the Harmonic encoding, that added 2 octaves, the frequency range was between 250 and 6,000 Hz which already contains the ILD frequency domain. The Noise encoding with the broadband sound allows both cues (ITD and ILD) to be fully used, which can theoretically improve azimuth localization accuracy in comparison with sounds with a lower spectral complexity, as previously shown in Morikawa and Hirahara (<xref ref-type="bibr" rid="B48">2013</xref>). However, in the current study, these drastic changes in the spectrum did not strongly affect the participants&#x00027; abilities, and the response patterns were similar. In other words, whatever the spectral complexity of the sound used in the encoding (white noise, complex tones or pure tones), binaural cues could be perceived and interpreted by the participants. It can be noticed that azimuth accuracy seems slightly higher with the two pitch-based encodings (the Harmonic and Monotonic encodings) in comparison with the spatialization-based only encoding (the Noise encoding). We did not find similar results in the scientific literature. This facilitation effect could result from a decrease in the cognitive load when the elevation is conveyed through the pitch modulation. As mentioned above, the pitch-based encodings seem more intuitive to localize the elevation, therefore it should globally decrease the cognitive load and thus facilitate the processing of the remaining dimension (i.e., the azimuth dimension). This effect does not seem to drastically shape the results and remains to be confirmed by other experiments.</p>
<p>The participants tended to overestimate the lateral position of the lateral targets with the three visual-to-auditory encodings: a shift to the left for targets on the left, and a shift to the right for targets on the right. Some studies also showed an overestimation pattern of lateral sound sources while using non-individualized HRTFs (Wenzel et al., <xref ref-type="bibr" rid="B67">1993</xref>), in a virtual environment while being blindfolded (Ahrens et al., <xref ref-type="bibr" rid="B2">2019</xref>), using ambisonics (Huisman et al., <xref ref-type="bibr" rid="B28">2021</xref>), and even with real sound sources (Oldfield and Parker, <xref ref-type="bibr" rid="B49">1984</xref>; Makous and Middlebrooks, <xref ref-type="bibr" rid="B40">1990</xref>). A possibility to decrease this overestimation might be to rescale the used HRTF positions to fit to the perceived ones. For example one could rescale the azimuth angles of the HRTFs database to compensate for the non-linear shape that was measured as the perceived ones and measure if it could linearize the response profile.</p>
<p>The participants also tended to localize the targets with a leftward bias between &#x02212;1.2 and &#x02212;5.03&#x000B0; in average. This systematic error might be due to a wrong auditory localization but also to a misperception of target distance. Geometrical considerations shows that an underestimation of the distance of the sound source would generate a leftward bias as we see in the current results. Since no indication concerning the sound distance was given, the participants could estimate that the sound sources were located closer than one meter. <xref ref-type="fig" rid="F7">Figure 7</xref> shows the effect of a misperception of target distance on the azimuth localization. However, for the same reason, a distance underestimation would have cause an overestimation of the elevation perception, which we did not measure.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Left shift scheme. If the participant perceived the distance of the target closer than the real target distance (<italic>d</italic><sub>1</sub> instead of <italic>d</italic> &#x0003D; 1m), it might induce an increase of the leftward bias (&#x003B5;<sub>1</sub>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1079998-g0007.tif"/>
</fig>
</sec>
<sec>
<title>4.3. A fast improvement in object localization performance</title>
<sec>
<title>4.3.1. A short but active familiarization method</title>
<p>After a first practice followed by a very short familiarization, participants&#x00027; abilities to localize an object with the visual-to-auditory SSD were improved. The elevation gains were improved for all the encodings (especially for pitch-based ones), and for the azimuth, the decrease in the lateral overshoot suggests that the interpretation of acoustic cues provided by the ILD and ITD for the azimuth was improved. Since no feedback was given during the first practice, it can be supposed that the familiarization session mainly contributed to acquire sensorimotor contingencies (Auvray, <xref ref-type="bibr" rid="B7">2004</xref>) through the mean of an audio-motor calibration (Aytekin et al., <xref ref-type="bibr" rid="B9">2008</xref>).</p>
<p>In order to avoid a too long experimental session, we used a short audio-motor familiarization session (60 s) during which participants were active by controlling the position of the target, which is known to improve the positive effect of the training (Aytekin et al., <xref ref-type="bibr" rid="B9">2008</xref>; H&#x000FC;g et al., <xref ref-type="bibr" rid="B27">2022</xref>). Other familiarization methods have been studied and have shown improvements in the use of SSDs. For example, prior to the experimental task, some studies simultaneously displayed to participants an image and its equivalent soundscape (Ambard et al., <xref ref-type="bibr" rid="B5">2015</xref>; Buchs et al., <xref ref-type="bibr" rid="B14">2021</xref>). In another study (Auvray et al., <xref ref-type="bibr" rid="B8">2007</xref>), participants were enrolled in an intensive training of 3 h. Using only a verbal explanation of the visual-to-auditory encoding scheme as been shown to be efficient to understand the main principles of the encoding scheme (Kim and Zatorre, <xref ref-type="bibr" rid="B30">2008</xref>; Buchs et al., <xref ref-type="bibr" rid="B14">2021</xref>; Scalvini et al., <xref ref-type="bibr" rid="B57">2022</xref>). The aim of the current study was not to directly investigate the effect of a short and active familiarization method on localization performance but it shows that a short practice might be sufficient to acquire the sensorimotor contingencies. The effect of the familiarization remains to be clearly assessed by comparing the efficiency of the existing methods with control conditions in order to optimize the SSD learning.</p>
</sec>
<sec>
<title>4.3.2. Calibration of the auditory space improves localization abilities</title>
<p>In the current study, participants were not aware of the size of the VAS neither that the head tracker was associated with a virtual camera capturing and converting into sounds a limited portion of the virtual scene in front of them. They only knew that the virtual target would appear at random locations in their front-field at different azimuth and elevation locations. As a consequence, they also did not know the spatial boundaries of the space where the target could be heard. After a short practice, the participants were able to build an accurate mental spatial representation of the virtual space where the visual-to-auditory encoding took place. For instance, the downward bias in elevation decreased after the familiarization session, suggesting that participants learned that the VAS was at a higher location. Also the decrease of the overestimation pattern in azimuth suggests that participants learned that the lateral VAS boundaries were closer.</p>
<p>It has to be noticed that the size of the VAS has an influence on the localization accuracy. The biggest the VAS is, the higher the localization error might be. Restricting the field of view of the camera would result in a smaller possible space in which an heard target could be placed, thus resulting in a lower angular error, but as a counterpart, it would cover a smaller subpart of the front-field without moving the head. For instance, for a target placed in a central position, a random pointing in a VAS with a field of view of 45 &#x000D7; 45&#x000B0; (azimuth &#x000D7; elevation) would result in an error in azimuth and elevation with a standard deviation 2 times lower than with a field of view of 90 &#x000D7; 90&#x000B0; while covering a space 4 times smaller. Studying the effect of various VAS sizes in a target localization task in which the user can freely move the head to point to a target as fast as possible would probably give some insights about the optimal VAS size. However, in ecological contexts, a large VAS size would have the advantage of providing auditory information about obstacles placed with a larger eccentricity with respect to the forward direction of the head. For this reason, in a real context of use, this parameter should probably be customizable according to the habit of use.</p>
</sec>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>Long trainings are required to master a visual-to-auditory SSD (Kristj&#x000E1;nsson et al., <xref ref-type="bibr" rid="B31">2016</xref>) because the used visual-to-auditory encodings are not enough intuitive (Hamilton-Fletcher et al., <xref ref-type="bibr" rid="B23">2016b</xref>). In our study, we investigated several visual-to-auditory encodings in order to develop a SSD whose auditory information could quickly be interpreted to localize obstacles. In line with previous studies, our results suggest that a visual-to-auditory SSD based on the creation of a VAS is efficient to convey visuo-spatial information about azimuth through soundscapes. Our study shows that a pitch-based elevation mapping can be easily learn to compensate for elevation localization impairments due to the use of non-individualized HRTFs in the creation process of the VAS. Despite a very short period of practice, the participants were able to improve their interpretation of the used acoustic cues both for the azimuth and the elevation encoding schemes.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation. The original contributions presented in the study will be made available in the following link: <ext-link ext-link-type="uri" xlink:href="http://leadserv.u-bourgogne.fr/en/members/maxime-ambard/pages/cross-modal-correspondance-enhances-elevation-localization">http://leadserv.u-bourgogne.fr/en/members/maxime-ambard/pages/cross-modal-correspondance-enhances-elevation-localization</ext-link>. Further questions should be directed to the corresponding author.</p>
</sec>
<sec sec-type="ethics-statement" id="s7">
<title>Ethics statement</title>
<p>The studies involving human participants were reviewed and approved by Comit&#x000E9; d&#x00027;Ethique pour la Recherche de Universit&#x000E9; Bourgogne Franche-Comt&#x000E9;. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>CB and MA contributed to conception and design of the study and interpreted the data. CB executed the study and was responsible for data analysis and wrote the first draft of the manuscript in closed collaboration with MA. FS, CM, and JD provided important feedback. All authors have read, approved the manuscript, and contributed substantially to it.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This research was funded by the Conseil R&#x000E9;gional de Bourgogne Franche-Comt&#x000E9; (2020_0335), France and the Fond Europ&#x000E9;en de D&#x000E9;veloppement R&#x000E9;gional (FEDER) (BG0027904).</p>

</sec>
<ack><p>Thanks to the Conseil R&#x000E9;gional de Bourgogne Franche-Comt&#x000E9;, France and the Fond Europ&#x000E9;en de D&#x000E9;veloppement R&#x000E9;gional (FEDER) for their financial support. We thank the Universit&#x000E9; de Bourgogne and le Centre National de la Recherche Scientifique (CNRS) for the providing administrative support and the infrastructure.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1079998/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1079998/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Image_1.jpg" id="SM1" mimetype="image/jpeg" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 1</label>
<caption><p>Elevation response position as a function of target elevation for each participant of the Monotonic group. Mean elevation response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the Noise (blue squares) and the Monotonic (orange circles) encodings. Error bars shows standard error of elevation response position. Solid lines represent the elevation gains with the Noise (blue) and Monotonic (orange) encodings. Black dashed lines indicate the optimal elevation gain 1.0.</p></caption> </supplementary-material>
<supplementary-material xlink:href="Image_2.jpg" id="SM2" mimetype="image/jpeg" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 2</label>
<caption><p>Elevation response position as a function of target elevation for each participant of the Harmonic group. Mean elevation response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the Noise (blue squares) and the Harmonic (red circles) encodings. Error bars shows standard error of elevation response position. Solid lines represent the elevation gains with the Noise (blue) and Harmonic (red) encodings. Black dashed lines indicate the optimal elevation gain 1.0.</p></caption> </supplementary-material>
<supplementary-material xlink:href="Image_3.jpg" id="SM3" mimetype="image/jpeg" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 3</label>
<caption><p>Azimuth response position as a function of target azimuth for each participant of the Monotonic group. Mean azimuth response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the Noise (blue squares) and the Monotonic (orange circles) encodings. Error bars shows standard error of azimuth response position. Solid lines represent the azimuth gains with the Noise (blue) and Monotonic (orange) encodings. Black dashed lines indicate the optimal azimuth gain 1.0.</p></caption> </supplementary-material>
<supplementary-material xlink:href="Image_4.jpg" id="SM4" mimetype="image/jpeg" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 4</label>
<caption><p>Azimuth response position as a function of target azimuth for each participant of the Harmonic group. Mean azimuth response positions (in degree) before <bold>(left)</bold> and after <bold>(right)</bold> are represented separately for the Noise (blue squares) and the Harmonic (red circles) encodings. Error bars shows standard error of azimuth response position. Solid lines represent the azimuth gains with the Noise (blue) and Harmonic (red) encodings. Black dashed lines indicate the optimal azimuth gain 1.0.</p></caption> </supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abboud</surname> <given-names>S.</given-names></name> <name><surname>Hanassy</surname> <given-names>S.</given-names></name> <name><surname>Levy-Tzedek</surname> <given-names>S.</given-names></name> <name><surname>Maidenbaum</surname> <given-names>S.</given-names></name> <name><surname>Amedi</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>EyeMusic: Introducing a &#x0201C;visual&#x0201D; colorful experience for the blind using auditory sensory substitution</article-title>. <source>Restor. Neurol Neurosci</source>. <volume>32</volume>, <fpage>247</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.3233/RNN-130338</pub-id><pub-id pub-id-type="pmid">24398719</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahrens</surname> <given-names>A.</given-names></name> <name><surname>Lund</surname> <given-names>K. D.</given-names></name> <name><surname>Marschall</surname> <given-names>M.</given-names></name> <name><surname>Dau</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Sound source localization with varying amount of visual information in virtual reality</article-title>. <source>PLoS ONE</source> <volume>14</volume>, <fpage>e0214603</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0214603</pub-id><pub-id pub-id-type="pmid">30925174</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Algazi</surname> <given-names>V. R, Duda, R.</given-names></name> <name><surname>Thompson</surname> <given-names>D.</given-names></name> <name><surname>Avendano</surname> <given-names>C.</given-names></name></person-group> (<year>2001a</year>). <article-title>&#x0201C;The CIPIC HRTF database,&#x0201D;</article-title> in <source>Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics</source> (<publisher-loc>New Platz, NY:L IEEE</publisher-loc>), <fpage>99</fpage>&#x02013;<lpage>102</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Algazi</surname> <given-names>V. R.</given-names></name> <name><surname>Avendano</surname> <given-names>C.</given-names></name> <name><surname>Duda</surname> <given-names>R. O.</given-names></name></person-group> (<year>2001b</year>). <article-title>Elevation localization and head-related transfer function analysis at low frequencies</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>109</volume>, <fpage>1110</fpage>&#x02013;<lpage>1122</lpage>. <pub-id pub-id-type="doi">10.1121/1.1349185</pub-id><pub-id pub-id-type="pmid">11303925</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ambard</surname> <given-names>M.</given-names></name> <name><surname>Benezeth</surname> <given-names>Y.</given-names></name> <name><surname>Pfister</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Mobile video-to-audio transducer and motion detection for sensory substitution</article-title>. <source>Front. ICT</source> <volume>2</volume>, <fpage>20</fpage>. <pub-id pub-id-type="doi">10.3389/fict.2015.00020</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Asano</surname> <given-names>F.</given-names></name> <name><surname>Suzuki</surname> <given-names>Y.</given-names></name> <name><surname>Sone</surname> <given-names>T.</given-names></name></person-group> (<year>1990</year>). <article-title>Role of spectral cues in median plane localization</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>88</volume>, <fpage>159</fpage>&#x02013;<lpage>168</lpage>. <pub-id pub-id-type="doi">10.1121/1.399963</pub-id><pub-id pub-id-type="pmid">2380444</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Auvray</surname> <given-names>M.</given-names></name></person-group> (<year>2004</year>). <source>Immersion et perception spatiale. L&#x00027;exemple des dispositifs de substitution sensorielle</source> (<publisher-loc>Ph.D. thesis</publisher-loc>). Ecole des Hautes Etudes en Sciences Sociales, Paris.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Auvray</surname> <given-names>M.</given-names></name> <name><surname>Hanneton</surname> <given-names>S.</given-names></name> <name><surname>O&#x00027;Regan</surname> <given-names>J. K.</given-names></name></person-group> (<year>2007</year>). <article-title>Learning to perceive with a visuo&#x02013;auditory substitution system: localisation and object recognition with &#x02018;the voice&#x00027;</article-title>. <source>Perception</source> <volume>36</volume>, <fpage>416</fpage>&#x02013;<lpage>430</lpage>. <pub-id pub-id-type="doi">10.1068/p5631</pub-id><pub-id pub-id-type="pmid">17455756</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aytekin</surname> <given-names>M.</given-names></name> <name><surname>Moss</surname> <given-names>C. F.</given-names></name> <name><surname>Simon</surname> <given-names>J. Z.</given-names></name></person-group> (<year>2008</year>). <article-title>A sensorimotor approach to sound localization</article-title>. <source>Neural Comput</source>. <volume>20</volume>, <fpage>603</fpage>&#x02013;<lpage>635</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2007.12-05-094</pub-id><pub-id pub-id-type="pmid">18045018</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bauer</surname> <given-names>R. W.</given-names></name> <name><surname>Matuzsa</surname> <given-names>J. L.</given-names></name> <name><surname>Blackmer</surname> <given-names>R. F.</given-names></name> <name><surname>Glucksberg</surname> <given-names>S.</given-names></name></person-group> (<year>1966</year>). <article-title>Noise localization after unilateral attenuation</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>40</volume>, <fpage>441</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1121/1.1910093</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Best</surname> <given-names>V.</given-names></name> <name><surname>Baumgartner</surname> <given-names>R.</given-names></name> <name><surname>Lavandier</surname> <given-names>M.</given-names></name> <name><surname>Majdak</surname> <given-names>P.</given-names></name> <name><surname>Kop&#x0010D;o</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>Sound externalization: a review of recent research</article-title>. <source>Trends Hear</source>. 24, 233121652094839. <pub-id pub-id-type="doi">10.1177/2331216520948390</pub-id><pub-id pub-id-type="pmid">32914708</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blauert</surname> <given-names>J.</given-names></name></person-group> (<year>1996</year>). <source>Spatial Hearing: The Psychophysics of Human Sound Localization, 6th Edn</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>D.</given-names></name> <name><surname>Macpherson</surname> <given-names>T.</given-names></name> <name><surname>Ward</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Seeing with sound? Exploring different characteristics of a visual-to-auditory sensory substitution device</article-title>. <source>Perception</source> <volume>40</volume>, <fpage>1120</fpage>&#x02013;<lpage>1135</lpage>. <pub-id pub-id-type="doi">10.1068/p6952</pub-id><pub-id pub-id-type="pmid">22208131</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buchs</surname> <given-names>G.</given-names></name> <name><surname>Haimler</surname> <given-names>B.</given-names></name> <name><surname>Kerem</surname> <given-names>M.</given-names></name> <name><surname>Maidenbaum</surname> <given-names>S.</given-names></name> <name><surname>Braun</surname> <given-names>L.</given-names></name> <name><surname>Amedi</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>A self-training program for sensory substitution devices</article-title>. <source>PLoS ONE</source> <volume>16</volume>, <fpage>e0250281</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0250281</pub-id><pub-id pub-id-type="pmid">33905446</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caraiman</surname> <given-names>S.</given-names></name> <name><surname>Morar</surname> <given-names>A.</given-names></name> <name><surname>Owczarek</surname> <given-names>M.</given-names></name> <name><surname>Burlacu</surname> <given-names>A.</given-names></name> <name><surname>Rzeszotarski</surname> <given-names>D.</given-names></name> <name><surname>Botezatu</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Computer vision for the visually impaired: the sound of vision system,&#x0201D;</article-title> in <source>2017 IEEE International Conference on Computer Vision Workshops (ICCVW)</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1480</fpage>&#x02013;<lpage>1489</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Comm&#x000E8;re</surname> <given-names>L.</given-names></name> <name><surname>Wood</surname> <given-names>S. U. N.</given-names></name> <name><surname>Rouat</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Evaluation of a vision-to-audition substitution system that provides 2D WHERE information and fast user learning</article-title>. <source>Techn. Rep. arXiv:2010.09041</source>, arXiv. arXiv:2010.09041 [cs]. <pub-id pub-id-type="doi">10.48550/arXiv.2010.09041</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deroy</surname> <given-names>O.</given-names></name> <name><surname>Fasiello</surname> <given-names>I.</given-names></name> <name><surname>Hayward</surname> <given-names>V.</given-names></name> <name><surname>Auvray</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Differentiated audio-tactile correspondences in sighted and blind individuals</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform</source>. <volume>42</volume>, <fpage>1204</fpage>&#x02013;<lpage>1214</lpage>. <pub-id pub-id-type="doi">10.1037/xhp0000152</pub-id><pub-id pub-id-type="pmid">26950385</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deroy</surname> <given-names>O.</given-names></name> <name><surname>Fernandez-Prieto</surname> <given-names>I.</given-names></name> <name><surname>Navarra</surname> <given-names>J.</given-names></name> <name><surname>Spence</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Unraveling the paradox of spatial pitch,&#x0201D;</article-title> in <source>Spatial Biases in Perception and Cognition, 1st Edn</source>, ed T. L. Hubbard (New York, NY: Cambridge University Press), <fpage>77</fpage>&#x02013;<lpage>93</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Evans</surname> <given-names>K. K.</given-names></name> <name><surname>Treisman</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Natural cross-modal mappings between visual and auditory features</article-title>. <source>J. Vis</source>. <volume>10</volume>, <fpage>6</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1167/10.1.6</pub-id><pub-id pub-id-type="pmid">20143899</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gardner</surname> <given-names>M. B.</given-names></name></person-group> (<year>1973</year>). <article-title>Some monaural and binaural facets of median plane localization</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>54</volume>, <fpage>1489</fpage>&#x02013;<lpage>1495</lpage>. <pub-id pub-id-type="doi">10.1121/1.1914447</pub-id><pub-id pub-id-type="pmid">4780802</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geronazzo</surname> <given-names>M.</given-names></name> <name><surname>Sikstrom</surname> <given-names>E.</given-names></name> <name><surname>Kleimola</surname> <given-names>J.</given-names></name> <name><surname>Avanzini</surname> <given-names>F.</given-names></name> <name><surname>de Gotzen</surname> <given-names>A.</given-names></name> <name><surname>Serafin</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;The impact of an accurate vertical localization with HRTFs on short explorations of immersive virtual reality scenarios,&#x0201D;</article-title> in <source>2018 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)</source> (<publisher-loc>Munich</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>90</fpage>&#x02013;<lpage>97</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hamilton-Fletcher</surname> <given-names>G.</given-names></name> <name><surname>Mengucci</surname> <given-names>M.</given-names></name> <name><surname>Medeiros</surname> <given-names>F.</given-names></name></person-group> (<year>2016a</year>). <source>Synaestheatre: Sonification of Coloured Objects in Space</source>. <publisher-loc>Brighton</publisher-loc>: <publisher-name>International Conference on Live Interfaces</publisher-name>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hamilton-Fletcher</surname> <given-names>G.</given-names></name> <name><surname>Obrist</surname> <given-names>M.</given-names></name> <name><surname>Watten</surname> <given-names>P.</given-names></name> <name><surname>Mengucci</surname> <given-names>M.</given-names></name> <name><surname>Ward</surname> <given-names>J.</given-names></name></person-group> (<year>2016b</year>). <article-title>&#x0201C;&#x00022;I always wanted to see the night sky&#x00022;: blind user preferences for sensory substitution devices,&#x0201D;</article-title> in <source>Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>San Jose, CA</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>2162</fpage>&#x02013;<lpage>2174</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanneton</surname> <given-names>S.</given-names></name> <name><surname>Auvray</surname> <given-names>M.</given-names></name> <name><surname>Durette</surname> <given-names>B.</given-names></name></person-group> (<year>2010</year>). <article-title>The Vibe: a versatile vision-to-audition sensory substitution device</article-title>. <source>Appl. Bionics Biomech</source>. <volume>7</volume>, <fpage>269</fpage>&#x02013;<lpage>276</lpage>. <pub-id pub-id-type="doi">10.1155/2010/282341</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hebrank</surname> <given-names>J.</given-names></name> <name><surname>Wright</surname> <given-names>D.</given-names></name></person-group> (<year>1974</year>). <article-title>Spectral cues used in the localization of sound sources on the median plane</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>56</volume>, <fpage>1829</fpage>&#x02013;<lpage>1834</lpage>. <pub-id pub-id-type="doi">10.1121/1.1903520</pub-id><pub-id pub-id-type="pmid">4443482</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howard</surname> <given-names>D. M.</given-names></name> <name><surname>Angus</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <source>Acoustics ans Psychoacoustics, 4th Edn</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Focal Press</publisher-name>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>H&#x000FC;g</surname> <given-names>M. X.</given-names></name> <name><surname>Bermejo</surname> <given-names>F.</given-names></name> <name><surname>Tommasini</surname> <given-names>F. C.</given-names></name> <name><surname>Di Paolo</surname> <given-names>E. A.</given-names></name></person-group> (<year>2022</year>). <article-title>Effects of guided exploration on reaching measures of auditory peripersonal space</article-title>. <source>Front. Psychol</source>. 13, 983189. <pub-id pub-id-type="doi">10.3389/fpsyg.2022.983189</pub-id><pub-id pub-id-type="pmid">36337523</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huisman</surname> <given-names>T.</given-names></name> <name><surname>Ahrens</surname> <given-names>A.</given-names></name> <name><surname>MacDonald</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>Ambisonics sound source localization with varying amount of visual information in virtual reality</article-title>. <source>Front. Virtual Real</source>. 2, 722321. <pub-id pub-id-type="doi">10.3389/frvir.2021.722321</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jicol</surname> <given-names>C.</given-names></name> <name><surname>Lloyd-Esenkaya</surname> <given-names>T.</given-names></name> <name><surname>Proulx</surname> <given-names>M. J.</given-names></name> <name><surname>Lange-Smith</surname> <given-names>S.</given-names></name> <name><surname>Scheller</surname> <given-names>M.</given-names></name> <name><surname>O&#x00027;Neill</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Efficiency of sensory substitution devices alone and in combination with self-motion for spatial navigation in sighted and visually impaired</article-title>. <source>Front. Psychol</source>. 11, 1443. <pub-id pub-id-type="doi">10.3389/fpsyg.2020.01443</pub-id><pub-id pub-id-type="pmid">32754082</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.-K.</given-names></name> <name><surname>Zatorre</surname> <given-names>R. J.</given-names></name></person-group> (<year>2008</year>). <article-title>Generalized learning of visual-to-auditory substitution in sighted individuals</article-title>. <source>Brain Res</source>. <volume>1242</volume>, <fpage>263</fpage>&#x02013;<lpage>275</lpage>. <pub-id pub-id-type="doi">10.1016/j.brainres.2008.06.038</pub-id><pub-id pub-id-type="pmid">18602373</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kristj&#x000E1;nsson</surname> <given-names>r.</given-names></name> <name><surname>Moldoveanu</surname> <given-names>A.</given-names></name> <name><surname>J&#x000F3;hannesson</surname> <given-names>m,. I.</given-names></name> <name><surname>Balan</surname> <given-names>O.</given-names></name> <name><surname>Spagnol</surname> <given-names>S.</given-names></name> <name><surname>Valgeirsd&#x000F3;ttir</surname> <given-names>V. V.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Designing sensory-substitution devices: Principles, pitfalls and potential1</article-title>. <source>Restor. Neurol Neurosci</source>. <volume>34</volume>, <fpage>769</fpage>&#x02013;<lpage>787</lpage>. <pub-id pub-id-type="doi">10.3233/RNN-160647</pub-id><pub-id pub-id-type="pmid">27567755</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>S.</given-names></name> <name><surname>Forster</surname> <given-names>H. M.</given-names></name> <name><surname>Bailey</surname> <given-names>P.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. D.</given-names></name></person-group> (<year>2008</year>). <article-title>Mapping unpleasantness of sounds to their auditory representation</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>124</volume>, <fpage>3810</fpage>&#x02013;<lpage>3817</lpage>. <pub-id pub-id-type="doi">10.1121/1.3006380</pub-id><pub-id pub-id-type="pmid">19206807</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumpik</surname> <given-names>D. P.</given-names></name> <name><surname>Kacelnik</surname> <given-names>O.</given-names></name> <name><surname>King</surname> <given-names>A. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Adaptive reweighting of auditory localization cues in response to chronic unilateral earplugging in humans</article-title>. <source>J. Neurosci</source>. <volume>30</volume>, <fpage>4883</fpage>&#x02013;<lpage>4894</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5488-09.2010</pub-id><pub-id pub-id-type="pmid">20371808</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuznetsova</surname> <given-names>A.</given-names></name> <name><surname>Brockhoff</surname> <given-names>P. B.</given-names></name> <name><surname>Christensen</surname> <given-names>R. H. B.</given-names></name></person-group> (<year>2017</year>). <article-title>lmertest package: tests in linear mixed effects models</article-title>. <source>J. Stat. Softw</source>. 82, 13. <pub-id pub-id-type="doi">10.18637/jss.v082.i13</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lenth</surname> <given-names>R. V.</given-names></name></person-group> (<year>2022</year>). <source>emmeans: Estimated Marginal Means, aka Least-Squares Means</source>. R package version 1.7.<fpage>4</fpage>&#x02013;<lpage>1</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levy-Tzedek</surname> <given-names>S.</given-names></name> <name><surname>Hanassy</surname> <given-names>S.</given-names></name> <name><surname>Abboud</surname> <given-names>S.</given-names></name> <name><surname>Maidenbaum</surname> <given-names>S.</given-names></name> <name><surname>Amedi</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Fast, accurate reaching movements with a visual-to-auditory sensory substitution device</article-title>. <source>Restor. Neurol Neurosci</source>. <volume>30</volume>, <fpage>313</fpage>&#x02013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.3233/RNN-2012-110219</pub-id><pub-id pub-id-type="pmid">22596353</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maidenbaum</surname> <given-names>S.</given-names></name> <name><surname>Abboud</surname> <given-names>S.</given-names></name> <name><surname>Amedi</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Sensory substitution: closing the gap between basic research and widespread practical visual rehabilitation</article-title>. <source>Neurosci. Biobehav. Rev</source>. <volume>41</volume>, <fpage>3</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2013.11.007</pub-id><pub-id pub-id-type="pmid">24275274</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maidenbaum</surname> <given-names>S.</given-names></name> <name><surname>Amedi</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Standardizing visual rehabilitation using simple virtual tests,&#x0201D;</article-title> in <source>2019 International Conference on Virtual Rehabilitation (ICVR)</source> (<publisher-loc>Tel Aviv</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Majdak</surname> <given-names>P.</given-names></name> <name><surname>Walder</surname> <given-names>T.</given-names></name> <name><surname>Laback</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>Effect of long-term training on sound localization performance with spectrally warped and band-limited head-related transfer functions</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>134</volume>, <fpage>2148</fpage>&#x02013;<lpage>2159</lpage>. <pub-id pub-id-type="doi">10.1121/1.4816543</pub-id><pub-id pub-id-type="pmid">23967945</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Makous</surname> <given-names>J. C.</given-names></name> <name><surname>Middlebrooks</surname> <given-names>J. C.</given-names></name></person-group> (<year>1990</year>). <article-title>Two-dimensional sound localization by human listeners</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>87</volume>, <fpage>2188</fpage>&#x02013;<lpage>2200</lpage>. <pub-id pub-id-type="doi">10.1121/1.399186</pub-id><pub-id pub-id-type="pmid">2348023</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meijer</surname> <given-names>P.</given-names></name></person-group> (<year>1992</year>). <article-title>An experimental system for auditory image representations</article-title>. <source>IEEE Trans. Biomed. Eng</source>. <volume>39</volume>, <fpage>112</fpage>&#x02013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.1109/10.121642</pub-id><pub-id pub-id-type="pmid">1612614</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mendon&#x000E7;a</surname> <given-names>C.</given-names></name> <name><surname>Campos</surname> <given-names>G.</given-names></name> <name><surname>Dias</surname> <given-names>P.</given-names></name> <name><surname>Santos</surname> <given-names>J. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Learning auditory space: generalization and long-term effects</article-title>. <source>PLoS ONE</source> <volume>8</volume>, <fpage>e77900</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0077900</pub-id><pub-id pub-id-type="pmid">24167588</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mhaish</surname> <given-names>A.</given-names></name> <name><surname>Gholamalizadeh</surname> <given-names>T.</given-names></name> <name><surname>Ince</surname> <given-names>G.</given-names></name> <name><surname>Duff</surname> <given-names>D. J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Assessment of a visual to spatial-audio sensory substitution system,&#x0201D;</article-title> in <source>2016 24th Signal Processing and Communication Application Conference (SIU)</source> (<publisher-loc>Zonguldak</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>245</fpage>&#x02013;<lpage>248</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Middlebrooks</surname> <given-names>J. C.</given-names></name></person-group> (<year>1999</year>). <article-title>Individual differences in external-ear transfer functions reduced by scaling in frequency</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>106</volume>, <fpage>1480</fpage>&#x02013;<lpage>1492</lpage>. <pub-id pub-id-type="doi">10.1121/1.427176</pub-id><pub-id pub-id-type="pmid">10489705</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Middlebrooks</surname> <given-names>J. C.</given-names></name> <name><surname>Green</surname> <given-names>D. M.</given-names></name></person-group> (<year>1990</year>). <article-title>Directional dependence of interaural envelope delays</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>87</volume>, <fpage>2149</fpage>&#x02013;<lpage>2162</lpage>. <pub-id pub-id-type="doi">10.1121/1.399183</pub-id><pub-id pub-id-type="pmid">2348020</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Middlebrooks</surname> <given-names>J. C.</given-names></name> <name><surname>Green</surname> <given-names>D. M.</given-names></name></person-group> (<year>1991</year>). <article-title>Sound localization by human listeners</article-title>. <source>Annu. Rev. Psychol</source>. <volume>42</volume>, <fpage>135</fpage>&#x02013;<lpage>159</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.ps.42.020191.001031</pub-id><pub-id pub-id-type="pmid">2018391</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>J.</given-names></name></person-group> (<year>1991</year>). <article-title>Channel interaction and the redundant-targets effect in bimodal divided attention</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform</source>. <volume>17</volume>, <fpage>160</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.17.1.160</pub-id><pub-id pub-id-type="pmid">1826309</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morikawa</surname> <given-names>D.</given-names></name> <name><surname>Hirahara</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>Effect of head rotation on horizontal and median sound localization of band-limited noise</article-title>. <source>Acoust. Sci. Technol</source>. <volume>34</volume>, <fpage>56</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1250/ast.34.56</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oldfield</surname> <given-names>S. R.</given-names></name> <name><surname>Parker</surname> <given-names>S. P. A.</given-names></name></person-group> (<year>1984</year>). <article-title>Acuity of sound localisation: a topography of auditory space. I. Normal hearing conditions</article-title>. <source>Perception</source> <volume>13</volume>, <fpage>581</fpage>&#x02013;<lpage>600</lpage>. <pub-id pub-id-type="doi">10.1068/p130581</pub-id><pub-id pub-id-type="pmid">6535983</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pourghaemi</surname> <given-names>H.</given-names></name> <name><surname>Gholamalizadeh</surname> <given-names>T.</given-names></name> <name><surname>Mhaish</surname> <given-names>A.</given-names></name> <name><surname>Duff</surname> <given-names>D. J.</given-names></name> <name><surname>Ince</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Real-time shape-based sensory substitution for object localization and recognition</article-title>. <source>Proceedings of the 11th International Conference on Advances in Computer-Human Interactions.</source></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Proulx</surname> <given-names>M. J.</given-names></name> <name><surname>Stoerig</surname> <given-names>P.</given-names></name> <name><surname>Ludowig</surname> <given-names>E.</given-names></name> <name><surname>Knoll</surname> <given-names>I.</given-names></name></person-group> (<year>2008</year>). <article-title>Seeing &#x02018;where through the ears: effects of learning-by-doing and long-term sensory deprivation on localization based on image-to-sound substitution</article-title>. <source>PLoS ONE</source> <volume>3</volume>, <fpage>e1840</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0001840</pub-id><pub-id pub-id-type="pmid">18364998</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Real</surname> <given-names>S.</given-names></name> <name><surname>Araujo</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>VES: a mixed-reality development platform of navigation systems for blind and visually impaired</article-title>. <source>Sensors</source> <volume>21</volume>, <fpage>6275</fpage>. <pub-id pub-id-type="doi">10.3390/s21186275</pub-id><pub-id pub-id-type="pmid">34577482</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Richardson</surname> <given-names>M.</given-names></name> <name><surname>Thar</surname> <given-names>J.</given-names></name> <name><surname>Alvarez</surname> <given-names>J.</given-names></name> <name><surname>Borchers</surname> <given-names>J.</given-names></name> <name><surname>Ward</surname> <given-names>J.</given-names></name> <name><surname>Hamilton-Fletcher</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>How much spatial information is lost in the sensory substitution process? Comparing visual, tactile, and auditory approaches</article-title>. <source>Perception</source> <volume>48</volume>, <fpage>1079</fpage>&#x02013;<lpage>1103</lpage>. <pub-id pub-id-type="doi">10.1177/0301006619873194</pub-id><pub-id pub-id-type="pmid">31547778</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romigh</surname> <given-names>G. D.</given-names></name> <name><surname>Simpson</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Specificity of adaptation to non-individualized head-related transfer functions</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>141</volume>, <fpage>3974</fpage>&#x02013;<lpage>3974</lpage>. <pub-id pub-id-type="doi">10.1121/1.4989065</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rouat</surname> <given-names>J.</given-names></name> <name><surname>Lescal</surname> <given-names>D.</given-names></name> <name><surname>Wood</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <source>Handheld Device for Substitution From Vision to Audition</source>. New York, NY: 20th International Conference on Auditory Display.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rusconi</surname> <given-names>E.</given-names></name> <name><surname>Kwan</surname> <given-names>B.</given-names></name> <name><surname>Giordano</surname> <given-names>B.</given-names></name> <name><surname>Umilta</surname> <given-names>C.</given-names></name> <name><surname>Butterworth</surname> <given-names>B.</given-names></name></person-group> (<year>2006</year>). <article-title>Spatial representation of pitch height: the SMARC effect</article-title>. <source>Cognition</source> <volume>99</volume>, <fpage>113</fpage>&#x02013;<lpage>129</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2005.01.004</pub-id><pub-id pub-id-type="pmid">15925355</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scalvini</surname> <given-names>F.</given-names></name> <name><surname>Bordeau</surname> <given-names>C.</given-names></name> <name><surname>Ambard</surname> <given-names>M.</given-names></name> <name><surname>Migniot</surname> <given-names>C.</given-names></name> <name><surname>Dubois</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Low-latency human-computer auditory interface based on real-time vision analysis,&#x0201D;</article-title> in <source>ICASSP 2022</source> - <italic>2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</italic> (Singapore: IEEE), <fpage>36</fpage>&#x02013;<lpage>40</lpage>.</citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shinn-Cunningham</surname> <given-names>B. G.</given-names></name> <name><surname>Durlach</surname> <given-names>N. I.</given-names></name> <name><surname>Held</surname> <given-names>R. M.</given-names></name></person-group> (<year>1998</year>). <article-title>Adapting to supernormal auditory localization cues. I. Bias and resolution</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>103</volume>, <fpage>3656</fpage>&#x02013;<lpage>3666</lpage>. <pub-id pub-id-type="doi">10.1121/1.423088</pub-id><pub-id pub-id-type="pmid">9637047</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sodnik</surname> <given-names>J.</given-names></name> <name><surname>Su&#x00161;nik</surname> <given-names>R.</given-names></name> <name><surname>&#x00160;tular</surname> <given-names>M.</given-names></name> <name><surname>Toma&#x0017D;i&#x0010D;</surname> <given-names>S.</given-names></name></person-group> (<year>2005</year>). <article-title>Spatial sound resolution of an interpolated HRIR library</article-title>. <source>Appl. Acoust</source>. <volume>66</volume>, <fpage>1219</fpage>&#x02013;<lpage>1234</lpage>. <pub-id pub-id-type="doi">10.1016/j.apacoust.2005.04.003</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spence</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>Crossmodal correspondences: a tutorial review</article-title>. <source>Attent. Percept. Psychophys</source>. <volume>73</volume>, <fpage>971</fpage>&#x02013;<lpage>995</lpage>. <pub-id pub-id-type="doi">10.3758/s13414-010-0073-7</pub-id><pub-id pub-id-type="pmid">21264748</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spence</surname> <given-names>C.</given-names></name> <name><surname>Deroy</surname> <given-names>O.</given-names></name></person-group> (<year>2013</year>). <article-title>How automatic are crossmodal correspondences?</article-title> <source>Conscious Cogn</source>. <volume>22</volume>, <fpage>245</fpage>&#x02013;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2012.12.006</pub-id><pub-id pub-id-type="pmid">23370382</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Steinmetz</surname> <given-names>C. J.</given-names></name> <name><surname>Reiss</surname> <given-names>J. D.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Pyloudnorm: a simple yet flexible loudness meter in python,&#x0201D;</article-title> in <source>150th AES Convention</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://csteinmetz1.github.io/pyloudnorm-eval/">https://csteinmetz1.github.io/pyloudnorm-eval/</ext-link></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stevens</surname> <given-names>S. S.</given-names></name> <name><surname>Volkmann</surname> <given-names>J.</given-names></name> <name><surname>Newman</surname> <given-names>E. B.</given-names></name></person-group> (<year>1937</year>). <article-title>A scale for the measurement of the psychological magnitude pitch</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>8</volume>, <fpage>185</fpage>&#x02013;<lpage>190</lpage>. <pub-id pub-id-type="doi">10.1121/1.1915893</pub-id><pub-id pub-id-type="pmid">8710456</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stiles</surname> <given-names>N. R. B.</given-names></name> <name><surname>Shimojo</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Auditory sensory substitution is intuitive and automatic with texture stimuli</article-title>. <source>Sci. Rep</source>. 5, 15628. <pub-id pub-id-type="doi">10.1038/srep15628</pub-id><pub-id pub-id-type="pmid">26490260</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Team</surname> <given-names>R. C.</given-names></name></person-group> (<year>2020</year>). <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Core Team</publisher-name>.</citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Voss</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Auditory spatial perception without vision</article-title>. <source>Front. Psychol</source>. 07, 01960. <pub-id pub-id-type="doi">10.3389/fpsyg.2016.01960</pub-id><pub-id pub-id-type="pmid">28066286</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wenzel</surname> <given-names>E. M.</given-names></name> <name><surname>Arruda</surname> <given-names>M.</given-names></name> <name><surname>Kistler</surname> <given-names>D. J.</given-names></name> <name><surname>Wightman</surname> <given-names>F. L.</given-names></name></person-group> (<year>1993</year>). <article-title>Localization using nonindividualized head-related transfer functions</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>94</volume>, <fpage>111</fpage>&#x02013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1121/1.407089</pub-id><pub-id pub-id-type="pmid">8354753</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Salvendy</surname> <given-names>G.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Individualization of head-related transfer function for three-dimensional virtual auditory display: a review,&#x0201D;</article-title> in <source>Proceedings of the 2nd International Conference on Virtual Reality, ICVR&#x00027;07</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>), <fpage>397</fpage>&#x02013;<lpage>407</lpage>.</citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zwicker</surname> <given-names>E.</given-names></name></person-group> (<year>1961</year>). <article-title>Subdivision of the audible frequency range into critical bands (Frequenzgruppen)</article-title>. <source>J. Acoust. Soc. Am</source>. <volume>33</volume>, <fpage>248</fpage>&#x02013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1121/1.1908630</pub-id></citation>
</ref>
</ref-list> 
</back>
</article> 