<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="brief-report" dtd-version="2.3">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2022.874411</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Learning to Recognize Unfamiliar Voices: An Online Study With 12- and 24-Month-Olds</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Orena</surname>
<given-names>Adriel John</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/718900/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mader</surname>
<given-names>Asia Sotera</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Werker</surname>
<given-names>Janet F.</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/9189/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Psychology, University of British Columbia</institution>, <addr-line>Vancouver, BC</addr-line>, <country>Canada</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Evaluation and Research Services, Fraser Health Authority</institution>, <addr-line>Surrey, BC</addr-line>, <country>Canada</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: S&#x00F3;nia Frota, University of Lisbon, Portugal</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Madeleine Yu, University of Toronto, Canada; Nicole Altvater-Mackensen, Johannes Gutenberg University Mainz, Germany</p></fn>
<corresp id="c001">&#x002A;Correspondence: Janet F. Werker, <email>jwerker@psych.ubc.ca</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Developmental Psychology, a section of the journal Frontiers in Psychology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>26</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>874411</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Orena, Mader and Werker.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Orena, Mader and Werker</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Young infants are attuned to the indexical properties of speech: they can recognize highly familiar voices and distinguish them from unfamiliar voices. Less is known about how and when infants start to recognize unfamiliar voices, and to map them to faces. This skill is particularly challenging when portions of the speaker&#x2019;s face are occluded, as is the case with masking. Here, we examined voice&#x2212;face recognition abilities in infants 12 and 24&#x2009;months of age. Using the online <italic>Lookit</italic> platform, children saw and heard four different speakers produce words with sonorous phonemes (high talker information), and words with phonemes that are less sonorous (low talker information). Infants aged 24&#x2009;months, but not 12&#x2009;months, were able to learn to link the voices to partially occluded faces of unfamiliar speakers, and only when the words were produced with high talker information. These results reveal that 24-month-old infants can encode and retrieve indexical properties of an unfamiliar speaker&#x2019;s voice, and they can access this information even when visual access to the speaker&#x2019;s mouth is blocked.</p>
</abstract>
<kwd-group>
<kwd>voice recognition</kwd>
<kwd>speaker perception</kwd>
<kwd>infancy</kwd>
<kwd>indexical information</kwd>
<kwd>online study</kwd>
</kwd-group>
<contract-num rid="cn1">435-2019-0306</contract-num>
<contract-num rid="cn1">895-2020-1004</contract-num>
<contract-sponsor id="cn1">Social Sciences and Humanities Research Council of Canada<named-content content-type="fundref-id">10.13039/501100000155</named-content></contract-sponsor>
<counts>
<fig-count count="3"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="35"/>
<page-count count="9"/>
<word-count count="6430"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>In face-to-face conversations, we can tell who is speaking by looking at whose mouth is moving. Given the intersensory redundancy between this visual information and the resulting vocal signals, it is tempting to dismiss voice recognition as a trivial skill. However, there are often situations in which a listener may not have access to visual information&#x2014;for example, during auditory-only telecommunication, when facing away from the speaker, or when the speaker is wearing a face mask. Moreover, familiarity with a speaker&#x2019;s voice has consequences for social communication and linguistic processing (see <xref ref-type="bibr" rid="ref3">Creel and Bregman, 2011</xref>). Thus, examining how listeners recognize voices when visual facial information is partially occluded is important for our understanding of how humans process speech. In the current study, we approached this question by examining whether infants can detect a speaker change when the speaker&#x2019;s face is partially occluded.</p>
<p>From an early age, humans are surprisingly adept at tracking highly familiar voices. Previous studies have reported that upon hearing their mother&#x2019;s voice (and relative to hearing an unfamiliar woman&#x2019;s voice), fetuses&#x2019; heart rate increased (<xref ref-type="bibr" rid="ref11">Hepper et al., 1993</xref>; <xref ref-type="bibr" rid="ref15">Kisilevsky et al., 2003</xref>), newborns preferentially sucked on a pacifier in a non-nutritive sucking procedure (<xref ref-type="bibr" rid="ref4">DeCasper and Fifer, 1980</xref>), and 4-month-olds looked toward their mother (live: <xref ref-type="bibr" rid="ref30">Spelke and Owsley, 1979</xref>; photograph: <xref ref-type="bibr" rid="ref21">Orena and Werker, 2021</xref>). Infants can also differentiate between their father&#x2019;s voice and an unfamiliar male voice (<xref ref-type="bibr" rid="ref4">DeCasper and Fifer 1980</xref>; <xref ref-type="bibr" rid="ref34">Ward and Cooper, 1999</xref>). These studies show that, in the absence of synchronous visual information, infants can match highly familiar voices to the identities of familiar individuals.</p>
<p>A related question is whether infants can discriminate and learn to recognize unfamiliar voices with the same ease. Indeed, processing voices from <italic>unfamiliar</italic> talkers is a separate and often more challenging task than processing <italic>highly familiar</italic> voices (<xref ref-type="bibr" rid="ref31">Stevenage, 2017</xref>). Infants show more attentive and mature processing of speech spoken by highly familiar voices versus an unfamiliar speaker (<xref ref-type="bibr" rid="ref24">Purhonen et al., 2004</xref>; <xref ref-type="bibr" rid="ref19">Naoi et al., 2012</xref>). Unsurprisingly, with increased exposure to a certain set of speakers, adult listeners similarly improve at differentiating the voices of those speakers (e.g., <xref ref-type="bibr" rid="ref17">Levi et al., 2011</xref>).</p>
<p>Research-to-date indicates that infants are highly adept at discriminating between unfamiliar voices, particularly when the pair of voices are highly distinct from each other (e.g., <xref ref-type="bibr" rid="ref8">Floccia et al., 2000</xref>). Moreover, infants can pair male and female voices with male and female faces, respectively, by 8&#x2009;months of age (<xref ref-type="bibr" rid="ref22">Patterson and Werker 2003</xref>). Infants can also successfully discriminate between voices of same-gender pairs of speakers. In a series of studies, researchers reported that, after being habituated to an unfamiliar female voice, infants as young as 4&#x2009;months of age dishabituated to a new unfamiliar female voice (<xref ref-type="bibr" rid="ref13">Johnson et al., 2011</xref>; <xref ref-type="bibr" rid="ref5">Fecher and Johnson; 2018</xref>, <xref ref-type="bibr" rid="ref6">2019</xref>).</p>
<p>Of note, much of the infant work on processing and learning unfamiliar voices has been limited to tests of discrimination, rather than recognition. A recent study by <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> revealed that even 16.5-month-old infants have difficulty in learning the voices of unfamiliar speakers. In their task, infants were shown pairings of two voices and their identities (either cartoon characters or talking human faces). When the pair of voices involved one male and one female speaker, infants showed learning of the two voices in a preferential looking procedure experiment. However, when the pair of voices involved two female speakers, infants showed no recognition of either voice, suggesting that learning to recognize unfamiliar face-voice pairings may be a challenging task.</p>
<p>Like adults, certain factors appear to modulate infants&#x2019; talker processing abilities. For instance, it is now well-established that listeners are better at learning others&#x2019; voices when they are speaking a familiar versus an unfamiliar language (<xref ref-type="bibr" rid="ref9">Goggin et al., 1991</xref>). Some studies indicate that this effect is a function of phonological processing (<xref ref-type="bibr" rid="ref23">Perrachione et al., 2011</xref>; <xref ref-type="bibr" rid="ref14">Kadam et al., 2016</xref>). Yet, even long-term systematic exposure to a language appears to facilitate talker processing (<xref ref-type="bibr" rid="ref20">Orena et al., 2015</xref>), suggesting that benefits to talker processing could emerge prior to comprehension of the language. Indeed, even infants show a language-familiarity benefit to voice discrimination (<xref ref-type="bibr" rid="ref13">Johnson et al., 2011</xref>). Recently, <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> found that 4-month-olds could discriminate between female voices speaking a familiar language, but not when speaking an unfamiliar language. These findings show that phonetic and indexical information is integrated early in speech processing.</p>
<p>In the current study, we examined the nature of early talker processing skills by tackling two research questions. First, we examined the voice recognition abilities of young children. We followed up on work by <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> to investigate whether young children will show voice recognition of unfamiliar voices when cognitive demands are eased. In <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref>, infants learned the face&#x2212;voice pairings during a six-trial training phase before being tested on their recognition of the pairings during a two-trial test phase. In our study, infants were taught the face&#x2212;voice pairings and then tested on their recognition of the face&#x2212;voice pairings within the same trial. Given that 16.5-month-olds in <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> still had difficulty in learning the voices of unfamiliar speakers, we chose a higher age range (i.e., 24-months-old) to examine if it is more stable at this age point. We also tested a young age range (i.e., 12-month-olds) to examine whether young infants would succeed in this modified task. Importantly, in our work, in the test phase the faces were partially occluded.</p>
<p>As a secondary question, we asked whether, like for adults, phonetic content influences the voice discrimination abilities of young children. Work by <xref ref-type="bibr" rid="ref1">Andics et al. (2007)</xref> found that segmental information contributes to ease of voice discrimination for adults. In their study, listeners heard blocks of various consonant-vowel-consonant words and had to decide whether the word they heard was produced by the same voice as the preceding word or by a different voice. Results indicated that certain segments&#x2014;particularly segments that were more sonorant (e.g., [m], [s])&#x2014;helped listeners discriminate voices more than other segments. Here, we examined whether certain phonetic segments in speech are helpful for infants to discriminate between voices.</p>
</sec>
<sec id="sec2">
<title>Experiment 1</title>
<sec id="sec3">
<title>Methods</title>
<sec id="sec4">
<title>Participants</title>
<p>We recruited families through two avenues. First, as part of the Early Development Research Group at the University of British Columbia, we contacted families in our database to take part in the online study. To supplement this recruitment effort, we also publicly posted our study on the <italic>LookIt</italic> website. Data collection occurred between August 2020 and July 2021.</p>
<p>We recruited infants from two age groups. Thirty-one 12-month-old infants participated in the study, but five infants were excluded from the final sample because of technical issues (1), being too fussy during the session (2), and parental interference (2). Thus, we analyzed data from 26 12-month-old infants (mean age&#x2009;=&#x2009;384&#x2009;days; age range&#x2009;=&#x2009;366&#x2013;396; 12 girls and 14 boys). In addition, 38 24-month-old infants participated in the study. Eight infants were excluded from the final sample because of technical issues (2), being too fussy during the session (3), and for being above our age criteria (3). We analyzed data from the remaining 30 24-month-old infants (mean age&#x2009;=&#x2009;784&#x2009;days; age range&#x2009;=&#x2009;732&#x2013;774; 16 girls and 14 boys). Based on parent-report measures, all infants were exposed to English at least 90% of the time and had no known hearing or language impairments.</p>
</sec>
<sec id="sec5">
<title>Study Platform</title>
<p>The study was conducted in participants&#x2019; homes, through the MIT-run online platform <italic>Lookit</italic> (<xref ref-type="bibr" rid="ref27">Scott and Schulz, 2017</xref>). Families were provided a website link and asked to participate in the study at a time of their choosing. Upon entering the study page, caregivers were guided on how to set up their webcam and speakers. They were given an opportunity to preview stimuli without their child present prior to beginning the task.</p>
</sec>
<sec id="sec6">
<title>Stimuli</title>
<p>Visual stimuli consisted of six different animated human characters and four animated animal characters (i.e., dogs, chicken, cat, and owl). The characters were created and animated using the mobile apps <italic>Zepeto, Talkr,</italic> and <italic>Animoji.</italic> The animations were created such that the character&#x2019;s mouth was moving when speech was playing.</p>
<p>Auditory stimuli for human trials consisted of nine different non-words (one for the task familiarization phase, and eight for the test phase). The phonemes in the non-words were selected carefully to reflect two types of words. Words with sonorous phonemes (henceforth referred to as <italic>high talker information</italic>) included /yom/, /yen/, /won/, and /wem/. Words with less sonorous phonemes (henceforth referred to as <italic>low talker information</italic>) included /gut/, /gip/, /dup/, /dit/. The selection of these words follows <xref ref-type="bibr" rid="ref1">Andics et al. (2007)</xref>, who found higher performance in talker discrimination with words that had consonants and vowels that were relatively higher in the sonority hierarchy. The auditory stimuli for animal trials consisted of recordings of animal vocalizations (i.e., sounds made by a dog, chicken, cat, and owl).</p>
<p>The word for the task familiarization phase was spoken by two female speakers, and the words for the test phase were spoken by four other female speakers. All speakers learned English from birth. Stimuli were recorded using the speakers&#x2019; mobile phones and edited through <italic>Praat</italic>. Audio files were edited to match average intensity (70&#x2009;dB). Each trial consisted of a 13-s videoclip, created using a combination of <italic>Keynote</italic> and <italic>iMovie</italic>.</p>
</sec>
<sec id="sec7">
<title>Procedure</title>
<p>The experiment was conducted on the family&#x2019;s home computer. Caregivers were instructed to either hold their child such that the caregiver&#x2019;s back and the infant&#x2019;s face is facing the computer screen, or to sit beside, or behind their child with their eyes closed. Prior to the start of the study, parents were asked to ensure that the child&#x2019;s face was visible through their webcam. The experiment used a preferential looking paradigm consisting of two phases: a task familiarization phase and a test phase (see <xref rid="tab1" ref-type="table">Table 1</xref> for order of task familiarization and test trials). In both phases, we included both animal and human trials. The animal trials were included to engage infants and sustain their attention. Though not the focus of the current study, infants&#x2019; performance in the animal test trials also gave us an opportunity to put their performance in human test trials into context. We predicted that infants would succeed at matching the animal sounds to the animal faces across both age groups.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption><p>Summary of notes for trials in the task familiarization and test phases. Each trial was 14 seconds long. See <xref ref-type="fig" rid="fig1">Figure 1</xref> for time course of audio.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top" colspan="2">Task familiarization phase</th>
</tr>
<tr>
<th align="left" valign="top">Trials</th>
<th align="left" valign="top">Notes</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">A1, A2</td>
<td align="left" valign="top">There was one <italic>Same Speaker</italic> and one <italic>Different speaker</italic> trial. For each trial:</td>
</tr>
<tr>
<td rowspan="3"/>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>In the first 2 seconds, two animal cartoons appeared silently side-by-side</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 2 second mark, one animal made a sound and wiggled</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 9 second mark, either the same or different animal makes a sound and wiggled. No ferns descended.</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">A3, A4</td>
<td align="left" valign="top">There was one <italic>Same Speaker</italic> and one <italic>Different speaker</italic> trial. For each trial:</td>
</tr>
<tr>
<td rowspan="3"/>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>In the first 2 seconds, two animals appeared silently side-by-side</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 2 second mark, one animal made a sound and wiggled</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 9 second mark, ferns descended. Then, either the same or different animal made a sound and wiggled.</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">A5, A6</td>
<td align="left" valign="top">There was one <italic>Same Speaker</italic> and one <italic>Different speaker</italic> trial. For each trial:</td>
</tr>
<tr>
<td rowspan="3"/>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>In the first 2 seconds, two human cartoons appeared silently side-by-side</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 2 second mark, one human made a sound and wiggles</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet"><list-item>
<p>At the 9 second mark, ferns descended. Then, either the same or different human made a sound and wiggled.</p>
</list-item>
</list>
</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2"><bold>Test Phase</bold>
</td>
</tr>
<tr>
<td align="left" valign="top"><bold>Trials</bold></td>
<td align="left" valign="top"><bold>Notes</bold></td>
</tr>
<tr>
<td align="left" valign="top">B1-B4</td>
<td align="left" valign="top">Each set of four trials had:</td>
</tr>
<tr>
<td align="left" valign="top">B6-B9</td>
<td align="left" valign="top" rowspan="2">
<list list-type="bullet"><list-item>
<p>Two <italic>Same Speaker</italic> and one <italic>Different speaker</italic> trial</p>
</list-item>
<list-item>
<p>Two <italic>High</italic> and two <italic>Low talker information</italic> trials</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="2">B11-B14</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">For each trial (see <xref ref-type="fig" rid="fig1">Figure 1</xref>):<break/>
<list list-type="bullet"><list-item>
<p>In the first 2 seconds, two human cartoons appeared silently side-by-side</p>
</list-item>
<list-item>
<p>At the 2 second mark, one human made a sound.</p>
</list-item>
<list-item>
<p>At the 9 second mark, ferns descended. Then, either the same or different human made a sound.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">B16-</td>
</tr>
<tr>
<td align="left" valign="top">B19</td>
</tr>
<tr>
<td align="left" valign="top">B5</td>
<td align="left" valign="top">Across the four animal trials, there were:</td>
</tr>
<tr>
<td align="left" valign="top">B15</td>
<td align="left" valign="top" rowspan="3">
<list list-type="bullet"><list-item>
<p>Two <italic>Same speaker</italic> and two <italic>Different speaker</italic> trialsFor each trial:</p>
</list-item>
<list-item>
<p>In the first 2 seconds, two animal cartoons appeared silently side-by-side</p>
</list-item>
<list-item>
<p>At the 2 second mark, one animal made a sound.</p>
</list-item>
<list-item>
<p>At the 9 second mark, ferns descended. Then, either the same or different animal made a sound.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">B20</td>
</tr>
<tr>
<td align="left" valign="top">B10</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The task familiarization phase consisted of four animal and two human trials, for a total of six trials. The task familiarization phase introduced the eventual task in a gradual manner. The first two trials (A1 and A2) of the task familiarization phase began with two cartoon animals&#x2014;one on each side of the screen. After a silent period (2&#x2009;s), one of the animals would make a sound three times for 7&#x2009;s, then wiggle slightly to indicate that they were the animal making the noise. Then, either the same or the other animal would make a sound once and wiggle (5&#x2009;s). In the following two trials (A3 and A4), participants once again saw two cartoon animals. One of the animals would once again make a sound three times and wiggle. This time, a set of ferns descended to cover the animals&#x2019; mouths. After the ferns descended, either the same or the other animal would make a sound once while wiggling. For the final two task familiarization trials (A5 and A6), two females, human animations appeared&#x2014;one on each side of the screen. These trials began with one of the speakers producing a one-syllable nonsense word three times while wiggling. Then, similar to the previous trials, a set of ferns descended to cover the mouths of both speakers. Finally, either the same or the other speaker would produce a one-syllable nonsense word once while wiggling. If they could learn the pairing, we expected that infants would look toward the speaker producing the word based on the vocal properties that they heard.</p>
<p>The test phase consisted of four sets of five trials (four human trials, followed by one animal trial), for a total of 20 trials. In the human trials (B1&#x2013;4), participants were once again presented with two female, animated humans&#x2014;one on each side of the screen&#x2014;and one of the speakers would produce a one-syllable nonsense word three times for 7&#x2009;s. In contrast to the task familiarization trials, the speakers did not wiggle during the test trials. Then, ferns would once again descend to cover the speakers&#x2019; mouths. Once their mouths were covered, and during the critical window of analysis (WoA; 2&#x2009;s), one of the speakers would produce the same one-syllable nonsense word once. The animal trials (B5) proceeded in the same way as the human trials, except that there were two cartoon animals instead of two human animations. See <xref rid="fig1" ref-type="fig">Figure 1</xref> for a visual time course of the test trials.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>Visual schematic of the two types of trials in the test phase of the experiment: <italic>Same Speaker</italic> trial, and <italic>Different Speaker</italic> trial. <bold><italic>WoA</italic></bold> represents window of analysis (WoA), which was 2&#x2009;s long.</p></caption>
<graphic xlink:href="fpsyg-13-874411-g001.tif"/>
</fig>
<p>Critically, test trials varied in two ways: trial type (<italic>Same</italic> vs. <italic>Different Speaker</italic>) and talker information (<italic>High</italic> vs. <italic>Low Talker Information</italic>). During <italic>Same Speaker</italic> trials, the speaker during the critical WoA was the same speaker as the one who spoke first during the trial. During <italic>Different Speaker</italic> trials, the other speaker would produce the same one-syllable nonsense word once. For some of the trials, the one-syllable nonsense word consisted of sonorous phonemes (<italic>high talker information</italic> trials). On other trials, the word consisted of less sonorous phonemes (<italic>low talker information</italic> trials). Each set of trials was counter-balanced, such that each set consisted of two Same Speaker and Different Speaker trials each, as well as two high talker information and low talker information trials each.</p>
<p>Several measures were put into place to minimize any bias that infants may have toward a particular character or stimuli. For both phases, the position of the speaker (i.e., left vs. right) was counterbalanced through the experiment. It was equally likely for each character to be the initial speaker. We also used different auditory tokens across a single trial for human trials.</p>
<p>Participants&#x2019; behavior was recorded <italic>via</italic> webcam for the duration of all trials. Participants were given two opportunities to take breaks, during which their webcam was not recording. The first break occurred after the task familiarization phase, and the second break occurred halfway through the testing phase. The breaks were not timed and the participants chose when to resume the study.</p>
</sec>
<sec id="sec8">
<title>Data Preparation and Predictions</title>
<p>Based on previous work (<xref ref-type="bibr" rid="ref21">Orena and Werker, 2021</xref>), we preset the critical WoA to be 367&#x2009;ms after the critical utterance (9,367&#x2013;11,367&#x2009;ms). Note that the WoA was offset by 367&#x2009;ms as this is the reported duration of time needed for infants to initiate an eye movement after hearing speech sounds (<xref ref-type="bibr" rid="ref32">Swingley, 2012</xref>).</p>
<p>For all trials, the dependent variable was proportion looking time to the target voice of the critical utterance. Thus, a proportion looking time above 0.5 indicates that the child was proportionally looking to the correct speaker. In contrast, a proportion looking time below 0.5 indicates that the child was proportionally looking to the other speaker. A proportion looking time of 0.5 indicates that the child was looking at both speakers at chance levels. Trials in which infants looked at the screen for less than 1 s were excluded from analysis. We expected that infants would continue to look at the correct speaker during the <italic>Same Speaker</italic> trials. If infants were able to learn aspects of the speakers&#x2019; voices during the trial, then they should switch to also look at the correct speaker during the <italic>Different Speaker</italic> trials.</p>
</sec>
</sec>
<sec id="sec9">
<title>Results and Discussion</title>
<sec id="sec10">
<title>12-Month-Old Infants</title>
<p>Before conducting the main analysis, we first examined 12-month-old infants&#x2019; looking behaviors during the <italic>Animal</italic> trials. Infants&#x2019; proportion looking to target voice during the WoA is plotted in <xref rid="fig2" ref-type="fig">Figure 2</xref>. Surprisingly, 12-month-old infants&#x2019; were at chance levels for both <italic>Same Speaker</italic> and <italic>Different Speaker</italic> trials [<italic>t</italic>(24)&#x2009;=&#x2009;0.33, <italic>p</italic>&#x2009;=&#x2009;0.74, <italic>d</italic>&#x2009;=&#x2009;0.07 and <italic>t</italic>(25)&#x2009;=&#x2009;&#x2212;0.36, <italic>p</italic>&#x2009;=&#x2009;0.72, <italic>d</italic>&#x2009;=&#x2009;0.07, respectively].</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p>Twelve-month-old infants&#x2019; looking data during the critical WoA in the test phase, separated by trial type (<italic>same speaker</italic> vs. <italic>different speaker</italic>) and talker information (high talker vs. low talker). The looking data for animal trials are also represented on this graph. A value above 0.5 represents proportionally longer looking time to the target voice. The dotted line at 0.5 refers to equal proportion looking to both faces on the screen. Errors bars represent standard error.</p></caption>
<graphic xlink:href="fpsyg-13-874411-g002.tif"/>
</fig>
<p>Next, we examined whether 12-month-old infants showed any pattern of voice recognition during the main trials. We submitted infants&#x2019; proportion looking to the target voice to a repeated-measures ANOVA, with trial type (same speaker vs. different speaker) and talker information (high talker information vs. low talker information) as within-subjects&#x2019; factors. There was a significant main effect of Trial Type [<italic>F</italic>(1,25)&#x2009;=&#x2009;17.27, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, <inline-formula>
<mml:math id="M1">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.41], suggesting that infants performed differently across the two trial types, but no main effect of Talker Information [<italic>F</italic>(1,25)&#x2009;=&#x2009;0.29, <italic>p</italic>&#x2009;=&#x2009;0.59, <inline-formula>
<mml:math id="M2">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.01], nor an interaction between the two factors [<italic>F</italic>(1,25)&#x2009;=&#x2009;0.36, <italic>p</italic>&#x2009;=&#x2009;0.55, <inline-formula>
<mml:math id="M3">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.01]. <italic>T</italic>-tests against chance levels (0.5) indicated that during <italic>Same Speaker</italic> trials, infants continued to look at the correct speaker during the final part of the trial when they heard speakers produce words with <italic>high talker information</italic> [<italic>t</italic>(25)&#x2009;=&#x2009;2.07, <italic>p</italic>&#x2009;=&#x2009;0.04, <italic>d</italic>&#x2009;=&#x2009;0.41], but not when they heard speakers produce words with <italic>low talker information</italic> [<italic>t</italic>(25)&#x2009;=&#x2009;0.71, <italic>p</italic>&#x2009;=&#x2009;0.48, <italic>d</italic>&#x2009;=&#x2009;0.14]. During <italic>different speaker</italic> trials, infants did not shift to look at the correct speaker during the final part of the trial. Instead, they continued to look at the first initial speaker&#x2014;whether they produced words with <italic>high or low talker information</italic> [<italic>t</italic>(25)&#x2009;=&#x2009;&#x2212;2.03, <italic>p</italic>&#x2009;=&#x2009;0.05, <italic>d</italic>&#x2009;=&#x2009;0.40 and <italic>t</italic>(25)&#x2009;=&#x2009;&#x2212;2.22, <italic>p</italic>&#x2009;=&#x2009;0.03, <italic>d</italic>&#x2009;=&#x2009;0.44, respectively].</p>
<p>Taken together, these looking patterns do not provide any evidence that infants were responding to the vocal stimuli during the critical part of the trial. Instead, the data suggest that infants were merely continuing to look at the initial speaker of the trial, regardless of whose voice they heard during the critical trial portion.</p>
</sec>
<sec id="sec11">
<title>24-Month-Old Infants</title>
<p>Next, we examined 24-month-old infants&#x2019; looking behaviors during the experiment. Infants&#x2019; proportion looking to target voice during the WoA is plotted in <xref rid="fig3" ref-type="fig">Figure 3</xref>. During the <italic>Animal</italic> trials, 24-month-old infants&#x2019; looked toward the target animal upon hearing the animal sounds&#x2014;and this was the case for both <italic>Same Speaker</italic> and <italic>Different Speaker</italic> trials [<italic>t</italic>(22)&#x2009;=&#x2009;2.01, <italic>p</italic>&#x2009;=&#x2009;0.05, <italic>d</italic>&#x2009;=&#x2009;0.42 and <italic>t</italic>(22)&#x2009;=&#x2009;3.12, <italic>p</italic>&#x2009;=&#x2009;0.01, <italic>d</italic>&#x2009;=&#x2009;0.65, respectively].</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p>Twenty-four-month-old infants&#x2019; looking data during the critical WoA in the test phase, separated by trial type (s<italic>ame speaker</italic> vs. <italic>different speaker</italic>) and talker information (high talker vs. low talker). The looking data for animal trials are also represented on this graph. A value above 0.5 represents proportionally longer looking time to the target voice. The dotted line at 0.5 refers to equal proportion looking to both faces on the screen. Errors bars represent standard error.</p></caption>
<graphic xlink:href="fpsyg-13-874411-g003.tif"/>
</fig>
<p>We then examined whether 24-month-old infants showed any pattern of voice recognition during the main trials. Similar to the earlier analysis with younger infants, we submitted 24-month-old infants&#x2019; proportion looking to the target voice to repeated-measures ANOVA, with trial type (Same Speaker vs. Different Speaker) and talker information (high talker information vs. Low Talker Information) as within-subjects factors. There were no main effects of either Trial Type [<italic>F</italic>(1,23)&#x2009;=&#x2009;3.51, <italic>p</italic>&#x2009;=&#x2009;0.07, <inline-formula>
<mml:math id="M4">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.13] or Talker Information [<italic>F</italic>(1,23)&#x2009;=&#x2009;2.23, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.15, <inline-formula>
<mml:math id="M5">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.09]. There was no interaction between the two factors [<italic>F</italic>(1,23)&#x2009;=&#x2009;1.35, <italic>p</italic>&#x2009;=&#x2009;0.26, <inline-formula>
<mml:math id="M6">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x03B7;</mml:mi>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>&#x2009;=&#x2009;0.06]. Nonetheless, we conducted planned comparisons against chance levels. During <italic>Same Speaker</italic> trials, infants continued to look at the correct speaker during the final part of the trial when they heard speakers produce words with <italic>high talker information</italic> [<italic>t</italic>(23)&#x2009;=&#x2009;2.76, <italic>p</italic>&#x2009;=&#x2009;0.01, <italic>d</italic>&#x2009;=&#x2009;0.56], as well as when they heard speakers produce words with <italic>low talker information [t</italic>(25)&#x2009;=&#x2009;2.84, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, <italic>d</italic>&#x2009;=&#x2009;0.58]. Intriguingly, during <italic>different speaker</italic> trials, infants shifted to look at the correct speaker during the final part of the trial when they heard speakers produce words with <italic>high talker information</italic> [<italic>t</italic>(23)&#x2009;=&#x2009;2.06, <italic>p</italic>&#x2009;=&#x2009;0.05, <italic>d</italic>&#x2009;=&#x2009;0.42]. However, when they heard speakers produce words with <italic>low talker information,</italic> their proportion looking to both speakers was at chance levels [<italic>t</italic>(23)&#x2009;=&#x2009;&#x2212;0.37, <italic>p</italic>&#x2009;=&#x2009;0.71, <italic>d</italic>&#x2009;=&#x2009;0.08].</p>
<p>These findings indicate that 24-month-old infants were able to learn the initial speaker&#x2019;s voice when producing words with &#x201C;high talker information.&#x201D; They continued to correctly look at the initial speaker even after mouth movements were occluded. They also disambiguated and looked at a different speaker when they heard another voice after the ferns covered the characters&#x2019; mouths. However, they did not show the same pattern of looking behaviors when the speakers were producing words with &#x201C;low talker information.&#x201D;</p>
</sec>
</sec>
</sec>
<sec id="sec12" sec-type="discussions">
<title>Discussion</title>
<p>In the current study, we examined the voice recognition skills of young children. Specifically, we tested infants&#x2019; ability to detect a speaker change while speakers&#x2019; faces were partially obscured. Infants&#x2019; performance in the task revealed at least two important findings.</p>
<p>First, our findings show that stable voice recognition skills&#x2014;especially for unfamiliar voices&#x2014;are present by 24 months of age. This older group of infants looked proportionally more at the target voices during the preset WoA for both <italic>Same Speaker</italic> and <italic>Different Speaker</italic> trials. These findings suggest that 24-month-old infants were able to learn some aspect of the first speaker&#x2019;s voice such that: (i) they continued looking at her when they heard her voice again, even after objects blocked infants&#x2019; visual access to the speakers&#x2019; faces, and (ii) they showed a disambiguation response and looked toward another speaker when they heard another speaker&#x2019;s voice.</p>
<p>Second, our secondary analyzes reveal that the level of unfamiliar voice learning depends, in part, on the phonetic content of the spoken words&#x2014;findings that mirror those found with adults in <xref ref-type="bibr" rid="ref1">Andics et al. (2007)</xref>. The 24-month-olds showed a disambiguation response in <italic>Different Speaker</italic> trials only during trials when speakers were saying words with <italic>high talker information</italic> (i.e., tokens with sonorous phonemes). These findings confirm that, even for infants, phonetic content affects infants&#x2019; ability to learn unfamiliar voices. The results further confirm that the speech processing system that is sensitive to the integration of indexical and linguistic information is in place by late infancy (e.g., <xref ref-type="bibr" rid="ref001">Mulak et al., 2017</xref>).</p>
<p>Interpreting data from the 12-month-old infants is less straightforward. Regardless of who was speaking during the critical period of the trial, 12-month-old infants continued to look at the first speaker of the trial. Interestingly, 12-month-old infants also did not show signs of recognition for animal sounds. One interpretation for these findings is that unfamiliar voice recognition is a challenging task for young infants. For example, <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> found that 16.5-month-old infants were able to learn the voices of two unfamiliar speakers, but only when the acoustic differences between the speakers were large (one male and one female speaker). When there were two speakers of the same gender, 16.5-month-old infants did not show signs of voice learning.</p>
<p>Why might 12-month-old infants have difficulty with the current task? Firstly, the occluding ferns may have disrupted infants&#x2019; learning of the unfamiliar voices. Indeed, visual information can help adults encode and retain aspects of a speaker&#x2019;s voice (<xref ref-type="bibr" rid="ref28">Sheffert and Olsen, 2004</xref>), and access to synchronized visual information can facilitate infants&#x2019; performance in other speech processing tasks (<xref ref-type="bibr" rid="ref12">Hollich et al., 2005</xref>). A real-world, similar context is the use of face masks, which cover the lower half of a speaker&#x2019;s face. Some have raised concerns about whether this visual occluder may compromise speech processing (<xref ref-type="bibr" rid="ref10">Green et al., 2021</xref>). As the current study did not have an extra condition without any visual occluders, further research is needed to address this question.</p>
<p>Alternatively, it is possible that learning to recognize an unfamiliar voice is a challenging task for all 12-month-olds (as in <xref ref-type="bibr" rid="ref7">Fecher et al., 2019</xref>). Indeed, voice recognition is demonstrably more difficult than voice discrimination (see <xref ref-type="bibr" rid="ref23">Perrachione et al., 2011</xref> for a review of talker processing tasks). The discrimination of two auditory tokens can rely on low-level mechanisms, but associating a novel face with a novel voice is more cognitively demanding. Moreover, in the current study, infants heard speakers produce only three instances of one token before being tested on their recognition of that voice. Infants may thus have needed more exposure to the initial speaker, or a less artificial experiment, to show recognition. Indeed, <xref ref-type="bibr" rid="ref7">Fecher et al. (2019)</xref> found that learning voices is facilitated when the naturalness and social relevance of the task is increased. Nonetheless, the current study suggests that unfamiliar voice recognition is more challenging at 12 months than at 24 months.</p>
<p>It is important to note that the current study was conducted <italic>via</italic> the online platform, <italic>Lookit</italic> (<xref ref-type="bibr" rid="ref27">Scott and Schulz, 2017</xref>). This platform is the first large-scale crowdsourcing platform for conducting online developmental studies. There are multitude benefits of an online platform, including the ability to efficiently run participants, to recruit more diverse participants, and to continue research during laboratory shut-downs, such as during the COVID-19 pandemic. There has been some success in validating the use of these platforms, including findings that experimental conclusions derived <italic>via</italic> this platform are comparable to those of in-lab studies (<xref ref-type="bibr" rid="ref26">Scott et al., 2017</xref>; <xref ref-type="bibr" rid="ref29">Smith-Flores et al. 2022</xref>). Note, however, that effect sizes appear to be smaller for online studies, which may also explain some of our smaller effects. Others have raised concerns about the replicability of online <italic>Lookit</italic> experiments (<xref ref-type="bibr" rid="ref16">Lapidow et al., 2021</xref>), potentially due to parental interference. These concerns further temper our interpretation of the 12-month-old data, especially given that they were not successful in identifying the animals from their sounds.</p>
<p>To conclude, this study examined infants&#x2019; ability to learn unfamiliar voices. We found that 24-month-old infants were able to learn an unfamiliar voice sufficiently well to detect a voice change when objects blocked visual access to the speaker&#x2019;s mouths. These findings are highly relevant to the current pandemic and the increased use of face masks. Particularly, these findings reveal that 24-month-old infants can encode indexical properties of an unfamiliar speaker&#x2019;s voice, and they can access this information even when visual access to the speaker&#x2019;s mouth is blocked. Certainly, there are important follow-ups to provide firmer conclusions. For example, face masks affect the acoustics of speech production (<xref ref-type="bibr" rid="ref18">Llamas et al., 2008</xref>), and it would be interesting to investigate whether children&#x2019;s perception of speech and voices are affected by these alterations. Moreover, one could ask whether exposure to more speakers&#x2014;including more speakers with masks&#x2014;might promote learning in this domain. Indeed, prior work has shown that speaker variability promotes learning in the linguistic domain (e.g., <xref ref-type="bibr" rid="ref25">Rost and McMurray, 2009</xref>; but also see <xref ref-type="bibr" rid="ref2">Bergmann and Cristia, 2018</xref>). Nonetheless, the successful evidence showed that infants at 24 months could recognize and learn speaker voices even when the face is partially obscured, complements the growing research showing that adults are able to adapt to mask wearing, with equal recognition of both speech and emotional expressions (<xref ref-type="bibr" rid="ref33">Trainin and Yeshurun, 2021</xref>). Research on these and related topics will help to improve our understanding of how infants make use of the multisensory information around them to adapt to different contexts.</p>
</sec>
<sec id="sec13" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="sec14">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by the University of British Columbia, Behavioural Research Ethics Board. Written informed consent to participate in this study was provided by the participants&#x2019; caregiver.</p>
</sec>
<sec id="sec15">
<title>Author Contributions</title>
<p>AO, AM, and JW contributed to the conception and design of the study. AM created the stimuli set, coded the looking data, with support from other research assistants, and wrote a section of the manuscript. AO and AM set up the study on <italic>Lookit</italic>. AO performed the statistical analysis, with guidance from JW, and wrote the first draft of the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec16" sec-type="funding-information">
<title>Funding</title>
<p>This research was funded by the Social Sciences and Humanities Research Council of Canada (grants 435-2019-0306 and 895-2020-1004) to JW, an NSERC Undergraduate Student Research Award to AM, and a Fonds de Recherche du Qu&#x00E9;bec&#x2014;Nature et Technologies Postdoctoral Fellowship to AO.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec18" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack>
<p>We thank the members of our research team, particularly, S. Nijeboer and J. Cloake for their assistance in study co-ordination, J. Beaudin, L. Caswell, A. Cui, T. Chong, S. Rivera, N. Ambareswari, L. Yam for their assistance in recruiting participants and coding the data, the <italic>LookIt</italic> staff for providing technical support and the online platform to conduct the study, and Linda Polka and Katherine White for their helpful discussions.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Andics</surname> <given-names>A.</given-names></name> <name><surname>McQueen</surname> <given-names>J. M.</given-names></name> <name><surname>Van Turennout</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). &#x201C;Phonetic content influences voice discriminability.&#x201D; in <italic>Proceedings of the 16th International Congress of Phonetic Sciences (ICPhS 2007)</italic>. eds. J. Trouvain and W. J. Barry, 1829&#x2212;1832.</citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergmann</surname> <given-names>C.</given-names></name> <name><surname>Cristia</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Environmental influences on infants&#x2019; native vowel discrimination: the case of talker number in daily life</article-title>. <source>Infancy</source> <volume>23</volume>, <fpage>484</fpage>&#x2013;<lpage>501</lpage>. doi: <pub-id pub-id-type="doi">10.1111/infa.12232</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Creel</surname> <given-names>S. C.</given-names></name> <name><surname>Bregman</surname> <given-names>M. R.</given-names></name></person-group> (<year>2011</year>). <article-title>How talker identity relates to language processing</article-title>. <source>Lang. Linguist. Compass</source> <volume>5</volume>, <fpage>190</fpage>&#x2013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1749-818X.2011.00276.x</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeCasper</surname> <given-names>A. J.</given-names></name> <name><surname>Fifer</surname> <given-names>W. P.</given-names></name></person-group> (<year>1980</year>). <article-title>Of human bonding: newborns prefer their mothers&#x2019; voices</article-title>. <source>Science</source> <volume>208</volume>, <fpage>1174</fpage>&#x2013;<lpage>1176</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.7375928</pub-id>, PMID: <pub-id pub-id-type="pmid">7375928</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fecher</surname> <given-names>N.</given-names></name> <name><surname>Johnson</surname> <given-names>E. K.</given-names></name></person-group> (<year>2018</year>). <article-title>The native-language benefit for talker identification is robust in 7.5-month-old infants</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>44</volume>, <fpage>1911</fpage>&#x2013;<lpage>1920</lpage>. doi: <pub-id pub-id-type="doi">10.1037/xlm0000555</pub-id>, PMID: <pub-id pub-id-type="pmid">29698034</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fecher</surname> <given-names>N.</given-names></name> <name><surname>Johnson</surname> <given-names>E. K.</given-names></name></person-group> (<year>2019</year>). <article-title>By 4.5 months, linguistic experience already affects infants&#x2019; talker processing abilities</article-title>. <source>Child Dev.</source> <volume>90</volume>, <fpage>1535</fpage>&#x2013;<lpage>1543</lpage>. doi: <pub-id pub-id-type="doi">10.1111/cdev.13280</pub-id>, PMID: <pub-id pub-id-type="pmid">31273757</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fecher</surname> <given-names>N.</given-names></name> <name><surname>Paquette-Smith</surname> <given-names>M.</given-names></name> <name><surname>Johnson</surname> <given-names>E. K.</given-names></name></person-group> (<year>2019</year>). <article-title>Resolving the (apparent) talker recognition paradox in developmental speech perception</article-title>. <source>Infancy</source> <volume>24</volume>, <fpage>570</fpage>&#x2013;<lpage>588</lpage>. doi: <pub-id pub-id-type="doi">10.1111/infa.12290</pub-id>, PMID: <pub-id pub-id-type="pmid">32677248</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Floccia</surname> <given-names>C.</given-names></name> <name><surname>Nazzi</surname> <given-names>T.</given-names></name> <name><surname>Bertoncini</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <article-title>Unfamiliar voice discrimination for short stimuli in newborns</article-title>. <source>Dev. Sci.</source> <volume>3</volume>, <fpage>333</fpage>&#x2013;<lpage>343</lpage>. doi: <pub-id pub-id-type="doi">10.1111/1467-7687.00128</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goggin</surname> <given-names>J. P.</given-names></name> <name><surname>Thompson</surname> <given-names>C. P.</given-names></name> <name><surname>Strube</surname> <given-names>G.</given-names></name> <name><surname>Simental</surname> <given-names>L. R.</given-names></name></person-group> (<year>1991</year>). <article-title>The role of language familiarity in voice identification</article-title>. <source>Mem. Cogn.</source> <volume>19</volume>, <fpage>448</fpage>&#x2013;<lpage>458</lpage>. doi: <pub-id pub-id-type="doi">10.3758/BF03199567</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Green</surname> <given-names>J.</given-names></name> <name><surname>Staff</surname> <given-names>L.</given-names></name> <name><surname>Bromley</surname> <given-names>P.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Petty</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>The implications of face masks for babies and families during the COVID-19 pandemic: a discussion paper</article-title>. <source>J. Neonatal Nurs.</source> <volume>27</volume>, <fpage>21</fpage>&#x2013;<lpage>25</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jnn.2020.10.005</pub-id>, PMID: <pub-id pub-id-type="pmid">33162776</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hepper</surname> <given-names>P. G.</given-names></name> <name><surname>Scott</surname> <given-names>D.</given-names></name> <name><surname>Shahidullah</surname> <given-names>S.</given-names></name></person-group> (<year>1993</year>). <article-title>Newborn and fetal response to maternal voice</article-title>. <source>J. Reprod. Infant Psychol.</source> <volume>11</volume>, <fpage>147</fpage>&#x2013;<lpage>153</lpage>. doi: <pub-id pub-id-type="doi">10.1080/02646839308403210</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollich</surname> <given-names>G.</given-names></name> <name><surname>Newman</surname> <given-names>R. S.</given-names></name> <name><surname>Jusczyk</surname> <given-names>P. W.</given-names></name></person-group> (<year>2005</year>). <article-title>Infants&#x2019; use of synchronized visual information to separate streams of speech</article-title>. <source>Child Dev.</source> <volume>76</volume>, <fpage>598</fpage>&#x2013;<lpage>613</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-8624.2005.00866.x</pub-id>, PMID: <pub-id pub-id-type="pmid">15892781</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>E. K.</given-names></name> <name><surname>Westrek</surname> <given-names>E.</given-names></name> <name><surname>Nazzi</surname> <given-names>T.</given-names></name> <name><surname>Cutler</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Infant ability to tell voices apart rests on language experience</article-title>. <source>Dev. Sci.</source> <volume>14</volume>, <fpage>1002</fpage>&#x2013;<lpage>1011</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-7687.2011.01052.x</pub-id>, PMID: <pub-id pub-id-type="pmid">21884316</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kadam</surname> <given-names>M. A.</given-names></name> <name><surname>Orena</surname> <given-names>A. J.</given-names></name> <name><surname>Theodore</surname> <given-names>R. M.</given-names></name> <name><surname>Polka</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Reading ability influences native and non-native voice recognition, even for unimpaired readers</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>139</volume>, <fpage>EL6</fpage>&#x2013;<lpage>EL12</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.4937488</pub-id>, PMID: <pub-id pub-id-type="pmid">26827051</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kisilevsky</surname> <given-names>B. S.</given-names></name> <name><surname>Hains</surname> <given-names>S. M. J.</given-names></name> <name><surname>Lee</surname> <given-names>K. L.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Ye</surname> <given-names>H. H.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Effects of experience on fetal voice recognition</article-title>. <source>Psychol. Sci.</source> <volume>14</volume>, <fpage>220</fpage>&#x2013;<lpage>224</lpage>. doi: <pub-id pub-id-type="doi">10.1111/1467-9280.02435</pub-id>, PMID: <pub-id pub-id-type="pmid">12741744</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lapidow</surname> <given-names>E.</given-names></name> <name><surname>Tandon</surname> <given-names>T.</given-names></name> <name><surname>Goddu</surname> <given-names>M.</given-names></name> <name><surname>Walker</surname> <given-names>C. M.</given-names></name></person-group> (<year>2021</year>). <article-title>A tale of three platforms: investigating preschoolers&#x2019; second-order inferences using in-person, zoom, and Lookit methodologies</article-title>. <source>Front. Psychol.</source> <volume>12</volume>:<fpage>731404</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2021.731404</pub-id>, PMID: <pub-id pub-id-type="pmid">34721195</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levi</surname> <given-names>S. V.</given-names></name> <name><surname>Winters</surname> <given-names>S. J.</given-names></name> <name><surname>Pisoni</surname> <given-names>D. B.</given-names></name></person-group> (<year>2011</year>). <article-title>Effects of cross-language voice training on speech perception: whose familiar voices are more intelligible?</article-title> <source>J. Acoust. Soc. Am.</source> <volume>130</volume>, <fpage>4053</fpage>&#x2013;<lpage>4062</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.3651816</pub-id>, PMID: <pub-id pub-id-type="pmid">22225059</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Llamas</surname> <given-names>C.</given-names></name> <name><surname>Harrison</surname> <given-names>P.</given-names></name> <name><surname>Donnelly</surname> <given-names>D.</given-names></name> <name><surname>Watt</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Effects of different types of face coverings on speech acoustics and intelligibility</article-title>. <source>York Papers Ling. Ser.</source> <volume>2</volume>, <fpage>80</fpage>&#x2013;<lpage>104</lpage>.</citation></ref>
<ref id="ref001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mulak</surname> <given-names>K. E.</given-names></name> <name><surname>Bonn</surname> <given-names>C. D.</given-names></name> <name><surname>Chl&#x00E1;dkov&#x00E1;</surname> <given-names>K.</given-names></name> <name><surname>Aslin</surname> <given-names>R. N.</given-names></name> <name><surname>Escudero</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Indexical and linguistic processing by 12-month-olds: Discrimination of speaker, accent and vowel differences</article-title>. <source>Plos One</source> <volume>12</volume>:<fpage>e0176762</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0176762</pub-id>, PMID: <pub-id pub-id-type="pmid">15094248</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naoi</surname> <given-names>N.</given-names></name> <name><surname>Minagawa-Kawai</surname> <given-names>Y.</given-names></name> <name><surname>Kobayashi</surname> <given-names>A.</given-names></name> <name><surname>Takeuchi</surname> <given-names>K.</given-names></name> <name><surname>Nakamura</surname> <given-names>K.</given-names></name> <name><surname>Yamamoto</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Cerebral responses to infant-directed speech and the effect of talker familiarity</article-title>. <source>NeuroImage</source> <volume>59</volume>, <fpage>1735</fpage>&#x2013;<lpage>1744</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neuroimage.2011.07.093</pub-id>, PMID: <pub-id pub-id-type="pmid">21867764</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orena</surname> <given-names>A. J.</given-names></name> <name><surname>Theodore</surname> <given-names>R. M.</given-names></name> <name><surname>Polka</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Language exposure facilitates talker learning prior to language comprehension, even in adults</article-title>. <source>Cognition</source> <volume>143</volume>, <fpage>36</fpage>&#x2013;<lpage>40</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cognition.2015.06.002</pub-id>, PMID: <pub-id pub-id-type="pmid">26113447</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orena</surname> <given-names>A. J.</given-names></name> <name><surname>Werker</surname> <given-names>J. F.</given-names></name></person-group> (<year>2021</year>). <article-title>Infants&#x2019; mapping of new faces to new voices</article-title>. <source>Child Dev.</source> <volume>92</volume>, <fpage>e1048</fpage>&#x2013;<lpage>e1060</lpage>. doi: <pub-id pub-id-type="doi">10.1111/cdev.13616</pub-id>, PMID: <pub-id pub-id-type="pmid">34156089</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Patterson</surname> <given-names>M. L.</given-names></name> <name><surname>Werker</surname> <given-names>J. F.</given-names></name></person-group> (<year>2003</year>). <article-title>Two-month-old infants match phonetic information in lips and voice</article-title>. <source>Dev. Sci.</source> <volume>6</volume>, <fpage>191</fpage>&#x2013;<lpage>196</lpage>. doi: <pub-id pub-id-type="doi">10.1111/1467-7687.00271</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perrachione</surname> <given-names>T. K.</given-names></name> <name><surname>Del Tufo</surname> <given-names>S. N.</given-names></name> <name><surname>Gabrieli</surname> <given-names>J. D. E.</given-names></name></person-group> (<year>2011</year>). <article-title>Human voice recognition depends on language ability</article-title>. <source>Science</source> <volume>333</volume>, <fpage>595</fpage>. doi: <pub-id pub-id-type="doi">10.1126/science.1207327</pub-id>, PMID: <pub-id pub-id-type="pmid">21798942</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Purhonen</surname> <given-names>M.</given-names></name> <name><surname>Kilpel&#x00E4;inen-Lees</surname> <given-names>R.</given-names></name> <name><surname>Valkonen-Korhonen</surname> <given-names>M.</given-names></name> <name><surname>Karhu</surname> <given-names>J.</given-names></name> <name><surname>Lehtonen</surname> <given-names>J.</given-names></name></person-group> (<year>2004</year>). <article-title>Cerebral processing of mother&#x2019;s voice compared to unfamiliar voice in 4-month-old infants</article-title>. <source>Int. J. Psychol.</source> <volume>52</volume>, <fpage>257</fpage>&#x2013;<lpage>266</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ijpsycho.2003.11.003</pub-id>, PMID: <pub-id pub-id-type="pmid">15094248</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rost</surname> <given-names>G. C.</given-names></name> <name><surname>McMurray</surname> <given-names>B.</given-names></name></person-group> (<year>2009</year>). <article-title>Speaker variability augments phonological processing in early word learning</article-title>. <source>Dev. Sci.</source> <volume>12</volume>, <fpage>339</fpage>&#x2013;<lpage>349</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-7687.2008.00786.x</pub-id>, PMID: <pub-id pub-id-type="pmid">19143806</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scott</surname> <given-names>K.</given-names></name> <name><surname>Chu</surname> <given-names>J.</given-names></name> <name><surname>Schulz</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Lookit (part 2): assessing the viability of online development research, results from three case studies</article-title>. <source>Open Mind</source> <volume>1</volume>, <fpage>15</fpage>&#x2013;<lpage>29</lpage>. doi: <pub-id pub-id-type="doi">10.1162/opmi_a_00001</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scott</surname> <given-names>K.</given-names></name> <name><surname>Schulz</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Lookit (part 1): a new online platform for developmental research</article-title>. <source>Open Mind</source> <volume>1</volume>, <fpage>4</fpage>&#x2013;<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1162/opmi_a_00002</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sheffert</surname> <given-names>S. M.</given-names></name> <name><surname>Olsen</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Audiovisual speech facilitates voice learning</article-title>. <source>Percept. Psychophys.</source> <volume>66</volume>, <fpage>352</fpage>&#x2013;<lpage>362</lpage>. doi: <pub-id pub-id-type="doi">10.3758/bf03194884</pub-id>, PMID: <pub-id pub-id-type="pmid">15129754</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith-Flores</surname> <given-names>A. S.</given-names></name> <name><surname>Perez</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>M. H.</given-names></name> <name><surname>Feigenson</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>Online measures of looking and learning in infancy</article-title>. <source>Infancy</source> <volume>27</volume>, <fpage>4</fpage>&#x2013;<lpage>24</lpage>. doi: <pub-id pub-id-type="doi">10.1111/infa.12435</pub-id>, PMID: <pub-id pub-id-type="pmid">34524727</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spelke</surname> <given-names>E. S.</given-names></name> <name><surname>Owsley</surname> <given-names>C. J.</given-names></name></person-group> (<year>1979</year>). <article-title>Intermodal exploration and knowledge in infancy</article-title>. <source>Infant Behav. Dev.</source> <volume>2</volume>, <fpage>13</fpage>&#x2013;<lpage>27</lpage>. doi: <pub-id pub-id-type="doi">10.1016/s0163-6383(79)80004-1</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stevenage</surname> <given-names>S. V.</given-names></name></person-group> (<year>2017</year>). <article-title>Drawing a distinction between familiar and unfamiliar voice processing: A review of neuropsychological, clinical and empirical findings</article-title>. <source>Neuropsychologia</source> <volume>116</volume>, <fpage>162</fpage>&#x2013;<lpage>178</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neuropsychologia.2017.07.005</pub-id>, PMID: <pub-id pub-id-type="pmid">28694095</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swingley</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>The looking-while-listening procedure</article-title>. <source>Res. Meth. Child Lang.</source> <fpage>29</fpage>&#x2013;<lpage>42</lpage>. doi: <pub-id pub-id-type="doi">10.1002/9781444344035.ch3</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trainin</surname> <given-names>N.</given-names></name> <name><surname>Yeshurun</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Reading the mind with a mask? Improvement in reading the mind in the eyes during the COVID-19 pandemic</article-title>. <source>Emotion</source> <volume>21</volume>, <fpage>1801</fpage>&#x2013;<lpage>1806</lpage>. doi: <pub-id pub-id-type="doi">10.1037/emo0001014</pub-id>, PMID: <pub-id pub-id-type="pmid">34793184</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ward</surname> <given-names>C. D.</given-names></name> <name><surname>Cooper</surname> <given-names>R. P.</given-names></name></person-group> (<year>1999</year>). <article-title>A lack of evidence in 4-month-old human infants for paternal voice preference</article-title>. <source>Dev. Psychobiol.</source> <volume>35</volume>, <fpage>49</fpage>&#x2013;<lpage>59</lpage>. doi: <pub-id pub-id-type="doi">10.1002/(SICI)1098-2302(199907)35:1&#x003C;49::AID-DEV7&#x003E;3.0.CO;2-3</pub-id>, PMID: <pub-id pub-id-type="pmid">10397896</pub-id></citation></ref>
</ref-list>
</back>
</article>