<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychology</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychology</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Research Foundation</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2012.00010</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Listening for the Norm: Adaptive Coding in Speech Categorization</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Huang</surname> <given-names>Jingyuan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001">&#x0002A;</xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Holt</surname> <given-names>Lori L.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Psychology, Center for the Neural Basis of Cognition, Carnegie Mellon University</institution> <country>Pittsburgh, PA, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Peter Neri, University of Aberdeen, UK</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Simon Baumann, Newcastle University, UK; Emily Myers, University of Connecticut, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Jingyuan Huang, Department of Psychology, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA 15213, USA. e-mail: <email>jingyuan&#x00040;andrew.cmu.edu</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Frontiers in Perception Science, a specialty of Frontiers in Psychology.</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>02</month>
<year>2012</year>
</pub-date>
<pub-date pub-type="collection">
<year>2012</year>
</pub-date>
<volume>3</volume>
<elocation-id>10</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>10</month>
<year>2011</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>01</month>
<year>2012</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2012 Huang and Holt.</copyright-statement>
<copyright-year>2012</copyright-year>
<license license-type="open-access" xlink:href="http://www.frontiersin.org/licenseagreement"><p>This is an open-access article distributed under the terms of the <uri xlink:href="http://creativecommons.org/licenses/by-nc/3.0/">Creative Commons Attribution Non Commercial License</uri>, which permits non-commercial use, distribution, and reproduction in other forums, provided the original authors and source are credited.</p></license>
</permissions>
<abstract>
<p>Perceptual aftereffects have been referred to as &#x0201C;the psychologist&#x02019;s microelectrode&#x0201D; because they can expose dimensions of representation through the residual effect of a context stimulus upon perception of a subsequent target. The present study uses such context-dependence to examine the dimensions of representation involved in a classic demonstration of &#x0201C;talker normalization&#x0201D; in speech perception. Whereas most accounts of talker normalization have emphasized talker-, speech-, or articulatory-specific dimensions&#x02019; significance, the present work tests an alternative hypothesis: that the long-term average spectrum (LTAS) of speech context is responsible for patterns of context-dependent perception considered to be evidence for talker normalization. In support of this hypothesis, listeners&#x02019; vowel categorization was equivalently influenced by speech contexts manipulated to sound as though they were spoken by different talkers and non-speech analogs matched in LTAS to the speech contexts. Since the non-speech contexts did not possess talker, speech, or articulatory information, general perceptual mechanisms are implicated. Results are described in terms of adaptive perceptual coding.</p>
</abstract>
<kwd-group>
<kwd>talker normalization</kwd>
<kwd>LTAS</kwd>
<kwd>speech perception</kwd>
</kwd-group>
<counts>
<fig-count count="3"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="32"/>
<page-count count="6"/>
<word-count count="4324"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction">
<title>Introduction</title>
<p>Perceptual systems adjust rapidly to changes in the environment, with neural and behavioral responses dynamically changing to mirror changes in the input (Sharpee et al., <xref ref-type="bibr" rid="B27">2006</xref>; Gutinsky and Dragoi, <xref ref-type="bibr" rid="B7">2008</xref>). Such adaptive codes provide efficient representations because they direct computational resources toward uncommon inputs and provide information about potentially important changes in the world (Barlow, <xref ref-type="bibr" rid="B1">1990</xref>).</p>
<p>Although adaptive coding is less-well-studied in audition than vision, behavioral demonstrations of context-dependence, particularly in speech perception, resonate with the perspective that context adaptively tunes perceptual codes. The identity (Ladefoged and Broadbent, <xref ref-type="bibr" rid="B15">1957</xref>) or accent (Evans and Iverson, <xref ref-type="bibr" rid="B3">2004</xref>) of a talker, the rate of the utterance (Liberman et al., <xref ref-type="bibr" rid="B17">1956</xref>), and the phonetic make-up of a preceding utterance (Mann, <xref ref-type="bibr" rid="B19">1986</xref>) all influence perception of subsequent speech targets. From an adaptive coding perspective, the perceptual system&#x02019;s response to preceding speech lingers to affect subsequent processing.</p>
<p>Significantly, however, this can only be true to the extent that context and target share common neural resources. In this way, perceptual aftereffects of context have been described as &#x0201C;the psychologist&#x02019;s microelectrode&#x0201D; (Frisby, <xref ref-type="bibr" rid="B6">1980</xref>) because they can expose neural coding related to perceptual experience through the residual effect of one stimulus upon perception of another (Clifford and Rhodes, <xref ref-type="bibr" rid="B2">2005</xref>). Thus, context-dependent speech perception perhaps can serve to reveal the underlying representation of speech.</p>
<p>A simple model clarifies the approach (Figure <xref ref-type="fig" rid="F1">1</xref>, after Rhodes et al., <xref ref-type="bibr" rid="B25">2005</xref>). Imagine two populations of units coding a dimension of auditory representational space (e.g., acoustic frequency, or a higher-order feature like talker identity). Pool 1 units best code below-average values along the dimension whereas Pool 2 units better code above-average values; within each pool, more extreme values are coded more robustly. The average value along the dimension is encoded implicitly in the neutral cross-over point at which pools respond equivalently. An input with a higher value along the dimension would result in strong Pool 2 response, thus reducing Pool 2 responsiveness in the short-term due to adaptation. In this way, the context stimulus &#x0201C;lingers&#x0201D; in the perceptual system to affect the resources available to process subsequent stimuli. As is evident in Figure <xref ref-type="fig" rid="F1">1</xref>, this reduction shifts the neutral point where the two pools respond equivalently toward higher values and the formerly &#x0201C;average&#x0201D; value is now more robustly coded by Pool 1 neurons. Overall, adaptive encoding results in a contrastive shift in target encoding as a function of whether the context was better-coded by Pool 1 or Pool 2 neurons.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Model of adaptive coding adapted from Rhodes et al. (<xref ref-type="bibr" rid="B25">2005</xref>)</bold>.</p></caption>
<graphic xlink:href="fpsyg-03-00010-g001.tif"/>
</fig>
<p>This toy model is simple, to be sure, but the general principle may extend to higher-dimensional perceptual spaces and more complex representations. In vision, adaptive coding has been successful in predicting patterns of interactions among low-level representations for brightness and hue (see Frisby, <xref ref-type="bibr" rid="B6">1980</xref>) as well as higher-level representations for faces and bodies (see Clifford and Rhodes, <xref ref-type="bibr" rid="B2">2005</xref>). What is implicit in this approach is the assumption that context and target stimuli are encoded along common dimension(s) of representation. Here, we pursue context-dependent speech categorization to examine significant representational dimensions of speech categories.</p>
<p>For these purposes, Ladefoged and Broadbent&#x02019;s (<xref ref-type="bibr" rid="B15">1957</xref>) classic demonstration of talker normalization is relevant. In their study, listeners heard a constant target-word with a relatively ambiguous vowel at the end of a context phrase &#x0201C;<italic>Please say what this word is</italic>&#x02026;.&#x0201D; The first (F1) and/or second (F2) formant frequencies (peaks in energy of a voice spectrum; Fant, <xref ref-type="bibr" rid="B4">1960</xref>) of the context phrase were increased or decreased. These shifts can be conceptualized, respectively, as decreasing and increasing the talker&#x02019;s vocal tract length and, correspondingly, as a change in talker. When the resulting phrases preceded the speech targets, listeners&#x02019; categorization shifted in a manner suggesting that they were compensating, or normalizing, for the change in vocal tract length or talker. A constant vowel was more often heard as &#x0201C;bit&#x0201D; when it followed a phrase synthesized as though spoken by a shorter vocal tract (higher formant frequencies in the phrase), but more often as &#x0201C;bet&#x0201D; following the same phrase modeling speech from a longer vocal tract (lower frequencies).</p>
<p>A central and enduring theoretical issue has been the representational dimension across which listeners &#x0201C;normalize&#x0201D; speech categorization in this way. Is the relevant representational dimension talker identity, vocal tract shape/anatomy, or acoustic phonetic space (Joos, <xref ref-type="bibr" rid="B13">1948</xref>; Ladefoged and Broadbent, <xref ref-type="bibr" rid="B15">1957</xref>; Halle and Stevens, <xref ref-type="bibr" rid="B8">1962</xref>; Nordstrom and Lindblom, <xref ref-type="bibr" rid="B23">1975</xref>; McGowan, <xref ref-type="bibr" rid="B20">1997</xref>; McGowan and Cushing, <xref ref-type="bibr" rid="B21">1999</xref>; Poeppel et al., <xref ref-type="bibr" rid="B24">2008</xref>)?</p>
<p>Here, we investigate the extent to which Ladefoged and Broadbent&#x02019;s classic results can be explained by adaptive coding along a representational dimension that has a general auditory, rather that talker-, or speech-specific basis: the long-term average spectrum (LTAS) of the preceding sound. Recent research has suggested that listeners are sensitive to the LTAS of sound stimuli and adjust perception of subsequent sounds contrastively opposing context LTAS. Holt (<xref ref-type="bibr" rid="B9">2005</xref>, <xref ref-type="bibr" rid="B10">2006a</xref>,<xref ref-type="bibr" rid="B11">b</xref>) found that sequences of 21 non-speech sine-wave tones, each with a unique frequency sampling a 1000&#x02009;Hz range affect speech categorization of /ga/&#x02013;/da/ as a function of the mean frequency of the tones forming the sequence. The influence of these contexts on speech categorization cannot be attributed to any particular acoustic segment of the sequences because tones were randomly ordered on a trial-by-trial basis. Instead, the pattern of context-dependent speech categorization is predicted only by the tone sequences&#x02019; LTAS. Perception of the subsequent speech targets was relative to, and contrastive with, the LTAS consistent with the adaptive coding scheme sketched above if LTAS serves as the common representational dimension linking non-speech contexts and speech targets.</p>
<p>Here, we seek to directly replicate the Ladefoged and Broadbent (<xref ref-type="bibr" rid="B15">1957</xref>) results with speech contexts and to explicitly test whether non-speech contexts modeling critical characteristics of the speech contexts&#x02019; LTAS produce the same effects on vowel categorization. Directly mimicking the methods of Ladefoged and Broadbent (<xref ref-type="bibr" rid="B15">1957</xref>), listeners categorized a series of speech sounds varying perceptually from &#x0201C;bet&#x0201D; to &#x0201C;but&#x0201D; in the context of a preceding phrase (&#x0201C;Please say what this word is&#x02026;&#x0201D;). This phrase was synthesized to model two different &#x0201C;talkers&#x0201D;: one with a larger vocal tract and the other with a smaller vocal tract. The same listeners also categorized the same &#x0201C;bet&#x0201D; to &#x0201C;but&#x0201D; targets preceded by sequences of non-speech sine-wave tones with frequencies modeling the LTAS of the context phrases. These extremely simple acoustic contexts carried no talker- or speech-specific information. Should the speech and non-speech contexts similarly influence vowel categorization, it would suggest that speech and non-speech contexts draw upon common neural resources. Rather than articulatory or talker-specific dimensions that typically have been proposed to account for talker normalization, listeners may rely on sounds&#x02019; LTAS to tune speech perception.</p>
</sec>
<sec sec-type="materials|methods">
<title>Materials and Methods</title>
<sec>
<title>Participants</title>
<p>Twenty-six adult native-English speakers from the Carnegie Mellon University campus with no reported speech or hearing disabilities were recruited for the experiment. All received written informed consent in accord with Carnegie Mellon University ethics approval, and course credit for their time.</p>
</sec>
<sec>
<title>Stimuli</title>
<p>Figure <xref ref-type="fig" rid="F2">2</xref> illustrates stimulus design. Each stimulus had a 1600&#x02009;ms context segment, followed by a 50&#x02009;ms silent interval and a 250&#x02009;ms speech categorization target drawn from a six-step series of speech syllables varying perceptually from /b&#x003B5;t/ to /b&#x02227;t/ (<italic>bet</italic> to <italic>but</italic>). The /bV/ segment was created by varying the second formant (F2) frequency of the main vowel portion in equal steps from 1300&#x02009;Hz/b&#x02227;/ to 1700&#x02009;Hz/b&#x003B5;/ using Klattworks (McMurray, in preparation). The onset F2 frequency was 1100&#x02009;Hz and gradually changed to the target frequency across 50&#x02009;ms. Similarly, the first formant was 150&#x02009;Hz and linearly transitioned to 600&#x02009;Hz over 50&#x02009;ms. The fundamental frequency and the third formant frequency were held constant at 120 and 2600&#x02009;Hz, respectively. The /t/ segment was taken from a natural utterance of &#x0201C;<italic>whit</italic>&#x0201D; recorded from a male native-English talker and appended to each speech target.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Stimulus design and results of the experiment</bold>. <bold>(A)</bold> Schematic illustration of stimulus components; <bold>(B)</bold> spectrogram in time x frequency dimensions for the high mean speech context (top panel) and mean percentage of &#x0201C;but&#x0201D; responses in speech contexts (bottom panel); <bold>(C)</bold> spectrogram in time x frequency dimensions for a representative high mean tone context (top panel) and mean percentage of &#x0201C;but&#x0201D; responses in tone context (bottom panel). Preceded by both speech <bold>(B)</bold> and tone <bold>(C)</bold> contexts, higher-frequency contexts led to more low-frequency target responses (&#x0201C;but&#x0201D;), and vice versa.</p></caption>
<graphic xlink:href="fpsyg-03-00010-g002.tif"/>
</fig>
<p>Two types of context preceded these speech targets, one speech and the other a sequence of tones. The speech context was generated by extracting formant frequencies and bandwidths from a recording a male voice uttering the sentence &#x0201C;<italic>Please say what this word is</italic>&#x02026;,&#x0201D; and using these values to synthetically reproduce the sentence in the parallel branch of the Klatt and Klatt (<xref ref-type="bibr" rid="B14">1990</xref>) synthesizer. Following the methods of Ladefoged and Broadbent with modern techniques, this 1600&#x02009;ms base phrase was spectrally manipulated by adjusting formant center frequencies and bandwidths to create different &#x0201C;talkers.&#x0201D; To mimic a longer vocal tract, a voice with relatively lower frequencies in the region of F2 was created (across the phrase, F2 frequencies ranged from 390 to 1868&#x02009;Hz with an average of 1300&#x02009;Hz). A voice with a relatively shorter vocal tract was mimicked by increasing these base frequencies by 400&#x02009;Hz. Thus, the mean acoustic energy in the range of F2 approximated the energy varying across the /b&#x02227;t/&#x02013;/b&#x003B5;t/ speech target stimuli.</p>
<p>The non-speech tone contexts were composed of a sequence of 16 repeated 70&#x02009;ms sine-wave tones (5&#x02009;ms linear amplitude onset/offset ramps) with 30&#x02009;ms silent intervals separating them as in Holt (2005; 1600&#x02009;ms total duration). The tone contexts modeled the mean F2 frequency of the speech contexts (1300 and 1700&#x02009;Hz for the long and short vocal tracts, respectively) and each tone was a single harmonic without variation. As such, the non-speech contexts did not sound like speech or possess information about talker identity, vocal tract anatomy, or phonetic space. Thus, these non-speech contexts eliminated shared information between context and target along talker- and speech-specific dimensions while preserving a similar frequency-specific peak in the LTAS. Stimuli were RMS-matched in amplitude to the &#x0201C;bet&#x0201D; endpoint of the target-word series. All stimuli were sampled at 11025&#x02009;Hz.</p>
</sec>
<sec>
<title>Procedure</title>
<p>Participants categorized each target as &#x0201C;bet&#x0201D; or &#x0201C;but&#x0201D; using labeled keyboard buttons across a 1&#x02009;h experiment under the control of E-prime (Schneider et al., <xref ref-type="bibr" rid="B26">2002</xref>). Participants first categorized 10 randomly ordered repetitions of each speech target in isolation to assure that targets were well-categorized as the intended vowels. They then categorized the same speech targets preceded by high and low versions of speech and non-speech contexts, blocked by context type with block order counterbalanced across participants. Each context/target pairing was presented 10 times in a random order. Sounds were presented diotically over linear headphones (Beyer Dt-150) at approximately 70&#x02009;dB SPL(A).</p>
</sec>
</sec>
<sec>
<title>Results</title>
<p>Participants&#x02019; categorization of speech targets in isolation was orderly, as indicated by a significant main effect target F2 frequency, <italic>F</italic> (5, 25)&#x02009;&#x0003D;&#x02009;139.43, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01. Individuals&#x02019; data conformed to this average pattern.</p>
<p>Figures <xref ref-type="fig" rid="F2">2</xref>B,C illustrate the influence of context on vowel categorization. A 2 (context type, speech/non-speech) X 2 (LTAS, low/high) X 6 (target F2 frequency) repeated-measures ANOVA of listeners&#x02019; percent &#x0201C;but&#x0201D; responses reveals that, as expected from vowel categorization in isolation, responses varied reliably as a function of the target F2 frequency, <italic>F</italic> (5, 25)&#x02009;&#x0003D;&#x02009;290.50, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001. Moreover, there was no main effect of context type indicating no overall bias in vowel categorizations a function of the type of context that preceded targets, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;0.274, <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.61.</p>
<p>Of greater interest, there was a significant main effect of LTAS on vowel categorization, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;27.05, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001. The vowels were categorized as /&#x02227;/ significantly more often following high-frequency contexts whereas the same vowels were more often labeled as /&#x003B5;/ following low-frequency contexts. Thus, the context-dependent effect was spectrally contrastive and consistent with previous studies of the influence of sentence-length contexts on speech categorization (Ladefoged and Broadbent, <xref ref-type="bibr" rid="B15">1957</xref>; Watkins and Makin, <xref ref-type="bibr" rid="B30">1994</xref>, <xref ref-type="bibr" rid="B31">1996</xref>; Holt, <xref ref-type="bibr" rid="B9">2005</xref>, <xref ref-type="bibr" rid="B10">2006a</xref>,<xref ref-type="bibr" rid="B11">b</xref>; Huang and Holt, <xref ref-type="bibr" rid="B12">2009</xref>). The interaction between LTAS and target F2 frequency was significant, <italic>F</italic> (5, 25)&#x02009;&#x0003D;&#x02009;11.65, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001, indicating that context had a greater influence on perceptually ambiguous targets.</p>
<p>This pattern of contrastive context-dependent vowel categorization was evident for both speech, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;21.62, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001, and non-speech contexts, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;17.38, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001. Of primary interest, there was no significant interaction between context type and LTAS, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;0.324, <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.574. It is interesting that although the LTAS contrast was larger in the non-speech contexts compared with speech context condition (Figure <xref ref-type="fig" rid="F3">3</xref>; see detail explanation in discussion), the magnitude of the influence of speech and non-speech contexts on speech target categorization was statistically indistinguishable. Neither the interaction between context type and target frequency, <italic>F</italic> (1, 25)&#x02009;&#x0003D;&#x02009;1.22, <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.30, nor the three-way interaction was significant, <italic>F</italic> (5, 25)&#x02009;&#x0003D;&#x02009;1.424, <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.22. In sum, &#x0201C;talker&#x0201D; is not an essential element of talker normalization as it appears even in the absence of a talker when context is merely a sequence of sine-wave tones.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Long-term average spectrum of target stimuli (A), speech contexts (B), and tone contexts (C)</bold>. The shaded area indicates F2 frequencies critical to distinguishing the speech targets and two context types (speech and tone) in present experiments.</p></caption>
<graphic xlink:href="fpsyg-03-00010-g003.tif"/>
</fig>
</sec>
<sec sec-type="discussion">
<title>Discussion</title>
<p>We exploited context-dependent speech categorization as &#x0201C;the psychologist&#x02019;s microelectrode&#x0201D; (Frisby, <xref ref-type="bibr" rid="B6">1980</xref>) to reveal characteristics of the underlying representation of speech. The residual effect of one stimulus upon perception of another demands that the two share common neural processing and/or representation. Thus, the comparable influence of speech and non-speech contexts on speech categorization indicates a common substrate of interaction. Importantly, since the two context types did not share linguistic, articulatory gestural, or talker-specific information, but yet produced equivalent effects on speech categorization, it does not appear that speech-, vocal tract-, or talker-specific information is essential in eliciting the patterns of context-dependent perception that have been described in the literature as &#x0201C;talker normalization.&#x0201D;</p>
<p>What the two context types in the present experiment shared was a similar pattern of spectral energy across their time course. Figure <xref ref-type="fig" rid="F3">3</xref> illustrates the LTAS of the speech-target endpoints (3A) and speech (3B) and non-speech contexts (3C). Gray shading highlights the region of acoustic energy that critically distinguishes the speech targets. We suggest that the context-dependent speech categorization observed here (and in Ladefoged and Broadbent, <xref ref-type="bibr" rid="B15">1957</xref>) arises because the auditory system is sensitive to the context LTAS and the speech target is encoded relative to, and contrastive with, that long-term average.</p>
<p>Specifically, we propose that speech is adaptively coded according to the LTAS of ambient sound. The simple model described in the introduction can clarify one means by which this might be accomplished. Imagine pools of neurons sensitive to energy in the range of the second formant with one pool better coding relatively lower frequencies and the other better coding higher frequencies. By this model, presentation of the speech context modeling a longer vocal tract with lower-frequency energy within this frequency range would result in greater activity among the pool of neurons that better code lower frequencies. Subsequent adaptation among this pool of responsive neurons would result in a shift toward the opposite, higher-frequency neural pool at the time of speech target presentation. The formerly &#x0201C;neutral&#x0201D; frequencies would be now more robustly encoded by the higher-frequency neural pool, shifting representation contrastively away from the lower-frequency context. This adaptive coding serves to exaggerate differences between the LTAS of the context and target (Holt, <xref ref-type="bibr" rid="B10">2006a</xref>).</p>
<p>By this model, &#x0201C;talker normalization&#x0201D; effects on speech targets are predicted and obtained even when no speech information is available in the context. Of note, the LTAS model makes no reference to specific linguistic units, such as phonemes. This generality makes the adaptive coding approach in general, and the LTAS model in particular, extend straightforwardly to other normalization phenomena in speech perception. Huang and Holt (<xref ref-type="bibr" rid="B12">2009</xref>) report that shifts in the peak energy of the LTAS in the region of the fundamental frequency (f0) of a Mandarin Chinese sentence predict patterns of context-dependent Mandarin lexical tone normalization (Leather, <xref ref-type="bibr" rid="B16">1983</xref>; Fox and Qi, <xref ref-type="bibr" rid="B5">1990</xref>; Moore and Jongman, <xref ref-type="bibr" rid="B22">1997</xref>). Further, non-speech precursors with matched LTAS produce the same effects. Similarly, although the adaptive coding approach to context-dependent speech categorization reveals LTAS as an important dimension of representation, the model&#x02019;s application is more general. Wade and Holt (<xref ref-type="bibr" rid="B29">2005</xref>), for example, investigated rate-dependent normalization effects whereby the rate of a precursor sentence affects how listeners categorize a rate-dependent speech distinction like /ba/ versus /wa/, finding that the rate of presentation of a sequence of non-speech tones evokes a similar contrastive influence on speech categorization.</p>
<p>If demonstrations of &#x0201C;talker normalization&#x0201D; can be accounted for by general perceptual processes that are not talker- or speech-specific then there remains the question of whether the context-dependent speech categorization taken as evidence of normalization really accommodates the acoustic variability in speech that arises from different talkers. We argue that it does. Across context sentences like those of the current study, speech maps the scope of a talker&#x02019;s articulatory space and the mean of that space resembles a talker&#x02019;s neutral vowel, the shape of the non-articulating vocal tract (Story, <xref ref-type="bibr" rid="B28">2005</xref>). This neutral vowel serves as an effective normalization referent because, as Story (<xref ref-type="bibr" rid="B28">2005</xref>) has demonstrated, most of the variability across talkers can be accounted for by differences in the shape the vocal air space of talkers&#x02019; neutral vowels. Thus, if listeners were able to extract an estimate of the neutral vowel, the mean of the articulatory space, they would have an excellent referent for talker normalization. However, tracking back from acoustics to articulator requires negotiating the inverse problem. The results we describe here suggest an alternative.</p>
<p>As talkers produce a variety of consonants and vowels, speech maps the articulatory space but it also produces a sound spectrum that maps the acoustic space and, across time, samples an LTAS. Instead of solving the inverse problem to recover the actual neutral vocal tract shape as an articulatory referent for normalization, listeners may use the average spectrum LTAS as an auditory referent. The observation that non-speech tones modeling the LTAS of a talker are as effective in shifting speech categorization as speech contexts supports the viability of a general auditory referent. However, it should be noted that LTAS is unlikely to be the only contributing factor in talker normalization. Speaker identities, for example, may mediate listeners&#x02019; attention and expectation and influence the way listeners tune their perception to the preceding contexts (Magnuson and Nusbaum, <xref ref-type="bibr" rid="B18">2007</xref>). Talker normalization is likely to be a multi-facet phenomenon. Nonetheless LTAS, which provides sufficient context information via general perceptual processes, is an important factor in the adaptive coding of speech perception processing that contributes to talker normalization.</p>
<p>The lingering influence of a sequence of non-speech tones on listeners&#x02019; response to speech targets indicates a shared substrate, the level of which remains to be evaluated. Relevant to this, Holt (<xref ref-type="bibr" rid="B9">2005</xref>) reported that non-speech contexts influence on speech categorization persists across 1300&#x02009;ms of silence. Moreover, the aftereffect of the tones was present even when 13 neutral-frequency tones intervened between tone contexts and speech targets. The influence of temporally non-adjacent tone sequences can even override the influence of temporally adjacent speech contexts on speech targets (Holt, <xref ref-type="bibr" rid="B11">2006b</xref>). These observations argue for a central, rather than peripheral, substrate.</p>
<p>In fitting the mind to the world, the ambient context plays a large role in how input is coded. The present results demonstrate that patterns of context-dependent speech categorization long taken to be evidence of talker-specific normalization for articulatory referents or rescaling of phonetic space may arise, instead, from general principles of adaptive perceptual coding. In &#x0201C;listening for the norm,&#x0201D; listeners appear to adjust perception to regularities of the ambient environment.</p>
</sec>
<sec>
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>This work was supported by NIH grants R01DC004674.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H. B.</given-names></name></person-group> (<year>1990</year>). <article-title>&#x0201C;A theory about the functional role and synaptic mechanism of visual aftereffects,&#x0201D;</article-title> in <source>Vision: Coding and Efficiency</source>, ed. <person-group person-group-type="editor"><name><surname>Blakemore</surname> <given-names>C.</given-names></name></person-group> (<publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>), <fpage>363</fpage>&#x02013;<lpage>375</lpage>.</citation></ref>
<ref id="B2"><citation citation-type="book"><person-group person-group-type="editor"><name><surname>Clifford</surname> <given-names>C. W. G.</given-names></name> <name><surname>Rhodes</surname> <given-names>G.</given-names></name></person-group> (eds). (<year>2005</year>). <source>Fitting the Mind to the World: Adaptation and After Effects in High-Level Vision</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Evans</surname> <given-names>B. G.</given-names></name> <name><surname>Iverson</surname> <given-names>P.</given-names></name></person-group> (<year>2004</year>). <article-title>Vowel normalization for accent: an investigation of best exemplar locations in northern and southern British English sentences</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>115</volume>, <fpage>352</fpage>&#x02013;<lpage>261</lpage>.<pub-id pub-id-type="doi">10.1121/1.1635413</pub-id><pub-id pub-id-type="pmid">14759027</pub-id></citation></ref>
<ref id="B4"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Fant</surname> <given-names>G.</given-names></name></person-group> (<year>1960</year>). <source>Acoustic Theory of Speech Production</source>. <publisher-loc>The Hague</publisher-loc>: <publisher-name>Mouton &#x00026; Co</publisher-name>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fox</surname> <given-names>R.</given-names></name> <name><surname>Qi</surname> <given-names>Y.</given-names></name></person-group> (<year>1990</year>). <article-title>Contextual effects in the perception of lexical tone</article-title>. <source>J. Chin. Ling.</source> <volume>18</volume>, <fpage>261</fpage>&#x02013;<lpage>283</lpage>.</citation></ref>
<ref id="B6"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Frisby</surname> <given-names>J. P.</given-names></name></person-group> (<year>1980</year>). <source>Seeing: Illusion, Mind and Brain</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>OUP</publisher-name>.</citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gutinsky</surname> <given-names>D. A.</given-names></name> <name><surname>Dragoi</surname> <given-names>V.</given-names></name></person-group> (<year>2008</year>). <article-title>Adaptive coding of visual information in neural populations</article-title>. <source>Nature</source> <volume>452</volume>, <fpage>220</fpage>&#x02013;<lpage>224</lpage>.<pub-id pub-id-type="doi">10.1038/nature06563</pub-id><pub-id pub-id-type="pmid">18337822</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Halle</surname> <given-names>M.</given-names></name> <name><surname>Stevens</surname> <given-names>K. N.</given-names></name></person-group> (<year>1962</year>). <article-title>Speech recognition: a model and a program for research</article-title>. <source>IEEE Trans. Inf. Theory</source> <volume>8</volume>, <fpage>155</fpage>&#x02013;<lpage>159</lpage>.<pub-id pub-id-type="doi">10.1109/TIT.1962.1057686</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holt</surname> <given-names>L. L.</given-names></name></person-group> (<year>2005</year>). <article-title>Temporally non-adjacent non-linguistic sounds affect speech categorization</article-title>. <source>Psychol. Sci.</source> <volume>16</volume>, <fpage>305</fpage>&#x02013;<lpage>312</lpage>.<pub-id pub-id-type="doi">10.1111/j.0956-7976.2005.01532.x</pub-id><pub-id pub-id-type="pmid">15828978</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holt</surname> <given-names>L. L.</given-names></name></person-group> (<year>2006a</year>). <article-title>The mean matters: effects of statistically-defined nonspeech spectral distributions on speech categorization</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>120</volume>, <fpage>2801</fpage>&#x02013;<lpage>2817</lpage>.<pub-id pub-id-type="doi">10.1121/1.2354071</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holt</surname> <given-names>L. L.</given-names></name></person-group> (<year>2006b</year>). <article-title>Speech categorization in context: joint effects of nonspeech and speech precursors</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>119</volume>, <fpage>4016</fpage>&#x02013;<lpage>4026</lpage>.<pub-id pub-id-type="doi">10.1121/1.2195119</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Holt</surname> <given-names>L. L.</given-names></name></person-group> (<year>2009</year>). <article-title>General perceptual contributions to lexical tone normalization</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>125</volume>, <fpage>3983</fpage>&#x02013;<lpage>3994</lpage>.<pub-id pub-id-type="doi">10.1121/1.3097690</pub-id><pub-id pub-id-type="pmid">19507980</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joos</surname> <given-names>M.</given-names></name></person-group> (<year>1948</year>). <article-title>Acoustic phonetics</article-title>. <source>Language</source> <volume>24</volume>, <fpage>1</fpage>&#x02013;<lpage>136</lpage>.<pub-id pub-id-type="doi">10.2307/522229</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klatt</surname> <given-names>D. H.</given-names></name> <name><surname>Klatt</surname> <given-names>L. C.</given-names></name></person-group> (<year>1990</year>). <article-title>Analysis, synthesis, and perception of voice quality variations among female and male talkers</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>87</volume>, <fpage>820</fpage>&#x02013;<lpage>857</lpage>.<pub-id pub-id-type="doi">10.1121/1.398894</pub-id><pub-id pub-id-type="pmid">2137837</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ladefoged</surname> <given-names>P.</given-names></name> <name><surname>Broadbent</surname> <given-names>D. E.</given-names></name></person-group> (<year>1957</year>). <article-title>Information conveyed by vowels</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>29</volume>, <fpage>98</fpage>&#x02013;<lpage>104</lpage>.<pub-id pub-id-type="doi">10.1121/1.1908694</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leather</surname> <given-names>J.</given-names></name></person-group> (<year>1983</year>). <article-title>Speaker normalization in perception of lexical tone</article-title>. <source>J. Phon.</source> <volume>11</volume>, <fpage>373</fpage>&#x02013;<lpage>382</lpage>.</citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liberman</surname> <given-names>A. M.</given-names></name> <name><surname>Delattre</surname> <given-names>P. C.</given-names></name> <name><surname>Gerstman</surname> <given-names>L. J.</given-names></name> <name><surname>Cooper</surname> <given-names>F. S.</given-names></name></person-group> (<year>1956</year>). <article-title>Tempo of frequency change as a cue for distinguishing classes of speech sounds</article-title>. <source>J. Exp. Psychol.</source> <volume>52</volume>, <fpage>127</fpage>&#x02013;<lpage>137</lpage>.<pub-id pub-id-type="doi">10.1037/h0041240</pub-id><pub-id pub-id-type="pmid">13345983</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Magnuson</surname> <given-names>J. S.</given-names></name> <name><surname>Nusbaum</surname> <given-names>H. C.</given-names></name></person-group> (<year>2007</year>). <article-title>Acoustic differences, listener expectations, and the perceptual accommodation of talker variability</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform.</source> <volume>33</volume>, <fpage>391</fpage>&#x02013;<lpage>409</lpage>.<pub-id pub-id-type="doi">10.1037/0096-1523.33.2.391</pub-id><pub-id pub-id-type="pmid">17469975</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mann</surname> <given-names>V. A.</given-names></name></person-group> (<year>1986</year>). <article-title>Distinguishing universal language-specific factors in speech perception: evidence from Japanese listeners&#x02019; perception of /1/ and /r/</article-title>. <source>Cognition</source> <volume>24</volume>, <fpage>169</fpage>&#x02013;<lpage>196</lpage>.<pub-id pub-id-type="doi">10.1016/0010-0277(86)90005-3</pub-id><pub-id pub-id-type="pmid">3816123</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGowan</surname> <given-names>R. S.</given-names></name></person-group> (<year>1997</year>). <article-title>Normalization for articulatory recovery</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>101</volume>, <fpage>3175</fpage>.<pub-id pub-id-type="doi">10.1121/1.418310</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGowan</surname> <given-names>R. S.</given-names></name> <name><surname>Cushing</surname> <given-names>S.</given-names></name></person-group> (<year>1999</year>). <article-title>Vocal tract normalization for midsagittal articulatory recovery with analysis-by-synthesis</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>106</volume>, <fpage>1090</fpage>&#x02013;<lpage>1105</lpage>.<pub-id pub-id-type="doi">10.1121/1.427117</pub-id><pub-id pub-id-type="pmid">10462814</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moore</surname> <given-names>C.</given-names></name> <name><surname>Jongman</surname> <given-names>A.</given-names></name></person-group> (<year>1997</year>). <article-title>Speaker normalization in the perception of Mandarin Chinese tones</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>102</volume>, <fpage>1864</fpage>&#x02013;<lpage>1877</lpage>.<pub-id pub-id-type="doi">10.1121/1.420350</pub-id><pub-id pub-id-type="pmid">9301064</pub-id></citation></ref>
<ref id="B23"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Nordstrom</surname> <given-names>P. E.</given-names></name> <name><surname>Lindblom</surname> <given-names>B.</given-names></name></person-group> (<year>1975</year>). <article-title>&#x0201C;A normalization procedure for vowel formant data,&#x0201D;</article-title> in <conf-name>Paper presented at the 8th International Congress of Phonetic Sciences</conf-name>, <conf-loc>Leed</conf-loc>, <fpage>17</fpage>&#x02013;<lpage>23</lpage>.</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poeppel</surname> <given-names>D.</given-names></name> <name><surname>Idsardi</surname> <given-names>W. J.</given-names></name> <name><surname>van Wassenhove</surname> <given-names>V.</given-names></name></person-group> (<year>2008</year>). <article-title>Speech perception at the interface of neurobiology and linguistics</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>363</volume>, <fpage>1071</fpage>&#x02013;<lpage>1086</lpage>.<pub-id pub-id-type="doi">10.1098/rstb.2007.2160</pub-id><pub-id pub-id-type="pmid">17890189</pub-id></citation></ref>
<ref id="B25"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Rhodes</surname> <given-names>G.</given-names></name> <name><surname>Robbins</surname> <given-names>R.</given-names></name> <name><surname>Jaquet</surname> <given-names>E.</given-names></name> <name><surname>McKone</surname> <given-names>E.</given-names></name> <name><surname>Jeffery</surname> <given-names>L.</given-names></name> <name><surname>Clifford</surname> <given-names>C. W. G.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;Adaptation and face perception &#x02013; how aftereffects implicate norm based coding of faces,&#x0201D;</article-title> in <source>Fitting the Mind to the World: Aftereffects in High-Level Vision</source>, eds <person-group person-group-type="editor"><name><surname>Clifford</surname> <given-names>C. W. G.</given-names></name> <name><surname>Rhodes</surname> <given-names>G.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>), <fpage>213</fpage>&#x02013;<lpage>240</lpage>.</citation></ref>
<ref id="B26"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Schneider</surname> <given-names>W.</given-names></name> <name><surname>Eschman</surname> <given-names>A.</given-names></name> <name><surname>Zuccolotto</surname> <given-names>A.</given-names></name></person-group> (<year>2002</year>). <source>E-Prime User&#x02019;s Guide</source>. <publisher-loc>Pittsburgh</publisher-loc>: <publisher-name>Psychology Software Tools Inc</publisher-name>.</citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharpee</surname> <given-names>T. O.</given-names></name> <name><surname>Sugihara</surname> <given-names>H.</given-names></name> <name><surname>Kurgansky</surname> <given-names>A. V.</given-names></name> <name><surname>Rebrik</surname> <given-names>S. P.</given-names></name> <name><surname>Stryker</surname> <given-names>M. P.</given-names></name> <name><surname>Miller</surname> <given-names>K. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Adaptive filtering enhances information transmission in visual cortex</article-title>. <source>Nature</source> <volume>439</volume>, <fpage>936</fpage>&#x02013;<lpage>942</lpage>.<pub-id pub-id-type="doi">10.1038/nature04519</pub-id><pub-id pub-id-type="pmid">16495990</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Story</surname> <given-names>B. H.</given-names></name></person-group> (<year>2005</year>). <article-title>A parametric model of the vocal tract area function for vowel and consonant simulation</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>117</volume>, <fpage>3231</fpage>&#x02013;<lpage>3234</lpage>.<pub-id pub-id-type="doi">10.1121/1.1869752</pub-id><pub-id pub-id-type="pmid">15957790</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wade</surname> <given-names>T.</given-names></name> <name><surname>Holt</surname> <given-names>L. L.</given-names></name></person-group> (<year>2005</year>). <article-title>Perceptual effects of preceding non-speech rate on temporal properties of speech categories</article-title>. <source>Percept. Psychophys.</source> <volume>67</volume>, <fpage>939</fpage>&#x02013;<lpage>950</lpage>.<pub-id pub-id-type="doi">10.3758/BF03193621</pub-id><pub-id pub-id-type="pmid">16396003</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watkins</surname> <given-names>A. J.</given-names></name> <name><surname>Makin</surname> <given-names>S. J.</given-names></name></person-group> (<year>1994</year>). <article-title>Perceptual compensation for speaker differences and for spectral-envelope distortion</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>96</volume>, <fpage>1263</fpage>&#x02013;<lpage>1282</lpage>.<pub-id pub-id-type="doi">10.1121/1.410275</pub-id><pub-id pub-id-type="pmid">7962994</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watkins</surname> <given-names>A. J.</given-names></name> <name><surname>Makin</surname> <given-names>S. J.</given-names></name></person-group> (<year>1996</year>). <article-title>Effects of spectral contrast on perceptual compensation for spectral-envelope distortions</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>99</volume>, <fpage>3749</fpage>&#x02013;<lpage>3757</lpage>.<pub-id pub-id-type="doi">10.1121/1.414981</pub-id><pub-id pub-id-type="pmid">8655806</pub-id></citation></ref>
</ref-list>
</back>
</article>