<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2023.1228506</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Sound category habituation requires task-relevant attention</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Moskowitz</surname>
<given-names>Howard S.</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2540364/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Sussman</surname>
<given-names>Elyse S.</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1049/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Otorhinolaryngology-Head and Neck Surgery, Albert Einstein College of Medicine</institution>, <addr-line>Bronx, NY</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Neuroscience, Albert Einstein College of Medicine</institution>, <addr-line>Bronx, NY</addr-line>, <country>Unites States</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Bruno L. Giordano, UMR7289 Institut de Neurosciences de la Timone (INT), France</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Justin Yao, The State University of New Jersey &#x2013; Busch Campus, United States; Giorgio Marinato, UMR7289 Institut de Neurosciences de la Timone (INT), France</p></fn>
<corresp id="c001">&#x002A;Correspondence: Elyse S. Sussman, <email>elyse.sussman@einsteinmed.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>24</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1228506</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Moskowitz and Sussman.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Moskowitz and Sussman</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Processing the wealth of sensory information from the surrounding environment is a vital human function with the potential to develop learning, advance social interactions, and promote safety and well-being.</p>
</sec>
<sec>
<title>Methods</title>
<p>To elucidate underlying processes governing these activities we measured neurophysiological responses to patterned stimulus sequences during a sound categorization task to evaluate attention effects on implicit learning, sound categorization, and speech perception. Using a unique experimental design, we uncoupled conceptual categorical effects from stimulus-specific effects by presenting categorical stimulus tokens that did not physically repeat.</p>
</sec>
<sec>
<title>Results</title>
<p>We found effects of implicit learning, categorical habituation, and a speech perception bias when the sounds were attended, and the listeners performed a categorization task (task-relevant). In contrast, there was no evidence of a speech perception bias, implicit learning of the structured sound sequence, or repetition suppression to repeated within-category sounds (no categorical habituation) when participants passively listened to the sounds and watched a silent closed-captioned video (task-irrelevant). No indication of category perception was demonstrated in the scalp-recorded brain components when participants were watching a movie and had no task with the sounds.</p>
</sec>
<sec>
<title>Discussion</title>
<p>These results demonstrate that attention is required to maintain category identification and expectations induced by a structured sequence when the conceptual information must be extracted from stimuli that are acoustically distinct. Taken together, these striking attention effects support the theoretical view that top-down control is required to initiate expectations for higher level cognitive processing.</p>
</sec>
</abstract>
<kwd-group>
<kwd>categorical perception</kwd>
<kwd>attention</kwd>
<kwd>speech</kwd>
<kwd>implicit learning</kwd>
<kwd>event-related brain potentials (ERPs)</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="47"/>
<page-count count="11"/>
<word-count count="8564"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Auditory Cognitive Neuroscience</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1.</label>
<title>Introduction</title>
<p>In the modern-day world we are constantly inundated by a cavalcade of sensory input from our surrounding environment. Our brains must process this information efficiently to facilitate appropriate reactions and responses. The ability to perceive and monitor sounds to which we are not specifically attending serves an important role for general functioning and safety. However, a gate or mechanism must exist to govern this important function. How and when we make the decision regarding which sounds meet a sufficient level of importance would promote welfare while also controlling utilization of important cognitive and behavioral functions that could be directed elsewhere.</p>
<p>The exact manner by which these processes are generated continues to be a source of research and controversy. The classical view of perception maintains a feedforward view: information is received from the environment, processed at higher brain levels and then a response to the input is generated (<xref ref-type="bibr" rid="ref26">Mumford, 1992</xref>). Any mismatch with the actual sensory input constantly updates the prediction based on this error signal (<xref ref-type="bibr" rid="ref35">Rao and Ballard, 1999</xref>). An alternative concept suggests that we predict the nature of incoming sensory information based on previous experiences, to process efficiently and to allocate resources to novel stimuli (<xref ref-type="bibr" rid="ref12">Friston, 2005</xref>). The concept governing this process is explained by predictive processing theories that suggest that the brain generates models that automatically anticipate and predict upcoming sensory input based on the recent history of the sensory input (<xref ref-type="bibr" rid="ref6">Clark, 2013</xref>). A predictive model is generated in higher cortical areas and is communicated through feedback connections to lower sensory areas (<xref ref-type="bibr" rid="ref35">Rao and Ballard, 1999</xref>; <xref ref-type="bibr" rid="ref12">Friston, 2005</xref>). Recently, the concept of predictive processing has been validated by several brain imaging studies investigating predictive feedback and the processing of prediction errors (<xref ref-type="bibr" rid="ref7">den Ouden et al., 2009</xref>; <xref ref-type="bibr" rid="ref13">Friston and Kiebel, 2009</xref>; <xref ref-type="bibr" rid="ref8">Egner et al., 2010</xref>; <xref ref-type="bibr" rid="ref19">Jiang et al., 2013</xref>; <xref ref-type="bibr" rid="ref2">Alink and Blank, 2021</xref>). However, these models do not take into account the precise nature of the stimulus input on which the predictions are based, or how they are established or maintained in memory. Accordingly, these issues are still being debated (<xref ref-type="bibr" rid="ref45">Walsh et al., 2020</xref>). There is an essential lack of understanding of (1) the role attention plays in forming the predictions themselves; and (2) how the predictions are instantiated, such as whether they are based on simple stimulus repetition, attentional control, or something else (<xref ref-type="bibr" rid="ref35">Rao and Ballard, 1999</xref>; <xref ref-type="bibr" rid="ref14">Friston et al., 2006</xref>; <xref ref-type="bibr" rid="ref39">Summerfield et al., 2008</xref>; <xref ref-type="bibr" rid="ref4">Bubic et al., 2010</xref>; <xref ref-type="bibr" rid="ref21">Larsson and Smith, 2012</xref>; <xref ref-type="bibr" rid="ref6">Clark, 2013</xref>; <xref ref-type="bibr" rid="ref45">Walsh et al., 2020</xref>).</p>
<p>The current study tests hypotheses that investigate these questions within predictive processing theories. Namely, we designed a study that dissociated stimulus-specific adaptation from semantic categorical repetition to determine whether higher-level, conceptual expectations can be encoded from stimulus repetition when the stimulus repetition is based on category membership and there is no repetition in the acoustic characteristics of the sound tokens. To do this, we implemented a novel paradigm for measuring repetition suppression to assess the reduction of neural activity to repeated stimuli (<xref ref-type="bibr" rid="ref25">Moskowitz et al., 2020</xref>). Participants heard sounds presented in groups consisting of stimuli by semantic category (spoken words, sounds of musical instruments, environmental sounds) (<xref rid="fig1" ref-type="fig">Figure 1</xref>). Each four-stimulus category group was randomly followed by another group of four sounds that were from a different sound category (switch) or from the same category (repeat). There were no repeated stimulus tokens: No sounds were physical repeats of any other sound in the stimulus blocks, only the category was repeated or switched. Thus, a change in response induced by repetition of the category could not be specifically due to token-specific sensory adaptation or to repetition suppression.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>Schematic of the experimental paradigm. Sounds from three categories &#x2013; music (M, red font), environment (E, green font), and speech (S, blue font) &#x2013; were presented in categorical groups of four stimuli (denoted by solid line in their respective color, and numbered 1&#x2013;4). Randomly, the category switched to a new category (indicated with an arrow and labeled &#x2018;Category Switch&#x2019;) or the category repeated (indicated with an arrow and labeled &#x2018;Category Repeat&#x2019;). When the category randomly repeated, eight successive tokens from the same category were presented (denoted by a solid line and numbered 5&#x2013;8 for the repeated category). A random sample of the sounds from each category are depicted with their respective spectrograms, and labeled above (see <xref rid="SM1" ref-type="supplementary-material">Appendix I</xref> for full list of the sounds). Each within-category token was unique in every four-token group. Stimuli were presented isochronously, one sound every 1.1&#x2009;s, in each stimulus block, with no demarcation of the switch or repeat trials.</p></caption>
<graphic xlink:href="fnins-17-1228506-g001.tif"/>
</fig>
<p>We used event-related brain potentials during passive auditory and active auditory listening conditions to measure the brain&#x2019;s response to the same categorical sounds when the categorical aspect of the sounds was relevant compared to when the categorical aspect of the task was irrelevant. The P3a component reflects involuntary orienting to a salient sound regardless of the direction of attention (<xref ref-type="bibr" rid="ref11">Friedman et al., 2001</xref>; <xref ref-type="bibr" rid="ref32">Polich, 2007</xref>; <xref ref-type="bibr" rid="ref10">Fonken et al., 2020</xref>), with its greatest amplitude at frontocentral locations (e.g., Cz electrode). The P3b component is a non-modality-specific index of task-related processing (<xref ref-type="bibr" rid="ref42">Sutton et al., 1965</xref>; <xref ref-type="bibr" rid="ref30">Picton, 1992</xref>) and generally has its largest amplitude over parietal scalp electrodes (e.g., Pz) (<xref ref-type="bibr" rid="ref32">Polich, 2007</xref>). The P3b is elicited when attention is focused on stimuli to identify a target (<xref ref-type="bibr" rid="ref33">Polich and Criado, 2006</xref>). Therefore, it is elicited by task-relevant but not task-irrelevant stimuli, and can reflect category perception (<xref ref-type="bibr" rid="ref23">Maiste et al., 1995</xref>). The different neural substrates of the P3a (at Cz) and P3b (at Pz) components reflect different aspects of attention (attentional orienting and target detection, respectively) (<xref ref-type="bibr" rid="ref10">Fonken et al., 2020</xref>). The sensory-specific N1 component is an obligatory response of the ERPs, elicited by sound onsets and its amplitude is reduced with stimulus repetition (<xref ref-type="bibr" rid="ref28">N&#x00E4;&#x00E4;t&#x00E4;nen and Picton, 1987</xref>; <xref ref-type="bibr" rid="ref5">Budd et al., 1998</xref>; <xref ref-type="bibr" rid="ref18">Hsu et al., 2016</xref>; <xref ref-type="bibr" rid="ref36">Rosburg and Mager, 2021</xref>) and increased with focused attention (<xref ref-type="bibr" rid="ref17">Hillyard et al., 1998</xref>). These primary dependent measures, the P3a, P3b, and N1 components, provided neural responses to the categorical sounds that indexed involuntary orienting to sounds during passive and active auditory task listening (P3a) observable at the Cz electrode, an index of task performance during active auditory task listening (P3b) observable at the Pz electrode, and an index of auditory-specific sensory process (N1) elicited during passive and active auditory tasks, observable at Cz.</p>
<p>In this approach, we measured the effects of conceptual &#x201C;repetition suppression&#x201D; to sound categories and not to individual repeating physical stimuli. Further, responses to the sounds when they were task-relevant were compared with responses to the sounds when they were task-irrelevant to further test the automaticity of conceptual category representation &#x2013; to evaluate the role of directed attention in maintaining predictions.</p>
<p>Our results demonstrated conceptual categorical expectation, implicit learning, and a speech perception bias only when the sounds were attended, and a sound categorization task was performed with them. These results indicate that attentional control is required to maintain semantic category identification, and that higher-level processes use that information to predict upcoming events.</p>
</sec>
<sec sec-type="materials|methods" id="sec2">
<label>2.</label>
<title>Materials and methods</title>
<sec id="sec3">
<label>2.1.</label>
<title>Participants</title>
<p>Ten adults ranging in age from 19&#x2013;37&#x2009;years (mean&#x2009;=&#x2009;28.5, SD&#x2009;=&#x2009;5.5) were paid to participate in the study. All participants passed a hearing screening at 20&#x2009;dB HL or better at 500, 1,000, 2,000, and 4,000&#x2009;Hz in the left and right ears and had no reported history of neurological or otologic disorders. Procedures were approved by the Institutional Review Board of the Albert Einstein College of Medicine (Bronx, NY) where the study was conducted. The examiner described the procedure to all participants in accordance with the Declaration of Helsinki who subsequently gave written consent and were paid for their participation.</p>
<p>Results of a power analysis conducted using Statistica software (Tibco), with a two-tailed <italic>t</italic>-test for dependent means, a medium effect size (<italic>d</italic>&#x2009;=&#x2009;0.50) and an alpha of 0.05. determined that a sample size of nine participants would yield power of 0.90 to detect differences. The total of ten participants included in the current study thus exceeds the number required to obtain sufficient statistical power.</p>
</sec>
<sec id="sec4">
<label>2.2.</label>
<title>Stimuli</title>
<p>Stimuli were naturally produced complex sounds (32-bit stereo; 44,100&#x2009;Hz digitization) obtained from free online libraries (<xref rid="SM1" ref-type="supplementary-material">Appendix I</xref>). Three categories of sounds were presented, of which there were 25 different tokens of spoken speech, 25 different tokens of musical instruments, and 57 different tokens of various environmental sounds. Speech sounds were naturally spoken words (e.g., &#x201C;hello,&#x201D; &#x201C;goodbye&#x201D;); music sounds were taken from various musical instruments (e.g., piano, flute, bass); and environmental sounds were taken from a range of sources, including nature (e.g., water dripping), vehicles (e.g., engine revving), household (e.g., phone ring), and animals (e.g., bird chirp). We modified all the sounds to be 500&#x2009;ms in duration, with an envelope rise and fall times of 7.5&#x2009;ms at onset and offset using Adobe Audition software (Adobe Systems, San Jose, CA). To verify that 500&#x2009;ms in sound duration was sufficient to identify and distinguish the sound categories (e.g., spoken word, instrumental, environmental), three lab members (who were not included in the study) categorized a set of 150 sounds. The final set of 107 sounds used in the study had unanimous agreement as belonging to a category of speech, music, or environment. All 107 stimuli were equated for loudness using the root mean square (RMS) amplitude with Adobe Audition software. Categorical sounds were calibrated with a sound pressure level meter in free field (Br&#x00FC;el and Kajaer, Denmark) and presented through speakers at 65&#x2009;dB SPL with a stimulus onset asynchrony (SOA) of 1.1&#x2009;s.</p>
</sec>
<sec id="sec5">
<label>2.3.</label>
<title>Procedures</title>
<p>Participants sat in a comfortable chair in an electrically shielded and sound-controlled booth (IAC Acoustics, Bronx, NY). Stimuli were presented via two speakers placed approximately 1.5&#x2009;m, 45&#x00B0; to the left of center and 1.5&#x2009;m, 45&#x00B0; to the right of center from the seated listener.</p>
<p>There were two task conditions: <italic>passive auditory</italic> and <italic>active auditory</italic>. During the <italic>passive auditory</italic> condition, participants had no specific task with the sounds. They watched a captioned silent movie of their choosing during the presentation of the sounds. The experimenter monitored the EEG to ensure that participants were reading the closed captions. In the <italic>active auditory</italic> condition, participants listened to the sounds and performed a three-alternative forced-choice task. Participants were instructed to listen to and classify each sequential sound by pressing one of three buttons labeled on a response keypad that uniquely corresponded to the sound category (speech, music, or environment). Participants were not provided with any information about the patterned structure of the stimulus sequence at any time. Thus, the patterned structure could be extracted by implicit learning regardless of the condition in which they were presented.</p>
<p>A total of 3,840 stimuli were presented in 16 separately randomized stimulus blocks (240 stimuli per block), eight stimulus blocks per condition. Stimuli were presented in continuous sequences of 240 stimuli, patterned by categorical groups of four stimuli (spoken words, musical instruments, environmental sounds), with an equal distribution of the categories in each condition (0.33 speech, 0.33 music, and 0.33 environmental). Category switches and repeats occurred randomly within each stimulus block. There were no repeated stimuli within any of the stimulus groups. Every sound token was unique in each categorical group (e.g., the sounds of the instruments harp, piano, clarinet, and guitar could be repetitions in the category of the music group, <xref rid="fig1" ref-type="fig">Figure 1</xref>). Categories switched randomly after four stimuli 70% of the time overall (336 switch trials per condition), and randomly repeated categories after four stimuli 30% of the time (144 repeat trials per condition). Presentation of sound groups was quasi-randomized such that categories could only repeat one time. Thus, sounds occurred in groups of either four or eight repetitions of any category. Participants were not informed about the structure of the sequences at any time and there was no demarcation to indicate when the category switched, or when the category repeated within a stimulus block; sound tokens were presented isochronously throughout every stimulus block. Thus, position #1 stimuli were only a &#x2018;first position&#x2019; stimulus based only with implicit detection of the patterned categorical grouping.</p>
<p>Task conditions were randomized across participants, with half of participants presented with the passive condition first and half presented with the active condition first. Recording time was approximately 35&#x2009;min per condition, with a snack break at the halfway point at which time the participant was unhooked from the amplifiers and could walk around. Total session time including cap placement, recording time, and breaks was approximately 2&#x2009;h.</p>
</sec>
<sec id="sec6">
<label>2.4.</label>
<title>Electroencephalogram recordings</title>
<p>A 32-channel electrode cap incorporating a subset of the International 10&#x2013;20 system was used to obtain EEG recordings. Additional electrodes were placed over the left and right mastoids (LM and RM, respectively). An external electrode placed at the tip of the nose was used as the reference electrode. Horizontal electro-oculogram (EOG) was monitored with the F7 and F8 electrode sites and vertical EOG was monitored using a bipolar configuration between FP1 and an external electrode placed below the left eye. Impedances were kept below 5&#x2009;k&#x03A9; at all electrodes throughout the recording session. The EEG and EOG were digitized (Neuroscan Synamps amplifier, Compumedics Corp., Texas, United States) at a sampling rate of 1,000&#x2009;Hz (0.05&#x2013;200&#x2009;Hz bandpass). EEG was filtered off-line with a lowpass of 30&#x2009;Hz (zero phase shift, 24&#x2009;dB rolloff).</p>
</sec>
<sec id="sec7">
<label>2.5.</label>
<title>Data analysis</title>
<p>This report includes data from all 10 participants in the study. There were no exclusions.</p>
<p><italic>Behavioral Data</italic>: Hit rate (HR) and reaction time (RT) were calculated for the responses to each of the sounds, separately by category (speech, music, and environment and by stimulus position (1&#x2013;8)). Hits were counted when responses occurred 100&#x2013;1,100&#x2009;ms from the onset of the stimulus. The mean HR was calculated as the total number of correctly identified stimuli divided by the number of stimuli in each category for each position. RT was calculated for each sound from sound onset. Means were derived for each stimulus category, in each position separately.</p>
<p><italic>ERP Data</italic>: The filtered EEG was segmented into 4,500&#x2009;ms epochs, starting from 200&#x2009;ms pre-stimulus and ending 4,300&#x2009;ms post-stimulus onset from position 1 for the switch category to display ERP responses consecutively in positions 1&#x2013;4, and from the onset of position 5 for the repeat category to display ERP responses consecutively in positions 5&#x2013;8. Due to the length of these epochs, ocular artifact reduction was performed on all participants using Neuroscan EDIT software. This Singular Value Decomposition transform method is used to identify the blink component. From the continuous EEG, a file was created that reflected the spatial distribution of the blink and then used to remove the blink. The blink-corrected data were then baseline-corrected across the whole epoch (the mean was subtracted at each point across the epoch). After baseline correction, artifact rejection criteria were set to &#x00B1;75&#x2009;mV. On average, 89% of all trials were included.</p>
<p>To measure mean amplitudes, the peak amplitude of each of the ERP components was visually identified in the grand-mean waveforms, in each condition separately, at the electrode with greatest expected signal-to-noise ratio for each component based on previous literature (<xref ref-type="bibr" rid="ref28">N&#x00E4;&#x00E4;t&#x00E4;nen and Picton, 1987</xref>; <xref ref-type="bibr" rid="ref11">Friedman et al., 2001</xref>; <xref ref-type="bibr" rid="ref10">Fonken et al., 2020</xref>). Thus, we used the Pz electrode to measure the P3b component, the Cz electrode for the P3a component, and the Cz electrode for the N1 component. The peak latency in the grand-averaged waveforms were used to obtain mean amplitudes for statistical comparison. Mean amplitudes were calculated using a 50&#x2009;ms interval centered on the grand-mean peak, for each ERP component, separately for each stimulus category and position, in each condition, for each individual participant.</p>
</sec>
<sec id="sec8">
<label>2.6.</label>
<title>Statistical analyses</title>
<p>For behavioral data (HR and RT), separate two-way repeated measures ANOVA with factors of category (speech/music/environment) and position (1&#x2013;8) to determine effects of category switch and category repetition. For event-related potentials (N1/P3a/P3b), separate two-way repeated measures ANOVA with factors of category (speech/music/environment) and position (1&#x2013;8) were calculated to determine effects of category switch and category repetition on the mean amplitude of the ERPs. In cases where data violated the assumption of sphericity, the Greenhouse&#x2013;Geisser estimates of sphericity were used to correct the degrees of freedom. Corrected <italic>p</italic> values are reported. Tukey&#x2019;s HSD for repeated measures was conducted on pairwise contrasts for <italic>post hoc</italic> analyses when the omnibus ANOVA was significant. Contrasts were reported as significantly different at <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05. Effect sizes were computed and reported as partial eta squared (<italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>). Statistical analyses were performed using Statistica 13.3 software (Tibco).</p>
</sec>
</sec>
<sec sec-type="results" id="sec9">
<label>3.</label>
<title>Results</title>
<sec id="sec10">
<label>3.1.</label>
<title>Passive auditory condition. Task: watch a movie</title>
<sec id="sec11">
<label>3.1.1.</label>
<title>P3a component</title>
<p>P3a amplitude did not differ as a function of position (<italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;1.5, <italic>p</italic>&#x2009;=&#x2009;0.25), or category (<italic>F</italic><sub>2,18</sub>&#x2009;=&#x2009;2.0, <italic>p</italic>&#x2009;=&#x2009;0.18), and there were no interactions (<italic>F</italic><sub>14,126</sub>&#x2009;=&#x2009;1.4, <italic>p</italic>&#x2009;=&#x2009;0.24). When the listener watched a movie, each sound engaged attention and elicited a P3a component with similar amplitudes across positions and sound categories (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, left panel, gray solid line; <xref rid="fig3" ref-type="fig">Figure 3A</xref>; <xref rid="fig4" ref-type="fig">Figure 4</xref>, passive auditory). The salient categorical stimuli elicited an orienting response during both tasks (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, left panel, compare gray and black traces).</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p>Event-related brain potentials. The grand-mean ERP waveforms elicited in each separate condition: active auditory (black traces) and passive auditory (gray traces) are displayed for the speech sounds (top row, blue), music sounds (middle row, red), and environmental sounds (bottom row, green). The Cz electrode (left panel) best displays the N1 and P3a components (labeled and with arrows for P3a). The Pz electrode (right panel) best displays the P3b component (labeled with arrows). The amplitude in microvolts is denoted along the <italic>y</italic>-axis. The colored squares displayed below the <italic>x</italic>-axis show the <italic>stimulus presentation</italic> rate, of one sound presented every 1.1&#x2009;s. Categorical habituation and a speech processing bias are clearly demonstrated in the P3b amplitudes only during active task performance.</p></caption>
<graphic xlink:href="fnins-17-1228506-g002.tif"/>
</fig>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p>Grand-mean ERP amplitudes. <bold>(A)</bold> P3a component. The mean amplitudes and standard errors (whiskers) for speech (blue circle), music (red square), and environmental (green diamond) categories are overlain showing each stimulus position: switch (1&#x2013;4) and repeat (5&#x2013;8) trials, measured at the Cz electrode in the passive auditory condition. The onset of the switch and repeat trials are labeled along the <italic>x</italic>-axis, and the amplitude is denoted in microvolts along the <italic>y</italic>-axis for all panels. P3a amplitude did not index either category or expectation effects. <bold>(B)</bold> P3b component. The mean amplitudes and standard errors (whiskers) for speech (blue circle), music (red square), and environmental (green diamond) categories are overlain showing each stimulus position: switch (1&#x2013;4) and repeat (5&#x2013;8) trials, measured at the Pz electrode in the active auditory condition. A clear speech effect (larger amplitude for P3b speech in position 1), and a clear expectation effect (larger amplitude P3b for all categories at the switch (position 1)) are demonstrated. <bold>(C)</bold> N1 component. Grand-mean amplitudes and standard errors (whiskers) are compared for passive auditory (gray, dashed line) and active auditory (black, solid line) conditions at each stimulus position: switch (1&#x2013;4) and repeat (5&#x2013;8) trials, measured at the Cz electrode. Larger (more negative) N1 amplitudes were elicited by stimuli when a task was performed with the sounds.</p></caption>
<graphic xlink:href="fnins-17-1228506-g003.tif"/>
</fig>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption><p>Scalp distribution maps. Scalp voltage distribution maps, from the grand-mean of all ten participants, show the P3b components (active auditory) and the P3a components (passive auditory), taken at their respective peak latencies, separately, by category for each position: speech (blue font), music (red font), and environment (green font). Small red dots denote each electrode. The Pz electrode, where the P3b amplitude is typically largest, is indicated with an arrow for the active auditory condition (top rows of each category). The Cz electrode, where the P3a amplitude is typically largest, is indicated with an arrow for the passive auditory condition (bottom rows of each category). Red indicates positive polarity; blue indicates negative polarity. The scale is 0.20&#x2009;&#x03BC;V/step.</p></caption>
<graphic xlink:href="fnins-17-1228506-g004.tif"/>
</fig>
</sec>
</sec>
<sec id="sec12">
<label>3.2.</label>
<title>Active auditory condition. Task: categorize the sounds</title>
<sec id="sec13">
<label>3.2.1.</label>
<title>P3b component</title>
<p>The P3b amplitude was larger (more positive) when elicited by the category switch stimulus for all three categories (position 1) compared to the category repetition stimuli (positions 2&#x2013;8) (main effect of position, <italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;10.33, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.24, <italic>p</italic>&#x2009;=&#x2009;0.002, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.53) (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Pz electrode, right panel, black solid lines; <xref rid="fig3" ref-type="fig">Figure 3B</xref>; <xref rid="fig4" ref-type="fig">Figure 4</xref>, active auditory). This shows a dramatic decrease in the magnitude of the P3b amplitude after a single repetition of a categorical stimulus during active task categorization (compare the delta in peak P3b amplitude of responses to position 1 and 2 stimuli in <xref rid="fig2" ref-type="fig">Figure 2</xref>, right panel, Pz electrode, black traces, and <xref rid="fig3" ref-type="fig">Figure 3B</xref>). <italic>Post hoc</italic> analyses showed that there were no mean amplitude differences in responses elicited by stimuli in positions 2&#x2013;8. The P3b amplitude remained attenuated for within-category repetitions, stimulus positions 2&#x2013;4 after switching to a new category, and in stimulus positions 5&#x2013;8 after repeating a category. The reduced P3b amplitude to category repeats in positions 2&#x2013;8 demonstrates conceptual category &#x201C;repetition suppression&#x201D; during active identification (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Pz electrode, right panel, black traces) that cannot be explained by stimulus-specific repetition suppression.</p>
<p>There was no main effect of category (<italic>F</italic><sub>2,18</sub>&#x2009;&#x003C;&#x2009;1, <italic>p</italic>&#x2009;=&#x2009;0.83). However, there was an interaction between category and position (<italic>F</italic><sub>14,126</sub>&#x2009;=&#x2009;3.7, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.27, <italic>p</italic>&#x2009;=&#x2009;0.015, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.29). <italic>Post hoc</italic> analyses revealed that P3b amplitudes elicited by position 1 stimuli were larger than positions 2&#x2013;8 for all categories, and that the P3b amplitude elicited by spoken words in position 1 was larger than the P3b response to music and environmental sounds elicited in position 1 (with no amplitude difference between music and environment for position 1) (<xref rid="fig2" ref-type="fig">Figure 2</xref>, black traces, Pz electrode; <xref rid="fig3" ref-type="fig">Figure 3B</xref>). These results demonstrate both category habituation (position 1 larger than position 2 for all categories) and a speech effect (position 1 larger for speech than position 1 for music and environmental stimuli).</p>
</sec>
</sec>
<sec id="sec14">
<label>3.3.</label>
<title>Sensory-specific processes</title>
<sec id="sec15">
<label>3.3.1.</label>
<title>N1 component</title>
<p>The sensory-specific N1 component, elicited during active and passive auditory tasks, did not clearly reflect categorical habituation or repetition suppression (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, left panel, black and gray traces). There was a main effect of position (<italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;4.8, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.49, <italic>p</italic>&#x2009;=&#x2009;0.006, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.35), with <italic>post hoc</italic> test showing that the N1 elicited in position 1 was more negative than the N1 in position 3 (but not with position 2) and position 1 did not differ in magnitude from any of the other N1 positions (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode). There was also a main effect of category (<italic>F</italic><sub>2,18</sub>&#x2009;=&#x2009;56.6, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.98, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.42), with <italic>post hoc</italic> analyses showing that the N1 elicited by the music stimuli was larger in magnitude than either speech or environment. There was an attention effect, reflecting an expected attentional gain when attending vs. ignoring sounds (<xref ref-type="bibr" rid="ref17">Hillyard et al., 1998</xref>). The N1 amplitude was larger (more negative amplitude) when the sounds were attended (main effect of attention, <italic>F</italic><sub>1,9</sub>&#x2009;=&#x2009;5.5, <italic>p</italic>&#x2009;=&#x2009;0.04, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.38) (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, left panel, compare gray and black traces; <xref rid="fig3" ref-type="fig">Figure 3C</xref>).</p>
</sec>
</sec>
<sec id="sec16">
<label>3.4.</label>
<title>Task performance: categorizing the sounds</title>
<sec id="sec17">
<label>3.4.1.</label>
<title>Behavioral results</title>
<p>Performance results indicate implicit learning, with overall performance poorer when a category switch occurred (position 1, <xref rid="fig5" ref-type="fig">Figure 5</xref>). Mean reaction time was longest at the category switch (position 1, <xref rid="fig5" ref-type="fig">Figure 5A</xref>) (main effect of position, <italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;29.32, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.32, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.0001, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.77). The switch stimulus (position 1) also had the lowest hit rates (main effect of position, <italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;23.47, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.19, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.0001, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.72) (<xref rid="fig5" ref-type="fig">Figure 5B</xref>). The longer RT and lower HR may reflect the expectation of a category switch, in that additional processing would be at-the-ready to &#x2018;re-identify&#x2019; the category after four sounds (i.e., in position 1). Once the category was identified, confirmation of category membership for stimuli 2&#x2013;4 would only be needed, reflected by the faster RT and higher HR in positions 2&#x2013;4 stimuli. Implicit learning is indicated by RT, which was, on average, 150&#x2009;ms shorter to the second token of the within-category repetition (main effect of position: <italic>F</italic><sub>7,63</sub>&#x2009;=&#x2009;29.3, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.32, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.77). <italic>Post hoc</italic> tests show that RT was slowest for position 1 stimuli. The faster responses time occurred for all within-category stimulus repetitions (positions 2&#x2013;8). After only one repetition of a categorical stimulus, there was a dramatic decrease in RT (<xref rid="fig5" ref-type="fig">Figure 5A</xref>, compare positions 1 and 2).</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption><p>Behavioral data. <bold>(A)</bold> Reaction time. The mean reaction time (in ms, <italic>y</italic>-axis) for the three categorical sounds are overlain and displayed separately for each position (represented along the <italic>x</italic>-axis) for speech sounds (blue circle), music sounds (red square), and environmental sounds (green diamond). Position 1 is a category switch (labeled) and position 5 is a category repeat (labeled). Whiskers show the standard error. The slower mean reaction times in position 1 show a clear switch effect (implicit learning). <bold>(B)</bold> Hit rate. The mean hit rate in percentage (<italic>y</italic>-axis) is displayed for speech sounds (blue circle), music sounds (red square), and environmental sounds (green diamond) for each position (represented along the <italic>x</italic>-axis). Position 1 is a category switch and position 5 is a category repeat. Whiskers show standard error. A speech effect is seen in the lower mean hit rates to music and environmental sounds compared to speech in positions 1 and 2.</p></caption>
<graphic xlink:href="fnins-17-1228506-g005.tif"/>
</fig>
<p>For categorical effects, mean RT was shorter to speech and music sounds than to environmental sounds (main effect of stimulus category, <italic>F</italic><sub>2,18</sub>&#x2009;=&#x2009;11.01, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.72, <italic>p</italic>&#x2009;=&#x2009;0.003, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.55). <italic>Post hoc</italic> calculations revealed the fastest reaction time to speech sounds, but not faster than music sounds when RT was collapsed across position (<italic>p</italic>&#x2009;=&#x2009;0.13) (<xref rid="fig5" ref-type="fig">Figure 5A</xref>). The main effect of category on HR (<italic>F</italic><sub>2,18</sub>&#x2009;=&#x2009;3.89, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.78, <italic>p</italic>&#x2009;=&#x2009;0.039, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.31) was due to a higher HR to speech than music sounds (<italic>p</italic>&#x2009;=&#x2009;0.04) and HR for speech trended toward being higher than environmental sounds (<italic>p</italic>&#x2009;=&#x2009;0.1). The significant interaction between sound category and position (<italic>F</italic><sub>14,126</sub>&#x2009;=&#x2009;11.38, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.22, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.0001, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.56) was due to a higher HR for speech than music and environmental sounds at positions 1 and 2, whereas HR was not different across any positions for the speech sounds (<xref rid="fig5" ref-type="fig">Figure 5B</xref>). There was a significant interaction between sound category and position (<italic>F</italic><sub>14,126</sub>&#x2009;=&#x2009;2.79, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.26, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.24). <italic>Post hoc</italic> calculations showing that in addition to a longer RT at position 1 across all sound categories, mean RT was longer in the repeat position 5 compared to position 4 for music and environmental sounds, but not for speech sounds. RT was not different between positions 4 and 5 for the speech sounds. There was an interaction between category and position (<italic>F</italic><sub>2,18</sub>&#x2009;=&#x2009;11.0, <italic>&#x03B5;</italic>&#x2009;=&#x2009;0.26, <italic>p</italic>&#x2009;=&#x2009;0.048, <italic>&#x03B7;</italic><sup>2</sup><italic><sub>p</sub></italic>&#x2009;=&#x2009;0.55), which was due to slower RT in position 5 for music and environment. This may suggest anticipation of a switch in position 5 but no enhancement for speech, which already had faster response times.</p>
</sec>
</sec>
</sec>
<sec sec-type="discussions" id="sec18">
<label>4.</label>
<title>Discussion</title>
<p>The key finding of our study, using a unique category repetition paradigm, was that extracting higher-level meaning from sound input requires specific task attention. This is the first study we know of showing effects of neural habituation for conceptual category repetition. We found three fundamental effects associated with actively categorizing sounds by speech, music, and environment: (1) categorical habituation; (2) implicit learning; and (3) a speech perception bias. None of these effects were observed when the stimulus sequences were presented, and the listener was watching a movie and had no specific task with the sounds. A crucial differentiating feature of our experimental design was that we dissociated stimulus repetition from category identification. Most previous studies that evaluate effects of predictive processing rely on repetition suppression where the same physical stimulus or pattern of stimuli are repeated. Higher-level conceptual effects can thus be conflated with stimulus-specific effects. In the current study, we uncoupled conceptual categorical effects from stimulus-specific effects by presenting categorical stimulus tokens that did not physically repeat. Using this experimental paradigm, the reduction of the P3b ERP amplitude that occurred at the first repetition of a category stimulus could not reflect stimulus-specific adaptation or &#x2018;repetition suppression&#x2019;. We found that categorical stimulus-repetition suppressed the brain response only when a task was being performed with the sounds. This suggests that category membership was only derived through task-based attentional processing.</p>
<sec id="sec19">
<label>4.1.</label>
<title>Categorical habituation is a top-down phenomenon</title>
<p>Habituation is defined as a reduction in response to repeated stimulation that is not due to physiological effects, such as neural fatigue or adaptation (<xref ref-type="bibr" rid="ref22">Magnussen and Kurtenbach, 1980</xref>; <xref ref-type="bibr" rid="ref34">Rankin et al., 2009</xref>; <xref ref-type="bibr" rid="ref37">Schmid et al., 2014</xref>). In the current study, we demonstrate a form of habituation that cannot be attributed to repetition suppression or neural fatigue. Conceptual categorical habituation was demonstrated by a dramatic decrease in the magnitude of the P3b amplitude after a single repetition of any of the categorical stimulus during active task categorization (<xref rid="fig2" ref-type="fig">Figures 2</xref>, <xref rid="fig3" ref-type="fig">3</xref>). The delta in P3b amplitude from position 1 to position 2 was remarkable considering that the pattern of categorical stimuli occurred within an ongoing sequence of sounds, with no demarcation of when the grouping of category repetitions or switches were occurring. Further, the sounds themselves did not repeat, precluding stimulus-driven factors that could drive the response reduction by sensory adaptation or neural fatigue. The reduction in the magnitude of the neural signal after one repetition is notable because no two identical stimulus tokens were presented successively; the category repeated but not the physical stimuli. Accordingly, the reduction cannot be explained by stimulus-specific repetition suppression and reflects a neural habituation to the repetition of a conceptual category.</p>
<p>There was no categorical habituation when participants were watching a movie. It is well documented that when successive stimuli are acoustically unique, neural repetition suppression would not be expected (<xref ref-type="bibr" rid="ref28">N&#x00E4;&#x00E4;t&#x00E4;nen and Picton, 1987</xref>; <xref ref-type="bibr" rid="ref5">Budd et al., 1998</xref>; <xref ref-type="bibr" rid="ref16">Grill-Spector et al., 2006</xref>). There was no conceptual &#x201C;repetition suppression&#x201D; when participants watched a movie. Thus, these results demonstrate that the conceptual categorical aspect of the stimuli was not automatically processed during passive listening. Habituation by repetition was only initiated when attention was focused on the sounds and a semantic categorical task was performed.</p>
</sec>
<sec id="sec20">
<label>4.2.</label>
<title>Evidence of implicit expectation only during active categorization of sounds</title>
<p>Expectation effects were observed only when the listener performed the categorization task with the sounds; not when they watched a movie. No explicit instructions were provided to participants about the stimulus structure, and the patterned structure was irrelevant to both tasks. However, implicit expectations could be formed by the regularity of the stimulus structure, in which the listener could expect a category switch after four successive categorical sounds most of the time. The larger P3b amplitude elicited by the categorical switch stimuli (position 1) demonstrates an implicit expectation that a category switch was likely to occur (a target switch). It should be noted that it was not possible to build up 100% expectation for a category switch because 30% of the time, rather than switching category, stimuli from the same category continued for a second successive group of four. Implicit category learning also influenced task performance. RT was slowest for position 1 stimuli: RT was 150&#x2009;ms shorter to the second token of the within-category repetition. The faster responses time occurred for all the within-category stimulus repetitions (positions 2&#x2013;8). The dramatic decrease in RT after only one repetition of a categorical stimulus is consistent with the substantial reduction in P3b amplitude after one categorical stimulus repetition (position 2) (<xref rid="fig3" ref-type="fig">Figure 3</xref>, compare positions 1 and 2). This is remarkable when considering that reaction time is faster to a repeated event than to a non-repeated event (<xref ref-type="bibr" rid="ref001">Smith, 1968</xref>); that is, when it is the same physical stimulus. Here we show a reduction in RT to a conceptual repetition. The reduced response for position 2 stimuli indicates that the pattern of category repetition in the structure of the sequence was implicitly learned while performing the task. Implicit learning led to the knowledge that the category of stimuli would repeat after the switch position, even though that it was not the same physical sound token. The slower RT in position 1 and faster RT for positions 2&#x2013;8 is consistent with modulation of the P3b amplitude, which was smaller after one category repetition and remained at the small amplitude until the next category switch. There was also an indication of implicit expectation in the longer RT at position 5, where a category switch may have been expected. However, this was not significantly reflected in the ERPs, likely due to the lower probability of a repeat than a switch.</p>
<p>In contrast, there was no evidence of implicit learning associated with the category switch when the listener watched a movie. The P3a amplitude did not differ as a function of position or category. The amplitude in position 1 was no different than that in any other position. Thus, a robust P3a was elicited by each successive stimulus token, with no indication by change in its magnitude that the brain detected a pattern of conceptual category repetitions. There was no categorical &#x201C;repetition suppression.&#x201D; Finding no index of implicit expectation, diverges from previous studies that have shown that stimulus repetition can build strong expectations and influence the brain response without attention focused on the sounds (<xref ref-type="bibr" rid="ref43">Todorovic et al., 2011</xref>). However, our stimulus design is unique and may explain the differences in our results. The current study design differs from previous studies in two important ways. The repetition pattern of four sound tokens from the same category (speech, music, or environment) was comprised of four unique sound tokens from the category. For example, the listener may have heard the spoken words &#x201C;peace&#x201D; &#x2013; &#x201C;hello&#x201D; &#x2013; &#x201C;yes&#x201D; &#x2013; &#x201C;wonder&#x201D; as the four-token repetition for one group in the speech category. All the sounds were different from each other. Therefore, identification of category repetition could not occur based on stimulus-driven features or acoustic characteristics of the sounds (<xref ref-type="bibr" rid="ref25">Moskowitz et al., 2020</xref>). Secondly, expectations were not 100% predictable, that is, the category switch after the presentation of four sounds was not fully predictable; 30% of the time the category repeated. Consequently, during the passive condition, while attention was focused on reading captions and watching a movie, there could be some uncertainty about the regularity of the categorical aspect in the stimulus sequence, especially because the stimulus tokens themselves were not repeated, and attention was not actively monitoring the structure of the sound presentation. In addition, the structure of the sound sequence was irrelevant to performing the task. Thus, we conclude that attention focused onto the sounds with the intention to identify category membership was a key factor enabling expectations to be implicitly derived from the stimulus sequence.</p>
</sec>
<sec id="sec21">
<label>4.3.</label>
<title>Speech effects were observed only when attention was focused on the sounds</title>
<p>A surprising result of the study was that a speech bias was observed only during active listening. Response times were faster and ERP amplitudes were larger to speech category tokens during task performance. There was no categorical effect when listeners were passively listening and watching a movie. The automatic involuntary orienting response (indexed by the P3a component) did not differentiate speech from the other categories at any position (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, gray traces), whereas the P3b amplitude did differentiate speech (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Pz electrode, black traces). Moreover, there was an attentional orienting response to the sounds in both the active auditory and the passive auditory conditions (<xref rid="fig2" ref-type="fig">Figure 2</xref>, Cz electrode, left panel, compare black and gray traces). However, with the active auditory task, there was an additional P3b component elicited consistent with target detection. There was no P3b elicited in the passive auditory when there was no auditory task. Thus, only with attention focused on a task with the sounds, was there evidence that the higher-level categorical aspects of signal differentiation. That is, differentiation of the speech signal from other music and environmental sounds was only evident when attention was used to categorize the sounds. This is notable because there is considerable evidence from infancy showing that speech is processed differently from other environmental sounds (<xref ref-type="bibr" rid="ref9">Eimas et al., 1971</xref>; <xref ref-type="bibr" rid="ref31">Pisoni, 1979</xref>; <xref ref-type="bibr" rid="ref27">Murray et al., 2006</xref>; <xref ref-type="bibr" rid="ref44">Vouloumanos and Werker, 2007</xref>; <xref ref-type="bibr" rid="ref1">Agus et al., 2012</xref>; <xref ref-type="bibr" rid="ref15">Gervain and Geffen, 2019</xref>). Recent evidence, however, has suggested that speech may only show an &#x2018;advantage&#x2019; under specific listening or task situations (<xref ref-type="bibr" rid="ref25">Moskowitz et al., 2020</xref>). In previous studies showing a speech bias, this issue of attention may have not come to light because stimulus categories were not separated by unique tokens. Certainly, one can detect speech passively and unattended speech can alert our attention (e.g., the sound of your name being called) (<xref ref-type="bibr" rid="ref29">Navon et al., 1987</xref>). However, the current results indicate that when attention is not directed towards the sounds, the acoustic characteristics that distinguish speech from other environmental sounds are not automatically discriminated as a special category when there is a complex mixture of sound categories occurring. Our results indicate that attentional control is required to process the higher-level aspects of the speech signal, to extract the conceptual category (speech, music, or environment) when there are a variety of complex sounds. Speech may not be treated as a distinct or separate category without an active task and attention to the sounds. A question that remains is how specific the task must be to the conceptual process for it to alter the neural response; would performing a task not involving categorization also show no category effects?</p>
</sec>
</sec>
<sec id="sec22">
<label>5.</label>
<title>Summary and conclusions</title>
<p>Our results address a fundamental controversy about the role of attention in higher level processing. We distinguished between repetition suppression and conceptual categorical habituation by repeating sounds that fit a sound category but never repeating the same physical sound tokens. Predictive processing theory suggests that brain processes are continually generating and updating a model of the environment (<xref ref-type="bibr" rid="ref46">Winkler et al., 1996</xref>; <xref ref-type="bibr" rid="ref12">Friston, 2005</xref>). This theory suggests that the brain automatically builds expectations (priors), derived by sound patterns extracted through stimulus statistics. Thus, our results diverge somewhat from this aspect of the predictive processing theory in that we found no reduction in the magnitude of the neural response to a repeated sound category unless attention was directed to the categorical aspect of the sounds. The theoretical perspective that the brain calculates and anticipates all stimulus patterns within a sound sequence and automatically sets up expectations, implicitly learned without attention, is not upheld for higher-level conceptual categories involving a mixture of complex sounds with the current experimental design. We found no evidence of implicit learning of the structured sound sequence when the listener was passively listening and watching a movie. Our results are consistent with previous studies showing that task goals, rather than stimulus statistics, have great influence on neural processing of auditory and visual patterns (<xref ref-type="bibr" rid="ref40">Sussman et al., 1998</xref>, <xref ref-type="bibr" rid="ref41">2002</xref>; <xref ref-type="bibr" rid="ref24">Max et al., 2015</xref>; <xref ref-type="bibr" rid="ref38">Solomon et al., 2021</xref>).</p>
<p>Overall, we found that top-down knowledge was required to set up expectations for higher-level processes (<xref ref-type="bibr" rid="ref35">Rao and Ballard, 1999</xref>; <xref ref-type="bibr" rid="ref39">Summerfield et al., 2008</xref>). Our results thus link in with the question of how much, or what type of, processing of the unattended, irrelevant sounds occurs when performing another task. It is generally thought that attention can &#x2018;leak&#x2019; or &#x2018;slip&#x2019; to the unattended stimuli while performing another task (<xref ref-type="bibr" rid="ref20">Lachter et al., 2004</xref>). Watching a movie is not considered a highly demanding task, and it may be argued that attentional slips could easily occur. However, remarkably, there was no evidence of implicit learning of the structured sound sequence, or of categorical perception, such as a speech bias during passive listening, when it would be more likely there would have been potential slips of attention to the unattended sounds. These findings are consistent with the theory of <xref ref-type="bibr" rid="ref3">Broadbent (1956)</xref>, who proposed that attention is a limited resource and therefore attention to one set of sounds limits available resources to process the unattended sounds, beyond the simple sound features (e.g., frequency, intensity, spatial location). We suggest that the limited resource here is higher-level conceptual category formation. We found that the &#x2018;slippage&#x2019; of attention to irrelevant sounds was not enough to induce higher-level processing, indicating that those higher-level processes that identify linguistic, semantic, or categorical aspects of stimuli require some form of active attention. Although humans are experts at detecting and finding patterns in sensory input, the extent of processing and the role of attention in processing irrelevant sounds, under various listening situations, is still yet to be fully resolved.</p>
</sec>
<sec sec-type="data-availability" id="sec23">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec sec-type="ethics-statement" id="sec24">
<title>Ethics statement</title>
<p>The studies involving humans were approved by Albert Einstein College of Medicine. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.</p>
</sec>
<sec sec-type="author-contributions" id="sec25">
<title>Author contributions</title>
<p>HM: conceptualization, methodology, formal analysis, writing &#x2013; original draft, and visualization. ES: conceptualization, methodology, formal analysis, writing &#x2013; review and editing, visualization, supervision, and funding acquisition. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="sec26">
<title>Funding</title>
<p>This work was supported by the National Institutes of Health (R01DC004263, to ES).</p>
</sec>
<ack>
<p>The authors thank Wei Wei Lee for assistance with data collection.</p>
</ack>
<sec sec-type="COI-statement" id="sec27">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="sec28">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fnins.2023.1228506/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fnins.2023.1228506/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Agus</surname> <given-names>T. R.</given-names></name> <name><surname>Suied</surname> <given-names>C.</given-names></name> <name><surname>Thorpe</surname> <given-names>S. J.</given-names></name> <name><surname>Pressnitzer</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Fast recognition of musical sounds based on timbre</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>131</volume>, <fpage>4124</fpage>&#x2013;<lpage>4133</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.3701865</pub-id>, PMID: <pub-id pub-id-type="pmid">22559384</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alink</surname> <given-names>A.</given-names></name> <name><surname>Blank</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Can expectation suppression be explained by reduced attention to predictable stimuli?</article-title> <source>NeuroImage</source> <volume>231</volume>:<fpage>117824</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neuroimage.2021.117824</pub-id>, PMID: <pub-id pub-id-type="pmid">33549756</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Broadbent</surname> <given-names>D. E.</given-names></name></person-group> (<year>1956</year>). <article-title>The concept of capacity and the theory of behavior</article-title>. <source>Information theory; papers read at a symposium on information theory held at the Royal Institution, London, September 12th to 16th, 1955</source>. (<fpage>354</fpage>&#x2013;<lpage>360</lpage>). <publisher-loc>Oxford, England</publisher-loc>: <publisher-name>Academic Press, Inc.</publisher-name>.</citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bubic</surname> <given-names>A.</given-names></name> <name><surname>von Cramon</surname> <given-names>D. Y.</given-names></name> <name><surname>Schubotz</surname> <given-names>R. I.</given-names></name></person-group> (<year>2010</year>). <article-title>Prediction, cognition and the brain</article-title>. <source>Front. Hum. Neurosci.</source> <volume>4</volume>:<fpage>25</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnhum.2010.00025</pub-id>, PMID: <pub-id pub-id-type="pmid">20631856</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Budd</surname> <given-names>T. W.</given-names></name> <name><surname>Barry</surname> <given-names>R. J.</given-names></name> <name><surname>Gordon</surname> <given-names>E.</given-names></name> <name><surname>Rennie</surname> <given-names>C.</given-names></name> <name><surname>Michie</surname> <given-names>P. T.</given-names></name></person-group> (<year>1998</year>). <article-title>Decrement of the N1 auditory event-related potential with stimulus repetition: habituation vs. refractoriness</article-title>. <source>Int. J. Psychophysiol.</source> <volume>31</volume>, <fpage>51</fpage>&#x2013;<lpage>68</lpage>. doi: <pub-id pub-id-type="doi">10.1016/s0167-8760(98)00040-3</pub-id>, PMID: <pub-id pub-id-type="pmid">9934621</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Whatever next? Predictive brains, situated agents, and the future of cognitive science</article-title>. <source>Behav. Brain Sci.</source> <volume>36</volume>, <fpage>181</fpage>&#x2013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1017/s0140525x12000477</pub-id>, PMID: <pub-id pub-id-type="pmid">23663408</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>den Ouden</surname> <given-names>H. E.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>McIntosh</surname> <given-names>A. R.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name></person-group> (<year>2009</year>). <article-title>A dual role for prediction error in associative learning</article-title>. <source>Cereb. Cortex</source> <volume>19</volume>, <fpage>1175</fpage>&#x2013;<lpage>1185</lpage>. doi: <pub-id pub-id-type="doi">10.1093/cercor/bhn161</pub-id>, PMID: <pub-id pub-id-type="pmid">18820290</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Egner</surname> <given-names>T.</given-names></name> <name><surname>Monti</surname> <given-names>J. M.</given-names></name> <name><surname>Summerfield</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Expectation and surprise determine neural population responses in the ventral visual stream</article-title>. <source>J. Neurosci.</source> <volume>30</volume>, <fpage>16601</fpage>&#x2013;<lpage>16608</lpage>. doi: <pub-id pub-id-type="doi">10.1523/jneurosci.2770-10.2010</pub-id>, PMID: <pub-id pub-id-type="pmid">21147999</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eimas</surname> <given-names>P. D.</given-names></name> <name><surname>Siqueland</surname> <given-names>E. R.</given-names></name> <name><surname>Jusczyk</surname> <given-names>P.</given-names></name> <name><surname>Vigorito</surname> <given-names>J.</given-names></name></person-group> (<year>1971</year>). <article-title>Speech perception in infants</article-title>. <source>Science</source> <volume>171</volume>, <fpage>303</fpage>&#x2013;<lpage>306</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.171.3968.303</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fonken</surname> <given-names>Y. M.</given-names></name> <name><surname>Kam</surname> <given-names>J. W. Y.</given-names></name> <name><surname>Knight</surname> <given-names>R. T.</given-names></name></person-group> (<year>2020</year>). <article-title>A differential role for human hippocampus in novelty and contextual processing: implications for P300</article-title>. <source>Psychophysiology</source> <volume>57</volume>:<fpage>e13400</fpage>. doi: <pub-id pub-id-type="doi">10.1111/psyp.13400</pub-id>, PMID: <pub-id pub-id-type="pmid">31206732</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>D.</given-names></name> <name><surname>Cycowicz</surname> <given-names>Y. M.</given-names></name> <name><surname>Gaeta</surname> <given-names>H.</given-names></name></person-group> (<year>2001</year>). <article-title>The novelty P3: an event-related brain potential (ERP) sign of the brain&#x2019;s evaluation of novelty</article-title>. <source>Neurosci. Biobehav. Rev.</source> <volume>25</volume>, <fpage>355</fpage>&#x2013;<lpage>373</lpage>. doi: <pub-id pub-id-type="doi">10.1016/s0149-7634(01)00019-7</pub-id>, PMID: <pub-id pub-id-type="pmid">11445140</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2005</year>). <article-title>A theory of cortical responses</article-title>. <source>Phil. Trans. R. Soc. B</source> <volume>360</volume>, <fpage>815</fpage>&#x2013;<lpage>836</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rstb.2005.1622</pub-id>, PMID: <pub-id pub-id-type="pmid">15937014</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Kiebel</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Predictive coding under the free-energy principle</article-title>. <source>Phil. Trans. R. Soc. B</source> <volume>364</volume>, <fpage>1211</fpage>&#x2013;<lpage>1221</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rstb.2008.0300</pub-id>, PMID: <pub-id pub-id-type="pmid">19528002</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Kilner</surname> <given-names>J.</given-names></name> <name><surname>Harrison</surname> <given-names>L.</given-names></name></person-group> (<year>2006</year>). <article-title>A free energy principle for the brain</article-title>. <source>J. Physiol. Paris</source> <volume>100</volume>, <fpage>70</fpage>&#x2013;<lpage>87</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jphysparis.2006.10.001</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gervain</surname> <given-names>J.</given-names></name> <name><surname>Geffen</surname> <given-names>M. N.</given-names></name></person-group> (<year>2019</year>). <article-title>Efficient neural coding in auditory and speech perception</article-title>. <source>Trends Neurosci.</source> <volume>42</volume>, <fpage>56</fpage>&#x2013;<lpage>65</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tins.2018.09.004</pub-id>, PMID: <pub-id pub-id-type="pmid">30297085</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grill-Spector</surname> <given-names>K.</given-names></name> <name><surname>Henson</surname> <given-names>R.</given-names></name> <name><surname>Martin</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Repetition and the brain: neural models of stimulus-specific effects</article-title>. <source>Trends Cogn. Sci.</source> <volume>10</volume>, <fpage>14</fpage>&#x2013;<lpage>23</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tics.2005.11.006</pub-id>, PMID: <pub-id pub-id-type="pmid">16321563</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hillyard</surname> <given-names>S. A.</given-names></name> <name><surname>Vogel</surname> <given-names>E. K.</given-names></name> <name><surname>Luck</surname> <given-names>S. J.</given-names></name></person-group> (<year>1998</year>). <article-title>Sensory gain control (amplification) as a mechanism of selective attention: electrophysiological and neuroimaging evidence</article-title>. <source>Phil. Trans. R. Soc. B</source> <volume>353</volume>, <fpage>1257</fpage>&#x2013;<lpage>1270</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rstb.1998.0281</pub-id>, PMID: <pub-id pub-id-type="pmid">9770220</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hsu</surname> <given-names>Y. F.</given-names></name> <name><surname>H&#x00E4;m&#x00E4;l&#x00E4;inen</surname> <given-names>J. A.</given-names></name> <name><surname>Waszak</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>The auditory N1 suppression rebounds as prediction persists over time</article-title>. <source>Neuropsychologia</source> <volume>84</volume>, <fpage>198</fpage>&#x2013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neuropsychologia.2016.02.019</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J.</given-names></name> <name><surname>Summerfield</surname> <given-names>C.</given-names></name> <name><surname>Egner</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>Attention sharpens the distinction between expected and unexpected percepts in the visual brain</article-title>. <source>J. Neurosci.</source> <volume>33</volume>, <fpage>18438</fpage>&#x2013;<lpage>18447</lpage>. doi: <pub-id pub-id-type="doi">10.1523/jneurosci.3308-13.2013</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lachter</surname> <given-names>J.</given-names></name> <name><surname>Forster</surname> <given-names>K. I.</given-names></name> <name><surname>Ruthruff</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Forty-five years after Broadbent (1958): still no identification without attention</article-title>. <source>Psychol. Rev.</source> <volume>111</volume>, <fpage>880</fpage>&#x2013;<lpage>913</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0033-295x.111.4.880</pub-id>, PMID: <pub-id pub-id-type="pmid">15482066</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larsson</surname> <given-names>J.</given-names></name> <name><surname>Smith</surname> <given-names>A. T.</given-names></name></person-group> (<year>2012</year>). <article-title>fMRI repetition suppression: neuronal adaptation or stimulus expectation?</article-title> <source>Cereb. Cortex</source> <volume>22</volume>, <fpage>567</fpage>&#x2013;<lpage>576</lpage>. doi: <pub-id pub-id-type="doi">10.1093/cercor/bhr119</pub-id>, PMID: <pub-id pub-id-type="pmid">21690262</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Magnussen</surname> <given-names>S.</given-names></name> <name><surname>Kurtenbach</surname> <given-names>W.</given-names></name></person-group> (<year>1980</year>). <article-title>Adapting to two orientations: disinhibition in a visual aftereffect</article-title>. <source>Science</source> <volume>207</volume>, <fpage>908</fpage>&#x2013;<lpage>909</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.7355271</pub-id>, PMID: <pub-id pub-id-type="pmid">7355271</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maiste</surname> <given-names>A. C.</given-names></name> <name><surname>Wiens</surname> <given-names>A. S.</given-names></name> <name><surname>Hunt</surname> <given-names>M. J.</given-names></name> <name><surname>Scherg</surname> <given-names>M.</given-names></name> <name><surname>Picton</surname> <given-names>T. W.</given-names></name></person-group> (<year>1995</year>). <article-title>Event-related potentials and the categorical perception of speech sounds</article-title>. <source>Ear Hear.</source> <volume>16</volume>, <fpage>68</fpage>&#x2013;<lpage>89</lpage>. doi: <pub-id pub-id-type="doi">10.1097/00003446-199502000-00006</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Max</surname> <given-names>C.</given-names></name> <name><surname>Widmann</surname> <given-names>A.</given-names></name> <name><surname>Schr&#x00F6;ger</surname> <given-names>E.</given-names></name> <name><surname>Sussman</surname> <given-names>E.</given-names></name></person-group> (<year>2015</year>). <article-title>Effects of explicit knowledge and predictability on auditory distraction and target performance</article-title>. <source>Int. J. Psychophysiol.</source> <volume>98</volume>, <fpage>174</fpage>&#x2013;<lpage>181</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ijpsycho.2015.09.006</pub-id>, PMID: <pub-id pub-id-type="pmid">26386396</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moskowitz</surname> <given-names>H. S.</given-names></name> <name><surname>Lee</surname> <given-names>W. W.</given-names></name> <name><surname>Sussman</surname> <given-names>E. S.</given-names></name></person-group> (<year>2020</year>). <article-title>Response advantage for the identification of speech sounds</article-title>. <source>Front. Psychol.</source> <volume>11</volume>:<fpage>1155</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2020.01155</pub-id>, PMID: <pub-id pub-id-type="pmid">32655436</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mumford</surname> <given-names>D.</given-names></name></person-group> (<year>1992</year>). <article-title>On the computational architecture of the neocortex. II. The role of cortico-cortical loops</article-title>. <source>Biol. Cybern.</source> <volume>66</volume>, <fpage>241</fpage>&#x2013;<lpage>251</lpage>. doi: <pub-id pub-id-type="doi">10.1007/bf00198477</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murray</surname> <given-names>M. M.</given-names></name> <name><surname>Camen</surname> <given-names>C.</given-names></name> <name><surname>Gonzalez Andino</surname> <given-names>S. L.</given-names></name> <name><surname>Bovet</surname> <given-names>P.</given-names></name> <name><surname>Clarke</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>Rapid brain discrimination of sounds of objects</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>1293</fpage>&#x2013;<lpage>1302</lpage>. doi: <pub-id pub-id-type="doi">10.1523/jneurosci.4511-05.2006</pub-id>, PMID: <pub-id pub-id-type="pmid">16436617</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>N&#x00E4;&#x00E4;t&#x00E4;nen</surname> <given-names>R.</given-names></name> <name><surname>Picton</surname> <given-names>T.</given-names></name></person-group> (<year>1987</year>). <article-title>The N1 wave of the human electric and magnetic response to sound: a review and an analysis of the component structure</article-title>. <source>Psychophysiology</source> <volume>24</volume>, <fpage>375</fpage>&#x2013;<lpage>425</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1469-8986.1987.tb00311.x</pub-id>, PMID: <pub-id pub-id-type="pmid">3615753</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Navon</surname> <given-names>D.</given-names></name> <name><surname>Sukenik</surname> <given-names>M.</given-names></name> <name><surname>Norman</surname> <given-names>J.</given-names></name></person-group> (<year>1987</year>). <article-title>Is attention allocation sensitive to word informativeness?</article-title> <source>Psychol. Res.</source> <volume>49</volume>, <fpage>131</fpage>&#x2013;<lpage>137</lpage>. doi: <pub-id pub-id-type="doi">10.1007/bf00308678</pub-id>, PMID: <pub-id pub-id-type="pmid">3671630</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Picton</surname> <given-names>T. W.</given-names></name></person-group> (<year>1992</year>). <article-title>The P300 wave of the human event-related potential</article-title>. <source>J. Clin. Neurophysiol.</source> <volume>9</volume>, <fpage>456</fpage>&#x2013;<lpage>479</lpage>. doi: <pub-id pub-id-type="doi">10.1097/00004691-199210000-00002</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pisoni</surname> <given-names>D. B.</given-names></name></person-group> (<year>1979</year>). <article-title>On the perception of speech sounds as biologically significant signals</article-title>. <source>Brain Behav. Evol.</source> <volume>16</volume>, <fpage>330</fpage>&#x2013;<lpage>350</lpage>. doi: <pub-id pub-id-type="doi">10.1159/000121875</pub-id>, PMID: <pub-id pub-id-type="pmid">399200</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polich</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Updating P300: an integrative theory of P3a and P3b</article-title>. <source>Clin. Neurophysiol.</source> <volume>118</volume>, <fpage>2128</fpage>&#x2013;<lpage>2148</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.clinph.2007.04.019</pub-id>, PMID: <pub-id pub-id-type="pmid">17573239</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polich</surname> <given-names>J.</given-names></name> <name><surname>Criado</surname> <given-names>J. R.</given-names></name></person-group> (<year>2006</year>). <article-title>Neuropsychology and neuropharmacology of P3a and P3b</article-title>. <source>Int. J. Psychophysiol.</source> <volume>60</volume>, <fpage>172</fpage>&#x2013;<lpage>185</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ijpsycho.2005.12.012</pub-id>, PMID: <pub-id pub-id-type="pmid">16510201</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rankin</surname> <given-names>C. H.</given-names></name> <name><surname>Abrams</surname> <given-names>T.</given-names></name> <name><surname>Barry</surname> <given-names>R. J.</given-names></name> <name><surname>Bhatnagar</surname> <given-names>S.</given-names></name> <name><surname>Clayton</surname> <given-names>D. F.</given-names></name> <name><surname>Colombo</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Habituation revisited: an updated and revised description of the behavioral characteristics of habituation</article-title>. <source>Neurobiol. Learn. Mem.</source> <volume>92</volume>, <fpage>135</fpage>&#x2013;<lpage>138</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.nlm.2008.09.012</pub-id>, PMID: <pub-id pub-id-type="pmid">18854219</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>R. P.</given-names></name> <name><surname>Ballard</surname> <given-names>D. H.</given-names></name></person-group> (<year>1999</year>). <article-title>Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects</article-title>. <source>Nat. Neurosci.</source> <volume>2</volume>, <fpage>79</fpage>&#x2013;<lpage>87</lpage>. doi: <pub-id pub-id-type="doi">10.1038/4580</pub-id>, PMID: <pub-id pub-id-type="pmid">10195184</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosburg</surname> <given-names>T.</given-names></name> <name><surname>Mager</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>The reduced auditory evoked potential component N1 after repeated stimulation: refractoriness hypothesis vs. habituation account</article-title>. <source>Hear. Res.</source> <volume>400</volume>:<fpage>108140</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.heares.2020.108140</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmid</surname> <given-names>S.</given-names></name> <name><surname>Wilson</surname> <given-names>D. A.</given-names></name> <name><surname>Rankin</surname> <given-names>C. H.</given-names></name></person-group> (<year>2014</year>). <article-title>Habituation mechanisms and their importance for cognitive function</article-title>. <source>Front. Integr. Neurosci.</source> <volume>8</volume>:<fpage>97</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnint.2014.00097</pub-id>, PMID: <pub-id pub-id-type="pmid">25620920</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Solomon</surname> <given-names>S. S.</given-names></name> <name><surname>Tang</surname> <given-names>H.</given-names></name> <name><surname>Sussman</surname> <given-names>E.</given-names></name> <name><surname>Kohn</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Limited evidence for sensory prediction error responses in visual cortex of macaques and humans</article-title>. <source>Cereb. Cortex</source> <volume>31</volume>, <fpage>3136</fpage>&#x2013;<lpage>3152</lpage>. doi: <pub-id pub-id-type="doi">10.1093/cercor/bhab014</pub-id>, PMID: <pub-id pub-id-type="pmid">33683317</pub-id></citation></ref>
<ref id="ref001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>M. C.</given-names></name></person-group> (<year>1968</year>). <article-title>Repetition effect and short-term memory</article-title>. <source>J Exp Psychol.</source> <volume>77</volume>, <fpage>435</fpage>&#x2013;<lpage>9</lpage>. doi: <pub-id pub-id-type="doi">10.1037/h0021293</pub-id>, PMID: <pub-id pub-id-type="pmid">17573239</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Summerfield</surname> <given-names>C.</given-names></name> <name><surname>Trittschuh</surname> <given-names>E. H.</given-names></name> <name><surname>Monti</surname> <given-names>J. M.</given-names></name> <name><surname>Mesulam</surname> <given-names>M. M.</given-names></name> <name><surname>Egner</surname> <given-names>T.</given-names></name></person-group> (<year>2008</year>). <article-title>Neural repetition suppression reflects fulfilled perceptual expectations</article-title>. <source>Nat. Neurosci.</source> <volume>11</volume>, <fpage>1004</fpage>&#x2013;<lpage>1006</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nn.2163</pub-id>, PMID: <pub-id pub-id-type="pmid">19160497</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussman</surname> <given-names>E.</given-names></name> <name><surname>Ritter</surname> <given-names>W.</given-names></name> <name><surname>Vaughan</surname> <given-names>H. G.</given-names></name></person-group> (<year>1998</year>). <article-title>Predictability of stimulus deviance and the mismatch negativity</article-title>. <source>Neuroreport</source> <volume>9</volume>, <fpage>4167</fpage>&#x2013;<lpage>4170</lpage>. doi: <pub-id pub-id-type="doi">10.1097/00001756-199812210-00031</pub-id>, PMID: <pub-id pub-id-type="pmid">9926868</pub-id></citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussman</surname> <given-names>E.</given-names></name> <name><surname>Winkler</surname> <given-names>I.</given-names></name> <name><surname>Huotilainen</surname> <given-names>M.</given-names></name> <name><surname>Ritter</surname> <given-names>W.</given-names></name> <name><surname>N&#x00E4;&#x00E4;t&#x00E4;nen</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>Top-down effects can modify the initially stimulus-driven auditory organization</article-title>. <source>Brain Res. Cogn. Brain Res.</source> <volume>13</volume>, <fpage>393</fpage>&#x2013;<lpage>405</lpage>. doi: <pub-id pub-id-type="doi">10.1016/s0926-6410(01)00131-8</pub-id>, PMID: <pub-id pub-id-type="pmid">11919003</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>S.</given-names></name> <name><surname>Braren</surname> <given-names>M.</given-names></name> <name><surname>Zubin</surname> <given-names>J.</given-names></name> <name><surname>John</surname> <given-names>E. R.</given-names></name></person-group> (<year>1965</year>). <article-title>Evoked-potential correlates of stimulus uncertainty</article-title>. <source>Science</source> <volume>150</volume>, <fpage>1187</fpage>&#x2013;<lpage>1188</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.150.3700.1187</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todorovic</surname> <given-names>A.</given-names></name> <name><surname>van Ede</surname> <given-names>F.</given-names></name> <name><surname>Maris</surname> <given-names>E.</given-names></name> <name><surname>de Lange</surname> <given-names>F. P.</given-names></name></person-group> (<year>2011</year>). <article-title>Prior expectation mediates neural adaptation to repeated sounds in the auditory cortex: an MEG study</article-title>. <source>J. Neurosci.</source> <volume>31</volume>, <fpage>9118</fpage>&#x2013;<lpage>9123</lpage>. doi: <pub-id pub-id-type="doi">10.1523/jneurosci.1425-11.2011</pub-id>, PMID: <pub-id pub-id-type="pmid">21697363</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vouloumanos</surname> <given-names>A.</given-names></name> <name><surname>Werker</surname> <given-names>J. F.</given-names></name></person-group> (<year>2007</year>). <article-title>Listening to language at birth: evidence for a bias for speech in neonates</article-title>. <source>Dev. Sci.</source> <volume>10</volume>, <fpage>159</fpage>&#x2013;<lpage>164</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-7687.2007.00549.x</pub-id>, PMID: <pub-id pub-id-type="pmid">17286838</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Walsh</surname> <given-names>K. S.</given-names></name> <name><surname>McGovern</surname> <given-names>D. P.</given-names></name> <name><surname>Clark</surname> <given-names>A.</given-names></name> <name><surname>O&#x2019;Connell</surname> <given-names>R. G.</given-names></name></person-group> (<year>2020</year>). <article-title>Evaluating the neurophysiological evidence for predictive processing as a model of perception</article-title>. <source>Ann. N. Y. Acad. Sci.</source> <volume>1464</volume>, <fpage>242</fpage>&#x2013;<lpage>268</lpage>. doi: <pub-id pub-id-type="doi">10.1111/nyas.14321</pub-id>, PMID: <pub-id pub-id-type="pmid">32147856</pub-id></citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Winkler</surname> <given-names>I.</given-names></name> <name><surname>Karmos</surname> <given-names>G.</given-names></name> <name><surname>N&#x00E4;&#x00E4;t&#x00E4;nen</surname> <given-names>R.</given-names></name></person-group> (<year>1996</year>). <article-title>Adaptive modeling of the unattended acoustic environment reflected in the mismatch negativity event-related potential</article-title>. <source>Brain Res.</source> <volume>742</volume>, <fpage>239</fpage>&#x2013;<lpage>252</lpage>. doi: <pub-id pub-id-type="doi">10.1016/s0006-8993(96)01008-6</pub-id>, PMID: <pub-id pub-id-type="pmid">9117400</pub-id></citation></ref>
</ref-list>
</back>
</article>
