<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Hum. Neurosci.</journal-id>
<journal-title>Frontiers in Human Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Hum. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5161</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnhum.2016.00679</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Visual Cortical Entrainment to Motion and Categorical Speech Features during Silent Lipreading</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>O&#x02019;Sullivan</surname> <given-names>Aisling E.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/382860/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Crosse</surname> <given-names>Michael J.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/295571/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Di Liberto</surname> <given-names>Giovanni M.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/182073/overview"/>
</contrib> 
<contrib contrib-type="author" corresp="yes">
<name><surname>Lalor</surname> <given-names>Edmund C.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/20763/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Engineering, Trinity College Dublin</institution> <country>Dublin, Ireland</country></aff>
<aff id="aff2"><sup>2</sup><institution>Trinity Centre for Bioengineering, Trinity College Dublin</institution> <country>Dublin, Ireland</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Pediatrics and Department of Neuroscience, Albert Einstein College of Medicine</institution> <country>Bronx, NY, USA</country></aff>
<aff id="aff4"><sup>4</sup><institution>Trinity College Institute of Neuroscience, Trinity College Dublin</institution> <country>Dublin, Ireland</country></aff>
<aff id="aff5"><sup>5</sup><institution>Department of Biomedical Engineering and Department of Neuroscience, University of Rochester</institution> <country>Rochester, NY, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Benjamin Morillon, Aix-Marseille University, France</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Sanne Ten Oever, Maastricht University, Netherlands; Peter W. Donhauser, McGill University, Canada</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Edmund C. Lalor <email>edmund_lalor&#x00040;urmc.rochester.edu</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>10</volume>
<elocation-id>679</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>12</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 O&#x02019;Sullivan, Crosse, Di Liberto and Lalor.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>O&#x02019;Sullivan, Crosse, Di Liberto and Lalor</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution and reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract><p>Speech is a multisensory percept, comprising an auditory and visual component. While the content and processing pathways of audio speech have been well characterized, the visual component is less well understood. In this work, we expand current methodologies using system identification to introduce a framework that facilitates the study of visual speech in its natural, continuous form. Specifically, we use models based on the unheard acoustic envelope (E), the motion signal (M) and categorical visual speech features (V) to predict EEG activity during silent lipreading. Our results show that each of these models performs similarly at predicting EEG in visual regions and that respective combinations of the individual models (EV, MV, EM and EMV) provide an improved prediction of the neural activity over their constituent models. In comparing these different combinations, we find that the model incorporating all three types of features (EMV) outperforms the individual models, as well as both the EV and MV models, while it performs similarly to the EM model. Importantly, EM does not outperform EV and MV, which, considering the higher dimensionality of the V model, suggests that more data is needed to clarify this finding. Nevertheless, the performance of EMV, and comparisons of the subject performances for the three individual models, provides further evidence to suggest that visual regions are involved in both low-level processing of stimulus dynamics and categorical speech perception. This framework may prove useful for investigating modality-specific processing of visual speech under naturalistic conditions.</p></abstract>
<kwd-group>
<kwd>EEG</kwd>
<kwd>visual speech</kwd>
<kwd>lipreading/speechreading</kwd>
<kwd>visemes</kwd>
<kwd>motion</kwd>
<kwd>temporal response function (TRF)</kwd>
<kwd>EEG prediction</kwd>
</kwd-group>
<counts>
<fig-count count="3"/>
<table-count count="0"/>
<equation-count count="2"/>
<ref-count count="55"/>
<page-count count="11"/>
<word-count count="8865"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction" id="s1">
<title>Introduction</title>
<p>It is well established that during face-to-face conversation visual speech cues play a prominent role in speech perception and comprehension (Summerfield, <xref ref-type="bibr" rid="B50">1992</xref>; Campbell, <xref ref-type="bibr" rid="B12">2008</xref>; Peelle and Sommers, <xref ref-type="bibr" rid="B40">2015</xref>). It has been shown that audiovisual (AV) speech processing benefits from the visual modality at several hierarchical levels of linguistic unit, including syllables (Bernstein et al., <xref ref-type="bibr" rid="B6">2004</xref>), words (Sumby and Pollack, <xref ref-type="bibr" rid="B49">1954</xref>) and sentences (Grant and Seitz, <xref ref-type="bibr" rid="B27">2000</xref>), and that this gain is present in both noisy (Sumby and Pollack, <xref ref-type="bibr" rid="B49">1954</xref>; Ross et al., <xref ref-type="bibr" rid="B44">2007</xref>; Crosse et al., <xref ref-type="bibr" rid="B16">2016b</xref>) and noise-free conditions (Reisberg et al., <xref ref-type="bibr" rid="B42">1987</xref>; Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>). Research on the anatomical organization of auditory speech processing has established a pathway of hierarchical processing, where each level encodes acoustic features of different complexity (Hickok and Poeppel, <xref ref-type="bibr" rid="B28">2007</xref>). And while several studies have reported auditory cortical activation to silent lipreading (Sams et al., <xref ref-type="bibr" rid="B45">1991</xref>; Calvert et al., <xref ref-type="bibr" rid="B11">1997</xref>; Pekkola et al., <xref ref-type="bibr" rid="B41">2005</xref>), the role of this activation remains unclear, i.e., whether it simply serves a modulatory function (Kayser et al., <xref ref-type="bibr" rid="B31">2008</xref>; Falchier et al., <xref ref-type="bibr" rid="B24">2010</xref>) or actually categorizes visual speech features. If the latter were true, one would expect auditory cortical activity during silent speech to track the visual speech features, yet there is a lack of strong evidence of sustained tracking by auditory regions to continuous visual speech (Crosse et al., <xref ref-type="bibr" rid="B17">2015b</xref>). This, coupled with reports of activation of high-level visual pathways during speech reading, has fueled the theory that visual cortex may be capable of processing and interpreting visual speech (for review see Bernstein and Liebenthal, <xref ref-type="bibr" rid="B5">2014</xref>).</p>
<p>Several recent studies have sought to further characterize the role of visual cortex in speech perception. Using cortical surface recordings, stronger visual cortical activity has been observed in response to silent word onsets than AV words (Schepers et al., <xref ref-type="bibr" rid="B46">2015</xref>). This may represent visual cortex accessing word meaning in the absence of an informative audio input. And fMRI research on AV speech perception has shown increases in connectivity strength between putatively multisensory regions and visual cortex when the visual modality is more reliable (Nath and Beauchamp, <xref ref-type="bibr" rid="B36">2011</xref>). In terms of continuous visual speech processing, recent MEG work found extensive (bilateral) entrainment of visual cortex to visual speech (lip movements) when the visual signal was relevant for speech comprehension (Park et al., <xref ref-type="bibr" rid="B38">2016</xref>). Importantly, this entrainment was restricted to a much smaller area in early visual cortex (left-lateralized) when the visual speech was irrelevant. Another study that reconstructed an estimate of the acoustic envelope from occipital EEG data recorded during silent lipreading found a strong correlation between reconstruction accuracy and lipreading ability, suggesting that visual cortex encodes high-level visual speech features (Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>). This is supported by behavioral research which has shown that visually presented syllables are categorically perceived (Weinholtz and Dias, <xref ref-type="bibr" rid="B52">2016</xref>).</p>
<p>Electrophysiological evidence of categorical processing in the context of natural visual speech is lacking. Part of the reason for this is that researchers have focused on studying brain responses to discrete visual syllables, audio-speech envelope entrainment measures, and responses to lip and facial movements. Although studying how the brain encodes features such as the speech envelope and lip movements can inform our understanding of visual speech processing, such simplified speech parameters overlook higher-level categorical processing which may be present. Efforts to further parameterize visual speech have involved the application of multiple sensors to the face and tongue of the speaker (Jiang et al., <xref ref-type="bibr" rid="B30">2002</xref>; Bernstein et al., <xref ref-type="bibr" rid="B9">2011</xref>), a method that is time consuming and yet still cannot fully capture the diverse array of complex motion involved in the production of speech. Here, we take a simplified approach to quantifying visual speech by characterizing the low-level temporal information in the form of the acoustic envelope (given its correlation with speech movements, Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>) and the frame-to-frame motion signal, as well as the higher-level linguistic information as groupings of visually similar phonemes, i.e., visemes (Fisher, <xref ref-type="bibr" rid="B25">1968</xref>). A system identification technique is employed to map these features to the subject&#x02019;s EEG by calculating the so-called temporal response function (TRF) of the system (Crosse et al., <xref ref-type="bibr" rid="B18">2016a</xref>). These TRFs are then tested in their ability to predict unseen EEG data using Pearson&#x02019;s correlation (<italic>r</italic>). The variation in these EEG prediction accuracies across different models is used as a dependent measure for assessing how well the EEG reflects the processing of lower- and higher-level visual speech features. The overarching hypothesis is that visual cortex encodes both the low-level, motion-related features of visual speech, as well as the higher-level, categorical articulatory features. In testing this hypothesis, we aim to establish a framework that facilitates the study of natural visual speech processing in line with methods previously used to characterize the hierarchical organization of speech in the auditory modality (Lalor and Foxe, <xref ref-type="bibr" rid="B32">2010</xref>; Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>).</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and Methods</title>
<p>The EEG data analyzed here were collected as part of a previous study. A more detailed account of the participants, stimuli and experimental procedure can be found in Crosse et al. (<xref ref-type="bibr" rid="B15">2015a</xref>).</p>
<sec id="s2-1">
<title>Subjects</title>
<p>Twenty-one native English speakers (8 females; age range: 19&#x02013;37 years), none of which were trained lipreaders, gave written informed consent. All participants were right-handed, free of neurological diseases, had self-reported normal hearing and normal or corrected-to-normal vision. This study was carried out in accordance with the Declaration of Helsinki. The protocol was approved by the Ethics Committee of the Health Sciences Faculty at Trinity College Dublin, Ireland.</p>
</sec>
<sec id="s2-2">
<title>Stimuli and Procedure</title>
<p>The speech stimuli were drawn from a collection of videos featuring a well-known male speaker. The videos consisted of the speaker&#x02019;s head, shoulders and chest, centered in the frame (see Figure <xref ref-type="fig" rid="F1">1</xref>). The speech was conversational like, and the linguistic content focused on political policy. Stimulus presentation and data recording took place in a dark, sound-attenuated room with participants seated at a distance of 70 cm from the visual display. Visual stimuli were presented on a 19&#x02033; CRT monitor operating at a refresh rate of 60 Hz. Fifteen 60-s videos were rendered into 1280 &#x000D7; 720-pixel movies in VideoPad Video Editor (NCH Software). Soundtracks were deleted from the 15 videos which had a frame rate of 30 frames per second. Participants were instructed to fixate on the speaker&#x02019;s mouth while minimizing eye blinking and all other motor activity during recording. The study from which the data are taken involved seven conditions, most of which included audio speech (Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>). Each of the 15 videos were presented seven times to each subject, once per condition. Presentation order was randomized across conditions and videos. This work examines EEG recordings from the visual-only condition.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Assessing the representation of visual speech in EEG (adapted from Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>).</bold> 128-channel EEG data were recorded while subjects watched videos of continuous, natural speech consisting of a well-known male speaker. Linear regression was used to fit multivariate temporal response functions (mTRFs) between the low-frequency (0.3&#x02013;15 Hz) EEG recordings and seven different representations of the speech stimulus (EV, MV and EM models are not shown). Each mTRF model was then tested for its ability to predict EEG using leave-one-out cross-validation.</p></caption>
<graphic xlink:href="fnhum-10-00679-g0001.tif"/>
</fig>
<p>To encourage active engagement with the video content, participants were required to respond to target words via button press. Before each trial, a target word was displayed on the monitor until the participant was ready to begin. A target word could occur between one and three times in a given 60-s trial. This allowed identification of whether subjects were successful at lipreading or not. A different set of target words was used for each condition to avoid familiarity, and assignment of target words to the seven conditions was counterbalanced across participants. This, combined with the randomized presentation order of the 15 videos, made it quite unlikely that subjects would be able to recognize the silent video from a previously heard audio/AV version.</p>
</sec>
<sec id="s2-3">
<title>Visual Speech Representations</title>
<p>To investigate mappings between different representations of visual speech and low-frequency (0.3&#x02013;15 Hz) EEG, we defined seven representations of visual speech (Figure <xref ref-type="fig" rid="F1">1</xref>): (1) the broadband amplitude envelope of the corresponding acoustic signal (E); (2) the frame-to-frame motion of the video (M); (3) a time aligned sequence of viseme occurrences (V), and (4&#x02013;7) respective combinations of each of the individual models, i.e., EV, MV, EM and EMV.</p>
<p>Previous work has shown that the motion of the mouth during speech is correlated with the acoustic speech envelope (Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>). Therefore, the speech envelope can be thought of as a proxy measure of the local motion related to the mouth area, even though the subjects were not actually presented with the acoustic speech. The broadband amplitude envelope representation was obtained by bandpass filtering the speech signal into 256 logarithmically-spaced frequency bands between 80 Hz and 3000 Hz using a gammachirp filterbank (Irino and Patterson, <xref ref-type="bibr" rid="B29">2006</xref>). The envelope at each of the 256 frequency bands was calculated using a Hilbert transform, and the broadband envelope was obtained by averaging over the 256 narrowband envelopes.</p>
<p>To more explicitly represent the motion in the videos, we calculated their frame-to-frame motion. For each frame, a matrix of motion vectors was calculated using an &#x0201C;Adaptive Rood Pattern Search&#x0201D; block matching algorithm (Barjatya, <xref ref-type="bibr" rid="B3">2004</xref>). A measure of global motion flow was obtained by calculating the sum of all motion vector lengths of each frame (Bartels et al., <xref ref-type="bibr" rid="B4">2008</xref>). This was then upsampled from 30 Hz to 128 Hz to match the rate of the EEG data.</p>
<p>Previous work involving visual speech identification tasks have demonstrated groupings of phonemes which, when presented visually were perceptually similar (consider that a /p/ and a /b/ cannot be distinguished by vision alone;Woodward and Barber, <xref ref-type="bibr" rid="B53">1960</xref>; Fisher, <xref ref-type="bibr" rid="B25">1968</xref>). Each class can thus be defined as the smallest perceptual unit of visual speech, i.e., a viseme. To derive a viseme representation from our videos, we first obtained a phonemic representation as in Di Liberto et al. (<xref ref-type="bibr" rid="B22">2015</xref>), and then converted that to visemes based on the mapping defined in Auer and Bernstein (<xref ref-type="bibr" rid="B2">1997</xref>). Combined models (e.g., EMV) were formed by concatenating the individual models stimuli, resulting in a model whose dimension is equal to the sum of the dimension of each individual model. The phoneme-to-viseme transformation means that timing of our viseme representation is actually tied to the acoustic boundaries rather than the visual. Since features of visual speech have a complex temporal relationship with the sound produced (Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>; Schwartz and Savariaux, <xref ref-type="bibr" rid="B47">2014</xref>), the time window that provided a qualitatively good alignment with the other models (E and M) was 150 ms earlier for the viseme model. As will become clear below, taking account of this fact also enabled us to use a consistent time window across all individual models and combined models in relating the speech stimulus to the EEG, thereby ensuring that our comparison across models was fair.</p>
</sec>
<sec id="s2-4">
<title>EEG Acquisition and Pre-Processing</title>
<p>Continuous EEG data were acquired using an ActiveTwo system (BioSemi) from 128 scalp electrodes. The data were low-pass filtered online below 134 Hz and digitized at a rate of 512 Hz. Triggers were sent by an Arduino Uno microcontroller which detected an audio click at the start of each soundtrack to indicate the start of each trial. Subsequent pre-processing was conducted offline in MATLAB; the data were bandpass filtered between 0.3 Hz and 15 Hz, then downsampled to 128 Hz and re-referenced to the average of all channels. To identify channels with excessive noise, the time series were visually inspected in Cartool (Brunet, <xref ref-type="bibr" rid="B10">1996</xref>), and the standard deviation of each channel was compared with that of the surrounding channels in MATLAB. Channels contaminated by noise were replaced by spline-interpolating the remaining clean channels with weightings based on their relative scalp location in EEGLAB (Delorme and Makeig, <xref ref-type="bibr" rid="B19">2004</xref>).</p>
</sec>
<sec id="s2-5">
<title>Temporal Response Function Estimation</title>
<p>In order to relate the continuous EEG to the various visual speech representations introduced above, we use a regression analysis that describes a mapping from one to the other. This mapping is known as a TRF and was computed using a custom-built toolbox in MATLAB (Crosse et al., <xref ref-type="bibr" rid="B18">2016a</xref>). A TRF can be thought of as a filter that describes how a particular stimulus feature (e.g., the acoustic envelope) is transformed into the continuous EEG at each channel. So if <italic>s</italic>(<italic>t</italic>) represents the stimulus feature at time <italic>t</italic>, the EEG response at channel <italic>n</italic>, <italic>r</italic>(<italic>t</italic>, <italic>n</italic>), can be modeled as a convolution with a to-be-estimated TRF, <italic>w</italic>(&#x003C4;, <italic>n</italic>).</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mi>r</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mrow><mml:mi>w</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mstyle></mml:mrow></mml:math></disp-formula>
<p>where <italic>&#x003B5;</italic>(<italic>t</italic>, <italic>n</italic>) is the residual response at each channel not explained by the model. Of course, the effect of a stimulus event is not seen in the EEG until several tens of milliseconds later and lasts for several hundred milliseconds. So, the TRF is defined across a certain set of time-lags between stimulus and response (<italic>T</italic><sub>min</sub> &#x02212; <italic>T</italic><sub>max</sub>). In our case, we fit TRFs for each 60-s trial using ridge regression expressed in the following matrix form:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mi>w</mml:mi><mml:mtext>&#x2009;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x2009;</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>S</mml:mi><mml:mtext>&#x2009;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x2009;</mml:mtext><mml:mi>&#x03BB;</mml:mi><mml:mi>I</mml:mi><mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x2009;</mml:mtext><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>r</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>where &#x003BB; is the ridge parameter, chosen to optimize the stimulus-response mapping, <italic>S</italic> is a matrix containing a time series of stimulus samples for the window of interest (i.e., the lagged time series), <italic>r</italic> is a matrix of all 128 channels of neural response data, and <italic>I</italic> is the identity matrix which provides regularization and prevents overfitting. For a more detailed explanation of this approach, see Crosse et al. (<xref ref-type="bibr" rid="B18">2016a</xref>).</p>
</sec>
<sec id="s2-6">
<title>EEG Prediction and Model Evaluation</title>
<p>We wished to use this TRF modeling approach to assess how well each of the abovementioned visual speech features was being encoded by visual cortex. To do this, we fit TRFs describing the mapping between each feature and the EEG. Then, using leave-one-out cross-validation, we assess how well we could predict unseen EEG data using the different models. If one can predict EEG with accuracy greater than chance using a particular model or combination of models, one can assert with some confidence that the EEG is reflecting the encoding of that particular feature or set of features. Because we had 15 trials for each subject, leave-one-out cross-validation meant that each TRF was fit to the data from 14 trials and then the average TRF across these 14 trials was used to predict the EEG in the remaining trial (Crosse et al., <xref ref-type="bibr" rid="B18">2016a</xref>).</p>
<p>Prediction accuracy was measured by calculating Pearson&#x02019;s (<italic>r</italic>) linear correlation coefficient between the predicted and original EEG responses at each electrode channel. The time window that best captures the stimulus-response mapping is used for EEG prediction (i.e., <italic>T</italic><sub>min</sub>, <italic>T</italic><sub>max</sub>). This is identified by examining the TRFs on a broad time window (e.g., &#x02212;200 ms to 500 ms) and then choosing the temporal region of the TRF that includes all relevant components that map the stimulus to the EEG with no evident response outside of this range (e.g., 30&#x02013;380 ms for E, M and V models). This time window is also used for the combined models so that differences in performance are not affected by the choice of time window. To optimize performance within each model, we conducted a parameter search (over the range 2<sup>&#x02212;20</sup>, 2<sup>&#x02212;16</sup>&#x02026; 2<sup>20</sup>) for the regularization parameter &#x003BB; that maximized the correlation between the predicted and recorded EEG. To prevent overfitting, the &#x003BB; values were chosen as the value corresponding to the highest mean prediction accuracy across the 15 trials for each subject. The cross-validation is then re-run for each model with a constricted range of &#x003BB; values, based on the range that includes the optimum &#x003BB; value for each subject. Since the cross-validation procedure takes the average performance across trials, the models are not biased towards the test data used for cross-validation. As a result, the TRF is more generalized and capable of predicting new unseen data with a similar accuracy. This procedure is explained in more detail in Crosse et al. (<xref ref-type="bibr" rid="B18">2016a</xref>). After model optimization, a set of 11 electrodes from the occipital region of the scalp (represented by black dots in Figure <xref ref-type="fig" rid="F1">1</xref>) were selected for calculating EEG prediction accuracy because of their consistently high prediction correlations. A nonparametric test was performed in Cartool to test for topographical differences in prediction accuracies across models (i.e., a T-ANOVA). Importantly, there was no statistical difference in the topographic distribution of these predictions between the models (<italic>p</italic> &#x0003E; 0.05), thus ensuring electrode selection did not bias any of the models.</p>
<p>All statistical analyses were conducted using one-way repeated-measures ANOVAs. <italic>Post hoc</italic> comparisons were conducted using two-tailed paired <italic>t</italic>-tests. The level of chance was obtained by calculating the correlation between the predicted EEG and five randomly selected EEG trials from the remaining fourteen. The averages of these predictions (for all models) were then pooled together and the chance level was taken as the 95th percentile of these values. All numerical values are reported as mean &#x000B1; SD.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>To identify neural indices of lower- and higher-level speech reading, we investigated the neural response functions that mapped different representations of visual speech to the low-frequency (0.3&#x02013;15 Hz) EEG from 11 bilateral occipital electrodes (Figure <xref ref-type="fig" rid="F1">1</xref>) of subjects attending to natural visual speech.</p>
<sec id="s3-1">
<title>Envelope, Motion and Visemes are Reflected in EEG</title>
<p>As mentioned above, the acoustic speech envelope can be thought of as a proxy measure for mouth movement (Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>), or may reflect the tracking of visual speech features (Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>). Thus, the use of the envelope model to predict EEG activity sought to investigate further the nature of its representation in visual cortex. The motion model (M) accounts for both local and global motion present (Bartels et al., <xref ref-type="bibr" rid="B4">2008</xref>) which may be related (e.g., cheek, jaw, eye movements) or unrelated (e.g., movement during pauses) to speech. Thus, this model serves to represent the low-level information received by visual cortex during lipreading. The relationship between low-frequency EEG and a categorical phoneme representation of speech has previously been examined for audio speech (Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>). However, no such relationship has been investigated for visual speech. Transforming phonemes into a lower-dimensional viseme representation (V), by grouping visually indistinguishable phonemes, allows us to explore the processing of these visual speech features using electrophysiology. Using these visual speech representations, we find that individually the envelope, motion and viseme models perform similarly at predicting EEG (E: 0.040 &#x000B1; 0.017, M: 0.046 &#x000B1; 0.015, V: 0.047 &#x000B1; 0.021; <italic>F</italic><sub>(2,40)</sub> = 1.42, <italic>p</italic> = 0.253; Figure <xref ref-type="fig" rid="F2">2A</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>(A)</bold> Grand-average (<italic>N</italic> = 21) EEG prediction correlations (Pearson&#x02019;s <italic>r</italic>) for the visual speech models (mean &#x000B1; SEM) for low-frequency EEG (0.3&#x02013;15 Hz). The &#x025B4; indicates the models other models (&#x025B4; <italic>p</italic> &#x0003C; 0.05), except for EV vs. M (<italic>p</italic> = 0.263). There is no difference in performance between these three models (<italic>p</italic> &#x0003E; 0.05). The dotted line represents the 95th percentile of chance-level prediction accuracy. <bold>(B)</bold> The prediction accuracy (<italic>N</italic> = 21) for normal and randomized visemes within their active time points. <bold>(C)</bold> Correlation values between recorded EEG and that predicted by each mTRF model for individual subjects. The subjects are sorted according to the prediction accuracies of the viseme model. (*<italic>p</italic> &#x0003C; 0.05, **<italic>p</italic> &#x0003C; 0.005, ***<italic>p</italic> &#x0003C; 0.001).</p></caption>
<graphic xlink:href="fnhum-10-00679-g0002.tif"/>
</fig>
<p>The successful performance of the motion and envelope models was unsurprising given previous work investigating their relationship with EEG (Goncalves et al., <xref ref-type="bibr" rid="B26">2014</xref>; Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>). But the fact that a model based on labeling the video with categorical visemic labels was as good at predicting EEG as the others was not trivial and suggests that EEG may be reflecting categorical speech processing. However, it may have been possible that its performance could be attributed to the timing of visual speech onsets irrespective of the particular visemes corresponding to those onsets. We sought to test this by randomizing the particular visemes in the speech stimuli while preserving their onset and offset time points. A significant reduction in performance (<italic>t</italic><sub>(20)</sub> = 7.99, <italic>p</italic> = 1.17 &#x000D7; 10<sup>&#x02212;7</sup>; Figure <xref ref-type="fig" rid="F2">2B</xref>) demonstrates that timing alone does not account for the performance of the viseme model. Still, further evidence is required to definitively prove that the viseme model indeed captures high-level, linguistic processing of visual speech.</p>
</sec>
<sec id="s3-2">
<title>Complementary Information Provided by Visual Speech Models</title>
<p>In an effort to reveal the encoding of complementary information between the individual models, we looked at the performance of different model combinations. The approach taken here can be explained in view of our understanding of audio speech processing. Phonemes are defined as the smallest unit of audio speech (Chomsky and Halle, <xref ref-type="bibr" rid="B14">1968</xref>), and despite spectro-temporal variations, different occurrences of the same phoneme are categorically perceived (Okada et al., <xref ref-type="bibr" rid="B37">2010</xref>). Similarly, during natural speech, the motion associated with particular visemes varies. This is largely dependent on the location of the viseme in the word, phrase or utterance as well as the speaker and language (Demorest and Bernstein, <xref ref-type="bibr" rid="B20">1992</xref>; Yakel et al., <xref ref-type="bibr" rid="B54">2000</xref>; Soto-Faraco et al., <xref ref-type="bibr" rid="B48">2007</xref>). Thus, the motion model (M) is expected to capture the variation across visemes, while the viseme model (V) categorically labels visually similar phonemes and so is ignorant of these variations. When these models are employed to predict EEG responses to speechreading, it is expected that individually, they should perform similarly, given their complementary strengths, whereas a combination of the two should result in an enhanced representation, thus improving model performance. Therefore, we derived a model based on combining the motion signal with the viseme representation (MV). In line with our hypothesis, we found a significant improvement in performance of this model over the individual models (MV vs. M: <italic>t</italic><sub>(20)</sub> = 2.7, <italic>p</italic> = 0.014, MV vs. V: <italic>t</italic><sub>(20)</sub> = 6.3, <italic>p</italic> = 3.88 &#x000D7; 10<sup>&#x02212;6</sup>) suggesting that EEG reflects the processing of both low-level motion fluctuations and higher-level visual speech features (Figure <xref ref-type="fig" rid="F2">2A</xref>).</p>
<p>In natural speech, it is known that bodily movements do not function independently of lip movements (Yehia et al., <xref ref-type="bibr" rid="B55">2002</xref>; Munhall et al., <xref ref-type="bibr" rid="B35">2004</xref>) and given the reported correlation between lip movements and the acoustic envelope (Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>), there exists a degree of redundancy between the motion (M) and envelope (E) models. Nonetheless, during pauses and silent periods we would expect the envelope model to provide a more accurate representation of visual speech (i.e., is zero) than the motion, since any motion at these times is unrelated to speech. However, during speech, we might expect the motion model to be more representative of the visual speech content, since it captures the full range of dynamic visual input present. Thus, combining the envelope and motion models (EM) should result in an improved prediction. As expected we found that EM has an improved prediction accuracy (E: <italic>t</italic><sub>(20)</sub> = 6.19, <italic>p</italic> = 4.83 &#x000D7; 10<sup>&#x02212;6</sup>, M: <italic>t</italic><sub>(20)</sub> = 5.03, <italic>p</italic> = 6.37 &#x000D7; 10<sup>&#x02212;5</sup>; Figure <xref ref-type="fig" rid="F2">2A</xref>), demonstrating that these models track complementary neural processes in visual regions.</p>
<p>As previously mentioned, the acoustic envelope represents a proxy measure of lip movements. Another approach to representing articulatory movements is according to the categorical speech units with which the lip movements are associated. Whereas the envelope model (E) tracks differences in lip movements for each particular utterance, the viseme model (V) captures their categorical nature. Based on this reasoning, we expect that these models represent distinct stages of visual speech perception and seek to quantify this. Thus, we formed a combined model (EV) and assessed its ability to predict EEG. And while E and V perform similarly, EV has an improved performance over both individual models (E: <italic>t</italic><sub>(20)</sub> = 2.42, <italic>p</italic> = 0.025, V: <italic>t</italic><sub>(20)</sub> = 3.69, <italic>p</italic> = 0.001; Figure <xref ref-type="fig" rid="F2">2A</xref>). Although the envelope model (E) represents a good correlate of lip movements, we might well expect the motion model (M) to more comprehensively represent the low-level speech content since it is a more direct measure and captures the full range of motion present, e.g., head, eye movements etc. This is supported by the finding that MV outperforms EV (<italic>t</italic><sub>(20)</sub> = 2.23, <italic>p</italic> = 0.037; Figure <xref ref-type="fig" rid="F2">2A</xref>) suggesting that the motion model may capture more of the low-level visual speech features than those that are captured by the envelope.</p>
<p>To continue with our reasoning that the E, M and V models all capture complementary information about visual speech (Figure <xref ref-type="fig" rid="F2">2C</xref>), we used a model involving the combination of all three representations. This combined model, EMV, outperforms each of the individual models (E: <italic>t</italic><sub>(20)</sub> = 3.95, <italic>p</italic> = 7.90 &#x000D7; 10<sup>&#x02212;4</sup>, M: <italic>t</italic><sub>(20)</sub> = 3.83, <italic>p</italic> = 0.001, V: <italic>t</italic><sub>(20)</sub> = 8.17, <italic>p</italic> = 8.40 &#x000D7; 10<sup>&#x02212;8</sup>). The combined EMV model also outperforms EV (<italic>t</italic><sub>(20)</sub> = 6.04, <italic>p</italic> = 6.68 &#x000D7; 10<sup>&#x02212;6</sup>) and MV (<italic>t</italic><sub>(20)</sub> = 3.26, <italic>p</italic> = 0.004), although, despite having the highest mean prediction accuracy, it was not significantly better than EM (<italic>t</italic><sub>(20)</sub> = 0.52, <italic>p</italic> = 0.610). This was somewhat surprising, especially given that there was no significant difference in performance between EM and either EV (<italic>t</italic><sub>(20)</sub> = 1.93, <italic>p</italic> = 0.069) or MV (<italic>t</italic><sub>(20)</sub> = 0.54, <italic>p</italic> = 0.596; Figure <xref ref-type="fig" rid="F2">2A</xref>).</p>
<p>Finally, we wished to investigate whether or not our cross validation approach had successfully insured us against the risk of improved performance coming about simply as a result of having more free parameters. We did this by comparing the envelope and viseme model (EV; 13 free parameters) with the motion model (M; 1 free parameter) and found no significant difference in performance (<italic>t</italic><sub>(20)</sub> = 1.15, <italic>p</italic> = 0.263; Figure <xref ref-type="fig" rid="F2">2A</xref>). This suggests that generation of a higher dimensional model is not guaranteed to give a significantly improved performance over lower dimensional models. In fact, it is possible that higher dimensional models (i.e., V, and any combination including V) would underperform due to the requirement for greater amounts of data to ensure the models are optimally fit.</p>
</sec>
<sec id="s3-3">
<title>Spatiotemporal Representation of Visual Speech</title>
<p>Since the frame-to-frame motion is precisely time-locked to the stimulus, its TRF has sharp positive and negative components (Figure <xref ref-type="fig" rid="F3">3B</xref>). In contrast with this, the envelope and viseme TRFs suffer from some smearing effects since the stimuli are aligned with the unheard acoustic signal and so have a complex temporal relationship with the visual input. As a result, they are not quite as precisely time-locked to the EEG (Figures <xref ref-type="fig" rid="F3">3A,C</xref>). Unsurprisingly, we found that the TRF amplitudes were largest over occipital scalp, suggesting that visual speech is preferentially processed in visual cortex (Figures <xref ref-type="fig" rid="F3">3D&#x02013;F</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Spatiotemporal analysis of mTRF models for natural visual speech.</bold> mTRFs plotted for envelope <bold>(A)</bold>, motion <bold>(B)</bold> and visemes <bold>(C)</bold> at peri-stimulus time lags from &#x02212;50 ms to 500 ms, for representative central and occipital electrode channels. The dotted line separates vowels and consonant visemes. The phonemes contained within each viseme category are shown on the left. <bold>(D)</bold> Topographies of P1 (140 ms) and N1 (234 ms) TRF components for the envelope model. <bold>(E)</bold> Topographies of P1 (125 ms) and N1 (210 ms) TRF components for the motion model. <bold>(F)</bold> Topographies for representative vowel and consonant visemes corresponding to their P1 time points. All visemes have a similar distribution of scalp activity. The black markers in the topographies <bold>(D&#x02013;F)</bold> indicate the channels for which the corresponding TRFs <bold>(A&#x02013;C)</bold> were plotted. The viseme mTRF and the topography colorbar are the same.</p></caption>
<graphic xlink:href="fnhum-10-00679-g0003.tif"/>
</fig>
<p>The viseme TRF also fits expectations in that visemes associated with clear and extended visibility are more strongly represented in the model weights, e.g., bilabials (/b/, /p/, /m/) and labiodentals (/f/, /v/). The response to these frontal consonants is also much sharper than for the other categories and is consistent with behavioral (Lidestam and Beskow, <xref ref-type="bibr" rid="B34">2006</xref>) and electrophysiological studies (van Wassenhove et al., <xref ref-type="bibr" rid="B51">2005</xref>) of viseme identification. The topographic distribution of these TRF weights also showed markedly different patterns for different classes of visemes (Figure <xref ref-type="fig" rid="F3">3F</xref>). However we must express caution when examining the mTRF since viseme occurrences are not equal across categories (for all trials: v1 = 4446, v2 = 3274, v3 = 11,197, v4 = 301, v5 = 16,552, v6 = 7416, v7 = 3722, v8 = 16,147, v9 = 20,092, v10 = 6789, v11 = 2421 and v12 = 2992). Furthermore, the delay between viseme onset and phoneme onset varies depending on their location within a particular utterance (e.g., start of word vs. middle of word). Given that our viseme timings were based on a transformation from phoneme timings, this variation may result in suppression as well as smearing of the TRF amplitudes, e.g., alveolar-fricatives (/d/, /t/, /s/, /z/).</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>In this work, we introduce a framework for investigating the cortical representation of natural visual speech. Specifically, we model how well low- and high-level representations of visual speech are reflected in EEG activity arising from visual cortex. Our results suggest that visual regions are involved in processing the physical stimulus dynamics as well as categorical visual speech features.</p>
<sec id="s4-1">
<title>Visual Cortical Entrainment to Envelope, Motion and Visemes Indices during Lipreading</title>
<p>Neural entrainment to continuous visual speech has been previously studied in the context of physical, low-level information through the use of the acoustic envelope (Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>) and lip movements (Park et al., <xref ref-type="bibr" rid="B38">2016</xref>). Park et al. (<xref ref-type="bibr" rid="B38">2016</xref>) showed that high-level visual regions as well as speech processing regions specifically entrained to the visual component of speech. This is consistent with theories of visual speech encoding through visual pathways (Bernstein and Liebenthal, <xref ref-type="bibr" rid="B5">2014</xref>). Here, we build on these findings to incorporate evidence from perceptual studies (Woodward and Barber, <xref ref-type="bibr" rid="B53">1960</xref>; Fisher, <xref ref-type="bibr" rid="B25">1968</xref>; Auer and Bernstein, <xref ref-type="bibr" rid="B2">1997</xref>), which have identified a basic unit of visual speech, to demonstrate that visually similar phonemes (i.e., visemes) are reflected in low-frequency EEG recordings among persons with normal hearing during lipreading.</p>
<p>Specifically, we provide objective evidence that these features present complementary information to the physical motion present (Figure <xref ref-type="fig" rid="F2">2A</xref>). Central to this, is the meaningful interpretation of visemes, demonstrated by the significant reduction in performance upon randomization of visemes within their active time points (Figure <xref ref-type="fig" rid="F2">2B</xref>). This is in line with the idea that a combination of bottom-up (extracting information from the visual speech signal, i.e., motion) and top-down processing (e.g., use of working memory and categorical perception) are involved in visual speech perception (Lidestam and Beskow, <xref ref-type="bibr" rid="B34">2006</xref>). In keeping with evidence that high-level visual speech is processed in visual cortical regions, the observed visemic entrainment was strongest over occipital electrodes (Figures <xref ref-type="fig" rid="F3">3D&#x02013;F</xref>). These results also align well with recent work modeling the hierarchical processing of acoustic speech in auditory cortex using analogous techniques (Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>), suggesting that visual cortex may indeed process visual speech in a similar hierarchical fashion (Bernstein and Liebenthal, <xref ref-type="bibr" rid="B5">2014</xref>). In addition, this framework facilitates a more detailed analysis of this notion of hierarchical processing of visual speech through analysis of the timing and distribution of the TRF weights across visemes (Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>). This allows one to examine the sensitivity of neural responses to different viseme categories as a function of response latency. Thus, for a hierarchical processing system, one would expect to see differences in viseme encoding according to the different articulatory features that produced them, and for these differences to be more pronounced at longer latencies. However, due to the method used for generating the viseme model (i.e., based on acoustic timings) it was not appropriate, in this case, to carry out further analysis into the viseme mTRF latencies. It would also be informative to compare our TRFs with ERP research on responses to different phonemes and syllables presented visually. For example, similar to our findings, previous ERP work has shown strong responses to labial consonants (e.g., van Wassenhove et al., <xref ref-type="bibr" rid="B51">2005</xref>; Bernstein et al., <xref ref-type="bibr" rid="B7">2008</xref>; Arnal et al., <xref ref-type="bibr" rid="B1">2009</xref>). One caveat here though is that directly comparing the TRF to ERPs is complicated by the fact that the TRFs are inherently different in terms of how they are derived (see Lalor et al., <xref ref-type="bibr" rid="B33">2009</xref>). Furthermore, our stimuli involve natural speech and so have important differences from repeated presentations of phonemes/syllables (e.g., coarticulation effects, rhythmic properties, complex statistical structure, etc.).</p>
<p>Although the idea of visually indistinguishable phonemes was first described over 50 years ago (Woodward and Barber, <xref ref-type="bibr" rid="B53">1960</xref>), there remains disagreement about how, and to what extent, these are actually perceived. Finding categorical responses to visemes using electrophysiology would be an objective way to provide evidence for high-level neural computations involving visual speech perception. However, the viseme model can also be thought of as shorthand labeling of the detailed motion associated with each viseme. In addition, the envelope and motion regressors surely do not capture all of the detailed, relevant motion patterns associated with these visemes. Hence, the viseme model may simply perform well because it leads to an improved measure of the motion in the video, rather than being a result of higher-level, categorical processing. However, this is unlikely given the finding that EM and EMV perform similarly. Indeed this finding raises the question as to whether or not the information represented by the viseme model is already captured by the combination of the envelope and motion. However, it is evident from the performances of the individual models across subjects that V does not correlate with E or M (Figure <xref ref-type="fig" rid="F2">2C</xref>), suggesting that it reflects a distinct process in the neural activity. Furthermore, since EMV contains 12 <italic>additional</italic> parameters to EM, it will require more data to ensure the model is optimally fit. This could be remedied by the collection of more data or the development of a generic model for predicting subject&#x02019;s EEG activity (Di Liberto and Lalor, <xref ref-type="bibr" rid="B21">2016</xref>). Another way to resolve this issue would be to examine the cortical responses to time-reversed visual speech i.e., the frames presented in reverse order, which would facilitate the isolation of motion responses from speech specific responses. In fact, time-reversed visual speech contains segments that are not different from forward speech, such as vowels and transitions into and out of consonants (Ronquest et al., <xref ref-type="bibr" rid="B43">2010</xref>; Bernstein and Liebenthal, <xref ref-type="bibr" rid="B5">2014</xref>). It would also maintain similar low-level information, such as variation in motion between frames and rhythmic pattern. However it would no longer be identified as speech due to removal of lexical information (Paulesu et al., <xref ref-type="bibr" rid="B39">2003</xref>; Ronquest et al., <xref ref-type="bibr" rid="B43">2010</xref>), thus depleting any speech-specific processing effects seen in the forward speech models. In this case, we would expect the performance of the motion model to remain similar, while the visemes performance should be significantly reduced.</p>
<p>Our finding that prediction accuracy of EEG activity using the acoustic envelope is similar to that of motion and visemes (Figure <xref ref-type="fig" rid="F2">2A</xref>) is in line with work showing that occipital channels best reflect the dynamics of the acoustic envelope during lipreading (Crosse et al., <xref ref-type="bibr" rid="B17">2015b</xref>). This may not reflect entrainment to the speech envelope <italic>per se</italic>, but perhaps to speech related movements which are highly correlated with the envelope (Chandrasekaran et al., <xref ref-type="bibr" rid="B13">2009</xref>). Thus, the improved performance demonstrated by the combination of the envelope with the motion model may be explained by reports that visual cortex processes localized and global motion from a natural scene at specialized regions (Bartels et al., <xref ref-type="bibr" rid="B4">2008</xref>). However, our finding that MV outperforms EV suggests there is a greater amount of mutual information between the envelope and viseme models and this may be explained by an analysis-by-synthesis perspective of visual speech encoding. Such a mechanism has previously been implicated to underpin envelope tracking in auditory cortex (Ding et al., <xref ref-type="bibr" rid="B23">2013</xref>) and may also be responsible for the observed entrainment in visual regions, reflecting an internal synthesis of visual speech features. Work from van Wassenhove et al. (<xref ref-type="bibr" rid="B51">2005</xref>) led to a proposal whereby an analysis-by-synthesis mechanism involves perceptual categorization of visual inputs which are used to evaluate auditory inputs. This mechanism has also been suggested by Crosse et al. (<xref ref-type="bibr" rid="B15">2015a</xref>), following the finding of a strong correlation between behavior and envelope tracking during lipreading, and is in line with results presented here. In contrast with this, work from Park et al. (<xref ref-type="bibr" rid="B38">2016</xref>) did not find the acoustic envelope to be coherent with MEG activity in visual cortex. This could be explained by the intrinsic difference between MEG and EEG recordings, where MEG measures current flow tangential to the scalp whereas EEG is sensitive to both tangential and radial components. An alternative explanation is due to differences in study design, since their task did not require subjects to concentrate on lipreading. Instead, subjects attended to audio speech whereby the visual speech was either informative (i.e., matched the attended audio speech) or distracting (i.e., unmatched).</p>
<p>The combined models used here provide us with a means to quantify the differential tracking of particular stimulus features in the EEG. However, the contribution from different stimuli to the EEG prediction could be more clearly defined by regressing out the common variability between the predictors, thus creating independent predictors. One way to achieve this is using partial coherence, which removes the linear contribution of one predictor (e.g., the motion signal) from another (e.g., visemes) in order to reveal entrainment to visemes which cannot be accounted for by the motion signal. This approach has been used previously to separate neural entrainment to lip movements from the speech envelope (Park et al., <xref ref-type="bibr" rid="B38">2016</xref>). Applying this method within the framework presented here could shed light on the unique variability captured by each stimulus in the EEG, and coupled with a high spatial resolution imaging technique, such as fMRI, one could also localize these entrained regions.</p>
</sec>
<sec id="s4-2">
<title>Limitations and Future Directions</title>
<p>It is important to consider some limitations of the current work. First, this experiment was not specifically designed as a visual-only speech experiment and so there are a couple of considerations with regard to how this particular paradigm may have influenced results seen here. In the original experiment, there were seven conditions, six of which contained audio speech. Thus, subjects may have become familiar with the audio content before viewing the visual speech condition. However, to minimize the effect of memory, presentation order of the 105 trials (15 stimuli &#x000D7; 7 speech conditions) was completely randomized within participants. And if subjects were able to relate the silent speech back to a previously heard audio trial, then we would expect to see a much improved target word detection. However, the consistently poor target word detection scores (36.8 &#x000B1; 18.1%, Crosse et al., <xref ref-type="bibr" rid="B15">2015a</xref>) suggests that subjects were not good at recognizing the speech from an earlier condition. Another issue is that the use of a well-known speaker may have aided subjects&#x02019; lipreading ability. However, in the present experiment, the subjects were not familiar with the content of the speech and, as mentioned above, their lipreading performance was relatively poor. Therefore, it is reasonable to assume that these factors did not have a major influence on the results seen here. Another issue with using a well-known speaker is that subjects may have strongly imagined the audio speech that accompanied our silent videos. If so, one might expect to see evidence of auditory cortical tracking of such imagined audition. Indeed, we have previously sought to uncover EEG evidence for such a phenomenon with well-known speech (Crosse et al., <xref ref-type="bibr" rid="B17">2015b</xref>). However, the evidence we have found for this has been weak at best. And, again, because in the present experiment the content was unfamiliar, we do not expect imagined audition will have played a significant role.</p>
<p>One other limitation of the present experiment is that the viseme representation of the speech used here is sub-optimal since the viseme stimulus is derived from a phoneme alignment of the speech signal before transformation into viseme groups. The accuracy of this alignment can be affected by different representations of the same phonemes. For example, we might expect that the alignment work better for the /b/ phoneme since this is associated with a large spike in the spectrogram and so is easier for the software to identify. This can in turn result in a more effective mapping of stimulus to EEG compared with phonemes which may be more difficult to accurately align (e.g., /s/ or /z/). Secondly, since it is essentially a phonetic mapping, the visemes are tied to the acoustic boundaries, and given the complex temporal relationship between audio and visual speech the viseme timings are imprecise. Thus, when these features are mapped to EEG, these slight variations can cause smearing of response peaks (as seen in Figure <xref ref-type="fig" rid="F3">3C</xref>). It is also important to keep in mind the poor lipreading ability of normal hearing subjects. We would expect to see a large boost in viseme model performance for trained lipreaders or subjects familiar with the speech content due to an improved ability to recognize the silent speech (Bernstein et al., <xref ref-type="bibr" rid="B8">2000</xref>). Nevertheless, our results are consistent with the notion that the visual system interprets visual speech and takes a first step to investigate how the visual system may represent the rich psycholinguistic structure of visual speech.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>Conclusion</title>
<p>In summary, we have presented a framework to objectively assess the possibility of speech-sensitive regions in visual cortex encoding high-level, categorical speech features. The results presented are akin to findings in the auditory domain and support the theory that visual regions are involved in categorical speech perception (Bernstein and Liebenthal, <xref ref-type="bibr" rid="B5">2014</xref>). Future work will seek to strengthen the evidence provided here, for example, by applying these models to subjects watching known vs. unknown speech, or by recruiting individuals with hearing impairments who have superior speechreading ability. In addition, the coupling of these models with recently developed representations of auditory speech (Di Liberto et al., <xref ref-type="bibr" rid="B22">2015</xref>) may shed light on how the brain weights sensory inputs from different modalities to form a multi-sensory percept.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>AEO and MJC contributed equally to the work. MJC and ECL designed research; MJC collected data; AEO performed research; AEO, MJC and GMDL analyzed data; AEO, MJC, GMDL and ECL interpreted data and wrote the article.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was supported by the Programme for Research in Third-Level Institutions and co-funded under the European Regional Development fund. Additional support was provided by the Irish Research Council&#x02019;s Government of Ireland Postgraduate Scholarship Scheme.</p>
</sec>
<sec id="s8">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arnal</surname> <given-names>L. H.</given-names></name> <name><surname>Morillon</surname> <given-names>B.</given-names></name> <name><surname>Kell</surname> <given-names>C. A.</given-names></name> <name><surname>Giraud</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2009</year>). <article-title>Dual neural routing of visual facilitation in speech processing</article-title>. <source>J. Neurosci.</source> <volume>29</volume>, <fpage>13445</fpage>&#x02013;<lpage>13453</lpage>. <pub-id pub-id-type="doi">10.1523/jneurosci.3194-09.2009</pub-id><pub-id pub-id-type="pmid">19864557</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Auer</surname> <given-names>E. T.</given-names> <suffix>Jr.</suffix></name> <name><surname>Bernstein</surname> <given-names>L. E.</given-names></name></person-group> (<year>1997</year>). <article-title>Speechreading and the structure of the lexicon: computationally modeling the effects of reduced phonetic distinctiveness on lexical uniqueness</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>102</volume>, <fpage>3704</fpage>&#x02013;<lpage>3710</lpage>. <pub-id pub-id-type="doi">10.1121/1.420402</pub-id><pub-id pub-id-type="pmid">9407662</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barjatya</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Block matching algorithms for motion estimation</article-title>. <source>IEEE Trans. Evol. Comput.</source> <volume>8</volume>, <fpage>225</fpage>&#x02013;<lpage>239</lpage>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bartels</surname> <given-names>A.</given-names></name> <name><surname>Zeki</surname> <given-names>S.</given-names></name> <name><surname>Logothetis</surname> <given-names>N. K.</given-names></name></person-group> (<year>2008</year>). <article-title>Natural vision reveals regional specialization to local motion and to contrast-Invariant, global flow in the human brain</article-title>. <source>Cereb. Cortex</source> <volume>18</volume>, <fpage>705</fpage>&#x02013;<lpage>717</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhm107</pub-id><pub-id pub-id-type="pmid">17615246</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>L. E.</given-names></name> <name><surname>Auer</surname> <given-names>E. T.</given-names> <suffix>Jr.</suffix></name> <name><surname>Takayanagi</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>Auditory speech detection in noise enhanced by lipreading</article-title>. <source>Speech Commun.</source> <volume>44</volume>, <fpage>5</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.specom.2004.10.011</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>L. E.</given-names></name> <name><surname>Auer</surname> <given-names>E. T.</given-names></name> <name><surname>Wagner</surname> <given-names>M.</given-names></name> <name><surname>Ponton</surname> <given-names>C. W.</given-names></name></person-group> (<year>2008</year>). <article-title>Spatiotemporal dynamics of audiovisual speech processing</article-title>. <source>Neuroimage</source> <volume>39</volume>, <fpage>423</fpage>&#x02013;<lpage>435</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.08.035</pub-id><pub-id pub-id-type="pmid">17920933</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>L. E.</given-names></name> <name><surname>Demorest</surname> <given-names>M. E.</given-names></name> <name><surname>Tucker</surname> <given-names>P. E.</given-names></name></person-group> (<year>2000</year>). <article-title>Speech perception without hearing</article-title>. <source>Percept. Psychophys.</source> <volume>62</volume>, <fpage>233</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.3758/bf03205546</pub-id><pub-id pub-id-type="pmid">10723205</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>L. E.</given-names></name> <name><surname>Jiang</surname> <given-names>J.</given-names></name> <name><surname>Pantazis</surname> <given-names>D.</given-names></name> <name><surname>Lu</surname> <given-names>Z. L.</given-names></name> <name><surname>Joshi</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Visual phonetic processing localized using speech and nonspeech face gestures in video and point-light displays</article-title>. <source>Hum. Brain Mapp.</source> <volume>32</volume>, <fpage>1660</fpage>&#x02013;<lpage> 1676</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.21139</pub-id><pub-id pub-id-type="pmid">20853377</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>L. E.</given-names></name> <name><surname>Liebenthal</surname> <given-names>E.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural pathways for visual speech perception</article-title>. <source>Front. Neurosci.</source> <volume>8</volume>:<fpage>386</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00386</pub-id><pub-id pub-id-type="pmid">25520611</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunet</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <source>Cartool 3.51 ed.: The Functional Brain Mapping Laboratory</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.fbmlab.com/cartool-software">http://www.fbmlab.com/cartool-software</ext-link></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Calvert</surname> <given-names>G. A.</given-names></name> <name><surname>Bullmore</surname> <given-names>E. T.</given-names></name> <name><surname>Brammer</surname> <given-names>M. J.</given-names></name> <name><surname>Campbell</surname> <given-names>R.</given-names></name> <name><surname>Williams</surname> <given-names>S. C.</given-names></name> <name><surname>Mcguire</surname> <given-names>P. K.</given-names></name> <etal/></person-group>. (<year>1997</year>). <article-title>Activation of auditory cortex during silent lipreading</article-title>. <source>Science</source> <volume>276</volume>, <fpage>593</fpage>&#x02013;<lpage>596</lpage>. <pub-id pub-id-type="doi">10.1126/science.276.5312.593</pub-id><pub-id pub-id-type="pmid">9110978</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>The processing of audio-visual speech: empirical and neural bases</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>363</volume>, <fpage>1001</fpage>&#x02013;<lpage>1010</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2007.2155</pub-id><pub-id pub-id-type="pmid">17827105</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chandrasekaran</surname> <given-names>C.</given-names></name> <name><surname>Trubanova</surname> <given-names>A.</given-names></name> <name><surname>Stillittano</surname> <given-names>S.</given-names></name> <name><surname>Caplier</surname> <given-names>A.</given-names></name> <name><surname>Ghazanfar</surname> <given-names>A. A.</given-names></name></person-group> (<year>2009</year>). <article-title>The natural statistics of audiovisual speech</article-title>. <source>PLoS Comput. Biol.</source> <volume>5</volume>:<fpage>e1000436</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000436</pub-id><pub-id pub-id-type="pmid">19609344</pub-id></citation></ref>
<ref id="B14"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Chomsky</surname> <given-names>N.</given-names></name> <name><surname>Halle</surname> <given-names>M.</given-names></name></person-group> (<year>1968</year>). <source>The Sound Pattern of English</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Harper &#x00026; Row</publisher-name>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crosse</surname> <given-names>M. J.</given-names></name> <name><surname>Butler</surname> <given-names>J. S.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2015a</year>). <article-title>Congruent visual speech enhances cortical entrainment to continuous auditory speech in noise-free conditions</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>14195</fpage>&#x02013;<lpage>14204</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1829-15.2015</pub-id><pub-id pub-id-type="pmid">26490860</pub-id></citation></ref>
<ref id="B17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Crosse</surname> <given-names>M. J.</given-names></name> <name><surname>ElShafei</surname> <given-names>H. A.</given-names></name> <name><surname>Foxe</surname> <given-names>J. J.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2015b</year>). &#x0201C;<article-title>Investigating the temporal dynamics of auditory cortical activation to silent lipreading</article-title>,&#x0201D; in <conf-name>2015 7th International IEEE/EMBS Conference on Neural Engineering (NER)</conf-name> (<publisher-loc>Montpellier, France</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>308</fpage>&#x02013;<lpage>311</lpage>.</citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crosse</surname> <given-names>M. J.</given-names></name> <name><surname>Di Liberto</surname> <given-names>G. M.</given-names></name> <name><surname>Bednar</surname> <given-names>A.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2016a</year>). <article-title>The multivariate temporal response function (mTRF) toolbox: a MATLAB toolbox for relating neural signals to continuous stimuli</article-title>. <source>Front. Hum. Neurosci.</source> <volume>10</volume>:<fpage>604</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2016.00604</pub-id><pub-id pub-id-type="pmid">27965557</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crosse</surname> <given-names>M. J.</given-names></name> <name><surname>Di Liberto</surname> <given-names>G. M.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2016b</year>). <article-title>Eye can hear clearly now: inverse effectiveness in natural audiovisual speech processing relies on long-term crossmodal temporal integration</article-title>. <source>J. Neurosci.</source> <volume>36</volume>, <fpage>9888</fpage>&#x02013;<lpage>9895</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1396-16.2016</pub-id><pub-id pub-id-type="pmid">27656026</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Delorme</surname> <given-names>A.</given-names></name> <name><surname>Makeig</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>EEGLAB: an open source toolbox for analysis of single-trial EEG dynamics including independent component analysis</article-title>. <source>J. Neurosci. Methods</source> <volume>134</volume>, <fpage>9</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1016/j.jneumeth.2003.10.009</pub-id><pub-id pub-id-type="pmid">15102499</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Demorest</surname> <given-names>M. E.</given-names></name> <name><surname>Bernstein</surname> <given-names>L. E.</given-names></name></person-group> (<year>1992</year>). <article-title>Sources of variability in speechreading sentences: a generalizability analysis</article-title>. <source>J. Speech Hear. Res.</source> <volume>35</volume>, <fpage>876</fpage>&#x02013;<lpage>891</lpage>. <pub-id pub-id-type="doi">10.1044/jshr.3504.876</pub-id><pub-id pub-id-type="pmid">1405543</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Di Liberto</surname> <given-names>G. M.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2016</year>). <article-title>Indexing cortical entrainment to natural speech at the phonemic level: Methodological considerations for applied research</article-title>. <source>In Review</source></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Di Liberto</surname> <given-names>G. M.</given-names></name> <name><surname>O&#x02019;Sullivan</surname> <given-names>J. A.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Low-frequency cortical entrainment to speech reflects phoneme-level processing</article-title>. <source>Curr. Biol.</source> <volume>25</volume>, <fpage>2457</fpage>&#x02013;<lpage>2465</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2015.08.030</pub-id><pub-id pub-id-type="pmid">26412129</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>N.</given-names></name> <name><surname>Chatterjee</surname> <given-names>M.</given-names></name> <name><surname>Simon</surname> <given-names>J. Z.</given-names></name></person-group> (<year>2013</year>). <article-title>Robust cortical entrainment to the speech envelope relies on the spectro-temporal fine structure</article-title>. <source>Neuroimage</source> <volume>88c</volume>, <fpage>41</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.10.054</pub-id><pub-id pub-id-type="pmid">24188816</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Falchier</surname> <given-names>A.</given-names></name> <name><surname>Schroeder</surname> <given-names>C. E.</given-names></name> <name><surname>Hackett</surname> <given-names>T. A.</given-names></name> <name><surname>Lakatos</surname> <given-names>P.</given-names></name> <name><surname>Nascimento-Silva</surname> <given-names>S.</given-names></name> <name><surname>Ulbert</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Projection from visual areas V2 and prostriata to caudal auditory cortex in the monkey</article-title>. <source>Cereb. Cortex</source> <volume>20</volume>, <fpage>1529</fpage>&#x02013;<lpage>1538</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhp213</pub-id><pub-id pub-id-type="pmid">19875677</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>C. G.</given-names></name></person-group> (<year>1968</year>). <article-title>Confusions among visually perceived consonants</article-title>. <source>J. Speech Hear. Res.</source> <volume>11</volume>, <fpage>796</fpage>&#x02013;<lpage>804</lpage>. <pub-id pub-id-type="doi">10.1044/jshr.1104.796</pub-id><pub-id pub-id-type="pmid">5719234</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goncalves</surname> <given-names>N. R.</given-names></name> <name><surname>Whelan</surname> <given-names>R.</given-names></name> <name><surname>Foxe</surname> <given-names>J. J.</given-names></name> <name><surname>Lalor</surname> <given-names>E. C.</given-names></name></person-group> (<year>2014</year>). <article-title>Towards obtaining spatiotemporally precise responses to continuous sensory stimuli in humans: a general linear modeling approach to EEG</article-title>. <source>Neuroimage</source> <volume>97</volume>, <fpage>196</fpage>&#x02013;<lpage>205</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2014.04.012</pub-id><pub-id pub-id-type="pmid">24736185</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grant</surname> <given-names>K. W.</given-names></name> <name><surname>Seitz</surname> <given-names>P. F.</given-names></name></person-group> (<year>2000</year>). <article-title>The use of visible speech cues for improving auditory detection of spoken sentences</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>108</volume>, <fpage>1197</fpage>&#x02013;<lpage>1208</lpage>. <pub-id pub-id-type="doi">10.1121/1.1288668</pub-id><pub-id pub-id-type="pmid">11008820</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hickok</surname> <given-names>G.</given-names></name> <name><surname>Poeppel</surname> <given-names>D.</given-names></name></person-group> (<year>2007</year>). <article-title>The cortical organization of speech processing</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>8</volume>, <fpage>393</fpage>&#x02013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2113</pub-id><pub-id pub-id-type="pmid">17431404</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Irino</surname> <given-names>T.</given-names></name> <name><surname>Patterson</surname> <given-names>R. D.</given-names></name></person-group> (<year>2006</year>). <article-title>A dynamic compressive &#x003B3; chirp auditory filterbank</article-title>. <source>IEEE Trans. Audio Speech Lang. Process.</source> <volume>14</volume>, <fpage>2222</fpage>&#x02013;<lpage>2232</lpage>. <pub-id pub-id-type="doi">10.1109/tasl.2006.874669</pub-id><pub-id pub-id-type="pmid">19330044</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J.</given-names></name> <name><surname>Alwan</surname> <given-names>A.</given-names></name> <name><surname>Keating</surname> <given-names>P.</given-names></name> <name><surname>Auer</surname> <given-names>E.</given-names> <suffix>Jr.</suffix></name> <name><surname>Bernstein</surname> <given-names>L.</given-names></name></person-group> (<year>2002</year>). <article-title>On the relationship between face movements, tongue movements and speech acoustics</article-title>. <source>EURASIP J. Adv. Signal Process.</source> <volume>2002</volume>, <fpage>1174</fpage>&#x02013;<lpage>1188</lpage>. <pub-id pub-id-type="doi">10.1155/s1110865702206046</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kayser</surname> <given-names>C.</given-names></name> <name><surname>Petkov</surname> <given-names>C. I.</given-names></name> <name><surname>Logothetis</surname> <given-names>N. K.</given-names></name></person-group> (<year>2008</year>). <article-title>Visual modulation of neurons in auditory cortex</article-title>. <source>Cereb. Cortex</source> <volume>18</volume>, <fpage>1560</fpage>&#x02013;<lpage>1574</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhm187</pub-id><pub-id pub-id-type="pmid">18180245</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lalor</surname> <given-names>E. C.</given-names></name> <name><surname>Foxe</surname> <given-names>J. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Neural responses to uninterrupted natural speech can be extracted with precise temporal resolution</article-title>. <source>Eur. J. Neurosci.</source> <volume>31</volume>, <fpage>189</fpage>&#x02013;<lpage>193</lpage>. <pub-id pub-id-type="doi">10.1111/j.1460-9568.2009.07055.x</pub-id><pub-id pub-id-type="pmid">20092565</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lalor</surname> <given-names>E. C.</given-names></name> <name><surname>Power</surname> <given-names>A. J.</given-names></name> <name><surname>Reilly</surname> <given-names>R. B.</given-names></name> <name><surname>Foxe</surname> <given-names>J. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Resolving precise temporal processing properties of the auditory system using continuous stimuli</article-title>. <source>J. Neurophysiol.</source> <volume>102</volume>, <fpage>349</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1152/jn.90896.2008</pub-id><pub-id pub-id-type="pmid">19439675</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lidestam</surname> <given-names>B.</given-names></name> <name><surname>Beskow</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Visual phonemic ambiguity and speechreading</article-title>. <source>J. Speech Lang. Hear. Res.</source> <volume>49</volume>, <fpage>835</fpage>&#x02013;<lpage>847</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2006/059)</pub-id><pub-id pub-id-type="pmid">16908878</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Munhall</surname> <given-names>K. G.</given-names></name> <name><surname>Jones</surname> <given-names>J. A.</given-names></name> <name><surname>Callan</surname> <given-names>D. E.</given-names></name> <name><surname>Kuratate</surname> <given-names>T.</given-names></name> <name><surname>Vatikiotis-Bateson</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Visual prosody and speech intelligibility: head movement improves auditory speech perception</article-title>. <source>Psychol. Sci.</source> <volume>15</volume>, <fpage>133</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1111/j.0963-7214.2004.01502010.x</pub-id><pub-id pub-id-type="pmid">14738521</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nath</surname> <given-names>A. R.</given-names></name> <name><surname>Beauchamp</surname> <given-names>M. S.</given-names></name></person-group> (<year>2011</year>). <article-title>Dynamic changes in superior temporal sulcus connectivity during perception of noisy audiovisual speech</article-title>. <source>J. Neurosci.</source> <volume>31</volume>, <fpage>1704</fpage>&#x02013;<lpage>1714</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4853-10.2011</pub-id><pub-id pub-id-type="pmid">21289179</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Okada</surname> <given-names>K.</given-names></name> <name><surname>Rong</surname> <given-names>F.</given-names></name> <name><surname>Venezia</surname> <given-names>J.</given-names></name> <name><surname>Matchin</surname> <given-names>W.</given-names></name> <name><surname>Hsieh</surname> <given-names>I. H.</given-names></name> <name><surname>Saberi</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Hierarchical organization of human auditory cortex: evidence from acoustic invariance in the response to intelligible speech</article-title>. <source>Cereb. Cortex</source> <volume>20</volume>, <fpage>2486</fpage>&#x02013;<lpage>2495</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhp318</pub-id><pub-id pub-id-type="pmid">20100898</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>H.</given-names></name> <name><surname>Kayser</surname> <given-names>C.</given-names></name> <name><surname>Thut</surname> <given-names>G.</given-names></name> <name><surname>Gross</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Lip movements entrain the observers&#x02019; low-frequency brain oscillations to facilitate speech intelligibility</article-title>. <source>Elife</source> <volume>5</volume>:<fpage>e14521</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.14521</pub-id><pub-id pub-id-type="pmid">27146891</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paulesu</surname> <given-names>E.</given-names></name> <name><surname>Perani</surname> <given-names>D.</given-names></name> <name><surname>Blasi</surname> <given-names>V.</given-names></name> <name><surname>Silani</surname> <given-names>G.</given-names></name> <name><surname>Borghese</surname> <given-names>N. A.</given-names></name> <name><surname>De Giovanni</surname> <given-names>U.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>A functional-anatomical model for lipreading</article-title>. <source>J. Neurophysiol.</source> <volume>90</volume>, <fpage>2005</fpage>&#x02013;<lpage>2013</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00926.2002</pub-id><pub-id pub-id-type="pmid">12750414</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peelle</surname> <given-names>J. E.</given-names></name> <name><surname>Sommers</surname> <given-names>M. S.</given-names></name></person-group> (<year>2015</year>). <article-title>Prediction and constraint in audiovisual speech perception</article-title>. <source>Cortex</source> <volume>68</volume>, <fpage>169</fpage>&#x02013;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.1016/j.cortex.2015.03.006</pub-id><pub-id pub-id-type="pmid">25890390</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pekkola</surname> <given-names>J.</given-names></name> <name><surname>Ojanen</surname> <given-names>V.</given-names></name> <name><surname>Autti</surname> <given-names>T.</given-names></name> <name><surname>J&#x000E4;&#x000E4;skel&#x000E4;inen</surname> <given-names>I. P.</given-names></name> <name><surname>M&#x000F6;tt&#x000F6;nen</surname> <given-names>R.</given-names></name> <name><surname>Tarkiainen</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Primary auditory cortex activation by visual speech: an fMRI study at 3 T</article-title>. <source>Neuroreport</source> <volume>16</volume>, <fpage>125</fpage>&#x02013;<lpage>128</lpage>. <pub-id pub-id-type="doi">10.1097/00001756-200502080-00010</pub-id><pub-id pub-id-type="pmid">15671860</pub-id></citation></ref>
<ref id="B42"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Reisberg</surname> <given-names>D.</given-names></name> <name><surname>Mclean</surname> <given-names>J.</given-names></name> <name><surname>Goldfield</surname> <given-names>A.</given-names></name></person-group> (<year>1987</year>). &#x0201C;<article-title>Easy to hear but hard to understand: a lip-reading advantage with intact auditory stimuli</article-title>,&#x0201D; in <source>The Psychology of Lip-Reading</source>, eds <person-group person-group-type="editor"><name><surname>Dodd</surname> <given-names>B.</given-names></name> <name><surname>Campbell</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Hillsdale</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>), <fpage>97</fpage>&#x02013;<lpage>114</lpage>.</citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ronquest</surname> <given-names>R. E.</given-names></name> <name><surname>Levi</surname> <given-names>S. V.</given-names></name> <name><surname>Pisoni</surname> <given-names>D. B.</given-names></name></person-group> (<year>2010</year>). <article-title>Language identification from visual-only speech signals</article-title>. <source>Atten. Percept Psychophys.</source> <volume>72</volume>, <fpage>1601</fpage>&#x02013;<lpage>1613</lpage>. <pub-id pub-id-type="doi">10.3758/APP.72.6.1601</pub-id><pub-id pub-id-type="pmid">20675804</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ross</surname> <given-names>L. A.</given-names></name> <name><surname>Saint-Amour</surname> <given-names>D.</given-names></name> <name><surname>Leavitt</surname> <given-names>V. M.</given-names></name> <name><surname>Javitt</surname> <given-names>D. C.</given-names></name> <name><surname>Foxe</surname> <given-names>J. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Do you see what I am saying? Exploring visual enhancement of speech comprehension in noisy environments</article-title>. <source>Cereb. Cortex</source> <volume>17</volume>, <fpage>1147</fpage>&#x02013;<lpage>1153</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhl024</pub-id><pub-id pub-id-type="pmid">16785256</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sams</surname> <given-names>M.</given-names></name> <name><surname>Aulanko</surname> <given-names>R.</given-names></name> <name><surname>H&#x000E4;m&#x000E4;l&#x000E4;inen</surname> <given-names>M.</given-names></name> <name><surname>Hari</surname> <given-names>R.</given-names></name> <name><surname>Lounasmaa</surname> <given-names>O. V.</given-names></name> <name><surname>Lu</surname> <given-names>S.-T.</given-names></name> <etal/></person-group>. (<year>1991</year>). <article-title>Seeing speech: visual information from lip movements modifies activity in the human auditory cortex</article-title>. <source>Neurosci. Lett.</source> <volume>127</volume>, <fpage>141</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1016/0304-3940(91)90914-f</pub-id><pub-id pub-id-type="pmid">1881611</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schepers</surname> <given-names>I. M.</given-names></name> <name><surname>Yoshor</surname> <given-names>D.</given-names></name> <name><surname>Beauchamp</surname> <given-names>M. S.</given-names></name></person-group> (<year>2015</year>). <article-title>Electrocorticography reveals enhanced visual cortex responses to visual speech</article-title>. <source>Cereb. Cortex</source> <volume>25</volume>, <fpage>4103</fpage>&#x02013;<lpage>4110</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhu127</pub-id><pub-id pub-id-type="pmid">24904069</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwartz</surname> <given-names>J. L.</given-names></name> <name><surname>Savariaux</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>No, there is no 150 ms lead of visual speech on auditory speech, but a range of audiovisual asynchronies varying from small audio lead to large audio lag</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003743</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003743</pub-id><pub-id pub-id-type="pmid">25079216</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soto-Faraco</surname> <given-names>S.</given-names></name> <name><surname>Navarra</surname> <given-names>J.</given-names></name> <name><surname>Weikum</surname> <given-names>W. M.</given-names></name> <name><surname>Vouloumanos</surname> <given-names>A.</given-names></name> <name><surname>Sebasti&#x000E1;n-Galles</surname> <given-names>N.</given-names></name> <name><surname>Werker</surname> <given-names>J. F.</given-names></name></person-group> (<year>2007</year>). <article-title>Discriminating languages by speech-reading</article-title>. <source>Percept. Psychophys.</source> <volume>69</volume>, <fpage>218</fpage>&#x02013;<lpage>231</lpage>. <pub-id pub-id-type="doi">10.3758/bf03193744</pub-id><pub-id pub-id-type="pmid">17557592</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sumby</surname> <given-names>W. H.</given-names></name> <name><surname>Pollack</surname> <given-names>I.</given-names></name></person-group> (<year>1954</year>). <article-title>Visual contribution to speech intelligibility in noise</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>26</volume>, <fpage>212</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1121/1.1907309</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Summerfield</surname> <given-names>Q.</given-names></name></person-group> (<year>1992</year>). <article-title>Lipreading and audio-visual speech perception</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>335</volume>, <fpage>71</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.1992.0009</pub-id><pub-id pub-id-type="pmid">1348140</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Wassenhove</surname> <given-names>V.</given-names></name> <name><surname>Grant</surname> <given-names>K. W.</given-names></name> <name><surname>Poeppel</surname> <given-names>D.</given-names></name></person-group> (<year>2005</year>). <article-title>Visual speech speeds up the neural processing of auditory speech</article-title>. <source>Proc. Natl. Acad. Sci. U S A</source> <volume>102</volume>, <fpage>1181</fpage>&#x02013;<lpage>1186</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0408949102</pub-id><pub-id pub-id-type="pmid">15647358</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weinholtz</surname> <given-names>C.</given-names></name> <name><surname>Dias</surname> <given-names>J. W.</given-names></name></person-group> (<year>2016</year>). <article-title>Categorical perception of visual speech information</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>139</volume>, <fpage>2018</fpage>&#x02013;<lpage>2018</lpage>. <pub-id pub-id-type="doi">10.1121/1.4949950</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woodward</surname> <given-names>M. F.</given-names></name> <name><surname>Barber</surname> <given-names>C. G.</given-names></name></person-group> (<year>1960</year>). <article-title>Phoneme perception in lipreading</article-title>. <source>J. Speech Hear. Res.</source> <volume>3</volume>, <fpage>212</fpage>&#x02013;<lpage>222</lpage>. <pub-id pub-id-type="doi">10.1044/jshr.0303.212</pub-id><pub-id pub-id-type="pmid">13845910</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yakel</surname> <given-names>D. A.</given-names></name> <name><surname>Rosenblum</surname> <given-names>L. D.</given-names></name> <name><surname>Fortier</surname> <given-names>M. A.</given-names></name></person-group> (<year>2000</year>). <article-title>Effects of talker variability on speechreading</article-title>. <source>Percept. Psychophys.</source> <volume>62</volume>, <fpage>1405</fpage>&#x02013;<lpage>1412</lpage>. <pub-id pub-id-type="doi">10.3758/bf03212142</pub-id><pub-id pub-id-type="pmid">11143452</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yehia</surname> <given-names>H. C.</given-names></name> <name><surname>Kuratate</surname> <given-names>T.</given-names></name> <name><surname>Vatikiotis-Bateson</surname> <given-names>E.</given-names></name></person-group> (<year>2002</year>). <article-title>Linking facial animation, head motion and speech acoustics</article-title>. <source>J. Phon.</source> <volume>30</volume>, <fpage>555</fpage>&#x02013;<lpage>568</lpage>. <pub-id pub-id-type="doi">10.1006/jpho.2002.0165</pub-id></citation></ref>
</ref-list>
</back>
</article>