<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2021.767840</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Mini Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Listening Effort Informed Quality of Experience Evaluation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Sun</surname> <given-names>Pheobe Wenyi</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1257986/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Hines</surname> <given-names>Andrew</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1462578/overview"/>
</contrib>
</contrib-group>
<aff><institution>QxLab</institution>, <addr-line>School of Computer Science, University College Dublin</addr-line>, <addr-line>Dublin</addr-line>, <country>Ireland</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Guangtao Zhai, Shanghai Jiao Tong University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Xiongkuo Min, University of Texas at Austin, United States; Yucheng Zhu, Shanghai Jiao Tong University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Pheobe Wenyi Sun <email>wenyi.sun&#x00040;ucdconnect.ie</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Perception Science, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>767840</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Sun and Hines.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Sun and Hines</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> 
</permissions>
<abstract><p>Perceived quality of experience for speech listening is influenced by cognitive processing and can affect a listener&#x00027;s comprehension, engagement and responsiveness. Quality of Experience (QoE) is a paradigm used within the media technology community to assess media quality by linking quantifiable media parameters to perceived quality. The established QoE framework provides a general definition of QoE, categories of possible quality influencing factors, and an identified QoE formation pathway. These assist researchers to implement experiments and to evaluate perceived quality for any applications. The QoE formation pathways in the current framework do not attempt to capture cognitive effort effects and the standard experimental assessments of QoE minimize the influence from cognitive processes. The impact of cognitive processes and how they can be captured within the QoE framework have not been systematically studied by the QoE research community. This article reviews research from the fields of audiology and cognitive science regarding how cognitive processes influence the quality of listening experience. The cognitive listening mechanism theories are compared with the QoE formation mechanism in terms of the quality contributing factors, experience formation pathways, and measures for experience. The review prompts a proposal to integrate mechanisms from audiology and cognitive science into the existing QoE framework in order to properly account for cognitive load in speech listening. The article concludes with a discussion regarding how an extended framework could facilitate measurement of QoE in broader and more realistic application scenarios where cognitive effort is a material consideration.</p></abstract>
<kwd-group>
<kwd>Quality of Experience (QoE)</kwd>
<kwd>cognitive load</kwd>
<kwd>listening effort</kwd>
<kwd>subjective test</kwd>
<kwd>QoE framework</kwd>
</kwd-group>
<contract-num rid="cn001">17/RC/2289_P2</contract-num>
<contract-sponsor id="cn001">Science Foundation Ireland<named-content content-type="fundref-id">10.13039/501100001602</named-content></contract-sponsor>
<counts>
<fig-count count="1"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="57"/>
<page-count count="7"/>
<word-count count="5483"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Quality of experience (QoE) is a paradigm that assesses media quality by mimicking human judgement. The goal is to understand and quantify how consumers perceive media quality. Instead of using the measurable signal parameters, QoE researchers evaluate the quality of a multimedia event based on reported quality ratings from participants in subjective experimental studies. To void the biases from the interpersonal differences, a mean opinion score (MOS) is used to represent an averaged perceived quality. The subjective ratings from experiments are also used to develop signal-based QoE prediction models (also called objective models). Such models are expected to predict quality judgements for multimedia application. Thus, the QoE evaluation approach has been widely adopted to rapidly test the perceptual effect of new products and services.</p>
<p>Despite the wide applicability of QoE evaluation methods, current QoE evaluations for naturalistic multimedia consumption scenarios, when a person is listening to podcasts while driving for example, are limited. They lack the consideration of a person&#x00027;s comprehension, engagement, effort, and other mental status. The current QoE framework, a conceptual model that characterizes how QoE forms, adopts a simple filtering structure that collapse all the interactions of different influencing factors to a single outcome&#x02014;people&#x00027;s internal comparison between their expectation of the signal properties and what they actually perceive&#x02014;which can be observed from the subjective quality judgement. Such framework has been widely adopted and works well for many scenarios. For instance, the telecommunication industry uses it to analyse the quality impact of a change in network capacity or system parameters. However, how the cognitive processes affect the multimedia QoE are not addressed by the framework nor by the evaluation methods.</p>
<p>As the multimedia consumption scenarios become more complex, the cognitive aspects of the experience need to be taken into account. QoE evaluation methods applicable to more natural scenarios are important to understand the impact of potential technological changes. Although cognitive aspects are highly personal and are hard to be modeled, the theories and the empirical studies in cognitive science can provide us with practical tools to systematically evaluate the impacts of the cognitive processes. This paper reviews the existing QoE framework as well as the cognitive listening methods and models from the audiology and cognitive psychology domains. The paper then discusses the potential ways to integrate cognitive effort into the existing QoE framework. While this paper uses listening effort as a focus, this review prompts consideration of broader and more realistic QoE framework for application scenarios where cognitive effort is a factor.</p>
</sec>
<sec id="s2">
<title>2. The Existing QoE Framework and Its Limits</title>
<sec>
<title>2.1. The QoE Framework</title>
<p>The QoE framework is a conceptual model that describes a QoE formation mechanism for any multimedia consumption scenario. It can be applied as a template to characterize a quality judgement formation for an experience. The QoE framework identifies the QoE formation pathways, the QoE observables, and the QoE influencing factors (see <xref ref-type="fig" rid="F1">Figure 1</xref>). Quality of Experience (QoE) describes a person&#x00027;s satisfactory level of a perceptual event (Brunnstr&#x000F6;m et al., <xref ref-type="bibr" rid="B5">2013</xref>). It results from the fulfillment of expectations. The satisfactory level of a perceptual experience can be reflected by people&#x00027;s quality judgement. Therefore descriptions and ratings are used as the observables to indicate the latent state of interest&#x02014;the perceived QoE.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The QoE framework adapted from the QoE whitepaper (Brunnstr&#x000F6;m et al., <xref ref-type="bibr" rid="B5">2013</xref>) where the QoE formation pathways (lines with arrows), the QoE observables (gray boxes), and the QoE influencing factors (orange boxes) are identified. The elements in the existing framework are denoted in black and the expanded parts are in blue. The existing model assumes that the QoE is the outcome of comparing the expected event and the perceived event (see the mechanistic diagrams in black). Both expectation and perception are influenced by different influencing factors. The influencing factors are grouped to four categories (orange boxes). The perceived quality is observed by the subjective rating and/or description of an event (gray box at the bottom).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-12-767840-g0001.tif"/>
</fig>
<p>Building on the QoE formation mechanism, <italic>influencing factors</italic> are classified that contribute to either the formation of one&#x00027;s expectation or the perceived event via <italic>formation pathways</italic> (the black lines with arrows in <xref ref-type="fig" rid="F1">Figure 1</xref>). For example, the context of media consumption can influence one&#x00027;s expectation (Sackl et al., <xref ref-type="bibr" rid="B46">2017</xref>), e.g., for a free vs. paid telephone call, or listening-only radio vs. conversational telephone call (Moller et al., <xref ref-type="bibr" rid="B36">2011</xref>). Other factors such as noise and network conditions also affect the perceived event. All the possible QoE influencing factors are grouped to four categories in the QoE framework: signal, context, system, and human factors (Brunnstr&#x000F6;m et al., <xref ref-type="bibr" rid="B5">2013</xref>), each has its own pathway that ultimately contributes to the formation of QoE (see the orange boxes in <xref ref-type="fig" rid="F1">Figure 1</xref>). The identified categories of the QoE influencing factors provide a structural guideline for researchers to analyse the quality impact of any factors of interest in a variety of scenarios. Together with the QoE formation pathways and the observables, researchers can design subjective experimental procedures that yield quantitative QoE measures.</p>
</sec>
<sec>
<title>2.2. QoE Evaluation in Practice</title>
<p>The two commonly used QoE evaluation approaches, the &#x0201C;descriptive&#x0201D; and the &#x0201C;integrated&#x0201D; (Katz and Nicol, <xref ref-type="bibr" rid="B28">2019</xref>) approaches, conform well with the observables in the QoE framework. The <italic>descriptive</italic> (or <italic>performance</italic>) approach uses the verbal descriptions as QoE evaluation. The focus of the experiential aspects will shift across different application scenarios using this approach. For example, descriptions of the noise and intelligibility levels are useful to evaluate the QoE of a voice call; comments regarding the perceived origin of a sound or how it blends with the rest of the environment are useful in a spatial sound scenario. The <italic>integrated</italic> approach, to the contrary, uses a single numerical value to represent the impression of an overall QoE. For instance, the basic audio quality (BAQ) test (ITU-R, <xref ref-type="bibr" rid="B24">2015a</xref>,<xref ref-type="bibr" rid="B23">b</xref>; Sch&#x000F6;ffler, <xref ref-type="bibr" rid="B48">2017</xref>) uses the mean opinion scores (MOS) for QoE. Using a uni-dimensional representation for QoE makes the comparison of different experiences easier, and hence, making it an efficient solution for rapid evaluations in industry. While acknowledging that experience is a high dimensional concept, the QoE framework provides guidelines to evaluate QoE that is repeatable experimentally and useful for media technology development and evaluation.</p>
</sec>
<sec>
<title>2.3. The Overlooked Impact of Cognitive Processes</title>
<p>The cognitive processes are modeled in the QoE framework through the pathways connecting the human influencing factors (orange box in bottom left of <xref ref-type="fig" rid="F1">Figure 1</xref>). The human influencing factors comprise factors such as mood, motivation, language, or prior experience (Brunnstr&#x000F6;m et al., <xref ref-type="bibr" rid="B5">2013</xref>). The human influencing factors only contribute to expectation formation, not the downstream QoE formation as human influencing factors are considered to be either temporarily volatile (such as mood and motivation) or personal (such as language proficiency or prior experience). In order to model a QoE evaluation that is representative and relevant for a large population, the effect of the transient factors needs to be dampened in the model. To realize this, QoE evaluation protocols (ITU-T, <xref ref-type="bibr" rid="B25">1996</xref>) recommend implementing a variety of mechanisms to minimize the effect of the human influencing factors such as accent familiarity, voice preferences, fatigue, or boredom. Studies in both audiology and cognitive neuroscience (Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>; Peelle, <xref ref-type="bibr" rid="B39">2018</xref>; Herrmann and Johnsrude, <xref ref-type="bibr" rid="B18">2020</xref>) show that the effort expended on our cognitive process has a substantial impact on perceived experience. Increased listening effort is found to reduce the ability to memorize (Murphy et al., <xref ref-type="bibr" rid="B37">2000</xref>; Rabbitt, <xref ref-type="bibr" rid="B43">2007</xref>; Heinrich et al., <xref ref-type="bibr" rid="B17">2008</xref>; Heinrich and Schneider, <xref ref-type="bibr" rid="B16">2011</xref>), and thereafter comprehension can be adversely affected (Piquado et al., <xref ref-type="bibr" rid="B42">2012</xref>; Ward et al., <xref ref-type="bibr" rid="B54">2016</xref>) due to less context information available from the memory to help decode the current information. A sustained high listening effort is found to lead to lower arousal levels (Aston-Jones and Cohen, <xref ref-type="bibr" rid="B2">2005</xref>) and reduced affective responses (Francis and Love, <xref ref-type="bibr" rid="B13">2020</xref>) such as fatigue (Hockey, <xref ref-type="bibr" rid="B19">2011</xref>) and boredom (Elpidorou, <xref ref-type="bibr" rid="B10">2018</xref>). The strenuous cognitive process is also found to have negative impact on behaviors such as slower response time (Phillips, <xref ref-type="bibr" rid="B40">2016</xref>), inferior task performance (Wingfield et al., <xref ref-type="bibr" rid="B56">2006</xref>; Hornsby, <xref ref-type="bibr" rid="B20">2013</xref>; Lemke and Besser, <xref ref-type="bibr" rid="B31">2016</xref>; Phillips, <xref ref-type="bibr" rid="B40">2016</xref>), or withdrawal from listening task (Lemke and Besser, <xref ref-type="bibr" rid="B31">2016</xref>; Herrmann and Johnsrude, <xref ref-type="bibr" rid="B18">2020</xref>) and social interactions (Mick et al., <xref ref-type="bibr" rid="B33">2014</xref>; Shukla et al., <xref ref-type="bibr" rid="B49">2020</xref>). Several neurological evidences [such as EEG (Hunter and Pisoni, <xref ref-type="bibr" rid="B22">2018</xref>), fMRI (Kuchinsky et al., <xref ref-type="bibr" rid="B30">2013</xref>), and pupil dilation (Aston-Jones and Cohen, <xref ref-type="bibr" rid="B2">2005</xref>; Adank, <xref ref-type="bibr" rid="B1">2012</xref>)] have showed distinct patterns when listeners are exposed to challenging auditory material, indicating the recruitment of different cognitive resources in astute listening scenarios. These findings indicate that the adverse effect of heavy auditory cognition is not only relevant to the population who are diagnosed with hearing impairment, but also relevant to anyone who needs to engage with listening in their day-to-day activities as the recruitment of other cognitive resources can directly affect the allocation of attention and therefore the task performance.</p>
<p>From a multimodal perspective, the existing pathways in the QoE framework are not exhaustive in modeling the effect of different source signals. The combined effect of audio and visual input signals have been shown to produce shifts in attention in various studies (Talsma et al., <xref ref-type="bibr" rid="B52">2006</xref>; Rapela et al., <xref ref-type="bibr" rid="B44">2012</xref>; Chao et al., <xref ref-type="bibr" rid="B6">2020</xref>). Although the multimodal integration is still an active area of study in neuroscience (Koelewijn et al., <xref ref-type="bibr" rid="B29">2010</xref>; Fu et al., <xref ref-type="bibr" rid="B14">2020</xref>), the consideration of audio-visual interaction is shown to be useful for attention and saliency modeling to improve existing QoE prediction (Min et al., <xref ref-type="bibr" rid="B34">2015</xref>, <xref ref-type="bibr" rid="B35">2020</xref>; Zhu et al., <xref ref-type="bibr" rid="B57">2020</xref>).</p>
<p>Attentional saliency, comprehension, fatigue level, task performance, and emotional status are important building blocks for understanding QoE in realistic listening scenarios, and these aspects cannot be captured and fully understood by the quality judgement alone via the standard QoE observable adopted by the community. The existing QoE framework lacks an explicit systematic model to guide effective studies exploring the impact of the cognitive processes on QoE. The attentional control can be influenced by the source signals (e.g., multimodal interaction) as well as by the human influencing factor (e.g., mental capacity). This study will focus on the latter and use the uni-modal input signal as an example to show how studies from cognitive hearing and perception theory could provide complementary learning to supplement the existing QoE framework.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Integrating Listening Effort Into Existing QoE Framework</title>
<p>To integrate listening effort into the QoE framework model, we consider three questions: (i) what contributes to the increase in the cognitive effort; (ii) how increased effort affects QoE; (iii) how to quantify the effect of effort on QoE. These questions correspond to the three core component in the QoE framework: influencing factors, QoE pathways, and the observables.</p>
<p>This section addresses each question and discuss how each component in the existing QoE framework can be adapted with reference to two cognitive hearing models: the Framework for understanding Effortful Listening (FUEL) (Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>) and the Model of Listening Engagement (MoLE) (Herrmann and Johnsrude, <xref ref-type="bibr" rid="B18">2020</xref>). They also draw on the more general cognitive load models (the load theory Murphy et al., <xref ref-type="bibr" rid="B38">2016</xref> and the mental capacity model Kahneman, <xref ref-type="bibr" rid="B27">1973</xref>).</p>
<sec>
<title>3.1. Influencing Factors</title>
<p>Listening effort increases along with the listening demand (McGarrigle et al., <xref ref-type="bibr" rid="B32">2014</xref>) as more attentional resources need to be allocated to meet the demand. The FUEL (Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>) model categorizes the sources of listening effort as source, transmission, listener, message, and context factors. These categories all have their counterparts in the QoE framework. <xref ref-type="table" rid="T1">Table 1</xref> illustrates how different sources of listening effort can be mapped to different influencing factor categories in the FUEL and the QoE framework. The middle column highlights that all four QoE influencing factor categories contribute to the effort formation. The overlapping factors of concern in both frameworks indicate that the existing QoE framework has already incorporated the main factors that lead to listening effort. The next step is to analyse whether the cognitive effect of these influencing factors can be modeled by the QoE formation pathways.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Sources of listening effort and their corresponding influencing factor categories in the QoE framework and the FUEL.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Factors</bold></th>
<th valign="top" align="left"><bold>QoE</bold></th>
<th valign="top" align="left"><bold>FUEL</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Voice degradation</td>
<td valign="top" align="left">System</td>
<td valign="top" align="left">Transmission</td>
</tr>
<tr>
<td valign="top" align="left">Bandwidth limit</td>
<td valign="top" align="left">System</td>
<td valign="top" align="left">Transmission</td>
</tr>
<tr>
<td valign="top" align="left">Noise</td>
<td valign="top" align="left">System</td>
<td valign="top" align="left">Transmission</td>
</tr>
<tr>
<td valign="top" align="left">Reverberation</td>
<td valign="top" align="left">System</td>
<td valign="top" align="left">Transmission</td>
</tr>
<tr>
<td valign="top" align="left">Multi-talker</td>
<td valign="top" align="left">Signal</td>
<td valign="top" align="left">Source &#x00026; context</td>
</tr>
<tr>
<td valign="top" align="left">Spatial separation</td>
<td valign="top" align="left">Signal</td>
<td valign="top" align="left">Source &#x00026; context</td>
</tr>
<tr>
<td valign="top" align="left">Synthesized voice</td>
<td valign="top" align="left">Signal</td>
<td valign="top" align="left">Source</td>
</tr>
<tr>
<td valign="top" align="left">Sustained speech</td>
<td valign="top" align="left">Context</td>
<td valign="top" align="left">Source</td>
</tr>
<tr>
<td valign="top" align="left">Voice similarity</td>
<td valign="top" align="left">Signal</td>
<td valign="top" align="left">Source</td>
</tr>
<tr>
<td valign="top" align="left">Foreign language</td>
<td valign="top" align="left">Signal &#x00026; context</td>
<td valign="top" align="left">Message &#x00026; context</td>
</tr>
<tr>
<td valign="top" align="left">Reward</td>
<td valign="top" align="left">Human</td>
<td valign="top" align="left">Motivation</td>
</tr>
<tr>
<td valign="top" align="left">Hearing loss</td>
<td valign="top" align="left">Human</td>
<td valign="top" align="left">Listener</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.2. Pathways</title>
<p>The formation pathways in a model identify the possible mechanisms through which the influencing factors can follow to impact an outcome. Although the formation pathways are not concrete, they are depicted in the models to guide research protocol designs wishing to evaluate the effect of factors of interest. The implications of increased listening effort are the result of complex combinations of interactions. The existing QoE formation pathways collapse the contributions of influencing factors to an internal comparison, which limits the capacity to capture the wider cognitive effects that make up our listening experience. Cognitive hearing studies (McGarrigle et al., <xref ref-type="bibr" rid="B32">2014</xref>; Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>; Herrmann and Johnsrude, <xref ref-type="bibr" rid="B18">2020</xref>) indicate that multiple effort formation pathways exist during speech listening. When a speech signal is being processed at an early stage, with presence of noise for instance, effort arises when listeners inhibit the irrelevant signals and keep attentive to the target signals. However, sometimes a higher load level helps people to concentrate (Mick et al., <xref ref-type="bibr" rid="B33">2014</xref>; Murphy et al., <xref ref-type="bibr" rid="B38">2016</xref>; Herrmann and Johnsrude, <xref ref-type="bibr" rid="B18">2020</xref>). At a later stage when the speech signal is being processed semantically, effort increases when the content topic is obscure and more context information needs to be recalled from memory to aid comprehension. Effort is also be influenced by the demands of concurrent tasks (Skowronek and Raake, <xref ref-type="bibr" rid="B50">2014</xref>) as attention needs to be constantly reallocated depending on the dynamics of a subtask. This pathway is particularly relevant to the design of technology and multimedia applications where people increasingly consume multimedia while multi-tasking in day-to-day scenarios.</p>
<p>It has yet to be shown whether the effect of multiple effort formation pathways can be simplified to a single pathway. Therefore, we show multiple potential effort formation pathways so that systematic investigations into the cognitive impact can be designed. Multiple pathways might result in different experiential implications in addition to the quality judgement, thus additional measurements that capture different aspects of an experience need to be recorded to compare the differences in the perceptual experiences.</p>
</sec>
<sec>
<title>3.3. Observables</title>
<p>The observables are used by researchers to infer the impact of influencing factors. The choice of the observables depends on the outcome of interest and the corresponding formation pathways. For instance, the corresponding observables for the percept (Johnsrude and Rodd, <xref ref-type="bibr" rid="B26">2016</xref>), cognitive activity, and the mental capacity as a result of listening effort can be the self-reported responses, neuroimaging, and concurrent task performance. As multiple listening effort formation pathways might exist, a single observable (i.e., a quality judgement) may not be sufficient to capture the QoE. Initiatives in the QoE domain (Engelke et al., <xref ref-type="bibr" rid="B11">2017</xref>) already attempt to use other observables to give a broader definition of QoE. We will next summarize the various listening effort observables in use and discuss how different types of observable account for different aspects of an experience.</p>
<p>The most direct observables for listening effort are the self-reported ratings or descriptions. Ratings are more commonly adopted as they are both scalable and easier to process. The NASA-TLX mental effort scale (Hart and Staveland, <xref ref-type="bibr" rid="B15">1988</xref>), for example, is a mature instrument that asks subjects to rate on different relevant aspects such as fatigue, stress, and task difficulty to gauge one&#x00027;s overall cognitive load (Rubio et al., <xref ref-type="bibr" rid="B45">2004</xref>). Another example of a self-reported measure asks subjects to estimate the duration they can sustain a task to gauge the cognitive load while listening (Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>). However, due to the retrospective nature of these self-reported measures, such measures are susceptible to memory and descriptive biases.</p>
<p>Behavioral responses are also used to indicate effort. These include the memory recall, speech comprehension (observed after the task), or attention-related task performance (observed during the task). The Span Test (Conway et al., <xref ref-type="bibr" rid="B7">2005</xref>) is a well established working memory test where participants are asked to read a series of sentences and to recall the last word from each sentence. It is used to indirectly evaluate listening effort based on the assumption of working memory capacity (Baddeley, <xref ref-type="bibr" rid="B3">2000</xref>). In a demanding listening scenario, an increase in the allocated cognitive resources to comprehend the signal will adversely impact information recall capacity. Another popular experimental paradigm is the dual-task method where participants conduct a parallel task simultaneously to force the division of attention. In this case, an increase in the listening effort is indicated by a performance reduction in the concurrent task (Hunter, <xref ref-type="bibr" rid="B21">2020</xref>). The dual-task paradigm is based on the assumption that attention allocated to one task will leave less spare cognitive capacity to process another task (Kahneman, <xref ref-type="bibr" rid="B27">1973</xref>; Beatty, <xref ref-type="bibr" rid="B4">1977</xref>; Sweller, <xref ref-type="bibr" rid="B51">1994</xref>; Schnotz and K&#x000FC;rschner, <xref ref-type="bibr" rid="B47">2007</xref>) leading to an observable reduced performances in the less attended task.</p>
<p>Psychophysiological changes are also used to indicate the effort involved in a listening task. Some physiological observables (e.g., pupil dilation, cardiac responses, skin conductance, and hormonal changes) are the result of sympathetic or parasympathetic responses to stress or effort (de Waard, <xref ref-type="bibr" rid="B8">1996</xref>; Peelle, <xref ref-type="bibr" rid="B39">2018</xref>). Thus, they are regarded as indirect measures for listening effort. Observables captured around the brain area (such as the activity intensity and the differences in the activated brain regions) are also used as indicators of listening effort. For example, an increase in the alpha band power in the electroencephalography signal can be observed when there is signal degradation or an increased demand for information storage (Piquado et al., <xref ref-type="bibr" rid="B42">2012</xref>; Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>; Hunter, <xref ref-type="bibr" rid="B21">2020</xref>). An increase in activity is found in the cingulo-opercular network from the functional magnetic resonance imaging when listeners are exposed to less intelligible signals (Wild et al., <xref ref-type="bibr" rid="B55">2012</xref>; Erb et al., <xref ref-type="bibr" rid="B12">2013</xref>; Vaden et al., <xref ref-type="bibr" rid="B53">2013</xref>; Eckert et al., <xref ref-type="bibr" rid="B9">2016</xref>). The psychophysiological observables are highly susceptible to many other internal and external factors such as environment temperature and mental status. Yet the high resolution in time makes them the preferred instruments for event-related analysis.</p>
<p>Identifying the potential and appropriate observables is critical in order to select the methods that will capture how effort affects different aspects of our experience. Using multiple observables is also recommended to reduce the structural interference in data analysis (Kahneman, <xref ref-type="bibr" rid="B27">1973</xref>; Pichora-Fuller et al., <xref ref-type="bibr" rid="B41">2016</xref>). The theoretical and empirical cognitive psychology literature provides a broad selection of observables to complement the commonly-used self-reported measures in the QoE community. It also prompts looking beyond the existing QoE framework to consider pathways to better capture different impacts of listening effort in naturalistic scenarios.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Conclusion and Future Direction</title>
<p>This review introduced the QoE framework model used by the media technology community to assign in designing and selecting the appropriate methods to empirically evaluate quality of experience. We introduced the rationale behind the framework and explained the structural influencing factors, pathways and observables. The limited capability within the framework to capture and quantify how effort interacts with QoE was highlighted. With a focus on listening effort, this paper reviewed multiple listening effort formation pathways from the cognitive science domain to complement the existing QoE formation pathway. A review of literature and methods drawn from the audiology and cognitive science domains, illustrated how the QoE framework could be expanded and QoE experimental methods could be applied to naturalistic listening scenarios where the cognitive process plays a significant part in QoE formation. Pathways and observables beyond self-reported quality ratings were reviewed. We believe the review warrants adding a cognitive dimension to QoE framework. It would allow for more direct comparisons of different subjective experiments. It would encourage the community to design subjective experiments that consider the impact of less explored cognitive processes. Furthermore, subjective experiments guided by such framework should provide new insights into the more nuanced experiential aspects of our multimedia consumption experience.</p>
<p>More generally, the review highlights the flexibility within the framework for extension and the potential to capture a better understanding of audio influence within wider QoE studies, e.g., listening effort impacting video or immersive QoE. This review also presents an opportunity to apply a similar approach beyond listening, identifying new pathways and observables within the QoE framework, for visual, haptic or multimodal interactions.</p>
</sec>
<sec id="s5">
<title>Author Contributions</title>
<p>PS and AH both contributed to writing, development, and editing. Both authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s6">
<title>Funding</title>
<p>This publication has emanated from research conducted with the financial support of Science Foundation Ireland (SFI) under Grant Number 12/RC/2289_P2.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s7">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> 
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adank</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <article-title>The neural bases of difficult speech comprehension and speech production: two activation likelihood estimation (ALE) meta-analyses</article-title>. <source>Brain Lang</source>. <volume>122</volume>, <fpage>42</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1016/j.bandl.2012.04.014</pub-id><pub-id pub-id-type="pmid">22633697</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aston-Jones</surname> <given-names>G.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>2005</year>). <article-title>An integrative theory of locus coeruleus-norepinephrine function: adaptive gain and optimal performance</article-title>. <source>Annu. Rev. Neurosci</source>. <volume>28</volume>, <fpage>403</fpage>&#x02013;<lpage>450</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.neuro.28.061604.135709</pub-id><pub-id pub-id-type="pmid">16022602</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baddeley</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>The episodic buffer: a new component of working memory?</article-title> <source>Trends Cogn. Sci</source>. <volume>4</volume>, <fpage>417</fpage>&#x02013;<lpage>423</lpage>. <pub-id pub-id-type="doi">10.1016/S1364-6613(00)01538-2</pub-id><pub-id pub-id-type="pmid">11058819</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Beatty</surname> <given-names>J.</given-names></name></person-group> (<year>1977</year>). <source>Activation and Attention</source>. <publisher-loc>Los Angeles, CA</publisher-loc>: <publisher-name>California Univ Los Angeles Dept of Psychology</publisher-name>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunnstr&#x000F6;m</surname> <given-names>K.</given-names></name> <name><surname>Beker</surname> <given-names>S. A.</given-names></name> <name><surname>de Moor</surname> <given-names>K.</given-names></name> <name><surname>Dooms</surname> <given-names>A.</given-names></name> <name><surname>Egger</surname> <given-names>S.</given-names></name> <name><surname>Garcia</surname> <given-names>M.-N.</given-names></name> <etal/></person-group>. (<year>2013</year>). <source>Qualinet White Paper on Definitions of Quality of Experience</source>. Technical report, Novi Sad.</citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chao</surname> <given-names>F. Y.</given-names></name> <name><surname>Ozcinar</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Zerman</surname> <given-names>E.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Hamidouche</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Audio-visual perception of omnidirectional video for virtual reality applications</article-title>, in <source>2020 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2020</source> (<publisher-loc>London</publisher-loc>: <publisher-name>IEEE</publisher-name>).</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conway</surname> <given-names>A. R.</given-names></name> <name><surname>Kane</surname> <given-names>M. J.</given-names></name> <name><surname>Bunting</surname> <given-names>M. F.</given-names></name> <name><surname>Hambrick</surname> <given-names>D. Z.</given-names></name> <name><surname>Wilhelm</surname> <given-names>O.</given-names></name> <name><surname>Engle</surname> <given-names>R. W.</given-names></name></person-group> (<year>2005</year>). <article-title>Working memory span tasks: a methodological review and user&#x00027;s guide</article-title>. <source>Psychonomic Bull. Rev</source>. <volume>12</volume>, <fpage>769</fpage>&#x02013;<lpage>786</lpage>. <pub-id pub-id-type="doi">10.3758/BF03196772</pub-id><pub-id pub-id-type="pmid">16523997</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>de Waard</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <source>The Measurement of Drivers&#x00027;Mental Workload</source> (Ph.D. thesis). <publisher-loc>University of Groningen</publisher-loc>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eckert</surname> <given-names>M. A.</given-names></name> <name><surname>Teubner-Rhodes</surname> <given-names>S.</given-names></name> <name><surname>Vaden</surname> <given-names>K. I.</given-names> <suffix>Jr</suffix></name></person-group>. (<year>2016</year>). <article-title>Is listening in noise worth it? The neurobiology of speech recognition in challenging listening conditions</article-title>. <source>Ear Hear</source>. <volume>37</volume>(<supplement>Suppl 1</supplement>):<fpage>101S</fpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000300</pub-id><pub-id pub-id-type="pmid">27355759</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elpidorou</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>The bored mind is a guiding mind: toward a regulatory theory of boredom</article-title>. <source>Phenomenol. Cogn. Sci</source> <volume>17</volume>, <fpage>455</fpage>&#x02013;<lpage>484</lpage>. <pub-id pub-id-type="doi">10.1007/s11097-017-9515-1</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Engelke</surname> <given-names>U.</given-names></name> <name><surname>Darcy</surname> <given-names>D. P.</given-names></name> <name><surname>Mulliken</surname> <given-names>G. H.</given-names></name> <name><surname>Bosse</surname> <given-names>S.</given-names></name> <name><surname>Martini</surname> <given-names>M. G.</given-names></name> <name><surname>Arndt</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Psychophysiology-Based QoE assessment: a survey</article-title>. <source>IEEE J. Select. Top. Signal Proc</source>. <volume>11</volume>, <fpage>6</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1109/JSTSP.2016.2609843</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Erb</surname> <given-names>J.</given-names></name> <name><surname>Henry</surname> <given-names>M. J.</given-names></name> <name><surname>Eisner</surname> <given-names>F.</given-names></name> <name><surname>Obleser</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>The brain dynamics of rapid perceptual adaptation to adverse listening conditions</article-title>. <source>J. Neurosci</source>. <volume>33</volume>, <fpage>10688</fpage>&#x02013;<lpage>10697</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4596-12.2013</pub-id><pub-id pub-id-type="pmid">23804092</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Francis</surname> <given-names>A. L.</given-names></name> <name><surname>Love</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Listening effort: are we measuring cognition or affect, or both?</article-title> <source>Wiley Interdiscip. Rev. Cogn. Sci</source>. <volume>11</volume>, <fpage>e1514</fpage>. <pub-id pub-id-type="doi">10.1002/wcs.1514</pub-id><pub-id pub-id-type="pmid">31381275</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>D.</given-names></name> <name><surname>Weber</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Kerzel</surname> <given-names>M.</given-names></name> <name><surname>Nan</surname> <given-names>W.</given-names></name> <name><surname>Barros</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>What can computational models learn from human selective attention? a review from an audiovisual unimodal and crossmodal perspective</article-title>. <source>Front. Integr. Neurosci</source>. <volume>14</volume>:<fpage>10</fpage>. <pub-id pub-id-type="doi">10.3389/fnint.2020.00010</pub-id><pub-id pub-id-type="pmid">32174816</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hart</surname> <given-names>S. G.</given-names></name> <name><surname>Staveland</surname> <given-names>L. E.</given-names></name></person-group> (<year>1988</year>). <article-title>Development of nasa-tlx (task load index): results of empirical and theoretical research</article-title>. <source>Adv. Psychol</source>. <volume>52</volume>, <fpage>139</fpage>&#x02013;<lpage>183</lpage>. <pub-id pub-id-type="doi">10.1016/S0166-4115(08)62386-9</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heinrich</surname> <given-names>A.</given-names></name> <name><surname>Schneider</surname> <given-names>B. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Elucidating the effects of ageing on remembering perceptually distorted word pairs</article-title>. <source>Q. J. Exp. Psychol</source>. <volume>64</volume>, <fpage>186</fpage>&#x02013;<lpage>205</lpage>. <pub-id pub-id-type="doi">10.1080/17470218.2010.492621</pub-id><pub-id pub-id-type="pmid">20694922</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heinrich</surname> <given-names>A.</given-names></name> <name><surname>Schneider</surname> <given-names>B. A.</given-names></name> <name><surname>Craik</surname> <given-names>F. I.</given-names></name></person-group> (<year>2008</year>). <article-title>Investigating the influence of continuous babble on auditory short-term memory performance</article-title>. <source>Q. J.Exp. Psychol</source>. <volume>61</volume>, <fpage>735</fpage>&#x02013;<lpage>751</lpage>. <pub-id pub-id-type="doi">10.1080/17470210701402372</pub-id><pub-id pub-id-type="pmid">17853231</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herrmann</surname> <given-names>B.</given-names></name> <name><surname>Johnsrude</surname> <given-names>I. S.</given-names></name></person-group> (<year>2020</year>). <article-title>A model of listening engagement (MoLE)</article-title>. <source>Hear Res</source>. <volume>397</volume>:<fpage>108016</fpage>. <pub-id pub-id-type="doi">10.1016/j.heares.2020.108016</pub-id><pub-id pub-id-type="pmid">32680706</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hockey</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <source>The Psychology of Fatigue: Work, Effort and Control</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <fpage>1</fpage>&#x02013;<lpage>272</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hornsby</surname> <given-names>B. W. Y.</given-names></name></person-group> (<year>2013</year>). <article-title>The effects of hearing aid use on listening effort and mental fatigue associated with sustained speech processing demands</article-title>. <source>Ear Hear</source>. <volume>34</volume>, <fpage>523</fpage>&#x02013;<lpage>534</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0b013e31828003d8</pub-id><pub-id pub-id-type="pmid">23426091</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hunter</surname> <given-names>C. R.</given-names></name></person-group> (<year>2020</year>). <article-title>Tracking cognitive spare capacity during speech perception with EEG/ERP: effects of cognitive load and sentence predictability</article-title>. <source>Ear Hear</source>. <volume>41</volume>, <fpage>1144</fpage>&#x02013;<lpage>1157</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000856</pub-id><pub-id pub-id-type="pmid">32282402</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hunter</surname> <given-names>C. R.</given-names></name> <name><surname>Pisoni</surname> <given-names>D. B.</given-names></name></person-group> (<year>2018</year>). <article-title>Extrinsic cognitive load impairs spoken word recognition in high-and low-predictability sentences</article-title>. <source>Ear Hear</source>. <volume>39</volume>, <fpage>378</fpage>&#x02013;<lpage>389</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000493</pub-id><pub-id pub-id-type="pmid">28945658</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><collab>ITU-</collab></person-group> (<year>2015b</year>). <source>BS.1534 Method for the subjective assessment of intermediate quality level of audio systems</source>. Technical report.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><collab>ITU-R</collab></person-group>. (<year>2015a</year>). <source>BS.1116-3 Methods for the subjective assessment of small impairments in audio systems</source>. Technical report, ITU.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><collab>ITU-T</collab></person-group>. (<year>1996</year>). <source>P.800: Methods for subjective determination of transmission quality</source>. Technical report, <publisher-name>Int. Telecomm. Union</publisher-name>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Johnsrude</surname> <given-names>I. S.</given-names></name> <name><surname>Rodd</surname> <given-names>J. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Chapter 40. Factors that increase processing demands when listening to speech</article-title>, in <source>Neurobiology of Language</source>, eds <person-group person-group-type="editor"><name><surname>Hickok</surname> <given-names>G.</given-names></name> <name><surname>Small</surname> <given-names>S. L.</given-names></name></person-group> (<publisher-name>Academic Press</publisher-name>), <fpage>491</fpage>&#x02013;<lpage>502</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-407794-2.00040-7</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kahneman</surname> <given-names>D.</given-names></name></person-group> (<year>1973</year>). <source>Attention and Effort</source>, Vol. <volume>1063</volume>. <publisher-loc>Englewood Cliffs, NJ</publisher-loc>: <publisher-name>Prentice-Hall Inc</publisher-name>.</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Katz</surname> <given-names>B. F.</given-names></name> <name><surname>Nicol</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>Binaural spatial reproduction</article-title>, in <source>Sensory Evaluation of Sound, Chapter 11</source>, ed N. Zacharov (<publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press Taylor &#x00026; Francis Group</publisher-name>), <fpage>349</fpage>&#x02013;<lpage>388</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koelewijn</surname> <given-names>T.</given-names></name> <name><surname>Bronkhorst</surname> <given-names>A.</given-names></name> <name><surname>Theeuwes</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Attention and the multiple stages of multisensory integration: a review of audiovisual studies</article-title>. <source>Acta Psychol</source>. <volume>134</volume>, <fpage>372</fpage>&#x02013;<lpage>384</lpage>. <pub-id pub-id-type="doi">10.1016/j.actpsy.2010.03.010</pub-id><pub-id pub-id-type="pmid">20427031</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuchinsky</surname> <given-names>S. E.</given-names></name> <name><surname>Ahlstrom</surname> <given-names>J. B.</given-names></name> <name><surname>Vaden</surname> <given-names>K. I.</given-names></name> <name><surname>Cute</surname> <given-names>S. L.</given-names></name> <name><surname>Humes</surname> <given-names>L. E.</given-names></name> <name><surname>Dubno</surname> <given-names>J. R.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Pupil size varies with word listening and response selection difficulty in older adults with hearing loss</article-title>. <source>Psychophysiology</source> <volume>50</volume>, <fpage>23</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1111/j.1469-8986.2012.01477.x</pub-id><pub-id pub-id-type="pmid">23157603</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lemke</surname> <given-names>U.</given-names></name> <name><surname>Besser</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Cognitive load and listening effort: concepts and age-related considerations</article-title>. <source>Ear Hear</source>. <volume>37</volume>, <fpage>77S</fpage>&#x02013;<lpage>84S</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000304</pub-id><pub-id pub-id-type="pmid">27355774</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGarrigle</surname> <given-names>R.</given-names></name> <name><surname>Munro</surname> <given-names>K. J.</given-names></name> <name><surname>Dawes</surname> <given-names>P.</given-names></name> <name><surname>Stewart</surname> <given-names>A. J.</given-names></name> <name><surname>Moore</surname> <given-names>D. R.</given-names></name> <name><surname>Barry</surname> <given-names>J. G.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Listening effort and fatigue: What exactly are we measuring? A british society of audiology cognition in hearing special interest group &#x00027;white paper&#x00027;</article-title>. <source>Int. J. Audiol</source>. <volume>53</volume>, <fpage>433</fpage>&#x02013;<lpage>445</lpage>. <pub-id pub-id-type="doi">10.3109/14992027.2014.890296</pub-id><pub-id pub-id-type="pmid">24673660</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mick</surname> <given-names>P.</given-names></name> <name><surname>Kawachi</surname> <given-names>I.</given-names></name> <name><surname>Lin</surname> <given-names>F. R.</given-names></name></person-group> (<year>2014</year>). <article-title>The association between hearing loss and social isolation in older adults</article-title>. <source>Otolaryngol. Head Neck Surg</source>. <volume>150</volume>, <fpage>378</fpage>&#x02013;<lpage>384</lpage>. <pub-id pub-id-type="doi">10.1177/0194599813518021</pub-id><pub-id pub-id-type="pmid">24384545</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Min</surname> <given-names>X.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>C.</given-names></name> <name><surname>Gu</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>Fixation prediction through multimodal analysis</article-title>, in <source>2015 Visual Communications and Image Processing (VCIP)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Min</surname> <given-names>X.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Member</surname> <given-names>S.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>X.-P.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A multimodal saliency model for videos with high audio-visual correspondence</article-title>. <source>IEEE Trans. Image Proc</source>. <volume>29</volume>:<fpage>2020</fpage>. <pub-id pub-id-type="doi">10.1109/TIP.2020.2966082</pub-id><pub-id pub-id-type="pmid">31976898</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moller</surname> <given-names>S.</given-names></name> <name><surname>Chan</surname> <given-names>W.-Y.</given-names></name> <name><surname>Cote</surname> <given-names>N.</given-names></name> <name><surname>Falk</surname> <given-names>T.</given-names></name> <name><surname>Raake</surname> <given-names>A.</given-names></name> <name><surname>Waltermann</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Speech quality estimation: models and trends</article-title>. <source>IEEE Signal Proc. Mag</source>. <volume>28</volume>, <fpage>18</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1109/MSP.2011.942469</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murphy</surname> <given-names>D. R.</given-names></name> <name><surname>Craik</surname> <given-names>F. I. M.</given-names></name> <name><surname>Li</surname> <given-names>K. Z. H.</given-names></name> <name><surname>Schneider</surname> <given-names>B. A.</given-names></name></person-group> (<year>2000</year>). <article-title>Comparing the effects of aging and background noise on short-term memory performance</article-title>. <source>Psychol. Aging</source> <volume>15</volume>, <fpage>323</fpage>&#x02013;<lpage>334</lpage>. <pub-id pub-id-type="doi">10.1037/0882-7974.15.2.323</pub-id><pub-id pub-id-type="pmid">10879586</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murphy</surname> <given-names>G.</given-names></name> <name><surname>Groeger</surname> <given-names>J. A.</given-names></name> <name><surname>Greene</surname> <given-names>C. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Twenty years of load theory&#x02013;Where are we now, and where should we go next?</article-title> <source>Psychonomic Bull. Rev</source>. <volume>23</volume>, <fpage>1316</fpage>&#x02013;<lpage>1340</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-015-0982-5</pub-id><pub-id pub-id-type="pmid">26728138</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peelle</surname> <given-names>J. E.</given-names></name></person-group> (<year>2018</year>). <article-title>Listening effort: How the cognitive consequences of acoustic challenge are reflected in brain and behavior</article-title>. <source>Ear Hear</source>. <volume>39</volume>, <fpage>204</fpage>&#x02013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000494</pub-id><pub-id pub-id-type="pmid">28938250</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Phillips</surname> <given-names>N. A.</given-names></name></person-group> (<year>2016</year>). <article-title>The implications of cognitive aging for listening and the framework for understanding effortful listening (FUEL)</article-title>. <source>Ear Hear</source>. <volume>37</volume>, <fpage>44S</fpage>&#x02013;<lpage>51S</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000309</pub-id><pub-id pub-id-type="pmid">27355769</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pichora-Fuller</surname> <given-names>M. K.</given-names></name> <name><surname>Kramer</surname> <given-names>S. E.</given-names></name> <name><surname>Eckert</surname> <given-names>M. A.</given-names></name> <name><surname>Edwards</surname> <given-names>B.</given-names></name> <name><surname>Hornsby</surname> <given-names>B. W.</given-names></name> <name><surname>Humes</surname> <given-names>L. E.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Hearing impairment and cognitive energy: the framework for understanding effortful listening (FUEL)</article-title>. <source>Ear Hear</source>. <volume>37</volume>, <fpage>5S</fpage>&#x02013;<lpage>27S</lpage>. <pub-id pub-id-type="doi">10.1097/AUD.0000000000000312</pub-id><pub-id pub-id-type="pmid">27355771</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Piquado</surname> <given-names>T.</given-names></name> <name><surname>Benichov</surname> <given-names>J. I.</given-names></name> <name><surname>Brownell</surname> <given-names>H.</given-names></name> <name><surname>Wingfield</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>The hidden effect of hearing acuity on speech recall, and compensatory effects of self-paced listening</article-title>. <source>Int. J. Audiol</source>. <volume>51</volume>, <fpage>576</fpage>&#x02013;<lpage>583</lpage>. <pub-id pub-id-type="doi">10.3109/14992027.2012.684403</pub-id><pub-id pub-id-type="pmid">22731919</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rabbitt</surname> <given-names>P. M. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Channel-capacity, intelligibility and immediate memory</article-title>. <source>Q. J. Exp. Psychol</source>. <volume>20</volume>, <fpage>241</fpage>&#x02013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1080/14640746808400158</pub-id><pub-id pub-id-type="pmid">5683763</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rapela</surname> <given-names>J.</given-names></name> <name><surname>Gramann</surname> <given-names>K.</given-names></name> <name><surname>Westerfield</surname> <given-names>M.</given-names></name> <name><surname>Townsend</surname> <given-names>J.</given-names></name> <name><surname>Makeig</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Brain oscillations in switching vs. focusing audio-visual attention</article-title>. <source>Annu. Int. Conf. IEEE Eng. Med. Biol. Soc</source>. <volume>2012</volume>, <fpage>352</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1109/EMBC.2012.6345941</pub-id><pub-id pub-id-type="pmid">23365902</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rubio</surname> <given-names>S.</given-names></name> <name><surname>D&#x000ED;az</surname> <given-names>E.</given-names></name> <name><surname>Mart&#x000ED;n</surname> <given-names>J.</given-names></name> <name><surname>Puente</surname> <given-names>J. M.</given-names></name></person-group> (<year>2004</year>). <article-title>Evaluation of subjective mental workload: a comparison of SWAT, NASA-TLX, and workload profile methods</article-title>. <source>Appl. Psychol</source>. <volume>53</volume>, <fpage>61</fpage>&#x02013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.1111/j.1464-0597.2004.00161.x</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sackl</surname> <given-names>A.</given-names></name> <name><surname>Schatz</surname> <given-names>R.</given-names></name> <name><surname>Raake</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>More than I ever wanted or just good enough? User expectations and subjective quality perception in the context of networked multimedia services</article-title>. <source>Quality User Exp</source>. <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1007/s41233-016-0004-z</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schnotz</surname> <given-names>W.</given-names></name> <name><surname>K&#x000FC;rschner</surname> <given-names>C.</given-names></name></person-group> (<year>2007</year>). <article-title>A reconsideration of cognitive load theory</article-title>. <source>Educ. Psychol. Rev</source>. <volume>19</volume>, <fpage>469</fpage>&#x02013;<lpage>508</lpage>. <pub-id pub-id-type="doi">10.1007/s10648-007-9053-4</pub-id><pub-id pub-id-type="pmid">25762908</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sch&#x000F6;ffler</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <source>Overall Listening Experience - a new Approach to Subjective Evaluation of Audio</source> (Ph.D. thesis).</citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shukla</surname> <given-names>A.</given-names></name> <name><surname>Harper</surname> <given-names>M.</given-names></name> <name><surname>Pedersen</surname> <given-names>E.</given-names></name> <name><surname>Goman</surname> <given-names>A.</given-names></name> <name><surname>Suen</surname> <given-names>J. J.</given-names></name> <name><surname>Price</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Hearing loss, loneliness, and social isolation: a systematic review</article-title>. <source>Otolaryngol. Head Neck Surg</source>. <volume>162</volume>, <fpage>622</fpage>&#x02013;<lpage>633</lpage>. <pub-id pub-id-type="doi">10.1177/0194599820910377</pub-id><pub-id pub-id-type="pmid">32151193</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skowronek</surname> <given-names>J.</given-names></name> <name><surname>Raake</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Assessment of cognitive load, speech communication quality and quality of experience for spatial and non-spatial audio conferencing calls</article-title>. <source>Speech Commun</source>. <volume>66</volume>, <fpage>154</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1016/j.specom.2014.10.003</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sweller</surname> <given-names>J.</given-names></name></person-group> (<year>1994</year>). <article-title>Cognitive load theory, learning difficulty, and instructional design</article-title>. <source>Learn. Instruct</source>. <volume>4</volume>, <fpage>295</fpage>&#x02013;<lpage>312</lpage>. <pub-id pub-id-type="doi">10.1016/0959-4752(94)90003-5</pub-id><pub-id pub-id-type="pmid">25993279</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Talsma</surname> <given-names>D.</given-names></name> <name><surname>Doty</surname> <given-names>T. J.</given-names></name> <name><surname>Woldorff</surname> <given-names>M. G.</given-names></name></person-group> (<year>2006</year>). <article-title>Selective attention and audiovisual integration: is attending to both modalities a prerequisite for early integration?</article-title> <source>Cereb. Cortex</source> <volume>17</volume>, <fpage>679</fpage>&#x02013;<lpage>690</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhk016</pub-id><pub-id pub-id-type="pmid">16707740</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaden</surname> <given-names>K. I.</given-names> <suffix>Jr</suffix></name> <name><surname>Kuchinsky</surname> <given-names>S. E.</given-names></name> <name><surname>Cute</surname> <given-names>S. L.</given-names></name> <name><surname>Ahlstrom</surname> <given-names>J. B.</given-names></name> <name><surname>Dubno</surname> <given-names>J. R.</given-names></name> <name><surname>Eckert</surname> <given-names>M. A.</given-names></name></person-group> (<year>2013</year>). <article-title>The cingulo-opercular network provides word-recognition benefit</article-title>. <source>J. Neurosci</source>. <volume>33</volume>, <fpage>18979</fpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1417-13.2013</pub-id><pub-id pub-id-type="pmid">24285902</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ward</surname> <given-names>C. M.</given-names></name> <name><surname>Rogers</surname> <given-names>C. S.</given-names></name> <name><surname>Van Engen</surname> <given-names>K. J.</given-names></name> <name><surname>Peelle</surname> <given-names>J. E.</given-names></name></person-group> (<year>2016</year>). <article-title>Effects of age, acoustic challenge, and verbal working memory on recall of narrative speech</article-title>. <source>Exp. Aging Res</source>. <volume>42</volume>, <fpage>97</fpage>&#x02013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1080/0361073X.2016.1108785</pub-id><pub-id pub-id-type="pmid">26683044</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wild</surname> <given-names>C. J.</given-names></name> <name><surname>Yusuf</surname> <given-names>A.</given-names></name> <name><surname>Wilson</surname> <given-names>D. E.</given-names></name> <name><surname>Peelle</surname> <given-names>J. E.</given-names></name> <name><surname>Davis</surname> <given-names>M. H.</given-names></name> <name><surname>Johnsrude</surname> <given-names>I. S.</given-names></name></person-group> (<year>2012</year>). <article-title>Effortful listening: the processing of degraded speech depends critically on attention</article-title>. <source>J. Neurosci</source>. <volume>32</volume>, <fpage>14010</fpage>&#x02013;<lpage>14021</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1528-12.2012</pub-id><pub-id pub-id-type="pmid">23035108</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wingfield</surname> <given-names>A.</given-names></name> <name><surname>McCoy</surname> <given-names>S. L.</given-names></name> <name><surname>Peelle</surname> <given-names>J. E.</given-names></name> <name><surname>Tun</surname> <given-names>P. A.</given-names></name> <name><surname>Cox</surname> <given-names>C. L.</given-names></name></person-group> (<year>2006</year>). <article-title>Effects of adult aging and hearing loss on comprehension of rapid speech varying in syntactic complexity</article-title>. <source>J. Am. Acad. Audiol</source>. <volume>17</volume>, <fpage>487</fpage>&#x02013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.3766/jaaa.17.7.4</pub-id><pub-id pub-id-type="pmid">16927513</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Min</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>The prediction of saliency map for head and eye movements in 360 degree images</article-title>. <source>IEEE Trans. Multimedia</source> <volume>22</volume>, <fpage>2331</fpage>&#x02013;<lpage>2344</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2019.2957986</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>