<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2022.884102</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The interaction of focus and phrasing with downstep and post-low-bouncing in Mandarin Chinese</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Bei</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1811706/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>K&#x00FC;gler</surname>
<given-names>Frank</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/101763/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Genzel</surname>
<given-names>Susanne</given-names>
</name>
<xref rid="aff3" ref-type="aff"><sup>3</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Key Laboratory of Language, Cognition and Computation, School of Foreign Languages, Beijing Institute of Technology</institution>, <addr-line>Beijing</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Linguistics, Goethe University Frankfurt</institution>, <addr-line>Frankfurt</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>i2x GmbH</institution>, <addr-line>Berlin</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Fang Liu, University of Reading, United Kingdom</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Albert Lee, The Education University of Hong Kong, Hong Kong SAR, China; Jiahong Yuan, Baidu, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Frank K&#x00FC;gler, <email>kuegler@em.uni-frankfurt.de</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Auditory Cognitive Neuroscience, a section of the journal Frontiers in Psychology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>30</day>
<month>09</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>884102</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>01</day>
<month>08</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Wang, K&#x00FC;gler and Genzel.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Wang, K&#x00FC;gler and Genzel</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>L(ow) tone in Mandarin Chinese causes both downstep and post-low-bouncing. Downstep refers to the lowering of a H(igh) tone after a L tone, which is usually measured by comparing the H tones in a &#x201C;H&#x2026;HLH&#x2026;H&#x201D; sentence with a &#x201C;H&#x2026;HHH&#x2026;H&#x201D; sentence (<italic>cross-comparison</italic>), investigating whether downstep sets a new pitch register for the scaling of subsequent tones. Post-low-bouncing refers to the raising of a H tone after a focused L tone. The current study investigates how downstep and post-low-bouncing interact with focus and phrasing in Mandarin Chinese. In the experiment, we systematically manipulated (a) the tonal environment by embedding two syllables with either LH or HH tone (syllable X and Y) sentence-medially in the same carrier sentences containing only H tones; (b) boundary strength between X and Y by introducing either a syllable boundary or a phonological phrase boundary; and (c) information structure by either placing a contrastive focus in the HL/HH word (XF), syllable Y (YF), or the sentence-final word (ZF). A wide-focus condition served as the baseline. With systematic control of focus and boundary strength around the L tone, the current study shows that the downstep effect in Mandarin is quite robust, lasting for 3&#x2013;5 H tones after the L tone, but eventually levelling back again to the register reference line of a H tone. The way how focus and phrasing interact with the downstep effect is unexpected. Firstly, sentence-final focus has no anticipatory effect on shortening the downstep effect; instead, it makes the downstep effect lasts longer as compared to the wide focus condition. Secondly, the downstep effect still shows when the H tone after the L tone is on-focus (YF), in a weaker manner than the wide focus condition, and is overridden by the post-focus-compression. Thirdly, the downstep effect gets greater when the boundary after the L tone is stronger, because the L tone is longer and more likely to be creaky. We further analyzed downstep by measuring the F0 drop between the two H tones surrounding the L tone (<italic>sequential-comparison</italic>). Comparing it with F0 drop in all-H sentences (i.e., declination), it showed that the downstep effect was much greater and more robust than declination. However, creaky voice in the L tone was not the direct cause of downstep. At last, when the L tone was under focus (XF), it caused a post-low-bouncing effect, which is weakened by a phonological phrase boundary. Altogether, the results showed that although intonation is largely controlled by informative functions, the physical-articulatory controls are relatively persistent, varying within the pitch range of 2.5 semitones. Downstep and post-low-bouncing in Mandarin Chinese thus seem to be mainly due to physical-articulatory movement on varying pitch, with the gradual tonal F0 change meeting the requirement of smooth transition across syllables, and avoiding confusion in informative F0 control.</p>
</abstract>
<kwd-group>
<kwd>downstep</kwd>
<kwd>post-low-bouncing</kwd>
<kwd>phrasing</kwd>
<kwd>focus</kwd>
<kwd>intonation</kwd>
<kwd>Mandarin Chinese</kwd>
</kwd-group>
<contract-num rid="cn1">18BYY079</contract-num>
<contract-sponsor id="cn1">Social Science Foundation of China</contract-sponsor>
<counts>
<fig-count count="10"/>
<table-count count="4"/>
<equation-count count="0"/>
<ref-count count="118"/>
<page-count count="21"/>
<word-count count="15618"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>Intonation carries communicative functions, such as focus and phrasing, but much of intonation variation also comes from tonal interactions. To better understand the interaction of tone and intonation, it is important to take into account of both informative and articulatory effects (<xref ref-type="bibr" rid="ref042">Xu et al., 2012</xref>). In Mandarin, for instance, L tone causes <italic>pre-low-raising, post-low-bouncing,</italic> and <italic>downstep</italic> in the surrounding H tones. Pre-low-raising refers to the pitch raising in the H tone preceding the L tone (<xref ref-type="bibr" rid="ref37">Lee et al., 2021</xref>). Post-low-bouncing is the phenomenon that F0 of the post-low syllables suddenly goes up first, then drops back gradually, in the condition that the following syllables carry neutral tones or the L tone is under focus (<xref ref-type="bibr" rid="ref58">Shen, 1994</xref>; <xref ref-type="bibr" rid="ref8">Chen and Xu, 2006</xref>; <xref ref-type="bibr" rid="ref21">Gu and Lee, 2009</xref>; <xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>). Downstep refers to the downtrend of F0 caused by L tones, in the way that the H tones after a L tone is with lower F0 than previous H tones (<xref ref-type="bibr" rid="ref74">Xu, 1997</xref>, <xref ref-type="bibr" rid="ref75">1999</xref>; <xref ref-type="bibr" rid="ref033">Shih, 2000</xref>; <xref ref-type="bibr" rid="ref35">Laniran and Clements, 2003</xref>; <xref ref-type="bibr" rid="ref9">Connell, 2011</xref>). The first two tonal effects have been extensively studied and well explained with articulatory movement of pitch control (<xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>; <xref ref-type="bibr" rid="ref37">Lee et al., 2021</xref>).</p>
<p>In this paper, our main goal is to study <italic>downstep</italic> in Mandarin, and its interaction with focus and phrasing. To be more specific, we aim to study how on-focus raising and post-focus compression (PFC) in F0 interact with downstep, and if a phrase boundary terminates downstep. Moreover, considering the presence of global F0 declination in all-H sentences (<xref ref-type="bibr" rid="ref033">Shih, 2000</xref>; <xref ref-type="bibr" rid="ref045">Yuan and Liberman, 2014</xref>), we investigate whether downstep and declination share the same pitch lowering mechanism. To go further, a L tone in Mandarin is commonly accompanied by creaky voice (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>, <xref ref-type="bibr" rid="ref020">2018</xref>), we thus investigate whether creaky voice may cause downstep. Thirdly, we aim to study the interaction of boundary with post-L-bouncing, which happens when a L tone is focused. Our investigation tackles the question whether a phonological phrase boundary cancels post-L-bouncing or not.</p>
<p>The next section starts with a review of downstep and declination followed by a review on post-low-bouncing, and then focus and phrasing. To better understand pitch from both linguistic and articulatory perspectives, we also briefly introduce a review on laryngeal movement of varying pitch. In the end of section &#x201C;Background,&#x201D; the research questions are summarized.</p>
</sec>
<sec id="sec2">
<title>Background</title>
<sec id="sec3">
<title>Downstep and declination</title>
<p>There is a global downtrend or declination in a sentence (e.g., <xref ref-type="bibr" rid="ref013">Gussenhoven, 2004</xref>). Articulatorily, declination is arguably caused by a decrease in subglottal pressure over time (<xref ref-type="bibr" rid="ref024">Lieberman, 1967</xref>; <xref ref-type="bibr" rid="ref006">Collier, 1975</xref>; <xref ref-type="bibr" rid="ref52">Pierrehumbert 1980</xref>; <xref ref-type="bibr" rid="ref010">Gelfer et al., 1983</xref>). Beside declination, lexical tones and tonal interactions also cause downtrend in F0 contours, e.g., downstep lowers the following H tones. Downstep has long been discussed in African languages (Yoruba (Niger-Congo): <xref ref-type="bibr" rid="ref72">Ward, 1952</xref>; <italic>cf.</italic> <xref ref-type="bibr" rid="ref12">Courtenay, 1971</xref>; Luo (Nilotic): <xref ref-type="bibr" rid="ref039">Tucker and Creider, 1975</xref>; Twi (Akan): <xref ref-type="bibr" rid="ref64">Stewart, 1965</xref>; <xref ref-type="bibr" rid="ref20">Genzel and K&#x00FC;gler, 2011</xref>; <xref ref-type="bibr" rid="ref32">K&#x00FC;gler 2017</xref>; Tswana (Southern Bantu): <xref ref-type="bibr" rid="ref84">Zerbian and K&#x00FC;gler, 2015</xref>, <xref ref-type="bibr" rid="ref85">2021</xref>; among many others). In the African linguistic tradition, downstep is distinguished from downdrift (<xref ref-type="bibr" rid="ref64">Stewart, 1965</xref>; <xref ref-type="bibr" rid="ref26">Hombert, 1974</xref>; see <xref ref-type="bibr" rid="ref29">Hyman and Leben, 2017</xref> for an overview), terrace (<xref ref-type="bibr" rid="ref12">Courtenay, 1971</xref>), or automatic and non-automatic downstep (see detailed discussion in <xref ref-type="bibr" rid="ref007">Connell, 2001</xref>; <xref ref-type="bibr" rid="ref031">Rialland and Som&#x00E9;, 2011</xref>; <xref ref-type="bibr" rid="ref36">Leben, 2014</xref>). Strictly speaking, downstep refers to a new register or ceiling established for subsequent H tones after a L tone (<xref ref-type="bibr" rid="ref61">Snider, 1990</xref>; <xref ref-type="bibr" rid="ref63">Snider and van der Hulst, 1993</xref>; <xref ref-type="bibr" rid="ref10">Connell, 2017</xref>; <italic>cf.</italic> <xref ref-type="bibr" rid="ref1">Akumbu, 2019</xref>). The differentiation between downstep and downdrift, or automatic and non-automatic downstep concerns the fact that in several African languages, both an overtly realized L tone and a floating L tone functions as the trigger of the lowering process. A floating L tone triggers non-automatic downstep, whereas a phonetically realized L tone triggers automatic downstep (<xref ref-type="bibr" rid="ref29">Hyman and Leben, 2017</xref>). Phonetically, no difference is found between these two types of downstep (e.g., <xref ref-type="bibr" rid="ref20">Genzel and K&#x00FC;gler, 2011</xref>; <xref ref-type="bibr" rid="ref32">K&#x00FC;gler 2017</xref> for Akan). Since there is no floating L tone in Mandarin Chinese, we do not need to differentiate them. We here take the broad definition of downstep as the lowering of F0 after a L tone, following <xref ref-type="bibr" rid="ref59">Shih (1988)</xref>; <xref ref-type="bibr" rid="ref75">Xu (1999)</xref>; <xref ref-type="bibr" rid="ref35">Laniran and Clements (2003)</xref>, and <xref ref-type="bibr" rid="ref19">Genzel (2013)</xref> among many others, see (1).<list list-type="order">
<list-item>
<p>Downstep: In a HLH tone sequence, the second H is realized with lowered F0 compared to the first H, due to the L tone (<italic>sequential-comparison</italic>). The size of downstep can be paradigmatically calculated as the difference of F0-maximum in the H tones after a L tone and in the corresponding H tones of an all-H tone phrase (<italic>cross-comparison</italic>).</p>
</list-item>
</list></p>
<p>In some West-African languages, downstep initiates a new pitch register to which subsequent tones are scaled, phonologically termed as register tones (e.g., <xref ref-type="bibr" rid="ref62">Snider, 1998</xref>) or register features (e.g., <xref ref-type="bibr" rid="ref1">Akumbu, 2019</xref>). In Mandarin, there is no study directly concerning the effect of downstep as setting up a new register tone or register line. We here introduce three studies, which suggest that downstep in Mandarin does not seem to set up a new register tone. First, it showed that several H tones after a L are lowered in F0 as compared to all H-tone sentences, then the pitch gradually reaches the target in the all-H sentence toward the end of the sentence (see Figure 4, pp. 66 in <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>). Second, <xref ref-type="bibr" rid="ref21">Gu and Lee (2009)</xref> found that the lowering effect in the H tones is greater when the preceding L tone is lower. Third, <xref ref-type="bibr" rid="ref70">Wang and Xu (2011)</xref> used sentence with HLHL&#x2026;HL and LHLH&#x2026;LH tone sequences, the sentence-medial H tones reach roughly the same height as the corresponding H tones in an all-H sentence, which is explained as the balance between pre-Low raising and downstep. Thus, downstep seems to be a tonal feature with gradual change in pitch. A terracing pattern of H tones in the LH sequence&#x2014;as found in the West-African pattern&#x2014;does not seem to exist in Mandarin Chinese.</p>
<p>What lacks in previous studies is that how downstep interacts with other informative functions, e.g., prosodic boundary and focus. The first question relates to the domain of downstep. The domain of downstp appears to vary across languages. In Kishamba, morpheme boundaries act as a trigger of downstep (<xref ref-type="bibr" rid="ref49">Odden, 1986</xref>). In Tswana (Southern Bantu), downstep occurs between prosodic words within a phonological phrase, whereas phonological phrase boundaries block downstep (<xref ref-type="bibr" rid="ref84">Zerbian and K&#x00FC;gler, 2015</xref>, <xref ref-type="bibr" rid="ref85">2021</xref>). In Yoruba, downstep applies across all boundaries within a breath group, which could roughly be interpreted as an intonation phrase (<xref ref-type="bibr" rid="ref12">Courtenay, 1971</xref>). In Japanese, only an accented word (H&#x002A;L) within a Major Phrase (MaP) triggers downstep (<xref ref-type="bibr" rid="ref53">Pierrehumbert and Beckman, 1988</xref>; <xref ref-type="bibr" rid="ref57">Selkirk and Tateishi, 1991</xref>). The downstep effect in Mandarin as reported in <xref ref-type="bibr" rid="ref75">Xu (1999)</xref> showed that a phrase boundary does not seem to block downstep, though no systematic data on this issue was provided.</p>
<p>As for the interaction of downstep and focus, we here introduce two studies. <xref ref-type="bibr" rid="ref30">Ishihara (2007)</xref> studied downstep systematically with sentences in the structure as N1&#x2009;+&#x2009;N2&#x2009;+&#x2009;N3&#x2009;+&#x2009;VP (N and VP are abbreviations of noun and verb phrase respectively). It showed that downstep between N2 and N3 is only partially reset, when N3 is focused and when the syntactic boundary is stronger between N2 and N3. It indicated that downstep is weakened by a strong phrase boundary, and a focused H tone after the L tone. <xref ref-type="bibr" rid="ref75">Xu (1999)</xref> has shown similar results in Mandarin that the size of downstep seems to be reduced when the H tone after the L tone is focused.</p>
<p>It has been also found that downstep can be canceled in yes/no questions in Hausa (<xref ref-type="bibr" rid="ref39">Lindau, 1986</xref>), meaning that final F0 raising may counter-balance the downstep effect. In Mandarin, however, it does not seem to be the case as shown in <xref ref-type="bibr" rid="ref75">Xu (1999)</xref>. It requires more systematic analysis on whether sentence final F0 raising interferes with the downstep effect.</p>
<p>As mentioned above, another term easy to be confused with downstep is declination, which refers to the F0 downtrend from the beginning through the end of an utterance. We can see that the crucial difference between declination and downstep is its scope. While declination is a gradual lowering of F0 within an intonation phrase, downstep is a local lowering of F0. Declination has been found in both non-tonal languages (<xref ref-type="bibr" rid="ref001">&#x2018;t Hart &#x0026; Cohen, 1973</xref>; <xref ref-type="bibr" rid="ref027">Maeda, 1976</xref>; <xref ref-type="bibr" rid="ref008">Cooper and Sorensen, 1977</xref>; <xref ref-type="bibr" rid="ref030">Pierrehumbert, 1979</xref>; <xref ref-type="bibr" rid="ref036">Sorensen and Cooper, 1980</xref>; <xref ref-type="bibr" rid="ref005">Cohen et al., 1982</xref>; <xref ref-type="bibr" rid="ref040">Umeda, 1982</xref>; <xref ref-type="bibr" rid="ref34">Ladd 1988</xref>) and tonal languages (Cantonese: <xref ref-type="bibr" rid="ref046">Zhang, 2017</xref>; <xref ref-type="bibr" rid="ref009">Ge and Li, 2018</xref>; Chinese: <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref033">Shih, 2000</xref>; <xref ref-type="bibr" rid="ref034">Shih and Lu, 2010</xref>). Some researchers argue that declination is a fundamental effect in human speech due to a drop in subglottal air pressure (<xref ref-type="bibr" rid="ref024">Lieberman, 1967</xref>; <xref ref-type="bibr" rid="ref006">Collier, 1975</xref>; <xref ref-type="bibr" rid="ref030">Pierrehumbert, 1979</xref>; <xref ref-type="bibr" rid="ref010">Gelfer et al., 1983</xref>; <xref ref-type="bibr" rid="ref013">Gussenhoven, 2004</xref>). However, other researchers stated that declination is a combined effect from different functions, e.g., sentence stress and terminal fall (<xref ref-type="bibr" rid="ref025">Lieberman and Tseng, 1980</xref>; <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref026">Liu and Xu, 2005</xref>), topic initial F0 raising (<xref ref-type="bibr" rid="ref040">Umeda, 1982</xref>; <xref ref-type="bibr" rid="ref70">Wang and Xu, 2011</xref>) and discourse structure (<xref ref-type="bibr" rid="ref014">Hirschberg and Pierrehumbert, 1986</xref>; <xref ref-type="bibr" rid="ref029">Nakajima and Allen, 1993</xref>; <xref ref-type="bibr" rid="ref035">Sluijter and Terken, 1993</xref>). Downstep and pre-low bouncing, as introduced earlier, also contribute to the overall declination (<xref ref-type="bibr" rid="ref023">Liberman and Pierrehumbert, 1984</xref>; <xref ref-type="bibr" rid="ref53">Pierrehumbert and Beckman, 1988</xref>; <xref ref-type="bibr" rid="ref59">Shih, 1988</xref>; <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>). <xref ref-type="bibr" rid="ref033">Shih (2000)</xref> used sentences with the tone sequence of LRH&#x2026;HN (L, R, H and N stands for low, rising, high and neutral tone respectively), and found that the H tones show declination in the way that the lowering slope is steeper in shorter sentences, after taking apart focus and final lowering. In <xref ref-type="bibr" rid="ref034">Shih and Lu (2010)</xref> the intonation of an all H tone digital string (338&#x2013;811-3783) drops from 300&#x2009;Hz to almost 100&#x2009;Hz. Similarly, <xref ref-type="bibr" rid="ref045">Yuan and Liberman (2014)</xref> found that shorter utterances have steeper declination in both the top line and the baseline, after excluding the initial rising and final lowering effects. They are in favor of the idea that declination is linguistically controlled, but not just a by-product of the physics and physiology of talking. It is possible that the declination in the previous three studies still involves some other unknown effects which are hidden by the regression model. In the current study, we calculated declination syllable-by-syllable, as the F0 drop between two adjacent H tones.</p>
</sec>
<sec id="sec4">
<title>Post-low-bouncing</title>
<p>A L tone could also cause F0 raising after it, especially when the following syllables carry the neutral tone, termed as post-low-bouncing (Mandarin and Cantonese: <xref ref-type="bibr" rid="ref6">Chao, 1968</xref>; <xref ref-type="bibr" rid="ref38">Lin and Yan, 1980</xref>; <xref ref-type="bibr" rid="ref59">Shih, 1988</xref>; <xref ref-type="bibr" rid="ref8">Chen and Xu, 2006</xref>; <xref ref-type="bibr" rid="ref21">Gu and Lee, 2009</xref>; <italic>cf.</italic> <xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>). As discussed in <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref>, post-low-bouncing has been considered mostly as an articulatory phenomenon, limited to the first neutral tone after the low tone. They emphasized that post-low F0 bouncing is different from a carryover effect, although it occurs between tones. The carryover effect shows in the way that the initial F0 of a syllable is heavily assimilated to the final F0 of the preceding tone, but over the course of the current syllable, F0 gradually approaches its own tonal target. To account for such assimilatory effect, <xref ref-type="bibr" rid="ref78">Xu and Wang (2001)</xref> proposed the Target Approximation model, which represents the production of successive tones as a process of asymptotically approaching each tonal target within the time interval of the respective syllable, starting from the offset F0 of the preceding syllable. Post-low-bouncing, instead, is the process that pitch increases first then drops back to the underlying target. They discussed the possible physical mechanism behind the low-bouncing effect and suggested a <italic>balance-perturbation hypothesis</italic>. In simple words, after producing a very low F0, the extrinsic laryngeal muscles, especially the sternohyoids (<xref ref-type="bibr" rid="ref51">Ohala, 1972</xref>; <xref ref-type="bibr" rid="ref2">Atkinson, 1978</xref>), stop contracting and thus temporarily tip the balance between the two antagonistic forces maintained by the intrinsic laryngeal muscles, resulting in a sudden increase of the vocal fold tension (<xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>, pp. 422). It still requires articulatory studies to verify the <italic>balance-perturbation hypothesis</italic>. From pitch analysis, one way to test it is to vary syllable duration in the L tone. A longer L tone may reduce post-low-bouncing as it gives more time for the muscles releasing the force. We hence predict that post-low-bouncing is weakened if the L tone is at a phrase boundary, as the L tone is with final lengthening.</p>
</sec>
<sec id="sec5">
<title>Focus and phrasing</title>
<p>There has been extensive research on how focus is realized prosodically in many languages (for an overview see <xref ref-type="bibr" rid="ref33">K&#x00FC;gler and Calhoun, 2020</xref>). In Mandarin, focus is realized by increasing the pitch range, intensity, duration and articulatory fullness of the focused word, and reducing the F<sub>0</sub> and intensity of the following words (post-focus-compression, PFC), while leaving the pre-focus words largely unchanged (<xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref7">Chen and Gussenhoven, 2008</xref>; <xref ref-type="bibr" rid="ref70">Wang and Xu, 2011</xref>). Although Mandarin is tonal, its prosodic focus pattern is very similar to English (<xref ref-type="bibr" rid="ref11">Cooper et al., 1985</xref>; <xref ref-type="bibr" rid="ref13">de Jong, 1995</xref>; <xref ref-type="bibr" rid="ref81">Xu and Xu, 2005</xref>), German (<xref ref-type="bibr" rid="ref18">F&#x00E9;ry and K&#x00FC;gler, 2008</xref>) and many other Indo-European languages (<xref ref-type="bibr" rid="ref042">Xu et al., 2012</xref>). A recent study found that PFC can go across a relative strong prosodic boundary in Mandarin (e.g., a boundary between to clauses), indicating that phrasing does not interfere with post-focus constituents (<xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). In other words, focus and phrasing are largely encoded in parallel in intonation, though focus may cause prosodic boundaries in some languages (e.g., <xref ref-type="bibr" rid="ref33">K&#x00FC;gler and Calhoun, 2020</xref>).</p>
<p>Prosodic boundaries are generally indicated by different phonetic cues such as pre-boundary lengthening, silent pause, F0 reset, phonological boundary tones and changes in voice quality (for detailed discussion, see <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). In Mandarin Chinese, boundary strength is realized with gradient means rather than categorical ones, differentiated mainly in pre-boundary lengthening and optional silent pause, but not F0 (<xref ref-type="bibr" rid="ref79">Xu and Wang, 2009</xref>; <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). Although pitch reset has been found at a strong boundary (Dutch: <xref ref-type="bibr" rid="ref14">de Pijper and Sandeman, 1994</xref>; <xref ref-type="bibr" rid="ref65">Swerts, 1997</xref>; English: <xref ref-type="bibr" rid="ref34">Ladd, 1988</xref>), F0 plays a limited role to distinguish boundary strength in Mandarin Chinese when tones and focus are carefully controlled (<xref ref-type="bibr" rid="ref79">Xu and Wang, 2009</xref>; <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). Minimum F0 is lowered at a strong boundary with a silent pause for about 200&#x2009;ms, but not at a phrase boundary within a sentence (<xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>).</p>
<p>It has been found in many languages that pre-boundary syllables are longer than non-final syllables (e.g., English: <xref ref-type="bibr" rid="ref5">Byrd, 2000</xref>; Finnish: <xref ref-type="bibr" rid="ref48">Nakai et al., 2009</xref>; Dutch: <xref ref-type="bibr" rid="ref66">Swerts and Geluykens, 1994</xref>), and articulatory gestures have slower velocity (<xref ref-type="bibr" rid="ref31">Krivokapic and Byrd 2012</xref>). The prolonged syllable might give rise to fully realized phonetic targets (<xref ref-type="bibr" rid="ref40">Lindblom, 1990</xref>; <xref ref-type="bibr" rid="ref15">DiCanio et al., 2021</xref>), e.g., tones in Mandarin (see Figure 4 in <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>, p. 36). Phrase-final tones carry both lexical tone and post-lexical tone, e.g., a pitch accent and a boundary tone (<xref ref-type="bibr" rid="ref002">Arvaniti and Fletcher, 2020</xref>). On the other hand, phrase-final position may also be the locus of glottalization (<xref ref-type="bibr" rid="ref017">Huffman, 2005</xref>), devoicing (<xref ref-type="bibr" rid="ref041">Wagner, 2002</xref>), and a gradual decay in intensity and F0 (<xref ref-type="bibr" rid="ref013">Gussenhoven, 2004</xref>; <xref ref-type="bibr" rid="ref021">Ladd, 2008</xref>; <italic>cf.</italic> <xref ref-type="bibr" rid="ref15">DiCanio et al., 2021</xref>).</p>
<p>Relating to the current study, we aim to find out whether a phonological phrase boundary reduces or even blocks downstep and post-low-bouncing effect, assuming that the pre-boundary syllable carrying a L tone is longer and hence the tone is fully realized with pitch raising toward the end of the syllable, since fall-rise is the citation form of the L tone in Mandarin. Another possibility is that pitch goes lower when the L tone is longer, thus makes a greater downstep effect.</p>
</sec>
<sec id="sec6">
<title>Laryngeal movement of varying pitch</title>
<p>After introducing the studies on the linguistic meaning of tone and intonation, we here would like to go back to articulatory studies on pitch control. It will help us to understand downstep and post-low-bouncing, since down to the bottom of the questions raised above, it is all about how the muscles, bones, vocal folds and brain cooperate to realize the pitch targets. The observed pitch contours reflect both linguistic meanings and articulatory constrains.</p>
<p><xref ref-type="bibr" rid="ref045">Yuan and Liberman (2014)</xref> discussed articulatory studies on how F0 is controlled. We here just briefly cite some most relevant studies. F0 is determined by the stiffness and effective mass of the vocal folds and the subglottal air pressure (<xref ref-type="bibr" rid="ref028">Murry, 1971</xref>; <xref ref-type="bibr" rid="ref015">Hollien, 1974</xref>, <xref ref-type="bibr" rid="ref016">1983</xref>; <xref ref-type="bibr" rid="ref003">Baer, 1979</xref>; <xref ref-type="bibr" rid="ref038">Titze, 1988</xref>; <xref ref-type="bibr" rid="ref037">Stevens, 2000</xref>; <xref ref-type="bibr" rid="ref047">Zhang, 2016</xref>). Intrinsic laryngeal muscles, especially the cricothyroid muscle (CT), are the main contributor to the adjustment of the stiffness and effective mass of the vocal folds. The contraction of CT raises F0; the relaxation of CT, along with the activity of other laryngeal muscles, lowers F0 (<xref ref-type="bibr" rid="ref006">Collier, 1975</xref>; <xref ref-type="bibr" rid="ref2">Atkinson, 1978</xref>). Extrinsic laryngeal muscles, which suspend and support the larynx, can also change the states of the vocal folds through vertical larynx movement (<xref ref-type="bibr" rid="ref51">Ohala 1972</xref>; <xref ref-type="bibr" rid="ref27">Honda 1995</xref>; <xref ref-type="bibr" rid="ref24">Hirose 1997</xref>), and F0 falls as the larynx moves down.</p>
<p>F0 lowering is not only accompanied by larynx lowering, relating to extrinsic laryngeal muscles (<xref ref-type="bibr" rid="ref27">Honda, 1995</xref>; <xref ref-type="bibr" rid="ref24">Hirose, 1997</xref>) but also involves the joint supraglottal action (<xref ref-type="bibr" rid="ref42">Lindqvist-Gauffin, 1969</xref>, <xref ref-type="bibr" rid="ref43">1972</xref>; <italic>cf.</italic> <xref ref-type="bibr" rid="ref41">Lindblom, 2009</xref>). In Mandarin, the basic role of larynx height in the execution of tone is complicated by the relationship of larynx height to the state of the larynx: constriction of the supra-glottal laryngeal structures is facilitated by raising the larynx (<xref ref-type="bibr" rid="ref17">Edmondson and Esling, 2006</xref>) and inhibited by lowering the larynx (<xref ref-type="bibr" rid="ref47">Moisik et al., 2014</xref>; <xref ref-type="bibr" rid="ref46">Moisik and Esling, 2014</xref>). <xref ref-type="bibr" rid="ref47">Moisik et al. (2014)</xref> shows that a L tone target can be reached either by lowering the larynx, or by combining the raise of larynx height and laryngeal constriction, which may lead to creakiness in the low tone. They show that producing the H tone requires any tone involving lowering in pitch is easily becoming creaky, especially the L tone (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>, <xref ref-type="bibr" rid="ref020">2018</xref>).</p>
<p><xref ref-type="bibr" rid="ref022">Ladefoged (1973</xref>, p. 75) suggested that the creaky voice phonation mechanism is that &#x201C;because the arytenoid cartilages move forward as they come together; the vocal cords tend to be less stretched in creaky voiced sounds; they are therefore likely to vibrate at a lower frequency. But the coming together of the arytenoids and the movements of the thyroid cartilage that stretch the vocal cords are independent laryngeal gestures, so that it is quite possible for creaky voiced sounds to occur on any pitch.&#x201D; Creaky voice in Mandarin L tone exhibits various laryngealization properties in acoustic waveforms, including aperiodicity, period doubling, or low-frequency pulse-like vibratory patterns (<xref ref-type="bibr" rid="ref012">Gerratt and Kreiman, 2001</xref>; <xref ref-type="bibr" rid="ref018">Keating et al., 2015</xref>). In Mandarin, creaky voice relates to the low target in pitch that the L tones are less creaky when the pitch range is raised, but creakier when the pitch range is lowered (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>, <xref ref-type="bibr" rid="ref020">2018</xref>). In previous studies on downstep and post-low-bouncing, creaky voice is usually not taken into account. A consequence of this discussion leads to the question whether creakiness causes downstep or not.</p>
</sec>
<sec id="sec7">
<title>Research questions and hypotheses</title>
<p>The main goal of the current study is to understand the property of downstep in Mandarin. The second goal is to provide some analysis on how post-low-bouncing interacts with boundary strength. These will lead us to better understand how intonation is shaped by both informative functions and articulatory constrains. The research questions and hypotheses are summarized as the following.</p>
<list list-type="order">
<list-item>
<p>How do focus and boundary interact with downstep? We divide this question into 6 sub-questions.</p>
<list list-type="simple">
<list-item>
<p>Q1: Does downstep set up a new register tone?</p>
</list-item>
<list-item>
<p>According to <xref ref-type="bibr" rid="ref75">Xu (1999)</xref>, we predict that downstep effect lasts for several syllables and approach the all-H reference line gradually in wide focus condition.</p>
</list-item>
<list-item>
<p>Q2: Does a sentence-final focus terminates downstep?</p>
</list-item>
<list-item>
<p>We predict that the answer is no because downstep is presumably local, and pitch target is realized syllable-by-syllable as stated in PENTA model (<xref ref-type="bibr" rid="ref043">Xu et al., 2022</xref>).</p>
</list-item>
<list-item>
<p>Q3: Is downstep eliminated by on-focus F0 raising and post-focus-compression?</p>
</list-item>
<list-item>
<p>We predict that informative functions of intonation may override an articulatory effect.</p>
</list-item>
<list-item>
<p>Q4: How does a phonological phrase boundary interact with downstep?</p>
</list-item>
<list-item>
<p>Given that pre-boundary L is lengthened, the tonal target is expected to be fully realized, and in turn, that may lead to greater downstep effect, since the L tone is lower or even being creaky.</p>
</list-item>
<list-item>
<p>Q5: Do declination and downstep share the same mechanism?</p>
</list-item>
<list-item>
<p>The answer to this question actually depends on how to measure declination and downstep. It also remains controversial whether there is any separate articulatory mechanism controlling declination. Our prediction is that downstep and declination may come from different articulatory control, since downstep is local whereas declination is global.</p>
</list-item>
<list-item>
<p>Q6: Is creaky voice the cause of downstep?</p>
</list-item>
<list-item>
<p>Downstep is caused by a L tone, which is usually creaky in Mandarin (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>). It is possible that creaky voice is the main cause of downstep.</p>
</list-item>
</list>
</list-item>
<list-item>
<p>When a L tone is under focus, post-low-bouncing is expected. Does a phrase boundary block post-low-bouncing (Q7)?</p>
<list list-type="simple">
<list-item>
<p>According to <italic>balance-perturbation hypothesis</italic> (<xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>), we predict that post-low-bouncing is weakened if the L tone is at a phrase boundary.</p>
</list-item>
</list>
</list-item>
</list>
</sec>
</sec>
<sec id="sec8" sec-type="materials|methods">
<title>Materials and methods</title>
<p>The experiment aimed to study the size and scope of downstep and post-low-bouncing in Mandarin Chinese, concerning its interaction with focus and phrasing. The size of downstep and post-low-bouncing effects was measured by comparing sentences with all H tones and a comparable sentence with a L tone inserted at the target position, while keeping the rest of the two sentences exactly the same. In this way, we can test whether downstep sets up a new pitch register, as taken the all-H sentence for reference. We named it as <italic>cross-comparison</italic> to answer Q1-Q4. Besides, we also calculated the F0 difference between the two H tones surrounds the L tone, and compared it with the F0 lowering in all-H sentences. We named it as <italic>sequential-comparison</italic> to answer Q5. Thus, the property of downstep and declination can be compared. Moreover, downstep effect caused by creaky and normal L tones were compared, to answer Q6. Post-low-bouncing only occured in the condition of the L tone being focused, thus focus condition is fixed. Only the boundary after the L tone was varied to test whether a strong boundary ends post-low-bouncing (Q7).</p>
<sec id="sec9">
<title>Reading materials</title>
<p>The carrier sentences contained only H tones, except for a neutral tone at sentence-final position. Two target words were embedded in the middle of the carrier sentence, one consisted of a LH word (named as syllable X and Y) triggering downstep and post-low-bouncing, and the other one consisted of a HH word, serving as the reference. The two sentences of each item were read in varied contexts eliciting 4 different focus, and 2 boundary conditions.</p>
<p>Three variables were independently manipulated in this experiment, that is, tone of syllable X (either H or L tone), boundary strength between syllable X and Y (syllable boundary or phrase boundary) and focus type (wide focus (WF), focus on syllable X (XF), on syllable Y (YF) and in sentence final position (ZF)). One set of the sentences in the condition of syllable and phrase boundary were provided in (1a) and (1b). Each sentence was with the syntactic structure as S-V1-O1-V2-O2, and the target words (syllables X and Y) were put in the O1 and V2 position, respectively. Here, by comparing the F0 of syllable Y and that of the following H tones between the two sentences (LH and HH), we can calculate the effect size and scope of downstep. The statistical analysis will then test for how many syllables after the L tone the downstep effect lasts, with consideration of boundary and focus conditions. <inline-graphic xlink:href="fpsyg-13-884102-igr0001.tif"/></p>
<p>For the two boundary conditions, a monosyllabic homophone of the target syllable Y was used to construct sentences with different syntactic boundaries. Based on the assumption of the syntax-phonology interface, prosodic boundaries, in particular in this experimental setting, are the result of matching syntactic constituents onto prosodic constituents (<xref ref-type="bibr" rid="ref032">Selkirk, 2011</xref>). Thus, in the syllable boundary condition (1a), the HL<sub>X</sub>H<sub>Y</sub> was one word, whereas in the phrase boundary condition (1b), the HL<sub>X</sub> was a word, and the following H<sub>Y</sub> was an adverb, phrased together with the following words as a verb phrase (VP). Thus, prosodic boundary in condition (1a) was weaker than that in (1b), named as a syllable boundary (SylB) and a phrase boundary (PhrB) respectively. In example (1a), the HHH sequence (yin1ou1dou1&#x6A31;&#x6B27;&#x515C;, Ying1ou1 bag1) meant a bag printed with ying1ou1 (a make-up word for an exotic plant), whereas in (1b), &#x2018;dou1&#x2019; in the HHH sequence (ying1ou1.dou1&#x6A31;&#x6B27;&#x90FD;, Ying1ou1 all1) was an adverb, meant &#x201C;all&#x201D; to modify the following verb &#x201C;lingchu (take-out).&#x201D; In this way, the two boundary conditions were clearly distinguished by using two different characters (&#x515C; vs. &#x90FD;, bag vs. all). It was the same construction for the HLH sequence, in which the HL tone word is ying1ou3 (&#x6A31;&#x85D5;Ying1ou3), which was also a make-up word for an exotic plant. Here, the contrast of syllable X, either being L or H toned, was straightforward by using the two different characters (&#x85D5;vs &#x6B27;, ou3 vs. ou1). In this way, no specific explanation of the material was necessary for the speakers. They were easily able to read the sentences with different tones and phrasing conditions in a natural way.</p>
<p>Focus was elicited by varying a preceding background sentence, which required a correction of the corresponding word in the target sentence. Taken the HL<sub>X</sub>H<sub>Y</sub> sentence in the syllable boundary condition (see 1a), the four focus conditions are presented in (2). Here, the H tone (syllable Y) is critical to test the effects of downstep and post-low-bouncing, and their interaction with focus. Thus, syllable Y was manipulated as either post-focus (focus on syllable X), on-focus (focus on syllable Y) or pre-focus (focus on syllable Z). A wide focus condition served as the baseline. Similar contexts were constructed for the other sentences, see <xref ref-type="sec" rid="sec230">Appendix I</xref> for the whole sentence sets.</p>
<p>The background sentences of the four focus conditions for the sentence (1a) are as follows.<list list-type="simple">
<list-item>
<p>Wide focus: &#x201C;ni3 ting1shuo1 le0 ma0?&#x201D; (Have you heard about it?)</p>
</list-item>
<list-item>
<p>X-focus: &#x201C;bu2shi4ying1an1&#x201D; (It is not &#x201C;Yingan.&#x201D;)</p>
</list-item>
<list-item>
<p>Y-focus: &#x201C;bu2shi4bao1&#x201D; (It is not the tote.)</p>
</list-item>
<list-item>
<p>Z-focus: &#x201C;bu2shi4lou2dao4&#x201D; (It is not the corridor.)</p>
</list-item>
</list></p>
<p>We constructed two sets of items. In total, 2 (tone of syllable X)&#x2009;&#x00D7;&#x2009;2 (boundary between X and Y)&#x2009;&#x00D7;&#x2009;4 (focus)&#x2009;&#x00D7;&#x2009;2 (sets)&#x2009;&#x00D7;&#x2009;3 (repetitions)&#x2009;&#x00D7;&#x2009;8 (speaker)&#x2009;=&#x2009;768 sentences were analyzed.</p>
</sec>
<sec id="sec10">
<title>Speakers</title>
<p>Eight native Mandarin speakers participated in the experiment at Minzu University of China (5 female and 3 male speakers), from the age of 20 to 28. They were born and brought up in Beijing, spoke no other Chinese dialects and reported no hearing or speaking impairments. They were paid with small amount of money for taking part in the experiment.</p>
</sec>
<sec id="sec11">
<title>Recording procedure</title>
<p>The subjects were recorded individually in the speech lab at Minzu University of China. They were asked to read aloud both the context and the target sentences at a normal speed and in a natural way. They sat before a computer monitor, on which the test sentences were displayed, using AudiRec, a custom-written recording program. To make the reading task a little easier for the speakers, the focused words were highlighted with color. A Shure 58 Microphone was placed about 10&#x2009;cm in front of the speaker. All sentences were digitized directly into a Thinkpad computer and saved as WAV files. The sampling rate was 48 KHz and the sampling format was one channel 6-bit linear. Each speaker repeated the whole set of sentences 3 times in different random order, with about 5 minutes break between sessions. Before the formal recording, they read the sentences silently to get familiar with them, and to make sure that they understood the meaning. The total recording time was about an hour.</p>
</sec>
<sec id="sec12">
<title>Acoustic measurements and statistical methods</title>
<p>The target sentences were extracted and saved as separate WAV files. ProsodyPro (<xref ref-type="bibr" rid="ref76">Xu, 2013</xref>) running under Praat (<xref ref-type="bibr" rid="ref4">Boersma and Weenink, 2013&#x2013;2022</xref>), was used to take F0 and duration of each syllable measurements from the target sentences, which were all segmented into syllables manually, and at the same time hand-checked vocal cycles markings generated for errors, such as double-marking and period skipping. ProsodyPro then generated syllable-by-syllable F0 contours that were either time-normalized or in the original time scale. At the same time, the script extracted various measurements, including maximum F0, minimum F0 and duration of each syllable. We could measure F0 at the offset of a syllable, however maximum F0 is toward the very end of the syllable (see <xref rid="fig5" ref-type="fig">Figure 5</xref>), it is highly probable that the two values are with very little difference. Maximum F0 is much more widely applied in previous studies (e.g., <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref20">Genzel and K&#x00FC;gler, 2011</xref>; <xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>). Thus, we choose maximum F0 to measure downstep effect.</p>
<p>The statistic tests were carried out in the R environment (<xref ref-type="bibr" rid="ref55">R Core Team, 2016</xref>) by using lme4 package Version 1.1&#x2013;18 (<xref ref-type="bibr" rid="ref3">Bates, et al., 2016</xref>) to estimate the effect of the fixed factors (tone, boundary and focus) and the random factors (speaker and sentence set) on the acoustic parameters, e.g., maximum F0 and duration. Regression coefficients (bs), standard errors (SEs) and <italic>t</italic>-values (<italic>t</italic>&#x2009;=&#x2009;b/SE) are reported, taken t&#x2009;&#x003E;&#x2009;2.0 as reaching the significant level at <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05 (<xref ref-type="bibr" rid="ref011">Gelman and Hill, 2006</xref>). In the results, we reported the best-fit model according to the model comparisons with the lowest AIC and BIC. For the fixed factors, we took the model with interaction only when there was significant interaction.</p>
<p>Creaky L tone was visually identified by checking the spectrum and the WAV files. <xref ref-type="bibr" rid="ref047">Zhang (2016)</xref> distinguished four types of creaky voice (see Figure 3 in that paper). We grouped all these types as creaky voice. Since F0 is the main concern in this paper, we here labeled the part with aperiodic pulses as creaky, see <xref rid="fig1" ref-type="fig">Figure 1</xref>. In this way, the part of the regular pulses was used to get the F0 values of the syllable. Most of the creaky L tone was similar to what <xref rid="fig1" ref-type="fig">Figure 1</xref> shows.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>An example of the syllable with creaky L tone. Here, L and c stand for the part with periodic and aperiodic pulses.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g001.tif"/>
</fig>
</sec>
</sec>
<sec id="sec13" sec-type="results">
<title>Results</title>
<p>In this section, the graphic analysis firstly shows how focus and phrasing are realized in intonation (<xref rid="fig2" ref-type="fig">Figures 2</xref>, <xref rid="fig3" ref-type="fig">3</xref>), followed by quantitative analysis of F0 and duration (<xref rid="fig4" ref-type="fig">Figure 4</xref> and <xref rid="tab1" ref-type="table">Table 1</xref>). These two sections serve to confirm that our results are largely consistent with previous studies on focus and phrasing, so that we are confident to further analyze their interaction with the tonal manipulation on intonation. To get an overview of the results, downstep and post-low-bouncing are firstly visually analyzed with intonation contours (<xref rid="fig5" ref-type="fig">Figure 5</xref>). Downstep is then quantitatively analyzed with two different methods, i.e., (a) the <italic>cross- comparison</italic> between LH and HH sentences to verify how many syllables it takes for the H tones after the L tone reaching the all-H sentences to answer Q1-Q4 (<xref rid="fig6" ref-type="fig">Figure 6</xref> and <xref rid="tab2" ref-type="table">Table 2</xref>); (b) the <italic>sequential-comparison</italic> between the H tones surrounding the L tone. By comparing the decrease of the H tones in the LH and HH sentence, we can tear apart the declination and the downstep effect to answer Q5 (<xref rid="fig7" ref-type="fig">Figures 7</xref> <xref rid="fig8" ref-type="fig">8</xref> and <xref rid="tab3" ref-type="table">Table 3</xref>). Thirdly, we noticed that L tones are mostly creaky, especially in the phrase-boundary condition. Therefore, we aim at answering the question whether the change of phonation type to creaky voice is a cause on downstep. We then analyzed the pitch height in the H tone after the L tone as compared between the creaky and normal L tones to answer Q6 (<xref rid="fig9" ref-type="fig">Figure 9</xref>). We can show that the change of phonation type is not the direct cause of downstep. For post-low-bouncing, it only happens when the L tone is focused (XF). In line with the findings in <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref>, we here provide further analysis on its interaction with boundary strength to answer Q7 (<xref rid="fig10" ref-type="fig">Figure 10</xref>). With this analysis, we can justify that the <italic>balance-perturbation hypothesis</italic> holds, which predicts weaker post-low-bouncing when the L tone is longer.</p>
<sec id="sec14">
<title>Graphic analysis on focus and phrasing</title>
<p>First, we present intonation contours to show how focus is encoded in intonation. <xref rid="fig2" ref-type="fig">Figure 2</xref> presents the HH and LH sentences in the condition of syllable boundary, with the 4 focus conditions overlaid in one figure. In each sentence, 10 time-normalized F0 points for each syllable were averaged across 48 observations (8 speakers&#x2009;&#x00D7;&#x2009;2 sets&#x2009;&#x00D7;&#x2009;3 repetitions).</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Time-normalized intonation contours of the HH (left) and LH (right) sentences in the conditions of the <italic>syllable</italic> boundary (between syllable X and Y), with the four focus conditions overlaid in one figure. Here XF, YF, ZF and WF stand for focus in word X, Y, Z and the wide focus condition. The <italic>x</italic>-axis are the syllable numbers. The vertical line indicates the critical boundary between X and Y.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g002.tif"/>
</fig>
<p>We can see in <xref rid="fig2" ref-type="fig">Figure 2</xref> that focus is realized as the tri-zone pattern as defined in <xref ref-type="bibr" rid="ref75">Xu (1999)</xref> and repetitively found in many other studies (e.g., <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). Looking at the HH tone sentences, we can clearly see that the on-focus syllables show raised F0 and expanded pitch range; the post-focus words exhibit lowered and compressed pitch; while the pre-focus words are similar to the wide focus condition. It holds in the LH sentences as well, except that when the L tone word (e.g., ying1ou3) is focused (XF), the pre-low H is raised. And in the YF condition, on-focus F0 raising still applies in the H tone after the L tone. Thus, downstep does not override (or cancel) on-focus F0 raising. The sentences in the phrase boundary show a very similar pattern, which is not presented here for the interest of space. A phrase boundary does not block post-focus F0 compression (PFC), as likewise reported in <xref ref-type="bibr" rid="ref71">Wang et al. (2018b)</xref>. In general, it confirms that tonal interactions and phrasing do not change how focus is realized, though the amount of focal raising appears to differ between tone conditions.</p>
<p>Secondly, <xref rid="fig3" ref-type="fig">Figure 3</xref> presents how boundary strength is encoded in intonation in the XF and YF conditions. No clear difference in F0 between the two boundary conditions can be seen here, in both the HH and LH (lower row) sentences. In WF and ZF conditions, the two boundary conditions do not show clear difference either, which is not presented here for the interest of space. It is in consistence with <xref ref-type="bibr" rid="ref71">Wang et al. (2018b)</xref> that F0 plays a limited role on phrasing, especially on boundaries within a sentence. Importantly, when pre- and post-boundary syllables are under focus (the X and Y focus condition), there is still no clear sign of using F0 to mark boundary strength. Thus, focus in Mandarin does not seem to invulnerably insert a prosodic boundary.</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>The time-normalized intonation contours of the two boundary conditions in the HH and the LH sentences under the XF and YF conditions. Here SylB and PhrB stand for syllable and phrase boundary between syllable X and Y. The <italic>x</italic>-axis are the syllable numbers.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g003.tif"/>
</fig>
<p>The above graphic observations show that F0 variation is mainly triggered by focus and tone, but not by prosodic boundaries. We further analyzed the nature of the boundary and whether speakers distinguished the two boundary conditions phonetically. The following analysis of syllable duration (see section &#x201C;Acoustic analysis on the interaction of focus and boundary,&#x201D; <xref rid="fig4" ref-type="fig">Figure 4</xref>) confirms that boundary strength was encoded mainly in pre-boundary lengthening, but not F0.</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption>
<p>Maximum F0 and duration of syllable X in the H<sub>X</sub>H<sub>Y</sub>(left) and L<sub>X</sub>H<sub>Y</sub>(right) sentences in the two boundary conditions (SylB and PhrB) and the four focus conditions.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g004.tif"/>
</fig>
</sec>
<sec id="sec15">
<title>Acoustic analysis on the interaction of focus and boundary</title>
<p>From the graphic analysis (see <xref rid="fig2" ref-type="fig">Figures 2</xref>, <xref rid="fig3" ref-type="fig">3</xref>), we can see that the intonation patterns of focus and phrasing are consistent with previous studies, e.g., <xref ref-type="bibr" rid="ref75">Xu (1999)</xref> and <xref ref-type="bibr" rid="ref71">Wang et al. (2018b)</xref>. Since focus and boundary effects have already been extensively studied, statistical analysis on all the syllables is not presented here. Statistic test on syllable X is of particular interest as it interacts with downstep and post-low-bouncing. To better understand the interaction between boundary and focus on syllable X, we present the boxplot of maximum F0 and duration of syllable X in <xref rid="fig4" ref-type="fig">Figure 4</xref>.</p>
<p>Linear-mixed-models on maximum F0 and duration in syllable X were carried out in HH and LH sentences separately, with focus and boundary as two non-interactive fixed factors, while speaker and sentence set as random factors (see <xref rid="tab1" ref-type="table">Table 1</xref>). Wide-focus and syllable-boundary were set as the baseline conditions. The LMM model was chosen to meet the criteria that (1) the model with presumed interaction did not show significant interactions, thus we took this model without interaction; and (2) it was with the lowest AIC and BIC while we tried different ways of setting the random effects.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>LMM analysis on maximum F0 and duration in syllable X, with HH and LH sentences separately tested taking focus and boundary as non-interactive fixed factors, whereas speaker and set as random factors in the equation as lmer(dv&#x2009;~&#x2009;focus+boundary+ (1|speaker)&#x2009;+&#x2009;(1|repetition)&#x2009;+&#x2009;(1|set), data&#x2009;=&#x2009;DT), here dv stands for dependent variable, which is MaxF0 and duration.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th/>
<th/>
<th align="center" valign="top">HH</th>
<th/>
<th/>
<th/>
<th align="center" valign="top">LH</th>
<th/>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" colspan="2">Random effects:</td>
<td align="center" valign="top" colspan="3">Num of Observations 429</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="bottom">MaxF0</td>
<td/>
<td align="left" valign="bottom">Var</td>
<td align="left" valign="bottom">SD</td>
<td/>
<td/>
<td align="left" valign="bottom">Var</td>
<td align="left" valign="bottom">SD</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Speaker</td>
<td align="center" valign="top">24.89</td>
<td align="center" valign="top">4.98</td>
<td/>
<td/>
<td align="center" valign="top">29.15</td>
<td align="center" valign="top">5.40</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Rep</td>
<td align="center" valign="top">0.06</td>
<td align="center" valign="top">0.25</td>
<td/>
<td/>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.21</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Set</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.13</td>
<td/>
<td/>
<td align="center" valign="top">0.24</td>
<td align="center" valign="top">0.49</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Res</td>
<td align="center" valign="top">1.10</td>
<td align="center" valign="top">1.05</td>
<td/>
<td/>
<td align="center" valign="top">1.22</td>
<td align="center" valign="top">1.10</td>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Duration</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Speaker</td>
<td align="center" valign="top">165.57</td>
<td align="center" valign="top">12.87</td>
<td/>
<td/>
<td align="center" valign="top">181.46</td>
<td align="center" valign="top">13.47</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Rep</td>
<td align="center" valign="top">2.14</td>
<td align="center" valign="top">1.46</td>
<td/>
<td/>
<td align="center" valign="top">6.07</td>
<td align="center" valign="top">2.46</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Set</td>
<td align="center" valign="top">553.76</td>
<td align="center" valign="top">23.53</td>
<td/>
<td/>
<td align="center" valign="top">387.53</td>
<td align="center" valign="top">19.93</td>
<td/>
<td/>
</tr>
<tr>
<td/>
<td align="left" valign="top">Res</td>
<td align="center" valign="top">1049.2</td>
<td align="center" valign="top">32.39</td>
<td/>
<td/>
<td align="center" valign="top">1305.3</td>
<td align="center" valign="top">36.12</td>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Fixed effects:</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td/>
<td/>
<td align="center" valign="top">Est</td>
<td align="center" valign="top">SE</td>
<td align="center" valign="top"><italic>df</italic></td>
<td align="center" valign="top"><italic>t</italic></td>
<td align="center" valign="top">Est</td>
<td align="center" valign="top">SE</td>
<td align="center" valign="top"><italic>df</italic></td>
<td align="center" valign="top"><italic>t</italic></td>
</tr>
<tr>
<td align="left" valign="top">MaxF0</td>
<td align="left" valign="top">Inter</td>
<td align="center" valign="top">92.0</td>
<td align="center" valign="top">1.67</td>
<td align="center" valign="top">8.22</td>
<td align="center" valign="top">54.91<sup>&#x002A;</sup></td>
<td align="center" valign="top">91.57</td>
<td align="center" valign="top">1.84</td>
<td align="center" valign="top">8.63</td>
<td align="center" valign="top">49.75</td>
</tr>
<tr>
<td/>
<td align="left" valign="top">XF</td>
<td align="center" valign="top">2.80</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">19.58<sup>&#x002A;</sup></td>
<td align="center" valign="top">2.05</td>
<td align="center" valign="top">0.15</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">13.54<sup>&#x002A;</sup></td>
</tr>
<tr>
<td/>
<td align="left" valign="top">YF</td>
<td align="center" valign="top">1.44</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">10.08<sup>&#x002A;</sup></td>
<td align="center" valign="top">0.35</td>
<td align="center" valign="top">0.15</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">2.29<sup>&#x002A;</sup></td>
</tr>
<tr>
<td/>
<td align="left" valign="top">ZF</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">0.70</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.15</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">0.45</td>
</tr>
<tr>
<td/>
<td align="left" valign="top">PhrB</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.08</td>
<td align="center" valign="top">0.11</td>
<td align="center" valign="top">413</td>
<td align="center" valign="top">0.75</td>
</tr>
<tr>
<td align="left" valign="top">Dur</td>
<td align="left" valign="top">Inter</td>
<td align="center" valign="top">165.92</td>
<td align="center" valign="top">17.55</td>
<td align="center" valign="top">1.22</td>
<td align="center" valign="top">9.46<sup>&#x002A;</sup></td>
<td align="center" valign="top">168.55</td>
<td align="center" valign="top">15.30</td>
<td align="center" valign="top">1.36</td>
<td align="center" valign="top">10.97</td>
</tr>
<tr>
<td/>
<td align="left" valign="top">XF</td>
<td align="center" valign="top">66.08</td>
<td align="center" valign="top">4.41</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">14.98<sup>&#x002A;</sup></td>
<td align="center" valign="top">61.93</td>
<td align="center" valign="top">4.92</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">12.60<sup>&#x002A;</sup></td>
</tr>
<tr>
<td/>
<td align="left" valign="top">YF</td>
<td align="center" valign="top">24.41</td>
<td align="center" valign="top">4.41</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">5.53<sup>&#x002A;</sup></td>
<td align="center" valign="top">30.09</td>
<td align="center" valign="top">4.92</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">6.12<sup>&#x002A;</sup></td>
</tr>
<tr>
<td/>
<td align="left" valign="top">ZF</td>
<td align="center" valign="top">&#x2212;7.19</td>
<td align="center" valign="top">4.41</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">&#x2212;1.63</td>
<td align="center" valign="top">6.62</td>
<td align="center" valign="top">4.92</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">1.35</td>
</tr>
<tr>
<td/>
<td align="left" valign="top">PhrB</td>
<td align="center" valign="top">20.06</td>
<td align="center" valign="top">3.12</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">6.44<sup>&#x002A;</sup></td>
<td align="center" valign="top">23.58</td>
<td align="center" valign="top">3.48</td>
<td align="center" valign="top">416</td>
<td align="center" valign="top">6.78<sup>&#x002A;</sup></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: &#x002A; stands for <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05.</p>
</table-wrap-foot>
</table-wrap>
<p>As for focus effect on syllable X, together with the observations in <xref rid="fig4" ref-type="fig">Figure 4</xref>, the statistical analysis in <xref rid="tab1" ref-type="table">Table 1</xref> shows that focus significantly increases both maximum F0 (about 2.8 st) and duration (about 66&#x2009;ms) of syllable X (see the line of XF in <xref rid="tab1" ref-type="table">Table 1</xref>). In the Y focus condition, the 3 syllables HXY (H means the high tone before syllable X, e.g., ying1ou1dou1 &#x2018;Yingou bag&#x2019;) is possibly grouped as one prosodic word, thus syllable X is also with increased maximum F0 (about 1.4 st) and duration (about 24&#x2009;ms; see the line of YF) in <xref rid="tab1" ref-type="table">Table 1</xref>, which is in consistent with the findings in <xref ref-type="bibr" rid="ref004">Chen (2006)</xref> on the durational domain of focus.</p>
<p>As for boundary effect on syllable X, the data in <xref rid="tab1" ref-type="table">Table 1</xref> (also see <xref rid="fig3" ref-type="fig">Figures 3</xref>, <xref rid="fig4" ref-type="fig">4</xref>) show that boundary does not have any effect in maximum F0 (92.7 st vs. 92.6 st), but only in duration of syllable X (192&#x2009;ms vs. 212&#x2009;ms). No interaction was found between focus and boundary in the duration of syllable X, meaning that the pre-boundary lengthening applies to roughly the same degree in all the focus conditions (see <xref rid="fig4" ref-type="fig">Figure 4</xref>), which is in consistent with <xref ref-type="bibr" rid="ref71">Wang et al. (2018b)</xref>. The above results hold for both HH and LH sentences. It leads us to conclude that focus and tone do not interfere with pre-boundary lengthening. Thus, durational adjustment due to focus, boundary and tone is also largely encoded in parallel. We can then further test whether the lengthened L tone decreases or increases the level of downstep and post-low-bouncing in the following sections.</p>
<p>From <xref rid="fig3" ref-type="fig">Figure 3</xref>, we can see that maximum F0 in the L tone is actually the end point of the preceding H tone, which does not show any difference between the two boundary conditions (see <xref rid="tab1" ref-type="table">Table 1</xref> and <xref rid="fig4" ref-type="fig">Figure 4</xref>). Does the minimum F0 in the L tone differ between the boundary conditions? With similar LMM tests in the LH and HH sentences separately, taken boundary and focus as two fixed factors with interaction, and speaker as the random factor, the minimum F0 of syllable X in the LH sentence showed no difference in the two boundary conditions either (86.4 st on average in both conditions) (Estimate&#x2009;=&#x2009;&#x2212;0.369, SE&#x2009;=&#x2009;0.426, <italic>df</italic>&#x2009;=&#x2009;352, <italic>t</italic>&#x2009;=&#x2009;&#x2212;0.866, <italic>p</italic>&#x2009;=&#x2009;0.387). However, there was an interaction between focus and phrasing, i.e., when the L tone is focused (XF), the minimum F0 in the phrase boundary condition is significantly lower than in the syllable boundary condition (Estimate&#x2009;=&#x2009;&#x2212;1.402, SE&#x2009;=&#x2009;0.598, <italic>df</italic>&#x2009;=&#x2009;352, <italic>t</italic>&#x2009;=&#x2009;&#x2212;2.344, <italic>p</italic>&#x2009;=&#x2009;0.0196). In the other three focus conditions, no difference in minimum F0 was found between the two boundary conditions.</p>
<p>When we labeled the speech data, we noticed that most of the L tones were creaky, that was 84.1% and 74.6% in the phrase and syllable boundary, conditions respectively. It is possible that creakiness is an additional feature of a stronger boundary, when minimum F0 cannot go any lower at a phrase boundary (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>).</p>
<p>To summarize, (1) focus is reliably realized in a tri-zone pattern, i.e., pre-focus F0 is largely intact, on-focus F0 is raised and post-focus F0 is lowered and compressed; in addition, focus increases duration of the focused syllable; (2) boundary strength has very little effect on maximum or minimum F0, but mainly realized by pre-boundary lengthening, which is independent from focus and tone; (3) The L tone is more likely to be creaky when it is before a phrase boundary than a syllable boundary.</p>
</sec>
<sec id="sec16">
<title>Graphic analysis on downstep and post-low-bouncing</title>
<p>The analysis on focus and boundary in section &#x201C;Graphic analysis on focus and phrasing&#x201D; and section &#x201C;Acoustic analysis on the interaction of focus and boundary&#x201D; shows that the current experiment is in agreement with previous findings on these two effects (<xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>). It validates the following analysis on the interaction of these two functional variations with the tonal effects, i.e., downstep and post-low-bouncing. As introduced in the beginning of the results section, we here firstly report the <italic>cross-comparison</italic> on assessing the downstep effect adopted from <xref ref-type="bibr" rid="ref75">Xu (1999)</xref> and <xref ref-type="bibr" rid="ref033">Shih (2000)</xref> among many others, by comparing crossly between the HH and LH sentences (see <xref rid="fig5" ref-type="fig">Figure 5</xref>).</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption>
<p>The comparison between the HH (black line) and LH (yellow line) sentences in four focus conditions (from left to right are the conditions of focus in syllable X, Y, Z and in wide-focus) under the condition that the boundary between X and Y is a syllable (SylB; upper row) or a phrase boundary (PhrB; lower row), as indicated by the vertical line. The downward arrow indicates where the downstep effect can be seen.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g005.tif"/>
</fig>
<p>In the wide- and Z focus sentences, we can see in <xref rid="fig5" ref-type="fig">Figure 5</xref> that F0 raises greatly in syllable Y in the LH sentence, which is the procedure of target approximation from a low starting point to the H target. As expected, F0 in syllable Y does not reach the height as the HH tone sentences in several H tones after the L tone, showing a clear downstep effect. We can also see that the downstep effect becomes weaker when the H tones are in a longer distance from the L tone. Five new findings are as below.</p>
<list list-type="order">
<list-item>
<p>The downstep effect also holds when the focused word is sentence final (ZF), indicating that on-focus F0 raising in word Z does not seem to have any anticipatory effect on downstep.</p>
</list-item>
<list-item>
<p>The above observations hold in both the syllable and phrase boundary conditions. Thus, a stronger phrase boundary does not block the downstep effect. Despite a longer duration in the L tone before a phrase boundary (see <xref rid="fig4" ref-type="fig">Figure 4</xref>), downstep still applies. The following analysis shows that this is because the L tone is with lower F0 and even becomes creaky at a phrase boundary.</p>
</list-item>
<list-item>
<p>When the H tone right after the L tone is focused (YF), the downstep effect still shows in syllable Y but not in the following H tones. Surprisingly, even on-focus F0 raising does not cancel the downstep effect. In other words, we can say that downstep does not cancel on-focus F0 raising. It further confirms that the downstep effect is relatively robust. However, post-focus-compression (PFC) seems to override the downstep effect since there is no clear difference in the H tones after syllable Y between the HH and LH sentences, which is statistically confirmed below in <xref rid="fig6" ref-type="fig">Figure 6</xref>.</p>
</list-item>
<list-item>
<p>Comparing the two boundary conditions it seems that downstep is greater in the phrase boundary condition, however, in the YF condition the downstep effect is weaker in the phrase boundary condition.</p>
</list-item>
<list-item>
<p>When the L tone is under focus (XF), instead of downstep, the post-low-bouncing effect shows in the adjacent H tones. Here, the H tone after the L tone goes up first, then drops gradually, as compared to the all-H sequence, as reported in <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref>. Note that the H tones in the baseline condition are realized lower, i.e., in a compressed pitch register (post-focal compression, <xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>; <xref ref-type="bibr" rid="ref042">Xu et al., 2012</xref>). The new finding is that the post-low-bouncing effect seems to be weaker in the phrase boundary condition than in the syllable boundary condition, which supports the <italic>balance-perturbation hypothesis</italic> as proposed in <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref>.</p>
</list-item>
</list>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption>
<p>The downstep size in the Wide-focus (WF), Z-focus(ZF), and Y-focus (YF) conditions, divided by syllable (SylB) and phrase boundary (PhrB) conditions. The significant downstep effects are marked with <sup>&#x002A;</sup> indicating that <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05. The <italic>x</italic>-axis shows syllable numbers, in which the 7th is syllable Y, the H tone right after the L tone.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g006.tif"/>
</fig>
<p>In summary, the graphic analysis of <xref rid="fig5" ref-type="fig">Figure 5</xref> shows that: (1) The downstep effect is relatively robust in varied focus and boundary conditions. More specifically, downstep is not blocked by a phrase boundary, neither is it overridden by on-focus F0 raising or phrase boundary. (2) When the L tone is under focus, post-low bouncing is found in the following H tones, and seems to be weakened by a phrase boundary.</p>
</sec>
<sec id="sec17">
<title>The <italic>cross-comparison</italic> of downstep effect</title>
<p>The main questions to be quantitatively analyzed are the size and the domain of downstep and post-low-bouncing effect, and their interactions with focus and phrase boundary.</p>
<p>Downstep is firstly analyzed by comparing the LH and the corresponding HH sentences in the WF, ZF and YF conditions. In the <italic>cross-comparison,</italic> the size of the downstep effect is calculated by the difference in maximum F0 between the H tones in the LH and HH sentence in syllable Y (syllable 7) and the following syllables (syllable 8 to 14). The post-low-bouncing effect is calculated in the X-focus condition in a similar way (see section &#x201C;F0 analysis on post-low-bouncing effect&#x201D;).</p>
<p><xref rid="fig6" ref-type="fig">Figure 6</xref> presents the size of downstep effect in the three focus and two boundary conditions. The mean values show how much F0 maximum is lowered in the LH sentence as compared to the HH sentence in the corresponding syllable. Paired-sample T tests were applied in each syllable to test whether the difference reached statistical significance at the level of <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05, which is marked by a &#x002A; in <xref rid="fig6" ref-type="fig">Figure 6</xref>.</p>
<p>To get an overall statistical analysis of the factors on the downstep effect, a LMM was applied, setting focus, boundary, and syllable as fixed factors with interactions presumed (WF, syllable boundary, and the 7<sup>th</sup> syllable are set as the base-line condition), while speaker is the random factor (see <xref rid="tab2" ref-type="table">Table 2</xref>). Putting it together with the t-test in <xref rid="fig6" ref-type="fig">Figure 6</xref>, the following findings are statistically supported: (1) The downstep effect in Y-focus condition is significantly smaller than that in wide-focus condition, while no difference is found between Z-focus and wide-focus condition. (2) The downstep effect decreases as the H tones are in longer distance from the L tone. (3) Unexpectedly, the downstep effect is greater in the phrase boundary condition than the syllable boundary condition, especially in the wide-focus conditions. It is probably because the L tone is with lower minimum F0 and with more creaky voice (see section &#x201C;Acoustic analysis on the interaction of focus and boundary&#x201D;), thus the following H tone is with a larger difference from the all-H reference, as compared to the syllable boundary condition. (4) In the Y-focus condition, the downstep effect interacts with focus and boundary, in the way that the downstep effect in the adjacent syllable of the L tone is greater in the syllable boundary condition than in the phrase boundary condition.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>LMM analysis on downstep size (difference of maximum F0 between LH and HH sentences in the H tones) with the equation as lmer(downstepsize&#x2009;~&#x2009;focus &#x002A; syllable &#x002A; boundary&#x2009;+&#x2009;(1 | speaker), data&#x2009;=&#x2009;DT2).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top" colspan="4">Number of observations: 2544</th>
</tr>
<tr>
<th align="left" valign="top" colspan="2">Random effects:</th>
<th align="center" valign="top">Variance</th>
<th align="center" valign="top">SD</th>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Speaker</td>
<td align="center" valign="top">(Intercept)</td>
<td align="left" valign="top">0.078</td>
<td align="center" valign="top">0.279</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">Residual</td>
<td/>
<td align="left" valign="top">1.218</td>
<td align="center" valign="top">1.104</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">Fixed effects:</td>
<td align="center" valign="top">Estimate</td>
<td align="left" valign="top">SE</td>
<td align="center" valign="top"><italic>df</italic></td>
<td align="left" valign="top"><italic>t</italic></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">2.84</td>
<td align="left" valign="top">0.27</td>
<td align="center" valign="top">435</td>
<td align="left" valign="top">10.47<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">YF</td>
<td align="left" valign="top">&#x2212;1.53</td>
<td align="left" valign="top">0.36</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;4.30<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">ZF</td>
<td align="left" valign="top">0.11</td>
<td align="left" valign="top">0.36</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">0.30</td>
</tr>
<tr>
<td align="left" valign="top">syllable</td>
<td align="left" valign="top">&#x2212;0.20</td>
<td align="left" valign="top">0.02</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;8.57<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">PhrB</td>
<td align="left" valign="top">0.79</td>
<td align="left" valign="top">0.36</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">2.21<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">YF:syllable</td>
<td align="left" valign="top">0.12</td>
<td align="left" valign="top">0.03</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">3.56<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">ZF:syllable</td>
<td align="left" valign="top">&#x2212;0.02</td>
<td align="left" valign="top">0.03</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;0.69</td>
</tr>
<tr>
<td align="left" valign="top">YF:PhrB</td>
<td align="left" valign="top">&#x2212;1.16</td>
<td align="left" valign="top">0.50</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;2.31<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">ZF:PhrB</td>
<td align="left" valign="top">&#x2212;0.55</td>
<td align="left" valign="top">0.50</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;1.08</td>
</tr>
<tr>
<td align="left" valign="top">syllable:PhrB</td>
<td align="left" valign="top">&#x2212;0.04</td>
<td align="left" valign="top">0.03</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">&#x2212;1.28</td>
</tr>
<tr>
<td align="left" valign="top">YF:syllable:PhrB</td>
<td align="left" valign="top">0.06</td>
<td align="left" valign="top">0.05</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">1.18</td>
</tr>
<tr>
<td align="left" valign="top">ZF:syllable:PhrB</td>
<td align="left" valign="top">0.02</td>
<td align="left" valign="top">0.05</td>
<td align="center" valign="top">2,522</td>
<td align="left" valign="top">0.45</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: &#x002A; stands for <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="sec18">
<title>The <italic>sequential-comparison</italic> on downstep and declination</title>
<p>Another way to analyze downstep is the degree of F0 lowering after a L tone. In this way, downstep effect can be compared with declination, which was analyzed by calculating the difference of the maximum F0 in the adjacent H tones in all-H sentences. Firstly, we just analyzed the two H tones surrounding the L tone, that is the difference of maximum F0 between syllable 7 and 5 (Difsy7sy5), presented in <xref rid="fig7" ref-type="fig">Figure 7</xref> as boxplots divided by focus conditions, with HH and LH sentences compared directly. Since the results already show that boundary has no effect on maximum F0 in syllable X (<xref rid="fig4" ref-type="fig">Figure 4</xref>), we here averaged the two boundary conditions in <xref rid="fig7" ref-type="fig">Figure 7</xref>.</p>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption>
<p>Maximum F0 difference between the two H tones surrounding syllable X (either L or H) in different focus conditions, while the two boundary conditions are averaged.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g007.tif"/>
</fig>
<p>To evaluate whether there is declination in all H tone sentence, we compared &#x201C;Difsy7sy5&#x201D; in the wide focus condition of the HH sentences with 0 in a one-sample <italic>t</italic>-test (<italic>t</italic>&#x2009;=&#x2009;&#x2212;3.359, <italic>df</italic>&#x2009;=&#x2009;107, <italic>p</italic>&#x2009;=&#x2009;0.001). The 95% confidence interval is &#x2212;0.3 to &#x2212;0.07. With the same analysis, however, declination is not found in the Z-focus condition (<italic>t</italic>&#x2009;=&#x2009;&#x2212;1.46, <italic>df</italic>&#x2009;=&#x2009;105, <italic>n.s.</italic>). Thus, declination is to a much less degree and vulnerable to be cancelled by a final focus.</p>
<p>The LMM model on &#x201C;Difsy7sy5,&#x201D; with focus, boundary and tone as fixed factors and speaker as random factor, showed a main effect in tone and focus, but not in boundary (see <xref rid="tab3" ref-type="table">Table 3</xref>). It further confirms that F0 plays a limited role on differentiating boundary degrees.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>LMM analysis on the difference of maximum F0 in the H tones before and after syllable X, with focus, boundary and tone as fixed factors, whereas speaker as a random factor in the formula as: difs7s5&#x2009;~&#x2009;tone &#x002A; focus + boundary + (1 | speaker).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top" colspan="4">Number of observations: 858</th>
</tr>
<tr>
<th align="left" valign="top">Random effects:</th>
<th/>
<th align="center" valign="top">Variance</th>
<th align="center" valign="top">SD</th>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Speaker</td>
<td align="center" valign="top">(Intercept)</td>
<td align="left" valign="top">0.3804</td>
<td align="center" valign="top">0.6167</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">Residual</td>
<td/>
<td align="left" valign="top">1.7171</td>
<td align="center" valign="top">1.3104</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">Fixed effects:</td>
<td align="center" valign="top">Estimate</td>
<td align="left" valign="top">SE</td>
<td align="center" valign="top"><italic>df</italic></td>
<td align="left" valign="top"><italic>t</italic></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#x2212;0.138</td>
<td align="left" valign="top">0.245</td>
<td align="center" valign="top">14.755</td>
<td align="left" valign="top">&#x2212;0.56</td>
</tr>
<tr>
<td align="left" valign="top">toneLH</td>
<td align="left" valign="top">&#x2212;1.655</td>
<td align="left" valign="top">0.179</td>
<td align="center" valign="top">841.013</td>
<td align="left" valign="top">&#x2212;9.214<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">XF</td>
<td align="left" valign="top">&#x2212;1.264</td>
<td align="left" valign="top">0.178</td>
<td align="center" valign="top">840.997</td>
<td align="left" valign="top">&#x2212;7.088<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">YF</td>
<td align="left" valign="top">1.724</td>
<td align="left" valign="top">0.178</td>
<td align="center" valign="top">841.000</td>
<td align="left" valign="top">9.649<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">ZF</td>
<td align="left" valign="top">0.101</td>
<td align="left" valign="top">0.179</td>
<td align="center" valign="top">841.011</td>
<td align="left" valign="top">0.566</td>
</tr>
<tr>
<td align="left" valign="top">boundaryPhrB</td>
<td align="left" valign="top">&#x2212;0.103</td>
<td align="left" valign="top">0.089</td>
<td align="center" valign="top">841.001</td>
<td align="left" valign="top">&#x2212;1.149</td>
</tr>
<tr>
<td align="left" valign="top">toneLH:XF</td>
<td align="left" valign="top">&#x2212;0.084</td>
<td align="left" valign="top">0.253</td>
<td align="center" valign="top">841.005</td>
<td align="left" valign="top">&#x2212;0.332</td>
</tr>
<tr>
<td align="left" valign="top">toneLH:YF</td>
<td align="left" valign="top">2.326</td>
<td align="left" valign="top">0.253</td>
<td align="center" valign="top">841.013</td>
<td align="left" valign="top">9.181<sup>&#x002A;</sup></td>
</tr>
<tr>
<td align="left" valign="top">toneLH:ZF</td>
<td align="left" valign="top">0.298</td>
<td align="left" valign="top">0.253</td>
<td align="center" valign="top">841.009</td>
<td align="left" valign="top">1.178</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: &#x002A; stands for <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05.</p>
</table-wrap-foot>
</table-wrap>
<p>In general, by comparing HH with LH in <xref rid="fig7" ref-type="fig">Figure 7</xref>, we can see that the difference on the degree of F0 drop in the all-H tone sentence is significantly less than the downstep effect in XF, ZF and WF conditions (<italic>p</italic>&#x2009;&#x003C;&#x2009;0.05). The interaction between focus and tone is not found in XF condition. We can see that Difsy7sy5 is greater in LH than in the HH sentence. Besides, Difsy7sy5 in the HH sentence is much smaller in the XF than in the WF condition, which reflects post-focus-compression (PFC) in F0. The new finding here is that downstep still shows aside from PFC. It means that the downstep effect is not just the general downtrend of F0. Declination and downstep are presumably not from the same articulatory mechanism.</p>
<p>When focus is on the H tone after the L tone (YF), the pitch difference between the two H tones (syl5 and syl7) is greater in the LH than the HH sentences. This comes from pre-low-raising (<xref ref-type="bibr" rid="ref37">Lee et al., 2021</xref>). Here, it also shows the pre-low-raising is independent of on-focus F0 raising.</p>
<p>Then, we further tested whether declination holds all along the sentence by comparing maximum F0 of each adjacent H tones, see <xref rid="fig8" ref-type="fig">Figure 8</xref>. We here only consider the wide-focus condition. The &#x002A;&#x002A;&#x002A; in the figure indicates that the F0 raise or drop between two adjacent H tones in all-H sentence is greater than 0 by on-sample <italic>t</italic>-test with <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, otherwise there is no difference between the two H tones. To put it in a simple way, the &#x002A;&#x002A;&#x002A; means that there is either F0 raising or declination in the current syllable. We can see that in the HH sentences, F0 goes up in the beginning of the sentence (increased 0.34 st), then drops gradually for about 3 syllables (decreased 0.25 st). However, between syllable 7 and 6 and between syllable 9 and 8, no significant difference is found in maximum F0. These two positions are phrase boundaries. It is possible that a phrase boundary cancels declination. Toward the end of the sentence, declination is absent as well. The last syllable is a neutral tone, which causes a sharp drop in F0. Thus, declination is with a very small pitch drop between two H tones, and can be easily cancelled due to topic, boundary, tone and other reasons.</p>
<fig position="float" id="fig8">
<label>Figure 8</label>
<caption>
<p>Boxplot of the maximum F0 difference between two adjacent H tones, as compared between HH and LH wide-focus sentences. The number in the <italic>x</italic>-axis (2&#x2013;14) means that this is the maximum F0 of the current syllable minus the preceding syllable. <sup>&#x002A;&#x002A;&#x002A;</sup> here indicates significant difference between 0 by one-sample <italic>t</italic>-test in the HH sentences.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g008.tif"/>
</fig>
<p>If we look at the LH sentences, we can see that the H tones around the L tone causes much greater F0 change than the all-H sentence. Toward the end of the sentence, the adjacent H tones do not differ in F0, which is similar to the all-H sentence. It is in agreement with the <italic>cross-comparison</italic> of the downstep effect, that downstep gets weaker as the H tones are further from the L tone.</p>
<p>All-together, both the <italic>cross-</italic> and <italic>sequential-comparison</italic> show that downstep effect is in a much greater degree than declination. Downstep effect is robust, lasting for about 2&#x2013;3 syllables. Downstep effect is not cancelled by focus or boundary, whereas declination can be cancelled by these two informative functions.</p>
</sec>
<sec id="sec19">
<title>Creaky L tone</title>
<p>The very last question concerning downstep is whether it is caused by creaky voice, or whether creaky L tones cause greater downstep effect. The number of creaky L tone in different conditions is presented in <xref rid="tab4" ref-type="table">Table 4</xref>. In line with previous studies, we also see that when the L tone is under focus and at a phrase boundary, it is more likely to be creaky. We do not go into detailed analysis on the acoustic parameters of the creaky L tone. Instead, we simply calculate the amount of creakiness in L tones to answer the question whether a creaky low tone causes greater downstep effect. In <xref rid="fig9" ref-type="fig">Figure 9</xref>, the maximum F0 of syllable Y is plotted against the duration of the creaky part in syllable X, with four focus conditions divided in different plots. When the creaky duration is 0, it means this is a normal L tone. Here we do not see any clear trend of a creaky L tone causes lower F0 in the following H tone, which is supported by the LMM model analysis with creaky, focus and gender as fixed factors and speaker as a random factor (lmer(maxF0syl7&#x2009;~ Creakylablel&#x002A;focus&#x002A;Gender+ (1|speaker), data&#x2009;=&#x2009;creaky)). The LMM shows significant effect in focus and gender, whereas creaky does not show any effect (Estimate&#x2009;=&#x2009;&#x2212;0.2455, SE&#x2009;=&#x2009;0.416, <italic>df</italic>&#x2009;=&#x2009;347, <italic>t</italic>&#x2009;=&#x2009;&#x2212;0.59, <italic>n.s.</italic>). Thus, creaky L tone is not the direct cause of downstep, but strengthens downstep. It then explains why downstep effect is greater after a phrase boundary (<xref rid="fig6" ref-type="fig">Figure 6</xref>).</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>The percentage of creaky L tone in different focus and boundary conditions (%).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">Syllable boundary</th>
<th align="center" valign="top">Phrase boundary</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">XF</td>
<td align="char" valign="top" char=".">95.8</td>
<td align="char" valign="top" char=".">91.6</td>
</tr>
<tr>
<td align="left" valign="top">YF</td>
<td align="char" valign="top" char=".">60.4</td>
<td align="char" valign="top" char=".">72.9</td>
</tr>
<tr>
<td align="left" valign="top">ZF</td>
<td align="char" valign="top" char=".">70.8</td>
<td align="char" valign="top" char=".">85.4</td>
</tr>
<tr>
<td align="left" valign="top">WF</td>
<td align="char" valign="top" char=".">64.5</td>
<td align="char" valign="top" char=".">83.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig9">
<label>Figure 9</label>
<caption>
<p>Scatter plot of the maximum F0 in syllable Y as functioned by duration of the creaky part in syllable X, with the color and shape differentiating speakers. The focus conditions are divided in each plot. The points with creaky duration of X being 0 means that this is a normal L tone.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g009.tif"/>
</fig>
</sec>
<sec id="sec20">
<title>F0 analysis on post-low-bouncing effect</title>
<p>The post-low-bouncing effect was calculated as the difference of maximum F0 in the H tones between the LH and HH sentences in the XF condition (the L tone is on-focus). As can be seen in <xref rid="fig7" ref-type="fig">Figures 7</xref>, <xref rid="fig10" ref-type="fig">10</xref>, F0 maximum is lower in the syllable right after the L tone (syllable 7), that is because the F0 maximum of syllable 7 in the HH sentence is the offset of the previous H tone (the maximum F0 hence appears at the onset of the syllable 7 representing the transition from the focused H tone to a post-focally H tone). Post-low-bouncing shows at the end of syllable 7, which can be observed in the maximum F0 of syllable 8 and 9, then the pitch gradually drops back. The linear-mixed-model analysis was carried out with syllable and boundary as two fixed factors (with interaction), while speaker and set are random factors. In this statistical test, we only considered syllable 8 to 10, since no difference is shown between HH and LH sentences after the 10<sup>th</sup> syllable (see <xref rid="fig10" ref-type="fig">Figure 10</xref>). The LMM analysis shows a significant effect in syllable (SE&#x2009;=&#x2009;0.093, <italic>t</italic>&#x2009;=&#x2009;&#x2212;5.087&#x002A;), boundary (SE&#x2009;=&#x2009;1.193, <italic>t</italic>&#x2009;=&#x2009;&#x2212;3.018&#x002A;) and the interaction (SE&#x2009;=&#x2009;0.132, <italic>t</italic>&#x2009;=&#x2009;2.899&#x002A;). Thus, the following observations in <xref rid="fig5" ref-type="fig">Figures 5</xref>, <xref rid="fig10" ref-type="fig">10</xref> are statistically supported: (1) The post-low-bouncing effect gradually decreases in the syllables after the L tone; (2) A phrase boundary weakens post-low-bouncing effect, especially in the second H tone after the L tone.</p>
<fig position="float" id="fig10">
<label>Figure 10</label>
<caption>
<p>Post-low-bouncing effect in the X-focus condition when the boundary between syllable X and Y is either a syllable (SylB) or a phrase (PhrB) boundary. The <italic>x</italic>-axis shows syllable numbers, in which the 7th is syllable Y, the H tone right after the L tone.</p>
</caption>
<graphic xlink:href="fpsyg-13-884102-g010.tif"/>
</fig>
</sec>
</sec>
<sec id="sec21">
<title>General discussion</title>
<p>The new contribution of the current study is on how pragmatic functions interacts with downstep and post-low-bouncing. With the control of focus, post-low-bouncing was brought in, which was mainly analyzed for neutral tones in previous studies (<xref ref-type="bibr" rid="ref8">Chen and Xu, 2006</xref>; <xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>). In the current study, it happened in the following H tones when the L tone is focused (XF in <xref rid="fig5" ref-type="fig">Figure 5</xref>). Although, a large part of intonation variation is informative, we want to emphasize that articulatory constrains on pitch change could not be neglected, since post-low-bouncing and downstep last for several syllables with decreasing in size from of 2.5 to 0.5 st, interacting actively with phrasing and focus. Our findings support the <italic>additive division hypothesis</italic> of pitch range, proposed in <xref ref-type="bibr" rid="ref45">Liu et al. (2021)</xref>. They found that pitch range of 5&#x2013;12 st above the baseline signals both focus and surprise, suggesting an overlap between different layers of meanings within this pitch range. In their study, to perceive focus, F0 needs to be raised about 3 st. We here show the downstep and post-low-bouncing are in a pitch range of less than 2.5 st (<xref rid="fig6" ref-type="fig">Figures 6</xref>, <xref rid="fig10" ref-type="fig">10</xref>), whereas on-focus F0 raising in syllable X is 2.8 st on average (<xref rid="tab1" ref-type="table">Table 1</xref>). In a rough sense, it explains why on-focus F0 raising needs to be about 3st, beneath which F0 variation reflects tone and articulatory constrains. It is possible that any sudden and great pitch raising may bring-in informative meaning, thus it takes several syllables for downstep and post-low-bouncing to go back to the reference line. The process of target approximation as proposed in PENTA model (<xref ref-type="bibr" rid="ref043">Xu et al., 2022</xref>) probably reflects both articulatory and perceptual constrains.</p>
<p>Relating to the tonal variation due to the L tone, the pre-low-raising was systematically studied in <xref ref-type="bibr" rid="ref37">Lee et al. (2021)</xref> in Thai and Cantonese, that is, the H tone is raised in pitch before a L tone. They discussed three possible explanations: (a) a velocity account, (b) a perceptual account, and (c) an anatomical account. More specifically, (a) the raising pitch in the preceding syllable may increase the distance of the downward movement toward the low tone; (b) pre-low-bouncing may enhance tonal contrasts to aid comprehension; (c) if pre-low-raising is not actively planned, it may be the direct result of intrinsic laryngeal muscle movement. Their analysis does not support (b), the perceptual account. Putting it together with <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref> and the current study, we can conclude that pitch movements caused by a L tone (pre-L-raising, downstep and post-low-bouncing) are largely the outcome of intrinsic and extrinsic laryngeal muscle movement. Below we will provide detailed discussion on the research questions.</p>
<sec id="sec22">
<title>How do focus and boundary interact with downstep (Q1-Q6)?</title>
<p>This is actually a very complicated question, since focus in Mandarin involves both on-focus raising and post-focus-compression in F0 (<xref ref-type="bibr" rid="ref59">Shih, 1988</xref>; <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref7">Chen and Gussenhoven, 2008</xref>; <xref ref-type="bibr" rid="ref70">Wang and Xu, 2011</xref>; <xref ref-type="bibr" rid="ref71">Wang et al., 2018b</xref>), see <xref rid="fig2" ref-type="fig">Figure 2</xref> in the current study. Besides, downstep refers to the relevant pitch height in the H tones, either as compared to all-H reference line (<italic>cross-comparison</italic> answering Q1-Q4), or as the F0 drop between the H tones before and after the L tone (s<italic>equential-comparison</italic> answering Q5). To make the question even more complicated, L tones usually become creaky (Q6). No previous study has considered the influence of creakiness on downstep (Q6). Thus, the first question is split to the following 6 sub-questions, aiming to fully understand the property of downstep, and to take apart declination and downstep. The results are interpretable and coherent to each other if we take the idea that downstep is mostly constrained by articulatory movement, instead of conveying linguistic meaning.</p>
<sec id="sec23">
<title>Q1: Does downstep set up a new register tone?</title>
<p>The answer is No. The original motivation of this study was whether downstep in Mandarin can be modelled as a phonetic or as a phonological tonal interaction. On the one hand, downstep was observed in West-African tone languages (<xref ref-type="bibr" rid="ref73">Welmers, 1959</xref>). The downstepped H tone defines a new ceiling for subsequent tones which was interpreted as a systematic, phonological effect, and downstep was phonologically modelled in terms of register tones (<xref ref-type="bibr" rid="ref62">Snider, 1998</xref>) or register features (<xref ref-type="bibr" rid="ref1">Akumbu, 2019</xref>). On the other hand, if downstep were a phonetic effect, the expectation is that the locally lowered F0 raises gradually back to its original register line. The present study suggests that downstep in Mandarin is indeed a phonetic tonal interaction. We observed that after a L tone, F0 does not raise back to the height of the all-H-tone sentences and lasts for several H tones decreasing in size, as has been repeatedly found in previous studies in Mandarin (<xref ref-type="bibr" rid="ref59">Shih, 1988</xref>; <xref ref-type="bibr" rid="ref75">Xu, 1999</xref>). The locally induced tonal interaction smoothly levels out such that the original reference line for a high tone in Mandarin is reached again (<xref rid="fig6" ref-type="fig">Figure 6</xref>). Thus, downstep in Mandarin is different from those in African languages. Moreover, our data showed that the effect size and the domain of the effect vary as a function of focus and prosodic boundary in Mandarin.</p>
</sec>
<sec id="sec24">
<title>Q2: Does a sentence-final focus ends downstep?</title>
<p>We predicted that the answer is no because downstep is presumably local, and pitch target of each tone is realized syllable-by-syllable as stated in PENTA model (<xref ref-type="bibr" rid="ref043">Xu et al., 2022</xref>). Indeed, we found that a late focus does not end downstep. Unexpectedly, downstep effect lasts longer in the Z-focus condition than in the wide focus condition (<xref rid="fig6" ref-type="fig">Figure 6</xref>). This is probably different from Hausa (<xref ref-type="bibr" rid="ref39">Lindau, 1986</xref>), in which downstep can be canceled in yes/no questions. It is possible that speakers try not to cause confusion, otherwise any pitch raising before the final word may increase the prominence level in that word, given that sentence-final focus is quite similar to wide focus intonation (<xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref44">Liu and Xu 2007</xref>; <xref ref-type="bibr" rid="ref042">Xu et al., 2012</xref>). Since the study on Hausa concerns question intonation, whereas ours is on final-focus, a controlled study of downstep in yes-no-questions in Mandarin would shed more light on this case.</p>
</sec>
<sec id="sec25">
<title>Q3: Is downstep eliminated by on-focus F0 raising and post-focus-compression?</title>
<p>We predicted that informative functions of intonation may override an articulatory effect. However, the results show that downstep is only weakened by on-focus F0 raising and post-focus-compression but not fully cancelled. This result is new. It indicates that downstep, as an articulatory pitch movement, is pretty robust. According to <xref ref-type="bibr" rid="ref044">Xu and Sun (2002)</xref>, the time of pitch rise can be estimated by <italic>t</italic>&#x2009;=&#x2009;89.6&#x2009;+&#x2009;8.7 <italic>d</italic> (here <italic>d</italic> stands for the change of pitch in semitone). Using this algorithm, we calculated the estimated time from the minimum F0 of the L tone to the maximum F0 of the following H tone. The exact duration of the H tone is actually longer than the estimated time (mean&#x2009;=&#x2009;33&#x2009;ms, sd&#x2009;=&#x2009;35.6). It means that the observed downstep is not because of time pressure. When the H tone is focused (YF), the exact H tone duration is 60&#x2009;ms (sd&#x2009;=&#x2009;47.8) longer than the estimated time, however downstep effect still shows (<xref rid="fig5" ref-type="fig">Figures 5</xref>, <xref rid="fig6" ref-type="fig">6</xref>). It further confirms that even in the condition of a longer H tone, downstep still applies. Thus, we draw the conclusion that informative intonation functions do not override downstep. The interaction between focus and downstep is gradual.</p>
</sec>
<sec id="sec26">
<title>Q4: How does a phonological phrase boundary interact with downstep?</title>
<p>As predicted, the pre-boundary L is lengthened at a phrase boundary (<xref rid="fig4" ref-type="fig">Figure 4</xref>), the tonal target is fully realized with higher frequency of being creaky (<xref rid="tab4" ref-type="table">Table 4</xref>), and in turn, it leads to greater downstep effect (see <xref rid="fig6" ref-type="fig">Figure 6</xref>). In wide focus condition, the L tone is lengthened about 14&#x2009;ms in the phrase boundary, with no difference in minimum F0 between the two boundary conditions (86.7 vs. 86.1 st). Instead, creaky L tone occurs more frequently in the phrase boundary condition than the syllable condition (84% vs. 67%). That might be the reason why the H tone is a little lower in the phrase boundary condition than in the syllable boundary condition (90.1 st vs. 90.4 st), showing as a greater downstep effect under the phrase-boundary condition. However, creakiness <italic>per se</italic> does not seem to cause downstep (see below, Q6).</p>
</sec>
<sec id="sec27">
<title>Q5: Do declination and downstep share the same mechanism?</title>
<p>The answer to this question actually depends on how to measure declination and downstep. It also remains controversial whether there is any separate articulatory mechanism of declination. We here take the <italic>sequential-comparison</italic> by calculating the difference of adjacent H tones (<xref rid="fig7" ref-type="fig">Figures 7</xref>, <xref rid="fig8" ref-type="fig">8</xref>). As predicted, we can see that downstep and declination come from different articulatory control. However, it is not because downstep is local whereas declination is global, rather downstep lasts for several syllables as well. It is because the downstep effect shows in a larger scale and in a more robust manner than declination. It is possible that there is some underlying articulatory control on declination, however, it is pretty weak and vulnerable to be overridden by varied reasons. We are in agreement with other studies (<xref ref-type="bibr" rid="ref75">Xu, 1999</xref>; <xref ref-type="bibr" rid="ref033">Shih, 2000</xref>; <xref ref-type="bibr" rid="ref045">Yuan and Liberman, 2014</xref>), showing that the general global downtrend, as modelled with a top and bottom regression line of intonation, is a combined effect from different functions. We further suggest not to just take the global downtrend in an abstract way, but to analyze it with full consideration of local tonal interactions.</p>
</sec>
<sec id="sec28">
<title>Q6: Is creaky voice the cause of downstep?</title>
<p>Downstep is caused by a L tone, which is usally creaky in Mandarin (<xref ref-type="bibr" rid="ref019">Kuang, 2017</xref>). Is it possible that creaky voice is the main cause of downstep? In our study we found that the L tone is more likely to be creaky when it is under focus and before a phrase boundary (<xref rid="tab4" ref-type="table">Table 4</xref>). It confirms the claim by <xref ref-type="bibr" rid="ref019">Kuang (2017)</xref> that creaky voice correlates with low pitch target. As discussed in Q4, more creaky L tones at a phrase boundary causes greater downstep effect. However, normal L tone causes roughly the same degree of downstep, as showed in the LMM that creakiness does not have any effect on the maximum F0 of the following H tone. No correlation is found between the duration of the creaky part in L tones and the pitch height in the following H tones (<xref rid="fig9" ref-type="fig">Figure 9</xref>). It indicates that creaky voice is probably not the direct cause of downstep. A normal L tone also causes downstep. However, a creaky L tone leads to a greater downstep effect.</p>
</sec>
</sec>
<sec id="sec29">
<title>Does a phrase boundary block post-low-bouncing (Q7)?</title>
<p>According to the <italic>balance-perturbation hypothesis</italic> (<xref ref-type="bibr" rid="ref54">Prom-on et al., 2012</xref>), we predicted that post-low-bouncing is weakened if the L tone is at a phrase boundary. It is indeed the case, as shown in <xref rid="fig7" ref-type="fig">Figure 7</xref>. They hypothesized that after producing a very low F0, the extrinsic laryngeal muscles (e.g., sternohyoids) stop contracting to maintain the balance between the two antagonistic forces in the intrinsic laryngeal muscles. When the L tone is focused, the extra force may cause a sudden increase of the vocal fold tension, resulting in the raise in F0 in the following H tone. We here see that when the L tone is before a prosodic phrase boundary, it then probably gives a little more time to release the tension between the extrinsic and intrinsic laryngeal muscles. This would explain the difference in size of post-low bouncing found in our data. In line with <xref ref-type="bibr" rid="ref54">Prom-on et al. (2012)</xref>, we also found that post-low-bouncing occurs in H tones when the L tone is under focus. In their study, neutral tones after a L tone show post-low-bouncing. The reason might lie in the fact that post-focal words are weakened in intensity and compressed in F0. The weakened H tones at post-focal position might share some similar mechanism with weak articulatory movement in the neutral tones.</p>
<p>At last, we here briefly introduce some preliminary findings in the current study, relating to <xref ref-type="bibr" rid="ref47">Moisik et al. (2014)</xref>. They have found that low F0 tone targets in Mandarin can not only be reached by lowering the larynx, but also by combining the raise of larynx height and laryngeal constriction, which may lead to creakiness in the low tone. In the L tone, the amount of F0 lowering correlates with larynx lowering in male speakers (r&#x2009;=&#x2009;0.73 and 0.86), while the female speaker uses larynx raising (<italic>r</italic>&#x2009;=&#x2009;0.13; Figures 11-13, pp. 39 in their study). In our study, the minimum F0 of the low tone (X) is positively correlated to the maximum F0 of the following H tone (Y) in the male speakers (wide focus: <italic>y</italic>&#x2009;=&#x2009;&#x2212;2&#x2009;+&#x2009;0.99<italic>x</italic>, <italic>r</italic><sup>2</sup>&#x2009;=&#x2009;0.66; X-focus condition: <italic>y</italic>&#x2009;=&#x2009;9.4&#x2009;+&#x2009;0.87<italic>x</italic>, r<sup>2</sup>&#x2009;=&#x2009;0.736), but not in the female speakers (wide focus: <italic>y</italic>&#x2009;=&#x2009;65&#x2009;+&#x2009;0.27<italic>x</italic>, r<sup>2</sup>&#x2009;=&#x2009;0.073; X-focus condition: <italic>y</italic>&#x2009;=&#x2009;79&#x2009;+&#x2009;0.12<italic>x</italic>, r<sup>2</sup>&#x2009;=&#x2009;0.01). To fully understand the anatomical process in downstep, articulatory studies considering gender difference are required.</p>
</sec>
</sec>
<sec id="sec30" sec-type="conclusions">
<title>Conclusion</title>
<p>To answer all the research questions concerning the interaction of focus/boundary with downstep/post-low-bouncing, we can draw the following conclusions.</p>
<p>In the wide focus condition, the downstep effect lasted for 3 syllables and gradually reached back to the all-H tone reference line. Downstep thus does not set up a new reference line in Mandarin (Q1). A sentence-final focus makes the downstep effect last for 5 syllables (Q2). When the H tone right after the low tone was focused (YF), on-focus F0 raising and post-focus-compression (PFC) weakened downstep (Q3). A phrase boundary strengthened downstep (Q4). We further analyzed downstep by measuring the F0 drop between the two H tones surrounding the L tone (<italic>sequential-comparison</italic>). Comparing it with F0 drop in all-H sentences, it showed that the downstep effect was much greater and more robust than declination (Q5). However, creaky voice in the L tone was not the direct cause of downstep (Q6). At last, when the L tone was under focus (XF), it caused a post-low-bouncing effect on the following H tones and lasted for about 3 syllables with F0 dropping back gradually. Moreover, post-low-bouncing is weakened by a phonological phrase boundary (Q7).</p>
<p>In general, this study showed that downstep and post-low-bouncing, as articulatory controls and local tonal interaction effects, interact with the execution of sentence-level pragmatic functions like focus and prosodic boundary. Pragmatic effects do not cancel or override articulatory effects, but affect the size and domain of the tonal interactions.</p>
</sec>
<sec id="sec31" sec-type="data-availability">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="sec32">
<title>Ethics statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="sec33">
<title>Author contributions</title>
<p>FK initiated the general research question. BW, FK, and SG designed the experiment together. BW checked the labeling of the wav files, wrote the paper, and finalized the data analysis, closely working together with FK. SG did a preliminary graphic and statistical analysis. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec34" sec-type="funding-information">
<title>Funding</title>
<p>The experiment was supported by SFB632 &#x201C;Information Structure&#x201D; at Potsdam University. The writing procedure was further supported by Social Science Foundation of China to BW (18BYY079). Goethe University Frankfurt provided financial support for open access support, and a DFG grant KU 2323-4/1 to FK supported further assistance.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>SG is employed by i2x GmbH.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The research was initiated while SG was affiliated with Potsdam University, and hence the research was conducted without a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="sec230" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2022.884102/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fpsyg.2022.884102/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
</body>
<back>
<ack>
<p>We thank Xiaxia Zhang and Qian Wu for running the experiment and labeling the speech data. We also thank Yi Xu for constructive discussions and suggestions. Part of the preliminary results were presented at the TAL conference in 2018 (<xref ref-type="bibr" rid="ref77">Wang et al., 2018a</xref>), we thank the audience.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Akumbu</surname> <given-names>P. W.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>A featural analysis of mid and downstepped high tone in Babanki</article-title>,&#x201D; in <source>Theory and Description in African Linguistics: Selected Papers From the 47th Annual Conference on African Linguistics</source>. eds. <person-group person-group-type="editor"><name><surname>Clem</surname> <given-names>E.</given-names></name> <name><surname>Jenks</surname> <given-names>P.</given-names></name> <name><surname>Sande</surname> <given-names>H.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Language Science Press</publisher-name>), <fpage>3</fpage>&#x2013;<lpage>20</lpage>.</citation></ref>
<ref id="ref002"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arvaniti</surname> <given-names>A.</given-names></name> <name><surname>Fletcher</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>The Autosegmental-Metrical theory of intonational phonology</article-title>. <source>Oxford Handbooks in linguistics. The Oxford handbook of language prosody</source>, <fpage>77</fpage>&#x2013;<lpage>95</lpage>.</citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atkinson</surname> <given-names>J. E.</given-names></name></person-group> (<year>1978</year>). <article-title>Correlation analysis of the physiological factors controlling fundamental voice frequency</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>63</volume>, <fpage>211</fpage>&#x2013;<lpage>222</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.381716</pub-id>, PMID: <pub-id pub-id-type="pmid">632414</pub-id></citation></ref>
<ref id="ref003"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baer</surname> <given-names>T.</given-names></name></person-group> (<year>1979</year>). <article-title>Reflex activation of laryngeal muscles by sudden induced subglottal pressure changes</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>65</volume>, <fpage>1271</fpage>&#x2013;<lpage>1275</lpage>.</citation></ref>
<ref id="ref3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bates</surname> <given-names>D.</given-names></name> <name><surname>Maechler</surname> <given-names>M.</given-names></name> <name><surname>Bolker</surname> <given-names>B.</given-names></name> <name><surname>Walker</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <source>lme4: Linear Mixed-Effects Models Using Eigen and S4 (R Package Version 1.112)</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name>.</citation></ref>
<ref id="ref4"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Boersma</surname> <given-names>P.</given-names></name> <name><surname>Weenink</surname> <given-names>D.</given-names></name></person-group> (<year>2013&#x2013;2022</year>). Available at: <ext-link xlink:href="http://www.fon.hum.uva.nl/praat/" ext-link-type="uri">http://www.fon.hum.uva.nl/praat/</ext-link></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Byrd</surname> <given-names>D.</given-names></name></person-group> (<year>2000</year>). <article-title>Articulatory vowel lengthening and coordination at phrasal junctures</article-title>. <source>Phonetica</source> <volume>57</volume>, <fpage>3</fpage>&#x2013;<lpage>16</lpage>. doi: <pub-id pub-id-type="doi">10.1159/000028456</pub-id>, PMID: <pub-id pub-id-type="pmid">10867568</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Chao</surname> <given-names>Y. R.</given-names></name></person-group> (<year>1968</year>). <source>A Grammar of Spoken Chinese</source>. <publisher-loc>Berkeley, CA</publisher-loc>: <publisher-name>University of California Press</publisher-name>.</citation></ref>
<ref id="ref004"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.-Y.</given-names></name></person-group> (<year>2006</year>). <article-title>Durational adjustment under corrective focus in Standard Chinese</article-title>. <source>Journal of Phonetics</source> <volume>34</volume>, <fpage>176</fpage>&#x2013;<lpage>201</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2005.05.002</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.-Y.</given-names></name> <name><surname>Gussenhoven</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>Emphasis and tonal implementation in standard Chinese</article-title>. <source>J. Phon.</source> <volume>36</volume>, <fpage>724</fpage>&#x2013;<lpage>746</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2008.06.003</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.-Y.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2006</year>). <article-title>Production of weak elements in speech -- evidence from f0 patterns of neutral tone in standard Chinese</article-title>. <source>Phonetica</source> <volume>63</volume>, <fpage>47</fpage>&#x2013;<lpage>75</lpage>. doi: <pub-id pub-id-type="doi">10.1159/000091406</pub-id>, PMID: <pub-id pub-id-type="pmid">16514275</pub-id></citation></ref>
<ref id="ref005"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>A.</given-names></name> <name><surname>Collier</surname> <given-names>R.</given-names></name> <name><surname>&#x2019;t Hart</surname> <given-names>J.</given-names></name></person-group> (<year>1982</year>). <article-title>Declination: Construct or Intrinsic Feature of Speech Pitch?</article-title> <source>Phonetica</source> <volume>39</volume>, <fpage>254</fpage>&#x2013;<lpage>273</lpage>.</citation></ref>
<ref id="ref006"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collier</surname> <given-names>R.</given-names></name></person-group> (<year>1975</year>). <article-title>Physiological correlates of intonation patterns</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>58</volume>, <fpage>249</fpage>&#x2013;<lpage>255</lpage>.</citation></ref>
<ref id="ref007"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Connell</surname> <given-names>B.</given-names></name></person-group> (<year>2001</year>). <source>Downdrift, downstep, and declination</source>. <publisher-loc>Paper presented at the Typology of African Prosodic Systems Workshop</publisher-loc>, <publisher-name>Bielefeld University, Germany</publisher-name>.</citation></ref>
<ref id="ref9"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Connell</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). &#x201C;<article-title>Downstep</article-title>&#x201D; in <source>Companion to Phonology</source>. eds. <person-group person-group-type="editor"><name><surname>Oostendorp</surname> <given-names>M. V.</given-names></name> <name><surname>Ewen</surname> <given-names>C. J.</given-names></name> <name><surname>Hume</surname> <given-names>E.</given-names></name> <name><surname>Rice</surname> <given-names>K.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Blackwell Publishing</publisher-name>), <fpage>824</fpage>&#x2013;<lpage>847</lpage>.</citation></ref>
<ref id="ref10"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Connell</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Tone and intonation in Mambila</article-title>,&#x201D; in <source>Intonation in African Tone Languages</source>. eds. <person-group person-group-type="editor"><name><surname>Downing</surname> <given-names>L. J.</given-names></name> <name><surname>Rialland</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>De Gruyter</publisher-name>), <fpage>132</fpage>&#x2013;<lpage>166</lpage>.</citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>W. E.</given-names></name> <name><surname>Eady</surname> <given-names>S. J.</given-names></name> <name><surname>Mueller</surname> <given-names>P. R.</given-names></name></person-group> (<year>1985</year>). <article-title>Acoustical aspects of contrastive stress in question-answer contexts</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>77</volume>, <fpage>2142</fpage>&#x2013;<lpage>2156</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.392372</pub-id>, PMID: <pub-id pub-id-type="pmid">4019901</pub-id></citation></ref>
<ref id="ref008"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>W. E.</given-names></name> <name><surname>Sorensen</surname> <given-names>J. M.</given-names></name></person-group> (<year>1977</year>). <article-title>Fundamental frequency contours at syntactic boundaries</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>62</volume>, <fpage>683</fpage>&#x2013;<lpage>692</lpage>.</citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Courtenay</surname> <given-names>K.</given-names></name></person-group> (<year>1971</year>). <article-title>Yoruba: A'terraced-level'language with three tonemes</article-title>. <source>Studies in African Linguistics</source> <volume>2</volume>, <fpage>239</fpage>&#x2013;<lpage>255</lpage>.</citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Jong</surname> <given-names>K. J.</given-names></name></person-group> (<year>1995</year>). <article-title>The supraglottal articulation of prominence in English: linguistic stress as localized hyperarticulation</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>97</volume>, <fpage>491</fpage>&#x2013;<lpage>504</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.412275</pub-id>, PMID: <pub-id pub-id-type="pmid">7860828</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Pijper</surname> <given-names>J. R.</given-names></name> <name><surname>Sandeman</surname> <given-names>A. A.</given-names></name></person-group> (<year>1994</year>). <article-title>On the perceptual strength of prosodic boundaries and its relation to suprasegmental cues</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>96</volume>, <fpage>2037</fpage>&#x2013;<lpage>2047</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.410145</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiCanio</surname> <given-names>C.</given-names></name> <name><surname>Benn</surname> <given-names>J.</given-names></name> <name><surname>Castillo Garc&#x00ED;a</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Disentangling the effects of position and utterance-level declination on the production of complex tones in Yolox&#x00F3;chitl Mixtec</article-title>. <source>Lang. Speech</source> <volume>64</volume>, <fpage>515</fpage>&#x2013;<lpage>557</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0023830920939132</pub-id>, PMID: <pub-id pub-id-type="pmid">32689854</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edmondson</surname> <given-names>J. A.</given-names></name> <name><surname>Esling</surname> <given-names>J. H.</given-names></name></person-group> (<year>2006</year>). <article-title>The valves of the throat and their functioning in tone, vocal register and stress: laryngoscopic case studies</article-title>. <source>Phonology</source> <volume>23</volume>, <fpage>157</fpage>&#x2013;<lpage>191</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S095267570600087X</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>F&#x00E9;ry</surname> <given-names>C.</given-names></name> <name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name></person-group> (<year>2008</year>). <article-title>Pitch accent scaling on given, new and focused constituents in German</article-title>. <source>J. Phon.</source> <volume>36</volume>, <fpage>680</fpage>&#x2013;<lpage>703</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2008.05.001</pub-id></citation></ref>
<ref id="ref009"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Declination and boundary effect in Cantonese declarative sentence</article-title>&#x201D; in <source>Paper presented at the 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP)</source></citation></ref>
<ref id="ref010"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Gelfer</surname> <given-names>C. E.</given-names></name> <name><surname>Harris</surname> <given-names>K. S.</given-names></name> <name><surname>Collier</surname> <given-names>R.</given-names></name> <name><surname>Baer</surname> <given-names>T.</given-names></name></person-group> (<year>1983</year>). &#x201C;<article-title>The Denver Center for the Performing Atrs, Inc.</article-title>&#x201D; in <source>Is declination actively controlled <italic>Vocal Fold Physiology: Biomechanics, Acoustics and Phonatory Control</italic></source> (<publisher-loc>Denver, Colorado</publisher-loc>), <fpage>113</fpage>&#x2013;<lpage>126</lpage>.</citation></ref>
<ref id="ref011"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Gelman</surname> <given-names>A.</given-names></name> <name><surname>Hill</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <source>Data analysis using regression and multilevel/hierarchical models</source>: Cambridge university press.</citation></ref>
<ref id="ref19"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Genzel</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <source>Lexical and post-lexical tones in Akan</source>. <comment>(PhD Thesis.)</comment>, <publisher-name>Universit&#x00E4;t Potsdam</publisher-name>, <publisher-loc>Potsdam</publisher-loc>.</citation></ref>
<ref id="ref20"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Genzel</surname> <given-names>S.</given-names></name> <name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name></person-group> (<year>2011</year>). &#x201C;Phonetic Realization of Automatic (Downdrift) and non-automatic Downstep in Akan.&#x201D; in <italic>Paper presented at the Proceedings of the XVII ICPhS</italic>, Hong Kong.</citation></ref>
<ref id="ref012"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gerratt</surname> <given-names>B. R.</given-names></name> <name><surname>Kreiman</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>Toward a taxonomy of nonmodal phonation</article-title>. <source>Journal of Phonetics</source> <volume>29</volume>, <fpage>365</fpage>&#x2013;<lpage>381</lpage>.</citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>W.</given-names></name> <name><surname>Lee</surname> <given-names>T.</given-names></name></person-group> (<year>2009</year>). <article-title>Effects of tone and emphatic focus on F0 contours of Cantonese speech: A comparison with standard Chinese</article-title>. <source>Chin. J. Phon.</source> <volume>2</volume>, <fpage>133</fpage>&#x2013;<lpage>147</lpage>.</citation></ref>
<ref id="ref013"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Gussenhoven</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <source>The Phonology of Tone and Intonation</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="ref24"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hirose</surname> <given-names>H.</given-names></name></person-group> (<year>1997</year>). &#x201C;<article-title>Investigating the physiology of laryngeal structures</article-title>,&#x201D; in <source>The handbook of phonetic sciences</source>. eds. <person-group person-group-type="editor"><name><surname>Hardcastle</surname> <given-names>W. J.</given-names></name> <name><surname>Laver</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Blackwell</publisher-name>), <fpage>116</fpage>&#x2013;<lpage>136</lpage>.</citation></ref>
<ref id="ref014"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hirschberg</surname> <given-names>J.</given-names></name> <name><surname>Pierrehumbert</surname> <given-names>J. B.</given-names></name></person-group> (<year>1986</year>). &#x201C;<article-title>The intonational structuring of discourse</article-title>&#x201D; in <source>Paper presented at the 24th Annual Meeting of the Association of Computational Linguistics</source></citation></ref>
<ref id="ref015"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollien</surname> <given-names>H.</given-names></name></person-group> (<year>1974</year>). <article-title>On vocal registers</article-title>. <source>Journal of Phonetics</source> <volume>2</volume>, <fpage>125</fpage>&#x2013;<lpage>143</lpage>.</citation></ref>
<ref id="ref016"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hollien</surname> <given-names>H.</given-names></name></person-group> (<year>1983</year>). &#x201C;<article-title>In search of vocal frequency control mechanisms</article-title>&#x201D; in <source>Vocal Fold Physiology: Comtemporary Research and Clinical Issues</source>. eds. <person-group person-group-type="editor"><name><surname>Bless</surname> <given-names>D. M.</given-names></name> <name><surname>Abbs</surname> <given-names>J. H.</given-names></name></person-group> (<publisher-name>College-Hill Press</publisher-name>), <fpage>361</fpage>&#x2013;<lpage>367</lpage>.</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hombert</surname> <given-names>J.-M.</given-names></name></person-group> (<year>1974</year>). <article-title>Universals of downdrift: their phonetic basis and significance for a theory of tone</article-title>. <source>Studies in African Linguistics</source> <volume>5</volume>, <fpage>169</fpage>&#x2013;<lpage>183</lpage>.</citation></ref>
<ref id="ref27"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Honda</surname> <given-names>K.</given-names></name></person-group> (<year>1995</year>). &#x201C;<article-title>Laryngeal and extra-laryngeal mechanisms of F0 control</article-title>,&#x201D; in <source>Producing Speech: Contemporary Issues</source>. eds. <person-group person-group-type="editor"><name><surname>Bell-Berti</surname> <given-names>F.</given-names></name> <name><surname>Raphael</surname> <given-names>L. J.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>American Institute of Physics</publisher-name>), <fpage>215</fpage>&#x2013;<lpage>232</lpage>.</citation></ref>
<ref id="ref017"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huffman</surname> <given-names>M. K.</given-names></name></person-group> (<year>2005</year>). <article-title>Segmental and prosodic effects on coda glottalization</article-title>. <source>Journal of Phonetics</source> <volume>33</volume>, <fpage>335</fpage>&#x2013;<lpage>362</lpage>.</citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hyman</surname> <given-names>L. M.</given-names></name> <name><surname>Leben</surname> <given-names>W. R.</given-names></name></person-group> (<year>2017</year>). <article-title>Word prosody II: tone systems</article-title>. <source>UC Berkeley PhonLab Annu. Rep.</source> <volume>13</volume>, <fpage>178</fpage>&#x2013;<lpage>209</lpage>. doi: <pub-id pub-id-type="doi">10.5070/P7131040752</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ishihara</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>Major phrase, focus intonation, multiple spell-out (MaP, FI, MSO)</article-title>. <source>Linguistic Rev.</source> <volume>24</volume>, <fpage>137</fpage>&#x2013;<lpage>167</lpage>. doi: <pub-id pub-id-type="doi">10.1515/TLR.2007.006</pub-id></citation></ref>
<ref id="ref018"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Keating</surname> <given-names>P. A.</given-names></name> <name><surname>Garellek</surname> <given-names>M.</given-names></name> <name><surname>Kreiman</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Acoustic properties of different kinds of creaky voice</article-title>&#x201D; in <source>Paper presented at the ICPhS</source></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krivokapic</surname> <given-names>J.</given-names></name> <name><surname>Byrd</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Prosodic boundary strength: an articulatory and perceptual study</article-title>. <source>J. Phon.</source> <volume>40</volume>, <fpage>430</fpage>&#x2013;<lpage>442</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2012.02.011</pub-id>, PMID: <pub-id pub-id-type="pmid">23441103</pub-id></citation></ref>
<ref id="ref019"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuang</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Covariation between voice quality and pitch: Revisiting the case of Mandarin creaky voice</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>142</volume>, <fpage>1693</fpage>&#x2013;<lpage>1706</lpage>.</citation></ref>
<ref id="ref020"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuang</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>The influence of tonal categories and prosodic boundaries on the creakiness in Mandarin</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>143</volume>:<fpage>EL509-EL515</fpage>.</citation></ref>
<ref id="ref32"><citation citation-type="book"><person-group person-group-type="author"><name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Tone and intonation in Akan</article-title>,&#x201D; in <source>Intonation in African Tone Languages</source>. eds. <person-group person-group-type="editor"><name><surname>Downing</surname> <given-names>L.</given-names></name> <name><surname>Rialland</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Mouton de Gruyter</publisher-name>), <fpage>89</fpage>&#x2013;<lpage>129</lpage>.</citation></ref>
<ref id="ref33"><citation citation-type="book"><person-group person-group-type="author"><name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name> <name><surname>Calhoun</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Prosodic encoding of information structure: a typological perspective</article-title>,&#x201D; in <source>The Oxford Handbook of Language Prosody</source>. eds. <person-group person-group-type="editor"><name><surname>Gussenhoven</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>).</citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ladd</surname> <given-names>D. R.</given-names></name></person-group> (<year>1988</year>). <article-title>Declination "reset" and the hierarchical organization of utterances</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>84</volume>, <fpage>530</fpage>&#x2013;<lpage>544</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.396830</pub-id></citation></ref>
<ref id="ref021"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ladd</surname> <given-names>D. R.</given-names></name></person-group> (<year>2008</year>). <source>Intonational phonology</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="ref022"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ladefoged</surname> <given-names>P.</given-names></name></person-group> (<year>1973</year>). <article-title>The features of the larynx</article-title>. <source>Journal of Phonetics</source> <volume>1</volume>, <fpage>73</fpage>&#x2013;<lpage>83</lpage>.</citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laniran</surname> <given-names>Y. O.</given-names></name> <name><surname>Clements</surname> <given-names>G. N.</given-names></name></person-group> (<year>2003</year>). <article-title>Downstep and hihg raising: interacting factors in Yoruba toe production</article-title>. <source>J. Phon.</source> <volume>31</volume>, <fpage>203</fpage>&#x2013;<lpage>250</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0095-4470(02)00098-0</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Leben</surname> <given-names>W. R.</given-names></name></person-group> (<year>2014</year>). The Nature(s) of Downstep. Paper Presented at the SLAO/1er Colloque International, Humboldt Kolleg Abidjan.</citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>A.</given-names></name> <name><surname>Prom-on</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Pre-low raising in Cantonese and Thai: effects of speech rate and vowel quantity</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>149</volume>, <fpage>179</fpage>&#x2013;<lpage>190</lpage>. doi: <pub-id pub-id-type="doi">10.1121/10.0002976</pub-id>, PMID: <pub-id pub-id-type="pmid">33514164</pub-id></citation></ref>
<ref id="ref023"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Liberman</surname> <given-names>M. Y.</given-names></name> <name><surname>Pierrehumbert</surname> <given-names>J.</given-names></name></person-group> (<year>1984</year>). &#x201C;<article-title>Intonational invariance under changes in pich range and length</article-title>&#x201D; in <source>Language sound structure</source>. eds. <person-group person-group-type="editor"><name><surname>Aronoff</surname> <given-names>M.</given-names></name> <name><surname>O</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT</publisher-name>), <fpage>157</fpage>&#x2013;<lpage>233</lpage>.</citation></ref>
<ref id="ref024"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Lieberman</surname> <given-names>P.</given-names></name></person-group> (<year>1967</year>). <source><italic>Intonation, perception, and language</italic>: Cambridge</source>, <publisher-loc>MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="ref025"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lieberman</surname> <given-names>P.</given-names></name> <name><surname>Tseng</surname> <given-names>C.</given-names></name></person-group> (<year>1980</year>). <article-title>On the fall of the declination theory: breath-group versus &#x201C;declination&#x201D; as the base form for intonation</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>67</volume>:<fpage>S63</fpage>.</citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>M.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name></person-group> (<year>1980</year>). <article-title>Beijinghua qingsheng de shengxue xingzhi (The acoustic nature of the mandarin neutral tone)</article-title>. <source>Fangyan (Dialect)</source> <volume>3</volume>, <fpage>166</fpage>&#x2013;<lpage>178</lpage>.</citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lindau</surname> <given-names>M.</given-names></name></person-group> (<year>1986</year>). <article-title>Testing a model of intonation in a tone language</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>80</volume>, <fpage>757</fpage>&#x2013;<lpage>764</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.393950</pub-id>, PMID: <pub-id pub-id-type="pmid">3760329</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Lindblom</surname> <given-names>B.</given-names></name></person-group> (<year>1990</year>). <source>Explaining phonetic variation: A sketch of the H&#x0026;H theory speech production and speech modelling</source>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>403</fpage>&#x2013;<lpage>439</lpage>.</citation></ref>
<ref id="ref41"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Lindblom</surname> <given-names>B.</given-names></name></person-group> (<year>2009</year>). F0 lowering, creaky voice, and glottal stop: Jan Gauf- fin's account of how the larynx works in speech. Paper presented at the Fonetik 2009, Stockholm.</citation></ref>
<ref id="ref42"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Lindqvist-Gauffin</surname> <given-names>J.</given-names></name></person-group> (<year>1969</year>). <source>Laryngeal Mechanisms in speech. Quarterly Progress and Status Report, Speech Transmission Laboratory</source>. <publisher-loc>Stockholm</publisher-loc>: <publisher-name>Royal Institute of Technology</publisher-name>, <fpage>26</fpage>&#x2013;<lpage>31</lpage>.</citation></ref>
<ref id="ref43"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Lindqvist-Gauffin</surname> <given-names>J.</given-names></name></person-group> (<year>1972</year>). <source>A Descriptive Model of Laryngeal Articulation in Speech. Quarterly Progress and Status Report, Speech Transmission Laboratory</source>. <publisher-loc>Stockholm</publisher-loc>: <publisher-name>Royal Institute of Technology</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="ref026"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2005</year>). <article-title>Parallel encoding of focus and interrogative meaning in Mandarin intonation</article-title>. <source>Phonetica</source> <volume>62</volume>, <fpage>70</fpage>&#x2013;<lpage>87</lpage>.</citation></ref>
<ref id="ref44"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2007</year>). &#x201C;Question intonation as affected by word stress and focus in English.&#x201D; in <italic>Paper presented at the 16th International Congress of Phonetic Sciences</italic>, Saarbrucken.</citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Tian</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Multiple prosodic meanings are conveyed through separate pitch ranges: evidence from perception of focus and surprise in mandarin Chinese</article-title>. <source>Cogn. Affect. Behav. Neurosci.</source> <volume>21</volume>, <fpage>1164</fpage>&#x2013;<lpage>1175</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13415-021-00930-9</pub-id></citation></ref>
<ref id="ref027"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Maeda</surname> <given-names>S.</given-names></name></person-group> (<year>1976</year>). <source>A characterization of American English Intonation</source>. <comment>(Doctoral Dissertation)</comment> <publisher-name>MIT</publisher-name>.</citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moisik</surname> <given-names>S. R.</given-names></name> <name><surname>Esling</surname> <given-names>J. H.</given-names></name></person-group> (<year>2014</year>). <article-title>Modeling the biomechanical influence of epilaryngeal stricture on the vocal folds: A low-dimensional model of vocal&#x2013;ventricular fold coupling</article-title>. <source>J. Speech Lang. Hear. Res.</source> <volume>57</volume>, <fpage>S687</fpage>&#x2013;<lpage>S704</lpage>. doi: <pub-id pub-id-type="doi">10.1044/2014_JSLHR-S-12-0279</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moisik</surname> <given-names>S. R.</given-names></name> <name><surname>Lin</surname> <given-names>H.</given-names></name> <name><surname>Esling</surname> <given-names>J. H.</given-names></name></person-group> (<year>2014</year>). <article-title>A study of laryngeal gestures in mandarin citation tones using simultaneous laryngoscopy and laryngeal ultrasound (SLLUS)</article-title>. <source>J. Int. Phon. Assoc.</source> <volume>44</volume>, <fpage>21</fpage>&#x2013;<lpage>58</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S0025100313000327</pub-id></citation></ref>
<ref id="ref028"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murry</surname> <given-names>T.</given-names></name></person-group> (<year>1971</year>). <article-title>Subglottal pressure and airflow measures during vocal fry phonation</article-title>. <source>Journal of Speech and Hearing Research</source> <volume>14</volume>, <fpage>544</fpage>&#x2013;<lpage>551</lpage>.</citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakai</surname> <given-names>S.</given-names></name> <name><surname>Kunnari</surname> <given-names>S.</given-names></name> <name><surname>Turk</surname> <given-names>A.</given-names></name> <name><surname>Suomi</surname> <given-names>K.</given-names></name> <name><surname>Ylitalo</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>Utterance-final lengthening and quantity in northern Finnish</article-title>. <source>J. Phon.</source> <volume>37</volume>, <fpage>29</fpage>&#x2013;<lpage>45</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2008.08.002</pub-id></citation></ref>
<ref id="ref029"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakajima</surname> <given-names>S.</given-names></name> <name><surname>Allen</surname> <given-names>J. F.</given-names></name></person-group> (<year>1993</year>). <article-title>A study on prosody and discourse structure in cooperative dialogues</article-title>. <source>Phonetica</source> <volume>50</volume>, <fpage>197</fpage>&#x2013;<lpage>210</lpage>.</citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Odden</surname> <given-names>D.</given-names></name></person-group> (<year>1986</year>). <article-title>On the role of the obligatory contour principle in phonological theory</article-title>. <source>Language</source> <volume>62</volume>, <fpage>353</fpage>&#x2013;<lpage>383</lpage>. doi: <pub-id pub-id-type="doi">10.2307/414677</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohala</surname> <given-names>J. J.</given-names></name></person-group> (<year>1972</year>). <article-title>How is pitch lowered?</article-title> <source>J. Acoust. Soc. Am.</source> <volume>52</volume>:<fpage>124</fpage>. doi: <pub-id pub-id-type="doi">10.1121/1.1981808</pub-id></citation></ref>
<ref id="ref030"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pierrehumbert</surname> <given-names>J.</given-names></name></person-group> (<year>1979</year>). <article-title>The perception of fundamental frequency declination</article-title>. <source>Journal of the Acoustics Society of America</source> <volume>66</volume>, <fpage>363</fpage>&#x2013;<lpage>379</lpage>.</citation></ref>
<ref id="ref52"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Pierrehumbert</surname> <given-names>J.</given-names></name></person-group> (<year>1980</year>). <source>The phonology and phonetics of english intonation</source>. (<comment>Ph. D. doctoral thesis</comment>), <publisher-name>Massachusetts Institute of Technology</publisher-name>, <publisher-loc>Cambridge</publisher-loc>.</citation></ref>
<ref id="ref53"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Pierrehumbert</surname> <given-names>J. B.</given-names></name> <name><surname>Beckman</surname> <given-names>M. E.</given-names></name></person-group> (<year>1988</year>). <source>Japanese Tone Structure</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prom-on</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Post-low bouncing in mandarin Chinese: acoustic analysis and computational modeling</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>132</volume>, <fpage>421</fpage>&#x2013;<lpage>432</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.4725762</pub-id>, PMID: <pub-id pub-id-type="pmid">22779489</pub-id></citation></ref>
<ref id="ref031"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Rialland</surname> <given-names>A.</given-names></name> <name><surname>Som&#x00E9;</surname> <given-names>A.-P.</given-names></name></person-group> (<year>2011</year>). &#x201C;<article-title>Downstep and linguistic scaling in Dagara-Wul&#x00E9;</article-title>&#x201D; in <source>Tones and features: Phonetic and Phonological Perspectives</source>. eds. <person-group person-group-type="editor"><name><surname>Goldsmith</surname> <given-names>J. A.</given-names></name> <name><surname>Hume</surname> <given-names>W. E.</given-names></name> <name><surname>Wetzels</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Berlin &#x0026; New York</publisher-loc>: <publisher-name>Mounton De Gruyter</publisher-name>), <fpage>108</fpage>&#x2013;<lpage>134</lpage>.</citation></ref>
<ref id="ref55"><citation citation-type="book"><person-group person-group-type="author"><collab id="coll1">R Core Team</collab></person-group>. (<year>2016</year>). <source>R: A language and environment for statistical computing.</source> <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation forStatisticalComputing</publisher-name>.</citation></ref>
<ref id="ref032"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Selkirk</surname> <given-names>E.</given-names></name></person-group> (<year>2011</year>). <article-title>The syntax-phonology interface</article-title>. In <person-group person-group-type="editor"><name><surname>Goldsmith</surname> <given-names>J. A.</given-names></name></person-group>, <person-group person-group-type="editor"><name><surname>Riggle</surname> <given-names>J.</given-names></name></person-group> and <person-group person-group-type="editor"><name><surname>Yu</surname> <given-names>A. C. L.</given-names></name></person-group> (Eds.), <source>The handbook of phonological theory</source> <edition>(2nd ed.)</edition> (Vol. <volume>2</volume>, pp. <fpage>435</fpage>&#x2013;<lpage>484</lpage>). <publisher-loc>Oxford</publisher-loc>: <publisher-name>Wiley-Blackwell</publisher-name>.</citation></ref>
<ref id="ref57"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Selkirk</surname> <given-names>E.</given-names></name> <name><surname>Tateishi</surname> <given-names>K.</given-names></name></person-group> (<year>1991</year>). &#x201C;<article-title>Syntax and downstep in Japanese</article-title>,&#x201D; in <source>Interdisciplinary Approaches to Language: Essays in Honor of S.-Y. Kuroda</source>. eds. <person-group person-group-type="editor"><name><surname>Georgopoulos</surname> <given-names>C.</given-names></name> <name><surname>Ishihara</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Kluwer Academic Publishers</publisher-name>), <fpage>519</fpage>&#x2013;<lpage>543</lpage>.</citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>J.</given-names></name></person-group> (<year>1994</year>). <article-title>Hanyu yudiao gouzao he yudiao leixing (intonation structures and patterns in mandarin) (in Chinese)</article-title>. <source>Fangyan (Dialect)</source> <volume>3</volume>, <fpage>221</fpage>&#x2013;<lpage>228</lpage>.</citation></ref>
<ref id="ref59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shih</surname> <given-names>C. L.</given-names></name></person-group> (<year>1988</year>). <article-title>Tone and intonation in mandarin. Working papers, Cornell phonetics</article-title>. <source>Laboratory</source> <volume>3</volume>, <fpage>83</fpage>&#x2013;<lpage>109</lpage>.</citation></ref>
<ref id="ref033"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Shih</surname> <given-names>C.</given-names></name></person-group> (<year>2000</year>). &#x201C;<article-title>A declination model of Mandarin Chinese</article-title>&#x201D; in <source>Intonation: Analysis, Modelling and Technology</source>. ed. <person-group person-group-type="editor"><name><surname>Botinis</surname> <given-names>A.</given-names></name></person-group> (<publisher-name>Kluwer Academic Publishers</publisher-name>), <fpage>243</fpage>&#x2013;<lpage>268</lpage>.</citation></ref>
<ref id="ref034"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Shih</surname> <given-names>C.</given-names></name> <name><surname>Lu</surname> <given-names>H.-Y. D.</given-names></name></person-group> (<year>2010</year>). &#x201C;<article-title>Prosody transfer and suppression: Stages of tone acquisition</article-title>&#x201D; in <source>Paper presented at the Speech Prosody 2010-Fifth International Conference</source></citation></ref>
<ref id="ref035"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sluijter</surname> <given-names>A.</given-names></name> <name><surname>Terken</surname> <given-names>J.</given-names></name></person-group> (<year>1993</year>). <article-title>Beyond sentence prosody: paragraph intonation in Dutch</article-title>. <source>Phonetica</source> <volume>50</volume>, <fpage>180</fpage>&#x2013;<lpage>188</lpage>.</citation></ref>
<ref id="ref61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Snider</surname> <given-names>K. L.</given-names></name></person-group> (<year>1990</year>). <article-title>Tonal upstep in Krachi: evidence for a register tier</article-title>. <source>Language</source> <volume>66</volume>, <fpage>453</fpage>&#x2013;<lpage>474</lpage>. doi: <pub-id pub-id-type="doi">10.2307/414608</pub-id></citation></ref>
<ref id="ref62"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Snider</surname> <given-names>K.</given-names></name></person-group> (<year>1998</year>). &#x201C;Tone and utterance length in Chumburung: an instrumental study.&#x201D; in <italic>Paper presented at the The 28th Colloquium on African Languages and Linguistics</italic>.</citation></ref>
<ref id="ref63"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Snider</surname> <given-names>K.</given-names></name> <name><surname>van der Hulst</surname> <given-names>H.</given-names></name></person-group> (<year>1993</year>). <source>Issues in the Representation of Tonal Register</source> <publisher-loc>Berlin</publisher-loc>: <publisher-name>Mouton de Gruyter</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>.</citation></ref>
<ref id="ref036"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Sorensen</surname> <given-names>J. M.</given-names></name> <name><surname>Cooper</surname> <given-names>W. E.</given-names></name></person-group> (<year>1980</year>). &#x201C;<article-title>Syntactic coding of fundamental frequency in speech production</article-title>&#x201D; in <source>Perception and production of fluent speech</source>. ed. <person-group person-group-type="editor"><name><surname>Cole</surname> <given-names>R. A.</given-names></name></person-group> (<publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Erlbaum</publisher-name>), <fpage>399</fpage>&#x2013;<lpage>440</lpage>.</citation></ref>
<ref id="ref037"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Stevens</surname> <given-names>K. N.</given-names></name></person-group> (<year>2000</year>). <source>Acoustic phonetics</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>MIT press</publisher-name>.</citation></ref>
<ref id="ref64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stewart</surname> <given-names>J. M.</given-names></name></person-group> (<year>1965</year>). <article-title>The typology of the Twi tone system</article-title>. <source>Bull. Inst. African Stud.</source> <volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>.</citation></ref>
<ref id="ref65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swerts</surname> <given-names>M.</given-names></name></person-group> (<year>1997</year>). <article-title>Prosodic features at discourse boundaries of different strength</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>101</volume>, <fpage>514</fpage>&#x2013;<lpage>521</lpage>. doi: <pub-id pub-id-type="doi">10.1121/1.418114</pub-id>, PMID: <pub-id pub-id-type="pmid">9000742</pub-id></citation></ref>
<ref id="ref66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swerts</surname> <given-names>M.</given-names></name> <name><surname>Geluykens</surname> <given-names>R.</given-names></name></person-group> (<year>1994</year>). <article-title>Prosody as a marker of information flow in spoken discourse</article-title>. <source>Lang. Speech</source> <volume>37</volume>, <fpage>21</fpage>&#x2013;<lpage>43</lpage>. doi: <pub-id pub-id-type="doi">10.1177/002383099403700102</pub-id></citation></ref>
<ref id="ref038"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Titze</surname> <given-names>I. R.</given-names></name></person-group> (<year>1988</year>). <article-title>The physics of small-amplitude oscillation of the vocal folds</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>83</volume>, <fpage>1536</fpage>&#x2013;<lpage>1552</lpage>.</citation></ref>
<ref id="ref001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x2018;t Hart</surname> <given-names>J.</given-names></name> <name><surname>Cohen</surname> <given-names>A.</given-names></name></person-group> (<year>1973</year>). <article-title>Intonation by rule: a perceptual quest</article-title>. <source>Journal of Phonetics</source> <volume>1</volume>, <fpage>309</fpage>&#x2013;<lpage>327</lpage>.</citation></ref>
<ref id="ref039"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Tucker</surname> <given-names>A. N.</given-names></name> <name><surname>Creider</surname> <given-names>C. A.</given-names></name></person-group> (<year>1975</year>). <article-title>Downdrift and downstep in Luo</article-title>. In <person-group person-group-type="editor"><name><surname>Herbert</surname> <given-names>R. K.</given-names></name></person-group> (Ed.), <source>Proceedings of the Sixth Conference on African Linguistics, OSU Working Papers in Linguistics</source> no. (pp. <fpage>125</fpage>&#x2013;<lpage>134</lpage>).</citation></ref>
<ref id="ref040"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Umeda</surname> <given-names>N.</given-names></name></person-group> (<year>1982</year>). <article-title>&#x201C;F0 declination&#x201D; is situation dependent</article-title>. <source>Journal of Phonetics</source> <volume>10</volume>, <fpage>279</fpage>&#x2013;<lpage>290</lpage>.</citation></ref>
<ref id="ref041"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wagner</surname> <given-names>M.</given-names></name></person-group> (<year>2002</year>). <article-title>The role of prosody in laryngeal neutralization. MIT Working Papers</article-title>. <source>Linguistics</source> <volume>42</volume>, <fpage>373</fpage>&#x2013;<lpage>392</lpage>.</citation></ref>
<ref id="ref77"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name> <name><surname>Genzel</surname> <given-names>S.</given-names></name></person-group> (<year>2018a</year>). <article-title>Downstep effect and the interaction with focus and prosodic boundary in Mandarin Chinese Paper presented at the Tonal Aspects of Languages (TAL)</article-title>. <publisher-name>Berlin, Germany</publisher-name>.</citation></ref>
<ref id="ref70"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). <article-title>Differential prosodic encoding of topic and focus in sentence-initial position in mandarin Chinese</article-title>. <source>J. Phon.</source> <volume>39</volume>, <fpage>595</fpage>&#x2013;<lpage>611</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2011.03.006</pub-id></citation></ref>
<ref id="ref71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Ding</surname> <given-names>Q.</given-names></name></person-group> (<year>2018b</year>). <article-title>Interactive prosodic marking of focus, boundary and newness in mandarin</article-title>. <source>Phonetica</source> <volume>75</volume>, <fpage>24</fpage>&#x2013;<lpage>56</lpage>. doi: <pub-id pub-id-type="doi">10.1159/000453082</pub-id>, PMID: <pub-id pub-id-type="pmid">28595174</pub-id></citation></ref>
<ref id="ref72"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ward</surname> <given-names>I. C.</given-names></name></person-group> (<year>1952</year>). <source>Introduction to the Yoruba Language</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>W. Heffer &#x0026; Sons Ltd.</publisher-name></citation></ref>
<ref id="ref73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Welmers</surname> <given-names>W. E.</given-names></name></person-group> (<year>1959</year>). <article-title>Tonemics, morphotonemics, and tonal morphemes</article-title>. <source>General Linguist.</source> <volume>4</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="ref74"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>1997</year>). <article-title>Contextual tonal variations in mandarin</article-title>. <source>J. Phon.</source> <volume>25</volume>, <fpage>61</fpage>&#x2013;<lpage>83</lpage>. doi: <pub-id pub-id-type="doi">10.1006/jpho.1996.0034</pub-id></citation></ref>
<ref id="ref75"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>1999</year>). <article-title>Effects of tone and focus on the formation and alignment of f0 contours</article-title>. <source>J. Phon.</source> <volume>27</volume>, <fpage>55</fpage>&#x2013;<lpage>105</lpage>. doi: <pub-id pub-id-type="doi">10.1006/jpho.1999.0086</pub-id></citation></ref>
<ref id="ref76"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). &#x201C;ProsodyPro&#x2014;a tool for large-scale systematic prosody analysis.&#x201D; in <italic>Paper Presented at the Proceedings of Tools and Resources for the Analysis of Speech Prosody (TRASP 2013)</italic>, Aix-en-Provence, France.</citation></ref>
<ref id="ref042"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>S.-w.</given-names></name> <name><surname>Wang</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <article-title>Prosodic focus with and without post-focus compression (PFC): A typological divide within the same language family?</article-title> <source>The Linguist. Rev.</source> <volume>29</volume>, <fpage>131</fpage>&#x2013;<lpage>147</lpage>. doi: <pub-id pub-id-type="doi">10.1515/tlr-2012-0006</pub-id></citation></ref>
<ref id="ref043"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Prom-on</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name></person-group> (<year>2022</year>). &#x201C;<article-title>The PENTA model: Concepts, use and implications</article-title>&#x201D; in <source>Prosodic Theory and Practice</source>. eds. <person-group person-group-type="editor"><name><surname>Shattuck-Hufnagel</surname> <given-names>S.</given-names></name> <name><surname>Barnes</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Cambridge</publisher-loc>: <publisher-name>The MIT Press</publisher-name>)</citation></ref>
<ref id="ref044"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name></person-group> (<year>2002</year>). <article-title>Maximum speed of pitch change and how it may relate to speech</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>111</volume>, <fpage>1399</fpage>&#x2013;<lpage>1413</lpage>.</citation></ref>
<ref id="ref78"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Q. E.</given-names></name></person-group> (<year>2001</year>). <article-title>Pitch targets and their realization: evidence from mandarin Chinese</article-title>. <source>Speech Comm.</source> <volume>33</volume>, <fpage>319</fpage>&#x2013;<lpage>337</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0167-6393(00)00063-7</pub-id></citation></ref>
<ref id="ref79"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>M. L.</given-names></name></person-group> (<year>2009</year>). <article-title>Organizing syllables into groups&#x2014;evidence from F0 and duration patterns in mandarin</article-title>. <source>J. Phon.</source> <volume>37</volume>, <fpage>502</fpage>&#x2013;<lpage>520</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2009.08.003</pub-id>, PMID: <pub-id pub-id-type="pmid">23482405</pub-id></citation></ref>
<ref id="ref81"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>C. X.</given-names></name></person-group> (<year>2005</year>). <article-title>Phonetic realization of focus in English declarative intonation</article-title>. <source>J. Phon.</source> <volume>33</volume>, <fpage>159</fpage>&#x2013;<lpage>197</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2004.11.001</pub-id></citation></ref>
<ref id="ref045"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>J.</given-names></name> <name><surname>Liberman</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>F0 declination in English and Mandarin Broadcast News Speech</article-title>. <source>Speech Communication</source> <volume>65</volume>, <fpage>67</fpage>&#x2013;<lpage>74</lpage>.</citation></ref>
<ref id="ref046"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Cantonese lexical tone and declination [in Chinese]</article-title>. <source>Language sciences Yuyan Kexue</source> <volume>2</volume>, <fpage>182</fpage>&#x2013;<lpage>191</lpage>.</citation></ref>
<ref id="ref047"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). <article-title>Mechanics of human voice production and control</article-title>. <source>The Journal of the Acoustical Society of America</source> <volume>140</volume>, <fpage>2614</fpage>&#x2013;<lpage>2635</lpage>.</citation></ref>
<ref id="ref84"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zerbian</surname> <given-names>S.</given-names></name> <name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). &#x201C;Downstep in Tswana (Southern Bantu).&#x201D; in <italic>Paper Presented at the ICPhS</italic>, Glasgow.</citation></ref>
<ref id="ref85"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zerbian</surname> <given-names>S.</given-names></name> <name><surname>K&#x00FC;gler</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Sequences of high tones across word boundaries in Tswana</article-title>. <source>J. Int. Phon. Assoc.</source> <fpage>1</fpage>&#x2013;<lpage>22</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S0025100321000141</pub-id></citation></ref>
</ref-list>
</back>
</article>