<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="Style/article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">2767-0279</journal-id>
<journal-title-group>
<journal-title>Glossa Psycholinguistics</journal-title>
</journal-title-group>
<issn pub-type="epub">2767-0279</issn>
<publisher>
<publisher-name>eScholarship Publishing</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5070/G6011.48921</article-id>
<article-categories>
<subj-group>
<subject>Regular article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>The prosodic perception-production link: Impact of auditory-perceptual acuity on prosodic cue production under cognitive load</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-2639-5499</contrib-id>
<name>
<surname>Hofmann</surname>
<given-names>Andrea</given-names>
</name>
<email>andhofma@uni-potsdam.de</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-8654-2446</contrib-id>
<name>
<surname>Tuomainen</surname>
<given-names>Outi</given-names>
</name>
<email>outi.tuomainen@uni-potsdam.de</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-5911-5572</contrib-id>
<name>
<surname>Hanne</surname>
<given-names>Sandra</given-names>
</name>
<email>hanne@uni-potsdam.de</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="fn" rid="affn1">*</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-1264-3017</contrib-id>
<name>
<surname>Ver&#237;ssimo</surname>
<given-names>Jo&#227;o</given-names>
</name>
<email>jlverissimo@edu.ulisboa.pt</email>
<xref ref-type="aff" rid="aff-2">2</xref>
<xref ref-type="fn" rid="affn1">*</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-5116-4441</contrib-id>
<name>
<surname>Wartenburger</surname>
<given-names>Isabell</given-names>
</name>
<email>isabell.wartenburger@uni-potsdam.de</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="fn" rid="affn1">*</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Cognitive Sciences, Department Linguistics, University of Potsdam</aff>
<aff id="aff-2"><label>2</label>Center of Linguistics, School of Arts and Humanities, University of Lisbon</aff>
<author-notes>
<fn id="affn1"><p>*Shared Senior Authorship</p></fn>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-07-06">
<day>06</day>
<month>07</month>
<year>2026</year>
</pub-date>
<pub-date pub-type="collection">
<year>2026</year>
</pub-date>
<volume>5</volume>
<issue>1</issue>
<elocation-id>12</elocation-id>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2026 The Author(s)</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://glossapsycholinguistics.journalpub.escholarship.org/articles/10.5070/G6011.48921/"/>
<abstract>
<p>Speech perception and production are linked through shared neural representations, such that individual differences in auditory-perceptual acuity predict articulatory precision for segmental contrasts. Whether similar perception-production links (PP-links) extend to suprasegmental prosodic features, however, remains unclear. In the present study, we investigated whether individual differences in auditory-perceptual acuity for prosodic boundary cues (pitch range, pause duration, final lengthening) predict how those same cues are produced, and whether this relationship is modulated by cognitive load. Using Bayesian distributional models, we examined how auditory-perceptual acuity modulated both the distinctiveness (mean cue differences between bracket and no-bracket productions) and trial-to-trial variability of prosodic boundary cue production. The results revealed cue-specific PP-links. For pitch, higher auditory-perceptual acuity was associated with stronger modulation of the produced pitch range, and it further modulated trial-to-trial variability in pitch range under cognitive load: under high cognitive load, speakers with higher pitch acuity maintained variable pitch range patterns across all productions, whereas speakers with lower pitch acuity showed reduced pitch range variability. For pause duration, higher auditory-perceptual acuity was associated with greater trial-to-trial variability in pause duration and, exploratorily, with reduced reliance on pauses for boundary marking. Final lengthening showed no relationship between auditory-perceptual acuity and production. Across cues, cognitive load systematically affected both production distinctiveness and trial-to-trial variability, interacting with auditory-perceptual acuity in a cue-specific manner. These findings demonstrate a PP-link at the prosodic level that is not uniform across prosodic boundary cues and suggest that, particularly for pitch, higher auditory-perceptual acuity supports flexible and adaptive prosodic cue production.</p>
</abstract>
</article-meta>
</front>
<body>
<sec>
<title>1. Introduction</title>
<p>Speech production relies on a tight interplay between perception and production within individual speakers. To produce speech that meets communicative goals, speakers continuously monitor their own output against internal auditory expectations, using sensorimotor feedback mechanisms to calibrate and refine articulatory movements (e.g., <xref ref-type="bibr" rid="B2">Arjmandi &amp; Behroozmand, 2024</xref>; <xref ref-type="bibr" rid="B38">Houde &amp; Jordan, 2002</xref>; <xref ref-type="bibr" rid="B74">Tourville et al., 2008</xref>). The perceptual and productive systems underlying this process are, thus, not independent, but rely on shared or tightly coupled representations (e.g., <xref ref-type="bibr" rid="B30">Guenther, 1995</xref>; <xref ref-type="bibr" rid="B80">Whalen, 2020</xref>). This coupling, commonly referred to as the <italic>perception-production link</italic> (PP-link) (e.g., <xref ref-type="bibr" rid="B21">Flege, 1995</xref>; <xref ref-type="bibr" rid="B54">Newman, 2003</xref>), has long been recognized in speech science and forms a foundation for theories of speech motor control.</p>
<sec>
<title>1.1 PP-link at the segmental level</title>
<p>Numerous studies investigating individual differences support the existence of a PP-link at the segmental level. These studies demonstrate that speakers who perceive phonemic contrasts more accurately also articulate those contrasts more distinctly. For example, English speakers who can discriminate between vowel contrasts (e.g., /&#593;/ &#8211; /&#652;/, /u/ &#8211; /&#650;/) or consonantal contrasts (e.g., /s/ &#8211; /&#643;/) more finely produce these categories with greater acoustic separation and reduced overlap, whereas poorer perceivers show smaller contrasts (<xref ref-type="bibr" rid="B62">Perkell, Guenther, et al., 2004</xref>; <xref ref-type="bibr" rid="B63">Perkell, Matthies, et al., 2004</xref>). Importantly, speakers with greater auditory-perceptual acuity can preserve such contrasts even under altered auditory feedback, indicating robust perception-linked control of articulation (e.g., <xref ref-type="bibr" rid="B8">Brunner et al., 2011</xref>; <xref ref-type="bibr" rid="B27">S. S. Ghosh et al., 2010</xref>). In such experiments, speakers receive real-time perturbations of their own speech signal via headphones, while producing differently sized speech units (see <xref ref-type="bibr" rid="B19">Elman, 1981</xref>, for one of the first studies on compensatory responses to real-time pitch perturbation). Moreover, PP-links have also been demonstrated at the subphonemic level: reaction times in cue-distractor tasks are reduced when the perceived distractor shares subphonemic features (e.g., vowel quality) with the target vowel to be produced (e.g., <xref ref-type="bibr" rid="B25">Ghaffarvand Mokari et al., 2020</xref>).</p>
<p>These behavioral findings can be explained by theoretical frameworks that posit tightly integrated sensory-motor mechanisms in speech planning and production. The Directions Into Velocities of Articulators (DIVA) model, for example, describes speech sounds as learned auditory-motor units, in which auditory targets in the phonetic space are closely linked to articulatory programs via shared neural mappings (see <xref ref-type="bibr" rid="B6">Bohland et al., 2010</xref>; <xref ref-type="bibr" rid="B30">Guenther, 1995</xref>; <xref ref-type="bibr" rid="B32">Guenther et al., 2013</xref>). Within this framework, speech sounds are acquired and planned in relation to sensory targets, and production accuracy is achieved through a combination of feedforward motor commands and auditory feedback-based error correction during development (<xref ref-type="bibr" rid="B31">Guenther, 2016</xref>; <xref ref-type="bibr" rid="B52">Meier &amp; Guenther, 2023</xref>). Over time, repeated feedback-based corrections are incorporated into feedforward commands, resulting in more precise and less variable segmental production, as well as reduced reliance on real-time feedback (e.g., <xref ref-type="bibr" rid="B70">Smith et al., 2020</xref>; <xref ref-type="bibr" rid="B77">Villacorta et al., 2007</xref>). As a consequence, individuals with finer auditory-perceptual acuity are expected to form more precise auditory targets, supporting more accurate error detection and more stable articulatory patterns. This architecture provides a principled explanation for observed PP-links at the segmental level, whereby auditory-perceptual acuity predicts production distinctiveness and precision (<xref ref-type="bibr" rid="B62">Perkell, Guenther, et al., 2004</xref>).</p>
</sec>
<sec>
<title>1.2 PP-link at the suprasegmental (prosodic) level</title>
<p>The evidence for a PP-link at the segmental level raises the question of whether a similar relationship exists within prosody. Prosody refers to the suprasegmental aspects of speech, realized through acoustic features such as pitch (the perceptual correlate of fundamental frequency, f0), duration, rhythm, and loudness. These features organize speech into hierarchical units and facilitate comprehension (e.g., <xref ref-type="bibr" rid="B23">Frazier et al., 2006</xref>; <xref ref-type="bibr" rid="B28">Gollrad et al., 2010</xref>). When one or more of these features are functionally coordinated to signal prosodic structure (e.g., stress, prominence, or prosodic boundaries), they are referred to as prosodic <italic>cues</italic>. A single cue may, therefore, be realized through multiple features, and individual features may contribute to different cues, depending on the context.</p>
<p>Empirical studies examining a potential PP-link at the prosodic level have adapted the perturbation paradigms used for segmental research, by observing how speakers handle perturbations of suprasegmental features. In the spectral domain, Patel et al. (<xref ref-type="bibr" rid="B59">2011</xref>) found that speakers compensated for attenuated stress cues by increasing f0 and intensity, though word duration remained unchanged. This suggests that duration was not recruited as a compensatory cue under f0 perturbation. In the temporal domain, compensatory responses to perturbed vowel durations are found to be hierarchically structured, depending on syllable position and stress. Oschkinat and Hoole (<xref ref-type="bibr" rid="B57">2022</xref>) perturbed the same syllable type in either an unstressed word-initial or a stressed word-medial position within a trisyllabic German word. Perturbation of the unstressed syllable caused a global slowing of all following segments, whereas perturbation of the stressed syllable elicited only local adjustment within the perturbed syllable. This discrepancy indicates that speakers do not merely correct individual segment durations in isolation; rather, their compensatory responses maintain prosodic timing relations at the word level by preserving the durational structure between stressed and unstressed syllables, as opposed to targeting sounds as independent units. Together with evidence for cue trading and adaptive re-weighting in prosodic boundary marking (e.g., <xref ref-type="bibr" rid="B15">Cole &amp; Shattuck-Hufnagel, 2016</xref>; <xref ref-type="bibr" rid="B43">Kentner et al., 2023</xref>; <xref ref-type="bibr" rid="B60">Patel &amp; Schell, 2008</xref>), these findings suggest that prosodic cue control involves coordinated adjustments across multiple dimensions of the speech signal.</p>
<p>A number of studies have also established a correlation between individual auditory-perceptual acuity and the magnitude of these compensatory and adaptive responses. In the spectral domain, finer-grained pitch discrimination has been associated with stronger adaptation to sustained pitch perturbations (e.g., <xref ref-type="bibr" rid="B49">Martin et al., 2018</xref>). Similarly, for prosodic timing, Oschkinat et al. (<xref ref-type="bibr" rid="B58">2022</xref>) found that speakers&#8217; responses to manipulated speech timing are jointly shaped by auditory-perceptual acuity and rhythmic abilities. However, the relative contributions of these factors vary according to the prosodic structure and whether responses signify online compensation or longer-term adaptation. These findings indicate that individuals differ systematically in how auditory feedback is integrated into prosodic motor control, paralleling individual-difference effects previously observed for segmental articulation.</p>
<p>The extension of PP-links from segments to prosody necessitates the introduction of additional distinctions, given the fundamental differences in the representation and control of suprasegmental features. The Gradient Order DIVA (GODIVA) model is an extension of the DIVA model that incorporates mechanisms for the planning and sequencing of multisyllabic utterances. These mechanisms include metrical structure, stress patterns, and the temporal coordination of speech units (see <xref ref-type="bibr" rid="B6">Bohland et al., 2010</xref>; <xref ref-type="bibr" rid="B73">Tourville &amp; Guenther, 2011</xref>). Recent advancements in the field have enabled the representation of spectral prosodic features, such as pitch, as auditory targets. This development constitutes a significant extension of the DIVA (LaDIVA) framework, as it now encompasses the modeling of pitch control and adaptation within an auditory feedback paradigm (<xref ref-type="bibr" rid="B79">Weerathunge et al., 2022</xref>). In contrast, temporal cues, such as pauses and the lengthening of segments, are not represented as explicit auditory targets. Instead, they are primarily implemented via GODIVA&#8217;s metrical and initiation maps, which govern the timing and release of speech units and reflect structural constraints of the prosodic system, rather than fine-grained auditory templates. Suprasegmental cues, thus, do not take the form of discrete symbolic representations. Rather, they are emergent control states that constrain the coordination of multiple acoustic features over time.</p>
<p>This representational asymmetry carries direct implications for the nature of PP-links. For spectral cues, such as pitch, higher auditory-perceptual acuity may facilitate more precise auditory goals and more flexible online correction. For temporal cues, however, the underlying mechanisms linking acuity and production are less well specified: these cues are primarily implemented via feedforward sequencing mechanisms rather than explicit auditory targets, and their production is further constrained by temporal irreversibility, placing greater emphasis on motor planning and stability (e.g., <xref ref-type="bibr" rid="B53">Miller &amp; Guenther, 2021</xref>; <xref ref-type="bibr" rid="B58">Oschkinat et al., 2022</xref>; <xref ref-type="bibr" rid="B57">Oschkinat &amp; Hoole, 2022</xref>). It is, therefore, hypothesized that prosodic PP-links will reflect an interaction between auditory-perceptual acuity, motor stability, and the structural properties of the prosodic system (e.g., prosodic hierarchy and cue-specific temporal constraints), and, consequently, that they are cue-specific in nature.</p>
</sec>
<sec>
<title>1.3 Prosody and syntactic disambiguation</title>
<p>Prosody plays a crucial role in speech comprehension by resolving structural ambiguities in sentences (e.g., <xref ref-type="bibr" rid="B23">Frazier et al., 2006</xref>). For example, in the structurally ambiguous sentence <italic>She will visit Berlin or Rome and Vienna</italic>, a prosodic break after <italic>Berlin</italic> leads listeners to interpret it as &#8216;She will visit (Berlin) or (Rome and Vienna)&#8217;, whereas a break after <italic>Rome</italic> yields &#8216;She will visit (Berlin or Rome) and (Vienna)&#8217;. Prosodic boundaries, thus, provide listeners with critical parsing information through variations in spectral and temporal boundary cues.</p>
<p>Coordinate name sequences have been used to systematically study how speakers use prosody to distinguish between different structural groupings (e.g., <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>; <xref ref-type="bibr" rid="B42">Kentner &amp; F&#233;ry, 2013</xref>; <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>; <xref ref-type="bibr" rid="B69">Schub&#246; et al., 2023</xref>). For instance, in response to a question like &#8220;Who is coming to the event?&#8221;, speakers might produce:</p>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(1)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p><italic>German: Bracket condition, with an internal prosodic boundary after Name2</italic></p></list-item>
<list-item><p>(Moni und Lilli) &#124;&#124; und Manu.</p></list-item>
<list-item><p>&#8216;Moni and Lilli, and Manu.&#8217;</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>&#160;</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>(Name1 und Name2) &#124;&#124; und Name3.</p></list-item>
<list-item><p>&#8216;Name1 and Name2, and Name3.&#8217;</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(2)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p><italic>German: No-bracket condition, without an internal prosodic boundary</italic></p></list-item>
<list-item><p>Moni und Lilli und Manu.</p></list-item>
<list-item><p>&#8216;Moni and Lilli and Manu.&#8217;</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>&#160;</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>Name1 und Name2 und Name3.</p></list-item>
<list-item><p>&#8216;Name1 and Name2 and Name3.&#8217;</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<p>In the <italic>bracket</italic> stimulus condition (1), a prosodic boundary following the second name, denoted here by &#124;&#124;, establishes a subgroup, suggesting that Moni and Lilli are arriving as a pair, separate from Manu. In the <italic>no-bracket</italic> stimulus condition (2), without this boundary, the implication is that all three are arriving collectively. The distinction in meaning relies entirely on the internal prosodic boundary, as the underlying words remain the same in both cases.</p>
<p>To mark prosodic boundaries in coordinated name sequences, German speakers reliably use a combination of prosodic boundary cues, most prominently pitch range, silent pause, and final lengthening (e.g., <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>; <xref ref-type="bibr" rid="B42">Kentner &amp; F&#233;ry, 2013</xref>; <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>; <xref ref-type="bibr" rid="B69">Schub&#246; et al., 2023</xref>). These cues differ systematically in their perceptual salience and functional role. Pauses typically elicit sharp, categorical boundary judgments and are particularly informative at strong boundaries, whereas pitch and final lengthening seem to have more gradient effects and interact with boundary strength and contextual expectations (e.g., <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>; <xref ref-type="bibr" rid="B65">Pijper &amp; Sanderman, 1994</xref>).</p>
<p>Cross-linguistic evidence further indicates that these cues are not functionally equivalent. Pitch cues demonstrate the highest degree of cross-linguistic variability, with their weighting and interpretation contingent on a language&#8217;s prosodic typology (e.g., stress-accented vs. tonal systems) and communicative demands (e.g., <xref ref-type="bibr" rid="B36">Holzgrefe-Lang et al., 2016</xref>; <xref ref-type="bibr" rid="B56">Ortega-Llebaria &amp; Nagao, 2025</xref>; <xref ref-type="bibr" rid="B85">Yang et al., 2014</xref>). In the German language, pitch is a primary but flexible boundary cue, contributing to both syntactic disambiguation and the expression of pragmatic or affective meaning. It is often integrated with temporal cues rather than used in isolation (e.g., <xref ref-type="bibr" rid="B36">Holzgrefe-Lang et al., 2016</xref>, <xref ref-type="bibr" rid="B35">2018</xref>). Pauses, by contrast, represent a highly salient and largely language-general segmentation cue, though their use is optional at weaker boundaries (e.g., <xref ref-type="bibr" rid="B48">M&#228;nnel &amp; Friederici, 2016</xref>; <xref ref-type="bibr" rid="B65">Pijper &amp; Sanderman, 1994</xref>), while final lengthening constitutes a near universal right-edge cue, robustly attested across languages and prosodic systems (e.g., <xref ref-type="bibr" rid="B10">Byrd et al., 2006</xref>; <xref ref-type="bibr" rid="B82">Wightman et al., 1992</xref>).</p>
</sec>
<sec>
<title>1.4 Individual differences in prosodic perception and production</title>
<p>Huttenlauch et al. (<xref ref-type="bibr" rid="B39">2021</xref>) demonstrated that while speakers show high intra-individual consistency in boundary marking, they differ from each other in their preferred prosodic boundary cue combinations. Despite the presence of these inter-individual differences, listeners demonstrated an ability to accurately recognize the intended prosodic groupings in over 95% of utterances, suggesting the presence of robust compensation for speaker variability.</p>
<p>In a complementary perception study, Hansen et al. (<xref ref-type="bibr" rid="B33">2023</xref>) found significant inter-individual variability in prosodic boundary perception. They presented the coordinated three-name stimuli incrementally, syllable by syllable (e.g., <italic>Mi, Mimmi, Mimmi und, Mimmi und Ne</italic>, and so forth) and listeners had to predict after each snippet if it would belong to a structure with or without a prosodic boundary after the second name. Some listeners could exploit subtle prosodic boundary cues very early &#8211; during the first name &#8211; to accurately predict upcoming syntactic grouping. Others required more prosodic information, only making accurate predictions after hearing the second name. This variation in early prosodic boundary cue sensitivity indicates that prosodic parsing depends not only on acoustic input, but also on listener-specific auditory-perceptual abilities.</p>
<p>These parallel findings of systematic inter-individual variation in both production and perception are consistent with broader evidence that speakers encode prosodic meaning through structured individual differences in prosodic boundary cue distributions. Xie et al. (<xref ref-type="bibr" rid="B84">2021</xref>) elicited question-statement productions from English speakers and found substantial structured variability in how different talkers used f0 and duration to mark the contrast: some speakers relied primarily on pitch rises, others used both pitch and lengthening, and, crucially, speakers also differed in how consistently they produced these cues. The authors employed ideal observer models that incorporated talker-specific distributional statistics of prosodic cues. These models, which reflected both typical realizations and their spread, exhibited substantially superior performance in comparison to models that only normalized for baseline pitch. In a subsequent perceptual study, listeners rapidly adapted to talker-specific &#8220;prosodic dialects,&#8221; shifting their categorization of identical utterance-final pitch contours as questions or statements after briefly learning a new speaker&#8217;s prosodic patterns. Together, these findings demonstrate that prosodic variability is structured and functionally meaningful: speakers produce systematic distributional patterns that listeners track and exploit for more accurate interpretation.</p>
<p>Building on this evidence of systematic individual variation in prosodic processing, our study investigates whether individual differences in perception predict individual differences in production. This leads to our central research question: Does a PP-link for prosodic boundary cues exist at the level of individual differences across participants? In other words, are participants who demonstrate heightened perceptual sensitivity to prosodic boundary cues also those who produce more distinctive and less variable prosodic boundaries?</p>
<p>Prior studies have demonstrated prosodic PP-links using perturbation paradigms (e.g., <xref ref-type="bibr" rid="B57">Oschkinat &amp; Hoole, 2022</xref>; <xref ref-type="bibr" rid="B59">Patel et al., 2011</xref>), which revealed compensatory mechanisms in prosodic control. However, because these paradigms rely on experimentally imposed feedback manipulations, they may induce heightened speaker awareness and reactive strategies that differ from those engaged during natural, communicative speech, potentially limiting their generalizability (e.g., <xref ref-type="bibr" rid="B16">Dahl et al., 2024</xref>; <xref ref-type="bibr" rid="B37">Houde &amp; Jordan, 1998</xref>; <xref ref-type="bibr" rid="B72">Tomassi et al., 2022</xref>). For instance, the magnitude of adaptation to f0 perturbations in linguistically neutral sustained vowels does not correlate with that during sentence production, suggesting distinct mechanisms at play in perturbation paradigms (<xref ref-type="bibr" rid="B16">Dahl et al., 2024</xref>), as compared to natural speech. Additionally, perturbation paradigms often focus on immediate compensation for specific cues (e.g., f0 shifts or temporal alterations), underexploring baseline variability across multiple prosodic features or the influence of moderators like attention. The current study, therefore, examines the PP-link in a more ecologically valid context: speakers&#8217; natural production of prosodic boundary cues (pitch range, pause duration, and final lengthening) during syntactic disambiguation tasks, with and without cognitive load. By incorporating all three primary prosodic boundary cues co-occurring in the same stimuli, the present design preserves their natural integration. At the same time, by analyzing each cue separately, we test how individual differences in auditory-perceptual acuity relate to prosodic boundary cue production distinctiveness and variability under varying cognitive load mimicking communicative demands (e.g., multitasking).</p>
</sec>
<sec>
<title>1.5 Testing the PP-link under cognitive load</title>
<p>While existing evidence suggests that prosodic PP-links are shaped by multiple interacting factors, there is currently no unified account specifying how individual differences in auditory-perceptual acuity map onto prosodic cue use. For segmental targets, the DIVA framework provides a mechanistic account of how auditory-perceptual acuity shapes production precision: finer auditory targets support more accurate error detection and more distinct articulatory patterns, a prediction that is consistently supported by empirical evidence involving individual differences (e.g., <xref ref-type="bibr" rid="B61">Perkell, 2012</xref>). The present study, therefore, adopts the segmental PP-link literature as a principled baseline. This allows us to formulate testable predictions about whether auditory-perceptual acuity similarly constrains the production of suprasegmental cues.</p>
<p>To test whether a PP-link exists at the prosodic level, we examined individual differences in auditory-perceptual acuity and in the production of prosodic boundaries in coordinated three-name sequences. In production, participants used prosody to disambiguate syntactic structure by marking or omitting an internal prosodic boundary after the second name (<italic>bracket</italic> vs. <italic>no-bracket</italic> stimulus condition). In perception, auditory-perceptual acuity was assessed independently, using Just-Noticeable-Difference (JND) thresholds for detecting graded changes in the same acoustic dimensions that function as prosodic boundary cues in production, namely, pitch range, pause duration, and final lengthening.</p>
<p>Perception and production were linked at the level of individual differences and analyzed using Bayesian distributional models that estimate both mean prosodic boundary cue realization (reflecting <italic>production distinctiveness</italic>) and trial-to-trial variability (captured by the <monospace>sigma</monospace> parameter, reflecting <italic>production variability</italic>). This modeling approach allows us to test not only whether participants with better perception produce stronger prosodic contrasts, but also whether they do so more or less variably across repetitions.</p>
<p>To examine individual differences in prosodic boundary production under varying processing demands, we manipulated cognitive load during speech production by introducing a dual-task condition, creating distinct low and high cognitive load conditions. According to Lindblom&#8217;s Hyper-Hypo theory (<xref ref-type="bibr" rid="B46">Lindblom, 1990</xref>), speakers continuously trade off communicative clarity against articulatory effort, depending on contextual demands. When cognitive resources are limited, speakers tend to economize speech production, resulting in reduced articulatory precision or attenuated prosodic marking, unless communicative demands require otherwise (e.g., <xref ref-type="bibr" rid="B46">Lindblom, 1990</xref>; <xref ref-type="bibr" rid="B83">Winkworth &amp; Davis, 1997</xref>). In the present study, the dual-task manipulation served a methodological purpose: it introduced controlled processing demands by drawing attentional resources away from speech planning and execution under high cognitive load, without inducing speech errors or breakdowns. Under these conditions, high cognitive load was expected to increase trial-to-trial variability within participants (captured by the <italic>&#963;</italic> parameter). High cognitive load may also amplify differences between participants in their mean prosodic boundary cue realization (captured by the <italic>&#956;</italic> parameter), thereby creating a critical testing ground for examining how auditory-perceptual acuity relates to production under varying speech processing demands. While differences in variability under high cognitive load could also reflect general cognitive advantages (e.g., superior multitasking ability), the present study focuses specifically on how auditory-perceptual acuity modulates these effects.</p>
<p>Cognitive load was expected to interact with auditory-perceptual acuity in a cue-specific manner. For spectral prosodic cues, such as pitch, which can be represented as auditory targets within the DIVA framework, higher auditory-perceptual acuity was predicted to support more precise auditory goal regions and more effective maintenance of informative prosodic marking under load (e.g., <xref ref-type="bibr" rid="B61">Perkell, 2012</xref>; <xref ref-type="bibr" rid="B73">Tourville &amp; Guenther, 2011</xref>). For temporal cues, such as pause and final lengthening, which are implemented via sequencing mechanisms rather than explicit auditory targets, the underlying mechanisms linking acuity and production are less well specified. Here, we provisionally extend the same prediction derived from segmental research and treat the presence or absence of PP-links for temporal cues as an empirical question, asking whether production effects associated with auditory-perceptual acuity for spectral prosodic boundary cues also appear for temporal prosodic boundary cues.</p>
<p>Finally, greater trial-to-trial variability in prosodic production would not necessarily reflect imprecision or noise. In prosody, production variability can also arise from flexible weighting of cues, in how they are combined, and from the fact that prosodic boundary cues serve multiple functions in speech (e.g., <xref ref-type="bibr" rid="B15">Cole &amp; Shattuck-Hufnagel, 2016</xref>; <xref ref-type="bibr" rid="B43">Kentner et al., 2023</xref>; <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>). From this perspective, individual differences in prosodic production variability may index differences in how participants dynamically deploy multiple prosodic boundary cues under varying contextual and processing demands, rather than differences in motor stability alone.</p>
</sec>
<sec>
<title>1.6 Aims and hypotheses</title>
<p>Building on prior evidence for individual differences in prosodic perception and production (<xref ref-type="bibr" rid="B33">Hansen et al., 2023</xref>; <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>), the present study investigates whether auditory-perceptual acuity for prosodic boundary cues predicts how those same cues are produced during syntactic disambiguation, and whether this relationship is modulated by cognitive load.</p>
<p>We address two primary hypotheses, each targeting a distinct aspect of prosodic production:</p>
<disp-quote>
<p><bold>Hypothesis 1 (H1): Effect of auditory-perceptual acuity on production distinctiveness.</bold> Our first hypothesis concerns production distinctiveness. Operationally, distinctiveness refers to the mean difference in prosodic boundary cue realization between stimulus conditions (<italic>bracket</italic> vs. <italic>no-bracket</italic>), estimated by the model&#8217;s <monospace>mean</monospace> (<italic>&#956;</italic>) parameter. For pitch range and final lengthening, derived from segmental perception-production evidence, we hypothesize that higher auditory-perceptual acuity is associated with greater production distinctiveness, that is, larger mean differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions, indicating stronger acoustic separation to convey the different groupings. Because pauses are almost exclusively realized in <italic>bracket</italic> productions, pause analyses focus on <italic>bracket</italic> stimuli only (see 2.4.4). For pauses, distinctiveness is, therefore, defined as pause duration within <italic>bracket</italic> stimuli rather than as a difference across stimulus conditions. We hypothesize that higher auditory-perceptual acuity is associated with longer pauses, particularly under high cognitive load. For all three prosodic boundary cues, we expect the relationship between auditory-perceptual acuity and production distinctiveness to be more pronounced under high cognitive load, where increased processing demands may amplify inter-individual differences.</p>
<p><bold>Hypothesis 2 (H2): Effect of auditory-perceptual acuity on production variability.</bold> Our second hypothesis concerns production variability. Operationally, production variability refers to the trial-to-trial spread of prosodic boundary cue realizations within a participant, captured by the model&#8217;s variability (<italic>&#963;</italic>) parameter. For all three prosodic boundary cues, we hypothesize that higher auditory-perceptual acuity is associated with lower production variability, that is, smaller <italic>&#963;</italic> values in prosodic boundary cue realizations across repeated productions. This variability effect is expected across both cognitive load conditions, but should be more pronounced under high cognitive load, where participants with lower auditory-perceptual acuity are predicted to show greater increases in production variability.</p>
</disp-quote>
<p>For pitch range and final lengthening, this hypothesis concerns whether auditory-perceptual acuity modulates the difference in production variability between <italic>bracket</italic> and <italic>no-bracket</italic> productions. For pause duration, it concerns production variability within <italic>bracket</italic> productions only.</p>
</sec>
</sec>
<sec>
<title>2. Methods and materials</title>
<p>This study was preregistered (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://osf.io/dvuw6">https://osf.io/dvuw6</ext-link>).</p>
<sec>
<title>2.1 Participants</title>
<p>Sixty native German speakers participated in the study, with a mean age of 24.78 years (standard deviation, SD = 6.07, range: 18&#8211;49). The sample comprised 48 females and 12 males. No formal a priori power analysis was conducted. The preregistered target sample size was determined based on the sample sizes used in previous work in prosodic perception-production research and was deemed sufficient to detect individual-differences effects (e.g., <xref ref-type="bibr" rid="B33">Hansen et al., 2023</xref>; <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>; <xref ref-type="bibr" rid="B49">Martin et al., 2018</xref>; <xref ref-type="bibr" rid="B57">Oschkinat &amp; Hoole, 2022</xref>). Participants were native speakers of German without a history of speech or language disorders, hearing impairments, or neurological or psychological conditions. They received either 40&#8364; or 5 course credits for completing two experimental sessions (approx. 2 hours each), scheduled on the same or separate days, based on availability. The two experiments reported here were both conducted during the first session.</p>
<p>No participants were excluded based on the preregistered criterion for adaptive staircase performance (<xref ref-type="bibr" rid="B58">Oschkinat et al., 2022</xref>). This criterion required JND thresholds to decrease below 70% of the initial cue difference, indicating appropriate task engagement (for JND task description and thresholds, see 2.3.2). However, following data inspection, some participants exhibited extreme JND values, suggesting atypical or invalid response patterns (e.g., one participant showed a JND for pitch that was more than 7 standard deviations above the group mean). To address this issue, we excluded participants whose JND scores fell below the first quartile (Q1) minus two times the interquartile range (IQR), or above the third quartile (Q3) plus two times the IQR. This led to the exclusion of three participants from the pitch model, leaving N = 57, and two from the pause model, leaving N = 58. No exclusions were made for the final lengthening model.</p>
</sec>
<sec>
<title>2.2 General procedure</title>
<p>Participants were tested individually in a sound-attenuated booth. Visual stimuli were presented on a 1080 &#215; 1920 pixel monitor, with keyboard input used to record responses for the perception task. For the production task, verbal responses were recorded at a 48 kHz sampling rate, using a Beyerdynamic DT-297 headset (80 Ohm headphones, 300 Ohm condenser mic, 5 cm distance from chin), connected to a Focusrite Scarlett 18i8 audio interface. All experimental procedures were controlled by custom Python 3.8 scripts (PyCharm, Windows 10).</p>
</sec>
<sec>
<title>2.3 Perception: JND task</title>
<sec>
<title>2.3.1 Stimuli</title>
<p>To assess auditory-perceptual acuity for acoustic dimensions that serve as prosodic boundary cues in production, we constructed three independent acoustic continua targeting pitch range, pause duration, and final lengthening. For brevity, we refer to these acoustic dimensions as prosodic boundary cues throughout, recognizing that their cue status depends on functional deployment in context. Crucially, while these cues are typically associated with boundary marking, the JND task itself was not designed to test boundary perception. Instead, it quantified listeners&#8217; sensitivity to graded changes in the underlying acoustic dimensions themselves.</p>
<p>All continua were derived from original speech recordings previously used in a perception study by de Beer et al. (<xref ref-type="bibr" rid="B17">2022</xref>). In that study, a phonetically trained female speaker produced coordinated three-name sequences with systematically varied prosodic realizations. From this corpus, we selected tokens that exhibited the strongest acoustic realization of each prosodic boundary cue. Specifically, we identified the (separate) instances in which (a) the pitch range between the stressed and unstressed syllables of Name2 was maximal, (b) a clear and long silent interval followed Name2, and (c) the final segment of Name2 showed maximal lengthening. Since the maximal realization of each prosodic boundary cue occurred in different three-name sequences within this corpus, the base stimulus for each continuum was drawn from a different three-name sequence.</p>
<p>These segments with maximal prosodic boundary cue realization were then extracted and served as the base stimuli for the construction of three JND continua, using custom Praat scripts (<xref ref-type="bibr" rid="B5">Boersma &amp; Weenink, 1992&#8211;2020</xref>). Each continuum ranged from the clearly (maximally) cued stimulus to a minimally cued reference stimulus, thereby spanning a graded decrease in the magnitude of the manipulated acoustic feature (and, thus, in the typical acoustic realization associated with boundary marking). Each continuum spanned from the original recording including the maximal realization of the targeted prosodic boundary cue to a version of the same recording manipulated to have the minimal realization of the targeted prosodic boundary cue. All intermediate steps consisted of systematically attenuated versions covering the range between these endpoints.</p>
<p><italic>Pitch range continuum</italic>: The pitch continuum was based on the <italic>Name2</italic> token <italic>Nelli</italic> (from the sequence <italic>Mimmi und Nelli &#124;&#124; und Lola</italic>.), produced with a pitch range of 13 semitones. This rising contour was gradually flattened from 13 to 0 semitones in steps of 0.005 semitones, yielding a fully level contour as the reference stimulus.</p>
<p><italic>Pause duration continuum</italic>: For the pause continuum, we used the phrase <italic>Name2 und Name3</italic> (<italic>Lilli [PAUSE] und Lisa</italic>, from the sequence <italic>Moni und Lilli &#124;&#124; und Lisa</italic>.), which contained a silent interval of 550 ms following <italic>Lilli</italic>. The pause duration was reduced in 1 ms increments from 550 ms to 0 ms, with the endpoint representing a stimulus without any pause.</p>
<p><italic>Final lengthening continuum</italic>: The final lengthening continuum was constructed from the <italic>Name2</italic> token <italic>Mimmi</italic> (from the sequence <italic>Leni und Mimmi &#124;&#124; und Manu</italic>.), in which the final vowel had a duration of 225 ms (approximately half of the total word duration). This segment was progressively shortened in steps of 0.3 ms until a duration of 61 ms was reached, resulting in approximately equal syllable durations and, thus, minimal perceptual lengthening.</p>
</sec>
<sec>
<title>2.3.2 Procedure</title>
<p>Auditory-perceptual acuity thresholds for each prosodic boundary cue continuum were obtained, using an AXB discrimination task combined with an adaptive staircase procedure (adapted from <xref ref-type="bibr" rid="B70">Smith et al., 2020</xref>). Each participant completed three separate JND tasks, one for each prosodic boundary cue (pitch range, pause duration, final lengthening). The task order was randomized across participants. At the beginning of each task, participants completed a brief practice phase with visual accuracy feedback until they produced four consecutive correct responses. No feedback was provided during the experimental phase.</p>
<p><bold>Trial structure:</bold> Each trial consisted of an AXB stimulus sequence (either AAB or ABB), with an inter-stimulus interval of 500 ms. Across trials, participants heard two types of tokens: a reference stimulus with minimal or absent cue expression (0 semitones pitch range, 0 ms pause duration, or 61 ms final lengthening) and a comparison stimulus drawn from the respective prosodic boundary cue continuum. The acoustic distance between reference and comparison, referred to here as the <italic>cue difference</italic>, determined task difficulty: larger cue differences yielded easier discrimination, while progressively smaller cue differences probed listeners&#8217; discrimination limits.</p>
<p>The order of reference and comparison tokens was randomized across trials, resulting in sequences containing either two identical reference tokens and one comparison token, or vice versa. Participants indicated which stimulus was different by pressing the right arrow key for AAB sequences and the left arrow key for ABB sequences. Visual response prompts remained visible until a response was made, and depicted arrow icons alongside schematic stimulus patterns (with &#8220;A&#8221; boxes shown in green and &#8220;B&#8221; boxes in black).</p>
<p><bold>Adaptive staircase:</bold> With the adaptive staircase procedure, we adjusted cue differences on a trial-by-trial basis to converge on each participant&#8217;s discrimination threshold, defined as the smallest reliably detectable acoustic difference. All staircases started with the maximum available cue difference, to ensure initial discriminability (13 semitones for pitch range, 550 ms for pause duration, and 164 ms for final lengthening). Following each response, cue differences were modified as a function of performance: correct responses reduced the cue difference, whereas incorrect responses increased it.</p>
<p><italic>Step sizes</italic>, defined as the amount by which the cue difference between reference and comparison stimuli was increased or decreased from one trial to the next, decreased progressively over the course of the staircase. Large initial step sizes enabled rapid movement toward the participant&#8217;s approximate threshold region, while smaller later step sizes allowed fine-grained estimation. Step sizes ranged from 0.75 to 0.005 semitones for pitch range, from 30 to 1 ms for pause duration, and from 7.5 to 0.3 ms for final lengthening.</p>
<p>At the beginning of each staircase, a 1-down-1-up adjustment rule was used to promote rapid convergence toward the threshold region. After the first incorrect response, the procedure switched to a 2-down-1-up rule, such that two consecutive correct responses were required to decrease the cue difference, whereas a single incorrect response was sufficient to increase it. This asymmetric update rule converges on a performance level of approximately 71% correct responses (<xref ref-type="bibr" rid="B44">Levitt, 1971</xref>), a threshold determined by the balance of probabilities of upward and downward step adjustments in a 2-down-1-up staircase: at equilibrium, the probability of decreasing the cue difference (requiring two consecutive correct responses, with probability <italic>p</italic><sup>2</sup>) equals the probability of increasing it (1 &#8211; <italic>p</italic><sup>2</sup>), yielding <inline-formula>
<alternatives>
<mml:math id="Eq005-mml">
<mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msqrt><mml:mo>&#x2248;</mml:mo><mml:mn>0.707</mml:mn></mml:mrow>
</mml:math>
<tex-math id="M5">
\documentclass[10pt]{article}
\usepackage{wasysym}
\usepackage[substack]{amsmath}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage[mathscr]{eucal}
\usepackage{mathrsfs}
\usepackage{pmc}
\usepackage[Euler]{upgreek}
\pagestyle{empty}
\oddsidemargin -1.0in
\begin{document}
\[
p = \sqrt {1/2} \approx 0.707
\]
\end{document}
</tex-math>
<graphic xlink:href="glossapx-5-1-48921-e5.gif"/>
</alternatives>
</inline-formula>. This value represents the cue difference, where discrimination is achieved with about 71% accuracy, providing a stable estimate of the perceptual threshold, while avoiding ceiling (near 100% performance) and floor (near chance) effects that could bias measurements. A reversal was defined as a change in the direction of cue difference adjustment (from increasing to decreasing or vice versa). Each task terminated after 120 trials or 18 reversals, whichever occurred first. Threshold estimation was based on the sequence of reversal points, as described below.</p>
</sec>
<sec>
<title>2.3.3 Data pre-processing</title>
<p>For each prosodic boundary cue, the JND threshold was computed as the mean of the six most stable consecutive reversal points, operationalized as the set of six adjacent reversals with the lowest standard deviation (<xref ref-type="bibr" rid="B8">Brunner et al., 2011</xref>; <xref ref-type="bibr" rid="B58">Oschkinat et al., 2022</xref>). These values represent the smallest acoustic differences participants could reliably discriminate, with lower raw thresholds indicating better auditory-perceptual acuity.</p>
<p>To facilitate interpretation and ensure consistency across analyses, raw JND thresholds were z-scored and sign-reversed, such that higher values correspond to better auditory-perceptual acuity. These transformed measures are referred to throughout as <italic>pitch acuity, pause acuity</italic>, and <italic>final lengthening acuity</italic>, reflecting participants&#8217; standardized auditory-perceptual acuity for each prosodic boundary cue. <xref ref-type="fig" rid="F1">Figure 1</xref> illustrates the distributions of raw JND thresholds in their original measurement units for each prosodic boundary cue.</p>
<fig id="F1">
<caption>
<p><bold>Figure 1:</bold> Distribution of individual JND thresholds for each prosodic boundary cue. Panels show raw JND thresholds in native units (ms for pause and final lengthening, semitones for pitch). Only data from analyzed participants are shown (n = 57 pitch, n = 58 pause, n = 60 final lengthening).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g1.png"/>
</fig>
</sec>
</sec>
<sec>
<title>2.4 Production: Cognitive load task</title>
<sec>
<title>2.4.1 Stimuli</title>
<p>The production task used 24 written name sequences arranged in coordinate structures that have been employed in prior production and perception research (<xref ref-type="bibr" rid="B17">de Beer et al., 2022</xref>; <xref ref-type="bibr" rid="B33">Hansen et al., 2023</xref>; <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>). Each sequence consisted of three disyllabic German names. The first two names consistently ended in the vowel /i/ (<italic>Moni, Lilli, Leni, Nelli, Mimmi, Manni</italic>), while the third name ended in either /u/ or /a/ (<italic>Manu, Nina, Lola</italic>).</p>
<p>Stimuli were balanced across two conditions. In the <italic>bracket</italic> stimulus condition (12 items), parentheses indicated a prosodic boundary after the second name (e.g., <italic>(Moni und Lilli) und Manu</italic>). In the <italic>no-bracket</italic> stimulus condition (12 items), the same sequences were presented without grouping markers (e.g., <italic>Moni und Lilli und Manu</italic>).</p>
</sec>
<sec>
<title>2.4.2 Procedure</title>
<p>Participants read aloud written name sequences under two cognitive load conditions: (i) a low cognitive load (single task) condition, involving only the reading task, and (ii) a high cognitive load (dual-task) condition, designed to increase cognitive load by adding two concurrent tasks.</p>
<p>Each trial began with the prompt <italic>Wer kommt?</italic> (&#8216;Who&#8217;s coming?&#8217;) displayed for 1 second, followed by a 1-second fixation cross. The name sequence then appeared for 5.8 seconds (350 frames), during which participants read the sequence aloud.</p>
<p>In the high cognitive load condition, participants simultaneously performed (a) a tone-counting task and (b) a motion-tracking task while reading. Each trial featured a sequence of approximately 23 tones (0.2 s duration, 0.58 s inter-stimulus-intervals) spanning the entire trial (~20 s). Tones were played before the name sequence appeared on screen, were then suspended during the reading phase (from 0.2 seconds before to 0.2 seconds after presenting the name sequence on the screen, to avoid auditory interference), and resumed afterward. During the reading phase, 500 moving dots (50% coherent motion in one of four directions: up, down, left, or right) were superimposed on the name sequence for 4.2 seconds. Participants had to track the motion direction while reading aloud.</p>
<p>Each trial, therefore, consisted of three phases: an initial tone-counting phase, a dual-task reading and motion-tracking phase, and a final tone-counting phase. The tone-counting task involved detecting and counting high-pitched deviant tones (octave 5 &#8220;A&#8221; notes) embedded among standard mid-pitched tones (octave 5 &#8220;C&#8221; notes), with 3&#8211;10 deviants randomly distributed across the initial and final phases of each trial (at least three overall). After each trial, participants provided two responses using arrow keys. First, they reported the perceived motion direction (selecting from four directional options). Second, they reported the cumulative count of deviant tones heard across both tone-counting phases, selecting from four numeric options (the correct count and three distractors within &#177;3).</p>
<p>Both cognitive load conditions included a practice phase. In each cognitive load condition, participants read aloud four practice stimuli: two <italic>bracket</italic> and two <italic>no-bracket</italic> items (featuring name sequences not used in the test phase). The low load practice phase served only to familiarize participants with the reading task and recording setup. No trial-by-trial performance feedback was provided. The high load practice phase used the same reading task, but additionally included the secondary tasks, that is, tone-counting and motion-tracking. For these secondary tasks, participants received automated on-screen accuracy feedback (&#8220;correct answer&#8221;/&#8220;incorrect answer&#8221;).</p>
<p>The main experiment comprised 24 trials under each cognitive load condition (12 <italic>bracket</italic>, 12 <italic>no-bracket</italic> stimuli each). Participants always completed the low cognitive load task first, followed by the high cognitive load task. No trial-by-trial performance feedback was provided during the main experimental blocks. Any feedback was restricted to the practice phase of the high cognitive load condition. Stimulus order was randomized. In high load trials, the name sequence onset, motion-tracking start time, deviant-tone distribution and number, and numeric response options were all randomized.</p>
<p>The two secondary tasks were selected to impose a significant cognitive load without inducing speech errors or articulatory distortions, as was verified during the pilot study. The motion-tracking task draws on visuospatial processing, rather than linguistic processing, with difficulty being parametrically controlled via dot motion coherence (e.g., <xref ref-type="bibr" rid="B47">Lively et al., 1993</xref>; <xref ref-type="bibr" rid="B55">Newsome et al., 1989</xref>; <xref ref-type="bibr" rid="B67">R&#233;v&#233;sz et al., 2016</xref>), while sustained auditory attention and working memory are taxed by the tone-counting task (e.g., <xref ref-type="bibr" rid="B7">Brown, 2025</xref>; <xref ref-type="bibr" rid="B34">Harmon et al., 2019</xref>). Together, these tasks draw on distinct resource pools that do not directly compete with phonological, prosodic or articulatory processing, which is consistent with the theory of multiple resources (<xref ref-type="bibr" rid="B81">Wickens, 2008</xref>).</p>
</sec>
<sec>
<title>2.4.3 Data pre-processing</title>
<p><bold>Data preparation and segmentation.</bold> Participant productions were automatically segmented using the Montreal Forced Aligner (<xref ref-type="bibr" rid="B50">McAuliffe et al., 2017</xref>) with a custom German dictionary. Trained research assistants then manually verified and corrected all segmentations. This manual verification process involved checking segmentation boundaries and correcting any mislabeled segments, including distinguishing between actual pauses and other phenomena, such as hesitations or stop closures. All pre-processing scripts are available in the Audio Analysis Pre-processing Scripts component of the OSF website for this article (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://osf.io/3ykhd/">https://osf.io/3ykhd/</ext-link>).</p>
<p><bold>Extraction of prosodic boundary cues.</bold> Segment durations and pitch information were extracted using Praat scripts (<xref ref-type="bibr" rid="B5">Boersma &amp; Weenink, 1992&#8211;2020</xref>) and logged in CSV files for further analysis:</p>
<p><italic>Pitch range</italic>: The f0 range between the two syllables of the second name was calculated based on f0 at the Center of Periodic Mass (CoM), extracted using the ProPer toolbox (see <xref ref-type="bibr" rid="B1">Albert, 2023</xref>). First, the f0 value at the CoM was identified for each syllable separately. Second, the larger value was indexed as f0<sub>Max</sub>, and the smaller as f0<sub>Min</sub>. <italic>Pitch range</italic> was then calculated as the ratio of f0<sub>Max</sub> to f0<sub>Min</sub> at CoM, and transformed into semitones using the formula:</p>
<disp-formula id="FD1">
<alternatives>
<mml:math id="Eq001-mml">
<mml:mrow><mml:mtext mathvariant="italic">Pitch&#x00A0;range</mml:mtext><mml:mo>=</mml:mo><mml:mn>12</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mrow><mml:mtext>log</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>f</mml:mi><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mtext mathvariant="italic">Max</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mtext mathvariant="italic">Min</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow>
</mml:math>
<tex-math id="M1">
\documentclass[10pt]{article}
\usepackage{wasysym}
\usepackage[substack]{amsmath}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage[mathscr]{eucal}
\usepackage{mathrsfs}
\usepackage{pmc}
\usepackage[Euler]{upgreek}
\pagestyle{empty}
\oddsidemargin -1.0in
\begin{document}
\[
Pitchrange = 12 \times {{\rm{log}}_2}\left({\frac{{f{0_{Max}}}}{{f{0_{Min}}}}} \right)
\]
\end{document}
</tex-math>
<graphic xlink:href="glossapx-5-1-48921-e1.gif"/>
</alternatives>
</disp-formula>
<p><italic>Pause duration</italic>: The duration of the silent interval after the second name was measured in milliseconds. For recordings without a pause, the duration was automatically coded as 0 ms. Hesitations and stop closures that were incorrectly labeled as pauses by the automatic aligner were manually recategorized during the verification process and not included in pause measurements.</p>
<p><italic>Final lengthening</italic>: The duration of the final vowel segment of the second name was measured in milliseconds.</p>
<p>We analyzed pause duration and final lengthening in milliseconds, rather than as ratios, as originally preregistered. This decision was based on three considerations: (i) both measures were highly correlated with their ratio-based counterparts (r &gt; .87), (ii) raw duration values can be directly analyzed without transformation by regression models that assume a <monospace>lognormal</monospace> response distribution, and (iii) raw values are more interpretable and can be more easily compared across participants and tasks. This change does not affect the core hypotheses or conclusions. Additionally, while we used the term <italic>f0 rise</italic> in the preregistration (following <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>), we use <italic>pitch range</italic> here to refer to the local f0 range between syllables, which better reflects the full range of contour shapes (rises, falls, mixed) used to signal prosodic boundaries (see <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>).</p>
</sec>
<sec>
<title>2.4.4 Data exclusion</title>
<p><bold>Data exclusion overall.</bold> From the total dataset, eight recordings (0.3%) were excluded, due to being incomplete, unintelligible, or of poor quality (e.g., background noise, excessive errors, or timing issues), or because reliable f0 measurements could not be obtained, due to issues like creaky voice or glottalization. These exclusion criteria were applied uniformly across all prosodic boundary cues and cognitive load task conditions, leaving 2872 recordings for the final analysis.</p>
<p><bold>Data exclusion on pause usage.</bold> As expected, we found that across stimulus conditions, pauses were primarily employed in the <italic>bracket</italic> stimulus condition, when needed to mark the prosodic boundary. In the <italic>no-bracket</italic> stimulus condition, only 14% of recordings (200 out of 1388) contained pauses. Conversely, in the <italic>bracket</italic> stimulus condition, 92.7% of recordings contained pauses, and only 7.3% (101 recordings) did not. Given this disparity, we focused our pause duration analysis exclusively on <italic>bracket</italic> stimuli, where pause usage serves a meaningful prosodic function. This approach mirrors the findings and analysis of Huttenlauch et al. (<xref ref-type="bibr" rid="B39">2021</xref>), who observed a similar pattern.</p>
</sec>
</sec>
<sec>
<title>2.5 Additional perception check for validation</title>
<p>To validate the perceptibility of the produced prosodic boundaries, we conducted a perception check using na&#239;ve listeners. Seven student labelers, unfamiliar with the study, listened to all recordings from the cognitive load production task in randomized order and categorized each as either <italic>bracket</italic> or <italic>no-bracket</italic> based on the perceived prosodic grouping. Labelers could replay recordings as needed and could also choose a &#8220;cannot determine&#8221; option for ambiguous cases. Response buttons were randomized but consistent for each labeler.</p>
<p>We modeled the perception data using a Bayesian multilevel logistic regression with a binomial outcome (correct/incorrect identification; &#8220;cannot determine&#8221; responses were excluded). The model included fixed effects for recording origin (high vs. low cognitive load), stimulus condition (<italic>bracket</italic> vs. <italic>no-bracket</italic>), and their interaction. We included by-subject and by-item random intercepts and slopes for all fixed effects. This approach allowed us to quantify how accurately listeners identified the presence vs. absence of an internal prosodic boundary and to determine whether identification was affected by cognitive load.</p>
<p>The perception check revealed high disambiguation accuracy across stimulus conditions, with na&#239;ve listeners correctly identifying the intended prosodic boundaries in over 90% of recordings. For recordings from the low cognitive load condition, accuracy was particularly high: 95.7% for <italic>bracket</italic> stimuli and 95.4% for <italic>no-bracket</italic> stimuli. While accuracy remained robust for recordings from the high cognitive load condition, it showed a slight decrease, to 90.3% for <italic>bracket</italic> stimuli and 93.0% for <italic>no-bracket</italic> stimuli. Ambiguity was rare across all stimulus conditions, with only 0.3&#8211;0.4% of recordings receiving &#8220;cannot determine&#8221; responses.</p>
<p>Bayesian multilevel logistic regression analysis revealed an effect of cognitive load on boundary identification accuracy (b = &#8211;0.456 log-odds, 95% credible interval (CI) [&#8211;0.774, &#8211;0.141]), indicating that prosodic boundaries were less accurately identified in recordings produced under high cognitive load than under low cognitive load. This corresponds to a 3.1 percentage point decrease in accuracy (low load: 97.8% vs. high load: 94.7%, estimated marginal means averaged across stimulus condition and random effects). We found no effect of stimulus condition (b = 0.0981 log-odds, 95% CI [&#8211;0.354, 0.546]), suggesting similar perceptibility for <italic>bracket</italic> and <italic>no-bracket</italic> productions. Also, no interaction between cognitive load and stimulus condition was observed (b = 0.012 log-odds, 95% CI [&#8211;0.503, 0.541]).</p>
<p>These findings confirm that participants successfully produced perceptually meaningful prosodic boundaries, with cognitive load creating detectable but relatively modest reductions in communicative effectiveness.</p>
</sec>
</sec>
<sec>
<title>3. Statistical modeling</title>
<p>The statistical models were fit within a Bayesian framework, which allows combining prior information with data to generate posterior distributions for each parameter (<xref ref-type="bibr" rid="B75">Vasishth et al., 2018</xref>; <xref ref-type="bibr" rid="B76">Ver&#237;ssimo, 2024</xref>). Bayesian mixed-effects distributional regression models were used via the <monospace>brms</monospace> package with <monospace>RStan</monospace> (<xref ref-type="bibr" rid="B9">Buerkner, 2018</xref>; <xref ref-type="bibr" rid="B71">Stan Development Team, 2020</xref>) to examine how the three prosodic boundary cues (pitch range, pause duration, final lengthening) related to auditory-perceptual acuity under varying cognitive loads. Unlike traditional regression approaches that focus solely on mean differences, distributional regression models estimate both mean prosodic boundary cue realization, modeled by a location parameter for the <monospace>mean</monospace> (<italic>&#956;</italic>), and trial-to-trial variability in prosodic boundary cue realization, modeled by a <monospace>sigma</monospace> parameter (<italic>&#963;</italic>). In the present study, <italic>&#956;</italic> captures production distinctiveness, whereas <italic>&#963;</italic> captures production variability, with lower <italic>&#963;</italic> indicating lower production variability. This approach aligns with recommendations to move beyond models that estimate only mean differences to capture how individuals differ not only in magnitude but also in consistency of cognitive processes (<xref ref-type="bibr" rid="B13">Ciaccio &amp; Ver&#237;ssimo, 2022</xref>).</p>
<p>For each prosodic boundary cue, we used a model tailored to the response variable&#8217;s distribution. Pitch range was modeled with a Gaussian distribution, appropriate for semitone-scale measurements. Pause duration (modeled only for <italic>bracket</italic> stimuli) required a hurdle-lognormal model: the <monospace>hurdle</monospace> component captured zero outcomes (no pauses, i.e., pause duration = 0), while the <monospace>lognormal</monospace> component modeled positive durations in <italic>bracket</italic> stimulus productions, where pauses serve a boundary-marking function (see 2.4.4). Final lengthening was modeled using a lognormal distribution. Both pause duration and final lengthening employed lognormal distributions, because their values are strictly positive and exhibit positively skewed distributions (<xref ref-type="bibr" rid="B13">Ciaccio &amp; Ver&#237;ssimo, 2022</xref>).</p>
<p>All dependent variables in the models are the acoustic feature measures of the prosodic boundary cues derived from the production task; auditory-perceptual acuity enters the models exclusively as a between-participants perception predictor derived from the independent JND perception tasks. The fixed effects in all models included stimulus condition (<italic>bracket</italic> vs. <italic>no-bracket</italic>) and cognitive load (high vs. low), both coded using sum contrasts (e.g., <italic>bracket</italic>/high load = +0.5; <italic>no-bracket</italic>/low load = &#8211;0.5). This contrast coding ensures that effects and interactions are interpreted at the mean of other predictors (i.e., averaged across their levels), rather than at a specific reference level. Additionally, we fitted models with nested contrasts to follow up on interactions between stimulus condition and cognitive load (<xref ref-type="bibr" rid="B68">Schad et al., 2020</xref>). All models included z-scored auditory-perceptual acuity scores (pitch acuity, pause acuity, final lengthening acuity).</p>
<p>We implemented a maximal random effects structure (<xref ref-type="bibr" rid="B3">Barr et al., 2013</xref>). Random effects included random intercepts for subjects and items, random slopes for stimulus condition and cognitive load for subjects and items, and random slopes for acuity scores for items only (as acuity is a between-participants measure). As preregistered, no random effects were included for the residual variability (<monospace>sigma</monospace>) component, because models with random effects on <monospace>sigma</monospace> take considerably longer to run, do not converge as easily, and we likely do not have sufficient data to estimate those parameters reliably (<xref ref-type="bibr" rid="B4">Bates et al., 2015</xref>). We implemented weakly-informative priors centered around zero on all parameters, following current recommendations (<xref ref-type="bibr" rid="B24">Gelman et al., 2008</xref>; J. <xref ref-type="bibr" rid="B26">Ghosh et al., 2018</xref>; <xref ref-type="bibr" rid="B51">McElreath, 2020</xref>; <xref ref-type="bibr" rid="B75">Vasishth et al., 2018</xref>). For intercept priors, we incorporated empirical data from Huttenlauch et al. (<xref ref-type="bibr" rid="B39">2021</xref>). The prior specifications were chosen to rule out only unreasonably extreme values, while still allowing large effects in either direction, as informed by the data. The appropriateness of these priors was confirmed through prior predictive checks. Prior specifications are provided in 1.1 (Prior distributions for all parameter estimates in the Bayesian distributional models) in the Supplementary Materials (page 2, <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
<p>The distributional models can be summarized schematically as follows, where &#215; denotes the full factorial expansion, including all lower-order terms:</p>
<p><bold>Pitch range model (Gaussian):</bold></p>
<disp-formula id="FD2">
<label>(E1)</label>
<alternatives>
<mml:math id="Eq002-mml">
<mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mtext>Pitch&#x00A0;range</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi mathvariant='script'>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Stimulus&#x00A0;condition</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pitch&#x00A0;acuity</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Stimulus&#x00A0;condition</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mtext>Pitch&#x00A0;acuity&#x00A0;&#x2223;&#x00A0;Participant</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Stimulus&#x00A0;condition</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pitch&#x00A0;acuity</mml:mtext><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mtext>Item</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mtext>Stimulus&#x00A0;condition</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pitch&#x00A0;acuity</mml:mtext></mml:mtd></mml:mtr></mml:mtable>
</mml:math>
<tex-math id="M2">
\documentclass[10pt]{article}
\usepackage{wasysym}
\usepackage[substack]{amsmath}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage[mathscr]{eucal}
\usepackage{mathrsfs}
\usepackage{pmc}
\usepackage[Euler]{upgreek}
\pagestyle{empty}
\oddsidemargin -1.0in
\begin{document}
\[
\begin{array}{l}
{\rm Pitch rang}{e_{ij}}\,\,\, \sim \;{\cal N}\left({{\mu _{ij}},{\sigma _{ij}}} \right)\\
\quad\quad\quad\quad\;\,{\mu _{ij}}\quad = \;{\beta _0} + {\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Pitch\,acuity}\\
\quad\quad\quad\quad\quad\quad\quad + \left({1 + {\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Pitch\,acuity}\mid {{\rm Participant}_i}} \right)\\
\quad\quad\quad\quad\quad\quad\,\;\;\; + \left({1 + {\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Pitch\,acuity}\mid {\rm Item}_{j}}\right)\\
\quad\quad\,{\rm{log}}\left({{\sigma _{ij}}} \right)\quad = \;{\gamma _0}\, + \,\,{\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Pitch\,acuity}
\end{array}
\]
\end{document}
</tex-math>
<graphic xlink:href="glossapx-5-1-48921-e2.gif"/>
</alternatives>
</disp-formula>
<p><bold>Pause duration model (hurdle-lognormal, <italic>bracket</italic> only, no stimulus condition contrast):</bold></p>
<disp-formula id="FD3">
<label>(E2)</label>
<alternatives>
<mml:math id="Eq003-mml">
<mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pr</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mtext>Pause</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mtext>logit</mml:mtext><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mtext>Cognitive</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pause&#x00A0;acuity</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mtext>Pause&#x00A0;duration</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mtext>Pause</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x223C;</mml:mo><mml:mtext>lognormal</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pause&#x00A0;acuity</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mtext>Cognitive&#x00A0;load&#x00A0;&#x2223;&#x00A0;Participant</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive&#x00A0;load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pause&#x00A0;acuity</mml:mtext><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mtext>Item</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Cognitive</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mtext>load</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>Pause</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mtext>acuity</mml:mtext></mml:mtd></mml:mtr></mml:mtable>
</mml:math>
<tex-math id="M3">
\documentclass[10pt]{article}
\usepackage{wasysym}
\usepackage[substack]{amsmath}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage[mathscr]{eucal}
\usepackage{mathrsfs}
\usepackage{pmc}
\usepackage[Euler]{upgreek}
\pagestyle{empty}
\oddsidemargin -1.0in
\begin{document}
\[
\begin{array}{l}
\quad\quad\quad\quad\quad\quad\,\,{\rm Pr}\left({{{\rm Pause}_{ij}} > 0} \right)\, = {\rm logit}^{-1}\left({{\alpha_0} + {\rm Cognitive\,load} \times {\rm Pause\,acuity}} \right)\\
{\rm Pause\,duration}_{ij}\mid {\rm Pause}_{ij} > 0\,\,\, \sim {\rm lognormal}\left( {{\mu _{ij}},{\sigma _{ij}}} \right)\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\;{\mu _{ij}} = {\beta _0} + \,\,{\rm Cognitive\,load} \times {\rm Pause\,acuity}\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\,\,+ \left({1 + {\rm Cognitive\,load}\mid {\rm Participant}_{i}} \right)\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\,\,+ \left({1 + {\rm Cognitive\,load} \times {\rm Pause\,acuity}\mid {\rm Item}_{j}} \right)\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad{\rm{log}}\left({{\sigma _{ij}}} \right) = {\gamma _0} + {\rm Cognitive\,load} \times {\rm Pause\,acuity}
\end{array}
\]
\end{document}
</tex-math>
<graphic xlink:href="glossapx-5-1-48921-e3.gif"/>
</alternatives>
</disp-formula>
<p><bold>Final lengthening model (lognormal):</bold></p>
<disp-formula id="FD4">
<label>(E3)</label>
<alternatives>
<mml:math id="Eq004-mml">
<mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mtext>Final&#x00A0;lengthening</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mtext>lognormal</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Stimulus&#x0020;condition</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Cognitive&#x0020;load</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Final&#x0020;lengthening&#x0020;acuity</mml:mi><mml:mo>&#x00A0;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Stimulus&#x0020;condition</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mtext>Cognitive&#x0020;load&#x00A0;&#x2223;&#x00A0;Participant</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Stimulus&#x0020;condition</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mtext>Cognitive&#x0020;load</mml:mtext><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Final&#x0020;lengthening&#x0020;acuity</mml:mi><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mtext>Item</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">ij</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Stimulus&#x0020;condition</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mtext>Cognitive&#x0020;load</mml:mtext><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>Final&#x0020;lengthening&#x0020;acuity</mml:mi></mml:mtd></mml:mtr></mml:mtable>
</mml:math>
<tex-math id="M4">
\documentclass[10pt]{article}
\usepackage{wasysym}
\usepackage[substack]{amsmath}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage[mathscr]{eucal}
\usepackage{mathrsfs}
\usepackage{pmc}
\usepackage[Euler]{upgreek}
\pagestyle{empty}
\oddsidemargin -1.0in
\begin{document}
\[
\begin{array}{l}
{\rm Final\,lengthening}_{ij}\, \sim {\rm lognormal}\left({{\mu _{ij}},{\sigma _{ij}}} \right)\\
\quad\quad\quad\quad\quad\quad\quad\;{\mu_{ij}}\;=\;{\beta_0}+\,\,{\rm Stimulus\,condition}\times{\rm Cognitive\,load}\times{\rm Final\,lengthening\,acuity}\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad + \left({1 + {\rm Stimulus\,condition} \times {\rm Cognitive\,load}\mid {\rm Participant}_{i}}\right)\\
\quad\quad\quad\quad\quad\quad\quad\quad\quad + \left({1 + {\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Final\,lengthening\,acuity}\mid {\rm Item}_{j}} \right)\\
\;\;\quad\quad\quad\quad\quad{\rm{log}}\left({{\sigma_{ij}}} \right)\, = {\gamma _0} + \,\,{\rm Stimulus\,condition} \times {\rm Cognitive\,load} \times {\rm Final\,lengthening\,acuity}
\end{array}
\]
\end{document}
</tex-math>
<graphic xlink:href="glossapx-5-1-48921-e4.gif"/>
</alternatives>
</disp-formula>
<p>The fixed effects in both <italic>&#956;</italic> and <italic>&#963;</italic> follow the same predictor structure, so effects on both parameters can be interpreted in parallel. A stimulus condition effect captures whether <italic>bracket</italic> and <italic>no-bracket</italic> productions differ in mean prosodic boundary cue realization (<italic>&#956;</italic>) or in trial-to-trial variability (<italic>&#963;</italic>, the spread of realizations across trials within <italic>bracket</italic> vs. within <italic>no-bracket</italic>). The interaction between stimulus condition and auditory-perceptual acuity shows whether this <italic>bracket/no-bracket</italic> difference varies with individual perceptual ability. Finally, the three-way interaction with cognitive load shows whether this modulation by auditory-perceptual acuity itself changes between low and high cognitive load.</p>
<p>Hypothesis testing was performed using Bayes factors, which quantify the strength of evidence for an effect by comparing an alternative model (including the effect) to a null model (excluding the effect). Since Bayes factors can be sensitive to prior specifications, we conducted a sensitivity analysis with five different prior configurations, ranging from narrower (moderately and strongly informative) to wider (moderately wide and wide) settings, compared to our default weakly-informative priors. The natural logarithm of the Bayes factor (lnBF<sub>10</sub>) was calculated using the Savage-Dickey method (<xref ref-type="bibr" rid="B18">Dickey &amp; Lientz, 1970</xref>; <xref ref-type="bibr" rid="B78">Wagenmakers et al., 2010</xref>). In this scale, natural-logged Bayes factors greater than 1 provide evidence for the alternative hypothesis, and values below &#8211;1 provide evidence for the null hypothesis, with values between &#8211;1 and 1 considered inconclusive or weak (<xref ref-type="bibr" rid="B41">Kass &amp; Raftery, 1995</xref>; <xref ref-type="bibr" rid="B76">Ver&#237;ssimo, 2024</xref>). Natural-logged Bayes factors greater than 3 are interpreted as strong evidence against the null hypothesis (see <xref ref-type="bibr" rid="B40">Jeffreys, 1991</xref>; <xref ref-type="bibr" rid="B41">Kass &amp; Raftery, 1995</xref>).</p>
<p>Model convergence was assessed using R-hat values, effective sample size, and visual inspection of trace plots. Additionally, posterior predictive checks were conducted, to evaluate model fit to the observed data distribution.</p>
</sec>
<sec>
<title>4. Results</title>
<p>We report the main perception-production analyses separately for each prosodic boundary cue: pitch range, pause duration, and final lengthening. We first report effects involving production distinctiveness (<monospace>mean</monospace>, <italic>&#956;</italic>; H1) and then effects involving production variability (<monospace>sigma</monospace>, <italic>&#963;</italic>; H2). For pitch range and final lengthening, <italic>production distinctiveness</italic> refers to the difference in mean cue realization between <italic>bracket</italic> and <italic>no-bracket</italic> productions. <italic>Production variability</italic> refers to the difference in trial-to-trial variability of cue realizations between <italic>bracket</italic> and <italic>no-bracket</italic> productions.</p>
<p>For pause duration, analyses focus on <italic>bracket</italic> productions only, because pauses were almost absent in <italic>no-bracket</italic> productions. Thus, pause distinctiveness (H1) refers to mean pause duration within <italic>bracket</italic> productions, and pause variability (H2) refers to the trial-to-trial variability of pause durations within <italic>bracket</italic> productions.</p>
<p>All <italic>&#963;</italic> estimates, their credible intervals, and Bayes factors are reported on the back-transformed SD scale of each dependent variable, while the raw log-scale estimates are provided in the Supplementary Materials (page 3 for pitch range, page 10 for pause duration, page 18 for final lengthening; <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
<p>Effects are reported when their natural-logged Bayes factors (lnBF<italic><sub>10</sub></italic>) exceed 1 in the base model with weakly-informative priors and show consistent patterns across the sensitivity analyses with different prior specifications. The evidence strength for such effects is characterized as moderate when lnBF<italic><sub>10</sub></italic> &gt; 1, and strong when lnBF<italic><sub>10</sub></italic> &gt; 3. Effects with lnBF<italic><sub>10</sub></italic> values between 0 and 1 are reported as suggestive but inconclusive, reflecting limited evidence that is insufficient to clearly favor either the null or the alternative hypothesis. Such effects indicate directional tendencies in the data, but should not be interpreted as robust support for the corresponding hypothesis. We, nevertheless, report such effects when they become stronger under narrower priors or when follow-up analyses reveal simple effects with lnBF<italic><sub>10</sub></italic> &gt; 1.</p>
<sec>
<title>4.1 Pitch range</title>
<p>Full fixed-effects tables, posterior distribution plots for both <monospace>mean</monospace> (<italic>&#956;</italic>) and variability (<monospace>sigma</monospace>, <italic>&#963;</italic>) parameters, and prior sensitivity analyses for the pitch range models are reported in 1.2 (Pitch range) in the Supplementary Materials (pages 3&#8211;8, <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
<sec>
<title>4.1.1 Pitch range distinctiveness</title>
<p>We found suggestive evidence that pitch acuity (perception) modulated pitch range distinctiveness, that is, how strongly participants used pitch range to distinguish <italic>bracket</italic> productions from <italic>no-bracket</italic> productions, depending on their pitch acuity and averaged across cognitive load conditions (stimulus condition &#215; pitch acuity interaction on the <monospace>mean</monospace>: b = 0.363 semitones, 95% CI [0.035, 0.688], lnBF<italic><sub>10</sub></italic> = 0.59). The evidence for this effect strengthened to moderate levels under narrower prior specifications (lnBF<italic><sub>10</sub></italic> up to 1.03). All participants used a higher pitch range in <italic>bracket</italic> productions than <italic>no-bracket</italic> productions, but the size of this difference varied with pitch acuity. As illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref>, participants at the upper end of the pitch acuity range (+1 SD) produced a <italic>bracket/no-bracket</italic> pitch range difference of 3.193 semitones (95% CI [2.710, 3.673], lnBF<italic><sub>10</sub></italic> &gt; 23), compared to 1.915 semitones (95% CI [1.057, 2.788], lnBF<italic><sub>10</sub></italic> = 5.80) for participants at the lower end (&#8211;2.5 SD).</p>
<p>To understand where this interaction originates, we examined the effect of pitch acuity separately within each stimulus condition. In <italic>bracket</italic> productions, the regression line has a positive slope: pitch range increased with pitch acuity, meaning that participants with higher pitch acuity used a higher pitch range when marking a boundary than participants with lower pitch acuity, though the evidence for this relationship was only suggestive (b = 0.407 semitones, 95% CI [&#8211;0.002, 0.817], lnBF<italic><sub>10</sub></italic> = 0.25). In <italic>no-bracket</italic> productions, by contrast, the regression line is nearly flat: pitch range was unrelated to pitch acuity (b = 0.044 semitones, 95% CI [&#8211;0.221, 0.313], lnBF<italic><sub>10</sub></italic> = &#8211;2.1), meaning that all participants produced similarly lower pitch ranges when no boundary was to be marked, regardless of their pitch acuity. This asymmetry in slopes is visually apparent in <xref ref-type="fig" rid="F2">Figure 2</xref>. However, given the limited evidence for the simple effect in <italic>bracket</italic> productions, the magnitude of this association is likely small and should be interpreted with caution.</p>
<fig id="F2">
<caption>
<p><bold>Figure 2:</bold> Predicted mean pitch range (in semitones) as a function of pitch acuity (z-scored; higher values indicate better auditory-perceptual discrimination), averaged across cognitive load conditions. The plot illustrates pitch range distinctiveness (mean), that is, how strongly participants used pitch range to distinguish <italic>bracket</italic> from <italic>no-bracket</italic> productions across the pitch acuity range. Regression lines represent posterior mean estimates, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals. Yellow lines correspond to <italic>bracket</italic> stimuli, and green lines correspond to <italic>no-bracket</italic> stimuli.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g2.png"/>
</fig>
<p>Finally, cognitive load did not modulate the relationship between pitch acuity and the pitch range difference between <italic>bracket</italic> and <italic>no-bracket</italic> productions (stimulus condition &#215; cognitive load &#215; pitch acuity interaction on the <monospace>mean</monospace>: b = 0.187 semitones, 95% CI [&#8211;0.097, 0.473], lnBF<italic><sub>10</sub></italic> = &#8211;1.09, evidence for the null), consistent with the visual pattern in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
</sec>
<sec>
<title>4.1.2 Pitch range variability</title>
<p>We found moderate evidence that pitch acuity (perception) modulated pitch range variability (<monospace>sigma</monospace>, <italic>&#963;</italic>) differently across cognitive load conditions (stimulus condition &#215; cognitive load &#215; pitch acuity interaction on <monospace>sigma</monospace>: b = 0.130 SD semitones, 95% CI [0.021, 0.238], lnBF<italic><sub>10</sub></italic> = 1.10). Here, variability refers to how much each individual participant&#8217;s own pitch range fluctuated from trial to trial. The three-way interaction indicates that how much more a given participant&#8217;s pitch range varied trial-to-trial in <italic>bracket</italic> productions compared to <italic>no-bracket</italic> productions, depended on both their pitch acuity and the cognitive load condition, as illustrated in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3">
<caption>
<p><bold>Figure 3:</bold> Predicted pitch range variability (sigma; SD of semitones, higher values indicate greater trial-to-trial variability) as a function of pitch acuity (z-scored; higher values indicate better auditory-perceptual acuity), stimulus condition (<italic>bracket</italic> vs. <italic>no-bracket</italic>), and cognitive load (low vs. high). The plot illustrates how the <italic>bracket/no-bracket</italic> difference in trial-to-trial pitch range variability scales with pitch acuity under each cognitive load condition. Regression lines represent posterior estimates of sigma, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals. Yellow lines correspond to <italic>bracket</italic> stimuli, and green lines correspond to <italic>no-bracket</italic> stimuli.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g3.png"/>
</fig>
<p>To understand where this interaction originates, we examined how pitch acuity related to the <italic>bracket/no-bracket</italic> difference in trial-to-trial variability separately under each cognitive load condition.</p>
<p>Under high cognitive load, pitch acuity strongly predicted differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions in trial-to-trial pitch range variability (b = 0.149 SD semitones, 95% CI [0.071, 0.228], lnBF<italic><sub>10</sub></italic> = 4.68). This is visible in the right panel of <xref ref-type="fig" rid="F3">Figure 3</xref> as a divergence between the two regression lines as pitch acuity increases: the <italic>bracket</italic> line has a positive slope, while the <italic>no-bracket</italic> line remains nearly flat. Participants with higher pitch acuity (+1 SD) showed considerably more trial-to-trial variability in pitch range during <italic>bracket</italic> productions than during <italic>no-bracket</italic> productions (b = 0.411 SD semitones, 95% CI [0.283, 0.543], lnBF<italic><sub>10</sub></italic> &gt; 23). This means, when marking a prosodic boundary under cognitive load, these participants drew on a wider range of pitch range realizations across repeated trials. Participants with lower pitch acuity (&#8211;2.5 SD), by contrast, showed no such difference in trial-to-trial variability between <italic>bracket</italic> and <italic>no-bracket</italic> productions (b = &#8211;0.034 SD semitones, 95% CI [&#8211;0.187, 0.120], lnBF<italic><sub>10</sub></italic> = &#8211;2.01). In the right panel of <xref ref-type="fig" rid="F3">Figure 3</xref>, this is visible at the lower end of the pitch acuity range, where the <italic>bracket</italic> and <italic>no-bracket</italic> regression lines converge and overlap, indicating that participants with lower pitch acuity showed similar trial-to-trial variability in both <italic>bracket</italic> and <italic>no-bracket</italic> productions, regardless of whether a boundary needed to be marked.</p>
<p>Under low cognitive load, pitch acuity did not predict a difference between <italic>bracket</italic> and <italic>no-bracket</italic> productions in trial-to-trial variability (b = 0.013 SD semitones, 95% CI [&#8211;0.068, 0.092], lnBF<italic><sub>10</sub></italic> = &#8211;2.07, evidence for the null). As visible in the left panel of <xref ref-type="fig" rid="F3">Figure 3</xref>, both regression lines have a positive slope and run in parallel. Participants with higher pitch acuity showed greater trial-to-trial variability in pitch range overall, in both <italic>bracket</italic> and <italic>no-bracket</italic> productions. The <italic>bracket/no-bracket</italic> difference in trial-to-trial variability was present across all pitch acuity levels and did not scale with pitch acuity, in contrast to the pattern observed under high cognitive load.</p>
</sec>
</sec>
<sec>
<title>4.2 Pause duration</title>
<p>Complete model outputs for pause duration, including fixed effects for the <monospace>lognormal</monospace> and <monospace>hurdle</monospace> components, posterior distributions, and prior sensitivity analyses, are reported in 1.3 (Pause duration) in the Supplementary Materials (pages 9&#8211;16, <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
<sec>
<title>4.2.1 Pause duration distinctiveness</title>
<p>We found no evidence that pause acuity (perception) modulated pause duration distinctiveness, that is, how strongly participants used pause duration to mark a prosodic boundary, depending on their pause acuity and averaged across cognitive load conditions. There was no main effect of pause acuity on pause duration (b = &#8211;0.108 log-ms, 95% CI [&#8211;0.269, 0.058], lnBF<italic><sub>10</sub></italic> = &#8211;0.96), and no interaction between cognitive load and pause acuity (b = &#8211;0.012 log-ms, 95% CI [&#8211;0.141, 0.117], lnBF<italic><sub>10</sub></italic> = &#8211;2.01). Participants produced pauses of similar duration, regardless of their pause acuity.</p>
<p>Cognitive load did, however, strongly modulate pause duration (b = &#8211;0.277 log-ms, 95% CI [&#8211;0.410, &#8211;0.143], lnBF<italic><sub>10</sub></italic> = 5.65). As illustrated in <xref ref-type="fig" rid="F4">Figure 4</xref>, participants with average pause acuity produced shorter pauses under high cognitive load than under low cognitive load. Given the absence of an interaction between cognitive load and pause acuity, this reduction was consistent across all pause acuity levels.</p>
<fig id="F4">
<caption>
<p><bold>Figure 4:</bold> Predicted mean pause duration (log-ms) as a function of cognitive load (low vs. high), averaged across pause acuity. The plot illustrates how strongly participants used pause duration to mark a prosodic boundary within <italic>bracket</italic> productions across cognitive load conditions. Points and intervals represent posterior mean estimates, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g4.png"/>
</fig>
</sec>
<sec>
<title>4.2.2 Pause duration variability</title>
<p>We found moderate evidence that pause acuity (perception) modulated pause duration variability (b = 0.027 SD log-ms, 95% CI [0.006, 0.049], lnBF<italic><sub>10</sub></italic> = 1.49). Here, variability refers to how much each individual participant&#8217;s own pause duration fluctuated from trial to trial in <italic>bracket</italic> productions only. As illustrated in panel A of <xref ref-type="fig" rid="F5">Figure 5</xref>, the regression line has a positive slope: participants with higher pause acuity showed greater trial-to-trial variability in pause duration than participants with lower pause acuity.</p>
<p>We also found strong evidence that cognitive load modulated pause duration variability (b = 0.093 SD log-ms, 95% CI [0.055, 0.131], lnBF<italic><sub>10</sub></italic> &gt; 23). As illustrated in panel B of <xref ref-type="fig" rid="F5">Figure 5</xref>, participants with average pause acuity showed greater trial-to-trial variability in pause duration under high cognitive load than under low cognitive load. Given the absence of an interaction between cognitive load and pause acuity, this pattern held across all pause acuity levels.</p>
<p>These two effects were additive, rather than interactive. We found no evidence for an interaction between cognitive load and pause acuity on trial-to-trial variability (b = 0.003 SD log-ms, 95% CI [&#8211;0.039, 0.045], lnBF<italic><sub>10</sub></italic> = &#8211;1.06). That is, higher pause acuity was associated with greater trial-to-trial variability in pause duration across both cognitive load conditions, and high cognitive load increased trial-to-trial variability across all pause acuity levels.</p>
<fig id="F5">
<caption>
<p><bold>Figure 5:</bold> Predicted pause duration variability (sigma; SD of log-ms, higher values indicate greater trial-to-trial variability) as a function of (A) pause acuity (z-scored; higher values indicate better auditory-perceptual acuity), averaged across cognitive load conditions, and (B) cognitive load (low vs. high), averaged across pause acuity. The plots illustrate how trial-to-trial variability in pause duration relates to pause acuity and cognitive load within <italic>bracket</italic> productions. Regression lines in Panel A and posterior intervals in Panel B represent posterior estimates of sigma, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g5.png"/>
</fig>
</sec>
<sec>
<title>4.2.3 Pause usage (hurdle parameter)</title>
<p>As an exploratory analysis, we examined whether pause acuity and cognitive load predicted whether participants used pauses at all to mark prosodic boundaries, independent of pause duration. This was assessed using the <monospace>hurdle</monospace> component of the pause model, which estimates the probability of producing no pause (i.e., a pause duration of zero). For ease of interpretation, the results are reported as pause usage probabilities (1 &#8211; <monospace>hurdle</monospace> probability). This analysis was not preregistered.</p>
<p>We found no evidence that the relationship between pause acuity and pause usage differed across cognitive load conditions (cognitive load &#215; pause acuity interaction on <monospace>hurdle</monospace>: b = &#8211;0.357 log-odds, 95% CI [&#8211;0.884, 0.147], lnBF<italic><sub>10</sub></italic> = &#8211;0.43, Bayes factors consistently favored, or tended toward, the null). Pause acuity did, however, predict pause usage. We found moderate evidence that higher pause acuity was associated with a higher probability of producing no pauses, and, thus, with reduced pause usage across all cognitive load conditions (b = 0.403 log-odds, 95% CI [0.146, 0.672], lnBF<italic><sub>10</sub></italic> = 2.56). As illustrated in Panel A of <xref ref-type="fig" rid="F6">Figure 6</xref>, the regression line has a negative slope: participants at the lower end of the pause acuity range (&#8211;2.3 SD) used pauses in 96.21% of <italic>bracket</italic> productions, compared to 87.18% for participants at the upper end (+2 SD).</p>
<fig id="F6">
<caption>
<p><bold>Figure 6:</bold> Predicted pause usage probability (expressed as the complement of the hurdle zero-pause probability) as a function of (A) pause acuity (z-scored; higher values indicate better auditory-perceptual acuity), averaged across cognitive load conditions, and (B) cognitive load (low vs. high), averaged across pause acuity. The plots illustrate how the probability of using a pause to mark a prosodic boundary within <italic>bracket</italic> productions relates to pause acuity and cognitive load. Points and intervals represent posterior predictions, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g6.png"/>
</fig>
<p>We also found strong evidence that cognitive load predicted pause usage (b = 1.209 log-odds, 95% CI [0.738, 1.711], lnBF<italic><sub>10</sub></italic> &gt; 23): high cognitive load increased the probability of producing no pauses, corresponding to reduced pause usage under load. As illustrated in Panel B of <xref ref-type="fig" rid="F6">Figure 6</xref>, participants with average pause acuity used pauses in 96.63% of <italic>bracket</italic> productions under low cognitive load, compared to 89.60% under high cognitive load. Given the absence of an interaction between cognitive load and pause acuity, this pattern held across all pause acuity levels.</p>
</sec>
</sec>
<sec>
<title>4.3 Final lengthening</title>
<p>Detailed fixed-effects estimates, posterior distributions for <monospace>mean</monospace> and <monospace>sigma</monospace> parameters, and prior sensitivity analyses for the final lengthening models are reported in 1.4 (Final lengthening) in the Supplementary Materials (pages 17&#8211;22, <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
<sec>
<title>4.3.1 Final lengthening distinctiveness</title>
<p>We found no evidence that final lengthening acuity modulated final lengthening production distinctiveness, that is, how strongly participants used final lengthening to distinguish <italic>bracket</italic> productions from <italic>no-bracket</italic> productions, depending on their final lengthening acuity and averaged across cognitive load conditions. There was no evidence for a three-way interaction between stimulus condition, cognitive load, and final lengthening acuity (b = 0.018 log-ms, 95% CI [&#8211;0.029, 0.066], lnBF<italic><sub>10</sub></italic> = &#8211;2.74), and no evidence for a two-way interaction between stimulus condition and final lengthening acuity (b = 0.013 log-ms, 95% CI [&#8211;0.057, 0.084], lnBF<italic><sub>10</sub></italic> = &#8211;2.57). Participants used final lengthening to distinguish <italic>bracket</italic> from <italic>no-bracket</italic> productions to a similar degree, regardless of their final lengthening acuity.</p>
</sec>
<sec>
<title>4.3.2 Final lengthening variability</title>
<p>We found no evidence that final lengthening acuity modulated final lengthening variability (<monospace>sigma</monospace>, <italic>&#963;</italic>). Here, variability refers to how much each individual participant&#8217;s own final lengthening fluctuated from trial to trial. There was no three-way interaction between stimulus condition, cognitive load, and final lengthening acuity (b = &#8211;0.010 SD log-ms, 95% CI [&#8211;0.030, 0.010], lnBF<italic><sub>10</sub></italic> = &#8211;1.51), an inconclusive two-way interaction between stimulus condition and final lengthening acuity (b = &#8211;0.011 SD log-ms, 95% CI [&#8211;0.021, &#8211;0.001], lnBF<italic><sub>10</sub></italic> = &#8211;0.29, all Bayes factors consistently favored, or tended toward, the null), and no main effect of final lengthening acuity (b = &#8211;0.002 SD log-ms, 95% CI [&#8211;0.007, 0.003], lnBF<italic><sub>10</sub></italic> = &#8211;2.97). Participants with higher and lower final lengthening acuity showed similar trial-to-trial variability in final lengthening across both <italic>bracket</italic> and <italic>no-bracket</italic> productions.</p>
<p>Cognitive load did, however, strongly modulate final lengthening variability (b = 0.028 SD log-ms, 95% CI [0.019, 0.037], lnBF<italic><sub>10</sub></italic> &gt; 23). As illustrated in <xref ref-type="fig" rid="F7">Figure 7</xref>, participants with average final lengthening acuity showed greater trial-to-trial variability in final lengthening under high cognitive load than under low cognitive load. Given the absence of an interaction between cognitive load and final lengthening acuity, this pattern held across all final lengthening acuity levels.</p>
<fig id="F7">
<caption>
<p><bold>Figure 7:</bold> Predicted final lengthening variability (sigma; SD of log-ms, higher values indicate greater trial-to-trial variability) as a function of cognitive load (low vs. high), averaged across final lengthening acuity and stimulus condition. The plot illustrates how trial-to-trial variability in final lengthening relates to cognitive load across all final lengthening acuity levels. Points and intervals represent posterior estimates of sigma, with shaded bands indicating 68% (darker) and 95% (lighter) credible intervals.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-48921-g7.png"/>
</fig>
</sec>
</sec>
<sec>
<title>4.4 Summary of PP-link findings</title>
<p><xref ref-type="table" rid="T1">Table 1</xref> provides an overview of our findings across all three prosodic boundary cues, summarizing the evidence for our two main hypotheses along with a short interpretation statement for each perception-production relationship examined.</p>
<table-wrap id="T1">
<caption>
<p><bold>Table 1:</bold> Summary of findings for prosodic boundary cue PP-links.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"><bold>Hypothesis</bold></td>
<td align="left" valign="top"><bold>Main findings</bold></td>
<td align="left" valign="top"><bold>Interpretation</bold></td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><bold>Pitch range</bold></td>
</tr>
<tr>
<td align="left" valign="top"><bold>H1 (mean, <italic>&#956;</italic>): Distinctiveness</bold></td>
<td align="left" valign="top">Higher pitch acuity &#8594; larger pitch range difference between <italic>bracket</italic> and <italic>no-bracket</italic> productions (<italic>&#956;</italic>), stronger under load.</td>
<td align="left" valign="top">Suggestive evidence for an acuity &#215; stimulus condition interaction on mean pitch range (<italic>&#956;</italic>), averaged across cognitive load conditions. No evidence for an acuity &#215; cognitive load &#215; stimulus condition interaction.</td>
<td align="left" valign="top">Pitch acuity shows a directionally consistent with, but inconclusive association with, pitch range distinctiveness (mean pitch range differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions); interpret cautiously. Partial support for H1 (<italic>&#956;</italic>).</td>
</tr>
<tr>
<td align="left" valign="top"><bold>H2 (sigma, <italic>&#963;</italic>): Variability</bold></td>
<td align="left" valign="top">Higher pitch acuity &#8594; lower trial-to-trial variability (smaller <italic>&#963;</italic>, less variable), esp. under load.</td>
<td align="left" valign="top">Moderate evidence for an acuity &#215; cognitive load &#215; stimulus condition interaction on pitch range variability (<italic>&#963;</italic>). Under high cognitive load, <italic>bracket</italic>/<italic>no-bracket</italic> differences in trial-to-trial variability (<italic>&#963;</italic>) are larger for participants with higher pitch acuity than for participants with lower pitch acuity.</td>
<td align="left" valign="top">Higher pitch acuity is associated with greater (not lower) pitch range variability, expressed as stronger <italic>bracket</italic>/<italic>no-bracket</italic> differences in trial-to-trial variability under high cognitive load, which contradicts the prediction. No support for H2 (<italic>&#963;</italic>).</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><bold>Pause duration (<italic>bracket</italic> only)</bold></td>
</tr>
<tr>
<td align="left" valign="top"><bold>H1 (mean, <italic>&#956;</italic>): Distinctiveness</bold></td>
<td align="left" valign="top">Higher pause acuity &#8594; longer pauses within <italic>bracket</italic> productions (<italic>&#956;</italic>), stronger under load.</td>
<td align="left" valign="top">Moderate evidence for a main effect of cognitive load, with higher load associated with shorter pauses. No evidence for an acuity &#215; cognitive load interaction or for a main effect of pause acuity on mean pause duration (<italic>&#956;</italic>).</td>
<td align="left" valign="top">Mean pause duration is primarily affected by cognitive load; pause acuity does not predict pause duration distinctiveness (mean pause duration) within <italic>bracket</italic> productions. No support for H1 (<italic>&#956;</italic>).</td>
</tr>
<tr>
<td align="left" valign="top"><bold>H2 (sigma, <italic>&#963;</italic>): Variability</bold></td>
<td align="left" valign="top">Higher pause acuity &#8594; lower trial-to-trial variability (smaller <italic>&#963;</italic>, less variable), esp. under load.</td>
<td align="left" valign="top">Moderate evidence for a main effect of pause acuity on pause duration variability (<italic>&#963;</italic>). Strong evidence for a main effect of cognitive load, with higher load increasing pause duration variability. No evidence for an acuity &#215; cognitive load interaction.</td>
<td align="left" valign="top">Higher pause acuity is associated with greater (not lower) trial-to-trial variability in pause duration; cognitive load further increases variability; contradicts the prediction. No support for H2 (<italic>&#963;</italic>).</td>
</tr>
<tr>
<td align="left" valign="top"><bold>Pause usage (exploratory; hurdle)</bold></td>
<td align="left" valign="top">Exploratory (not preregistered).</td>
<td align="left" valign="top">Moderate evidence for a main effect of pause acuity and strong evidence for a main effect of cognitive load on pause usage probability. No evidence for a cognitive load &#215; pause acuity interaction.</td>
<td align="left" valign="top">Higher pause acuity is associated with reduced pause usage across both cognitive load conditions. High cognitive load further reduces pause usage independently of pause acuity. Warrants further investigation. Exploratory.</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><bold>Final lengthening</bold></td>
</tr>
<tr>
<td align="left" valign="top"><bold>H1 (mean, <italic>&#956;</italic>): Distinctiveness</bold></td>
<td align="left" valign="top">Higher final lengthening acuity &#8594; larger final lengthening difference between <italic>bracket</italic> and <italic>no-bracket</italic> productions (<italic>&#956;</italic>), stronger under load.</td>
<td align="left" valign="top">No evidence for an acuity &#215; cognitive load &#215; stimulus condition interaction, an acuity &#215; stimulus condition interaction, no main effect of final lengthening acuity on mean final lengthening (<italic>&#956;</italic>).</td>
<td align="left" valign="top">Final lengthening acuity does not predict final lengthening distinctiveness (mean final lengthening differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions). No support for H1 (<italic>&#956;</italic>).</td>
</tr>
<tr>
<td align="left" valign="top"><bold>H2 (sigma, <italic>&#963;</italic>): Variability</bold></td>
<td align="left" valign="top">Higher final lengthening acuity &#8594; lower trial-to-trial variability (smaller <italic>&#963;</italic>, less variable), esp. under load.</td>
<td align="left" valign="top">Strong evidence for a main effect of cognitive load, with higher load increasing variability. No evidence for an acuity &#215; cognitive load &#215; stimulus condition interaction, an acuity &#215; stimulus condition interaction, or a main effect of final lengthening acuity on trial-to-trial variability in final lengthening (<italic>&#963;</italic>).</td>
<td align="left" valign="top">Final lengthening production variability is primarily modulated by cognitive load. Final lengthening acuity does not predict trial-to-trial variability in final lengthening. No support for H2 (<italic>&#963;</italic>).</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Note:</italic> Perception = auditory-perceptual acuity (z-scored, sign-reversed JND thresholds; higher values = better acuity). Production distinctiveness = model mean: for pitch range and final lengthening as the difference between <italic>bracket</italic> and <italic>no-bracket</italic> productions; for pause duration (<italic>bracket</italic> only) as mean pause duration within <italic>bracket</italic> productions. Production variability = model sigma: larger sigma indicates greater trial-to-trial variability.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec>
<title>5. Discussion</title>
<p>Our study examined whether individual differences in prosodic boundary cue perception predict how those same cues are produced, focusing on two separable aspects of production: distinctiveness and variability. Specifically, we tested whether auditory-perceptual acuity predicts (H1) the strength with which a prosodic boundary cue is used to distinguish <italic>bracket</italic> productions from <italic>no-bracket</italic> productions, and (H2) the trial-to-trial variability with which that prosodic boundary cue is produced, and whether these relationships are modulated by cognitive load.</p>
<p>Perception and production were assessed in separate tasks and linked at the level of individual differences, using distributional models that estimated both production distinctiveness (<monospace>mean</monospace>, <italic>&#956;</italic>, reflecting inter-individual differences) and trial-to-trial variability (<monospace>sigma</monospace>, <italic>&#963;</italic>, reflecting intra-individual differences) separately for each prosodic boundary cue. A summary of findings is provided in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<sec>
<title>5.1 Pitch range</title>
<p>We found suggestive, rather than strong, evidence for the first hypothesis (H1) for pitch range: participants with higher pitch acuity produced larger pitch range differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions (averaged across cognitive load conditions), but the evidence for this effect did not reach the threshold for strong support across prior specifications. This pattern is, nevertheless, directionally consistent with our predictions. The qualification &#8220;suggestive&#8221; reflects statistical strength, rather than a contradiction of the hypothesized effect.</p>
<p>In contrast, we found moderate evidence for a three-way interaction indicating that higher pitch acuity was associated with greater trial-to-trial variability in pitch range production, contradicting our second hypothesis (H2) that higher auditory-perceptual acuity would lead to less variable production. Participants with higher pitch acuity were considerably more variable in <italic>bracket</italic> productions than in <italic>no-bracket</italic> productions under both cognitive load conditions, and both the size of this difference and the variability level for each stimulus condition remained comparable across low and high cognitive load. Participants with lower pitch acuity also showed greater trial-to-trial variability in <italic>bracket</italic> productions than in <italic>no-bracket</italic> productions under low cognitive load, though at lower overall levels than participants with higher pitch acuity. Under high cognitive load, however, this difference disappeared: participants with lower pitch acuity showed the same amount of trial-to-trial variability in <italic>bracket</italic> and <italic>no-bracket</italic> productions, as their variability in <italic>bracket</italic> productions reduced to the level observed in <italic>no-bracket</italic> productions. Thus, while pitch acuity showed directional alignment with H1 at the level of production distinctiveness, it did not support the prediction of reduced trial-to-trial variability formulated in H2.</p>
<p>It could be argued that the finer resolution of the pitch range JND continuum, relative to the pause and final lengthening continua, might have increased the sensitivity of the pitch acuity measure, and thereby influenced the PP-link modeling. To address this concern, we conducted a post-hoc simulation analysis in which pitch range JND thresholds (and, thereafter, acuity scores) were recomputed from coarser continua. Refitting the pitch range production model with these simulated scores yielded virtually identical parameter estimates and Bayes factors, providing no evidence that continuum resolution biased our findings (see 1.5, Just-Noticeable-Difference (JND) task sensitivity analysis, in the Supplementary Materials for details; pages 23&#8211;24, <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>).</p>
</sec>
<sec>
<title>5.2 Pause duration</title>
<p>We found no evidence that pause acuity predicts pause production distinctiveness (H1). Because pauses are almost exclusively used in <italic>bracket</italic> productions, distinctiveness for pauses was operationalized as mean pause duration within <italic>bracket</italic> productions rather than as a <italic>bracket/no-bracket</italic> contrast. Cognitive load had a strong main effect on mean pause duration: participants produced shorter pauses under high cognitive load than under low cognitive load, across all pause acuity levels.</p>
<p>In contrast, we found moderate evidence for a relationship between pause acuity and pause duration variability, again providing no support for H2. Participants with higher pause acuity showed greater trial-to-trial variability in pause duration, whereas participants with lower pause acuity showed less trial-to-trial variability in pause duration. This difference in trial-to-trial variability associated with pause acuity was present across both cognitive load conditions. Thus, as for pitch, higher pause acuity was associated with increased, rather than reduced, trial-to-trial variability, contradicting the prediction that higher acuity would lead to more stable production patterns. Exploratory analyses further indicated that pause acuity predicted pause usage. However, because this analysis was not preregistered and the pause acuity effect on usage probability was only moderate, we treat this pattern as tentative.</p>
</sec>
<sec>
<title>5.3 Final lengthening</title>
<p>Neither hypothesis was supported for final lengthening: final lengthening acuity did not predict final lengthening distinctiveness (H1) or trial-to-trial variability (H2) in final lengthening. In contrast, cognitive load had a strong main effect on final lengthening variability, increasing trial-to-trial variability across all final lengthening acuity levels. Thus, while final lengthening production was sensitive to increased processing demands, this sensitivity did not interact with final lengthening acuity. This absence of a PP-link for final lengthening stands in contrast to the patterns observed for pitch and pause, and is taken up in 5.4.</p>
</sec>
<sec>
<title>5.4 Cue-specific patterns</title>
<p>Our results reveal a clear cue-dependent PP-link pattern. Auditory-perceptual acuity related most robustly to pitch range, more weakly and differently to pause duration, and not reliably to final lengthening. Importantly, for the prosodic boundary cues showing evidence of a PP-link, this link emerged primarily in production variability (<italic>&#963;</italic>) rather than in production distinctiveness (<italic>&#956;</italic>).</p>
<p>This pattern suggests that individual differences in auditory-perceptual acuity primarily govern the trial-to-trial variability and, thus, flexibility with which speakers produce prosodic boundary cues across contexts, rather than the magnitude of their productions. This interpretation aligns with Xie et al.&#8217;s (<xref ref-type="bibr" rid="B84">2021</xref>) demonstration that prosodic production variability is functionally meaningful: their talker-specific ideal observer models outperformed generic models precisely because they captured the full distributional statistics of individual speakers&#8217; cue production patterns, and listeners actively learned and exploited these talker-specific patterns to improve their categorization accuracy. From this perspective, speakers with better auditory-perceptual acuity may more flexibly adapt their coordination and weighting of prosodic cues to context. This is precisely the type of systematic distributional information that listeners are attuned to track. Thus, the variability effects we observed likely reflect communicative strategies, with production flexibility serving as a learnable signal that facilitates comprehension.</p>
<p>Pitch range showed the strongest and most selective PP-link, emerging for both production distinctiveness and trial-to-trial variability. Participants with higher pitch acuity produced larger pitch range differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions across both cognitive load conditions, though this distinctiveness effect was only suggestive, and maintained a consistent <italic>bracket/no-bracket</italic> differentiation in trial-to-trial pitch range variability across both cognitive load conditions, showing comparably high variability in both <italic>bracket</italic> and <italic>no-bracket</italic> productions even under high cognitive load. Participants with lower pitch acuity, by contrast, produced smaller pitch range differences between <italic>bracket</italic> and <italic>no-bracket</italic> productions across both cognitive load conditions, and converged on more uniform pitch ranges under high cognitive load, showing reduced trial-to-trial variability in both <italic>bracket</italic> and <italic>no-bracket</italic> productions compared to low cognitive load. This divergence suggests that higher pitch acuity supports continued and more flexible modulation of pitch range when processing demands increase. Evidence from pitch perturbation studies points in the same direction: individuals with better pitch discrimination show stronger adaptive responses to sustained pitch shifts (<xref ref-type="bibr" rid="B49">Martin et al., 2018</xref>), and speakers actively regulate pitch to preserve communicative contrasts under challenging conditions (<xref ref-type="bibr" rid="B59">Patel et al., 2011</xref>). Together, these findings are consistent with pitch functioning as a flexible and central prosodic boundary cue in German (<xref ref-type="bibr" rid="B36">Holzgrefe-Lang et al., 2016</xref>; <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>; <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>).</p>
<p>Pause duration showed a partially parallel profile: higher pause acuity was associated with greater trial-to-trial variability in pause duration production, but, unlike pitch, showed no corresponding modulation of mean pause duration. This dissociation suggests that, for pauses, higher pause acuity relates to how pauses are integrated into prosodic control, rather than to their average magnitude. Our exploratory analyses further indicated that participants with higher pause acuity relied less categorically on pause usage, showing greater flexibility in when they deployed pauses for marking a boundary. Importantly, reduced pause usage among participants with higher pause acuity did not reflect a general absence of pauses: all participants used pauses in the majority of <italic>bracket</italic> stimulus productions. Rather, pause omission occurred selectively and disproportionately among speakers with higher pause acuity, suggesting their reduced reliance on this categorical boundary cue and a greater ability to exploit alternative prosodic boundary cues when available. This pattern likely reflects pauses&#8217; dual function in cognition and communication: marking boundaries for listeners and providing planning time for speakers (e.g., <xref ref-type="bibr" rid="B20">Ferreira &amp; Karimi, 2015</xref>). Moreover, since pauses function as categorical perceptual cues once they exceed a perceptual threshold (e.g., <xref ref-type="bibr" rid="B64">Petrone et al., 2017</xref>), higher pause acuity may afford greater flexibility in pause usage, without requiring systematic scaling of pause duration in production.</p>
<p>Final lengthening showed no evidence for a PP-link, despite trial-to-trial variability being affected by cognitive load. This null result suggests that final lengthening may be governed primarily by automatic timing and sequencing processes that respond in a more uniform way to processing demands, rather than by perceptually guided adjustments. Consistent with this view, final lengthening exhibits non-monotonic scaling with boundary strength, and often decreases at strong boundaries where pauses take over (e.g., <xref ref-type="bibr" rid="B43">Kentner et al., 2023</xref>). It is also the least consistently used prosodic boundary cue across participants (e.g., <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>) and rarely functions as a reliable boundary signal in isolation (e.g., <xref ref-type="bibr" rid="B36">Holzgrefe-Lang et al., 2016</xref>), indicating that it plays a supportive rather than primary role in German prosodic boundary marking.</p>
<p>Overall, this cue-specific gradient aligns with perturbation evidence showing differential control mechanisms for spectral versus temporal cues. In particular, Patel et al. (<xref ref-type="bibr" rid="B59">2011</xref>) showed that when pitch (f0) is perturbed during sentence production, speakers compensate primarily by adjusting spectral properties, such as pitch and intensity, while durational properties remain comparatively stable. Likewise, timing-perturbation studies demonstrate that compensatory responses to temporal perturbations are organized at the level of prosodic structure, rather than at the level of individual segments, revealing that speakers maintain word-level timing relations, rather than, correcting segment durations in isolation (<xref ref-type="bibr" rid="B57">Oschkinat &amp; Hoole, 2022</xref>). Moreover, this prosodic boundary cue gradient also matches broader observations. The use of pitch for boundary marking shows substantial cross-linguistic variability in both weighting and interpretation, reflecting its flexible role across prosodic systems, whereas temporal boundary cues, such as pauses and final lengthening, are more widely used across languages and tend to be more constrained in how flexibly they are deployed (e.g., <xref ref-type="bibr" rid="B10">Byrd et al., 2006</xref>; <xref ref-type="bibr" rid="B56">Ortega-Llebaria &amp; Nagao, 2025</xref>; <xref ref-type="bibr" rid="B82">Wightman et al., 1992</xref>; <xref ref-type="bibr" rid="B85">Yang et al., 2014</xref>).</p>
</sec>
<sec>
<title>5.5 Segmental versus suprasegmental PP-links</title>
<p>Our hypotheses were informed by the segmental PP-link literature, which has consistently shown that participants with higher auditory-perceptual acuity tend to produce more precise and less variable segmental contrasts, even under challenging conditions (e.g., <xref ref-type="bibr" rid="B27">S. S. Ghosh et al., 2010</xref>; <xref ref-type="bibr" rid="B62">Perkell, Guenther, et al., 2004</xref>; <xref ref-type="bibr" rid="B77">Villacorta et al., 2007</xref>). We extended this logic to prosody, predicting that higher auditory-perceptual acuity for prosodic boundary cues would yield more distinct and more stable boundary marking, though, for temporal cues, this prediction was treated as an empirical question, rather than a direct extension of segmental evidence. Our findings only partially aligned with these predictions: distinctiveness effects were limited to pitch, while trial-to-trial variability effects emerged for both pitch and pause, but in the opposite direction: higher auditory-perceptual acuity was associated with greater, rather than lower, trial-to-trial variability. These patterns indicate that segmental perception-production predictions do not straightforwardly generalize to prosodic boundary marking.</p>
<p>A key reason may lie in how prosodic boundary cues are represented. The DIVA model formalizes segmental PP-links in terms of tightly coupled auditory-motor units, in which each speech sound is encoded as both an auditory target and a corresponding motor program (e.g., <xref ref-type="bibr" rid="B32">Guenther et al., 2013</xref>). This architecture supports direct mappings between auditory-perceptual discriminability and segmental production precision (e.g., <xref ref-type="bibr" rid="B61">Perkell, 2012</xref>). Prosodic boundary cues differ fundamentally in the nature of their representations: spectral cues, such as pitch, can be modeled as explicit auditory targets in a framework like DIVA, whereas durational and temporal cues, such as pause and final lengthening, are primarily represented as motor timing parameters, rather than as perceptual targets (e.g., <xref ref-type="bibr" rid="B53">Miller &amp; Guenther, 2021</xref>; <xref ref-type="bibr" rid="B86">Zhang et al., 2015</xref>). The GODIVA extension addresses this by incorporating metrical and initiation maps that control the timing and sequencing of speech chunks (e.g., <xref ref-type="bibr" rid="B6">Bohland et al., 2010</xref>; <xref ref-type="bibr" rid="B53">Miller &amp; Guenther, 2021</xref>), but these mechanisms are architecturally distinct from the auditory target representations available for spectral cues. This distinction has a plausible neurobiological basis: the right hemisphere&#8217;s longer temporal integration windows make it well-suited to tracking slow modulations, such as F0 contours and prosodic structure, whereas the faster left-hemisphere processing supports the sequencing and timing mechanisms underlying durational control (e.g., <xref ref-type="bibr" rid="B2">Arjmandi &amp; Behroozmand, 2024</xref>; <xref ref-type="bibr" rid="B22">Floegel et al., 2020</xref>; <xref ref-type="bibr" rid="B66">Poeppel, 2003</xref>). An additional factor is that prosodic boundaries are realized through the graded modulation of multiple cues over extended temporal spans, and this redundancy allows speakers to achieve the same communicative goal through different cue combinations (e.g., <xref ref-type="bibr" rid="B11">Cho, 2016</xref>; <xref ref-type="bibr" rid="B39">Huttenlauch et al., 2021</xref>). PP-links at the level of prosodic boundary marking are, therefore, more likely to reflect shared control processes (i.e., how multiple cues are coordinated, weighted, and adapted) than the shared symbolic representations that characterize segmental auditory-motor target mappings.</p>
<p>Our findings partially align with these theoretical predictions, while also revealing aspects not yet captured by DIVA/GODIVA. The clearest acuity-related PP-link emerged for pitch, the prosodic boundary cue with explicit auditory target representation, consistent with the prediction that tonal cues benefit from acuity-tuned auditory error maps supporting continuous feedback correction. Pause showed partial linking only (pause acuity predicting trial-to-trial variability, but not mean duration), and final lengthening showed no linking at all, consistent with durational cues relying primarily on feedforward sequencing mechanisms less directly modulated by auditory-perceptual acuity (e.g., <xref ref-type="bibr" rid="B14">Civier et al., 2013</xref>; <xref ref-type="bibr" rid="B73">Tourville &amp; Guenther, 2011</xref>). However, the directionality of the variability effect contradicts straightforward DIVA predictions. In segmental motor control, trial-to-trial variability would conventionally be interpreted as instability or motor noise resulting from poorly specified targets (<xref ref-type="bibr" rid="B61">Perkell, 2012</xref>). Our findings suggest a different interpretation for the prosodic domain: rather than reflecting imprecision, greater variability in speakers with higher auditory-perceptual acuity may reflect an expanded adaptive range, and, thus, the ability to flexibly modulate cue realization across varying contexts. One possibility is that this flexibility reflects the coordination of multiple cues: speakers with higher auditory-perceptual acuity for a given prosodic boundary cue may have access to a wider repertoire of cue combinations for signaling prosodic boundaries. This form of acuity-linked flexibility, relating to cue coordination rather than target precision, is &#8211; to the best of our knowledge &#8211; not yet captured in current neurocomputational models, which focus primarily on error minimization around fixed targets.</p>
</sec>
<sec>
<title>5.6 Limitations and future directions</title>
<p>Although we interpret the observed perception-production patterns primarily in terms of auditory-perceptual acuity interacting with cue-specific control demands, these relationships are likely shaped by additional cognitive and structural factors. For temporal prosody, in particular, perturbation work indicates that production responses are jointly influenced by auditory-perceptual acuity and rhythmic abilities, with their relative contributions varying as a function of prosodic structure and whether responses reflect online compensation or longer-term adaptation (e.g., <xref ref-type="bibr" rid="B58">Oschkinat et al., 2022</xref>). This suggests that temporal PP-links may depend on a broader set of moderators and task distinctions than perceptual resolution alone.</p>
<p>Beyond perceptual and motor constraints, individual differences in memory and attentional control may further shape how reliably participants produce distinctions between stimuli with and without a boundary, especially under cognitive load. Prosodic structure has been shown to influence how listeners allocate attention and encode material in memory, including benefits from pitch rises and boundary tones in serial recall tasks, pointing to systematic interactions between prosody, memory, and attention (e.g., <xref ref-type="bibr" rid="B20">Ferreira &amp; Karimi, 2015</xref>; <xref ref-type="bibr" rid="B29">Grice et al., 2024</xref>; <xref ref-type="bibr" rid="B45">Lialiou et al., in press</xref>). From this perspective, PP-links may reflect learning, adaptation, and executive control processes in addition to auditory-perceptual resolution.</p>
<p>These considerations point to several concrete directions. First, the theoretical accounts advanced here could be tested more directly by incorporating independent measures of (i) working memory or serial recall, (ii) sustained attention or executive functioning, (iii) rhythmic or musical experience, and (iv) sensorimotor adaptation to auditory perturbations, and testing whether these variables explain variance in <italic>&#956;</italic> and/or <italic>&#963;</italic> beyond auditory-perceptual acuity. Second, PP-links for additional prosodic dimensions, such as intensity, have yet to be examined. Domain-initial strengthening is a particularly promising candidate, given that segments following a boundary show enhanced articulation (e.g., <xref ref-type="bibr" rid="B12">Cho et al., 2007</xref>). Third, presenting the secondary tasks in separate blocks would allow the independent contribution of each to the observed cognitive load effects to be formally tested. Finally, because prosodic boundary cue weighting differs across languages and prosodic systems, an important direction for future work is to test whether the cue-specific PP-link pattern observed here generalizes beyond German, particularly for pitch (e.g., <xref ref-type="bibr" rid="B56">Ortega-Llebaria &amp; Nagao, 2025</xref>; <xref ref-type="bibr" rid="B85">Yang et al., 2014</xref>).</p>
</sec>
<sec>
<title>5.7 Conclusion</title>
<p>Cognitive load served as an effective stress test for prosodic production, revealing individual differences that were not apparent under low cognitive load. The data indicate a prosodic PP-link for pitch and pause, but not for final lengthening. Higher auditory-perceptual acuity was associated with greater production variability for pitch and pause, suggesting that for prosodic boundary marking, variability reflects adaptive flexibility in prosodic boundary cue production, rather than imprecision in motor control. Although this direction of the variability effects runs counter to our preregistered hypothesis (which predicted lower trial-to-trial variability with higher auditory-perceptual acuity), the pattern is consistent with variability reflecting adaptive flexibility in prosodic cue deployment, rather than motor imprecision. These cue-specific patterns align with theoretical distinctions between spectral and temporal cues in neurocomputational models, and suggest that prosodic PP-links emerge from the interaction of perceptual precision, control mechanisms, and contextual demands.</p>
</sec>
</sec>
</body>
<back>
<sec>
<title>Abbreviations</title>
<p>BF &#8211; Bayes factor</p>
<p>CI &#8211; Credible interval</p>
<p>CoM &#8211; Center of periodic Mass</p>
<p>CSV &#8211; Comma-separated values</p>
<p>DIVA &#8211; Directions Into Velocities of Articulators (model)</p>
<p>f0 &#8211; Fundamental frequency</p>
<p>GODIVA &#8211; Gradient Order DIVA (model)</p>
<p>IQR &#8211; Interquartile range</p>
<p>JND &#8211; Just-Noticeable-Difference</p>
<p>lnBF<sub>10</sub> &#8211; Natural logarithm of Bayes factor</p>
<p>ms &#8211; Milliseconds</p>
<p>OSF &#8211; Open Science Framework</p>
<p>PP-link &#8211; Perception-production link</p>
<p>ProPer &#8211; Prosodic analysis with Periodic energy (toolbox)</p>
<p>SD &#8211; Standard deviation</p>
</sec>
<sec>
<title>Data accessibility statement</title>
<p>All materials, data, and reproducible analysis code are available through the Open Science Framework: experimental code for production task (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/N3Z2K">https://doi.org/10.17605/OSF.IO/N3Z2K</ext-link>), experimental code for perception task (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/MQY2P">https://doi.org/10.17605/OSF.IO/MQY2P</ext-link>), and analysis scripts with data (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.17605/OSF.IO/KQHMT">https://doi.org/10.17605/OSF.IO/KQHMT</ext-link>). Audio recordings are available upon request to the corresponding author.</p>
</sec>
<sec>
<title>Ethics and consent</title>
<p>The study was conducted in accordance with the Declaration of Helsinki and approved by the University of Potsdam Ethics Committee (approval code: 99/2020). Informed consent was obtained from all participants.</p>
</sec>
<sec>
<title>Acknowledgements</title>
<p>Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) &#8211; Project-ID 317633480 &#8211; SFB 1287. Jo&#227;o Ver&#237;ssimo has been funded by the Funda&#231;&#227;o para a Ci&#234;ncia e a Tecnologia (FCT, Foundation for Science and Technology), grant UID/00214/2025 to the Center of Linguistics of the University of Lisbon.</p>
</sec>
<sec>
<title>Competing interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<sec>
<title>Author contributions</title>
<p>Andrea Hofmann: conceptualization, methodology, software, investigation, data curation, formal analysis, visualization, writing &#8211; original draft, writing &#8211; review &amp; editing</p>
<p>Outi Tuomainen: conceptualization, funding acquisition, advice</p>
<p>Sandra Hanne: conceptualization, funding acquisition, advice</p>
<p>Jo&#227;o Ver&#237;ssimo: senior author, supervision, formal analysis, regular consultation, writing &#8211; review &amp; editing</p>
<p>Isabell Wartenburger: senior author, supervision, conceptualization, funding acquisition, regular consultation, writing &#8211; review &amp; editing</p>
</sec>
<ref-list>
<ref id="B1"><mixed-citation publication-type="book"><string-name><surname>Albert</surname>, <given-names>A.</given-names></string-name> (<year>2023</year>). <source>A model of sonority based on pitch intelligibility</source>. <publisher-name>Zenodo</publisher-name>. <pub-id pub-id-type="doi">10.5281/zenodo.7837175</pub-id></mixed-citation></ref>
<ref id="B2"><mixed-citation publication-type="journal"><string-name><surname>Arjmandi</surname>, <given-names>M. K.</given-names></string-name>, &amp; <string-name><surname>Behroozmand</surname>, <given-names>R.</given-names></string-name> (<year>2024</year>). <article-title>On the interplay between speech perception and production: Insights from research and theories</article-title>. <source>Frontiers in Neuroscience</source>, <volume>18</volume>, <elocation-id>1347614</elocation-id>. <pub-id pub-id-type="doi">10.3389/fnins.2024.1347614</pub-id></mixed-citation></ref>
<ref id="B3"><mixed-citation publication-type="journal"><string-name><surname>Barr</surname>, <given-names>D. J.</given-names></string-name>, <string-name><surname>Levy</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Scheepers</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Tily</surname>, <given-names>H. J.</given-names></string-name> (<year>2013</year>). <article-title>Random effects structure for confirmatory hypothesis testing: Keep it maximal</article-title>. <source>Journal of Memory and Language</source>, <volume>68</volume>(<issue>3</issue>). <pub-id pub-id-type="doi">10.1016/j.jml.2012.11.001</pub-id></mixed-citation></ref>
<ref id="B4"><mixed-citation publication-type="webpage"><string-name><surname>Bates</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Kliegl</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Vasishth</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>H.</given-names></string-name> (<year>2015</year>). <source>Parsimonious mixed models</source>. <uri>http://arxiv.org/pdf/1506.04967</uri></mixed-citation></ref>
<ref id="B5"><mixed-citation publication-type="webpage"><string-name><surname>Boersma</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Weenink</surname>, <given-names>D.</given-names></string-name> (<year>1992&#8211;2020</year>). <source>Praat: Doing phonetics by computer [computer program]</source>. <uri>http://www.praat.org/</uri></mixed-citation></ref>
<ref id="B6"><mixed-citation publication-type="journal"><string-name><surname>Bohland</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Bullock</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2010</year>). <article-title>Neural representations and mechanisms for the performance of simple speech sequences</article-title>. <source>Journal of Cognitive Neuroscience</source>, <volume>22</volume>(<issue>7</issue>), <fpage>1504</fpage>&#8211;<lpage>1529</lpage>. <pub-id pub-id-type="doi">10.1162/jocn.2009.21306</pub-id></mixed-citation></ref>
<ref id="B7"><mixed-citation publication-type="journal"><string-name><surname>Brown</surname>, <given-names>V. A.</given-names></string-name> (<year>2025</year>). <article-title>Measuring the dual-task costs of audiovisual speech processing across levels of background noise</article-title>. <source>Journal of Experimental Psychology. General</source>, <volume>154</volume>(<issue>12</issue>), <fpage>3428</fpage>&#8211;<lpage>3449</lpage>. <pub-id pub-id-type="doi">10.1037/xge0001826</pub-id></mixed-citation></ref>
<ref id="B8"><mixed-citation publication-type="journal"><string-name><surname>Brunner</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ghosh</surname>, <given-names>S. S.</given-names></string-name>, <string-name><surname>Hoole</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Matthies</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name> (<year>2011</year>). <article-title>The influence of auditory acuity on acoustic variability and the use of motor equivalence during adaptation to a perturbation</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>54</volume>(<issue>3</issue>), <fpage>727</fpage>&#8211;<lpage>739</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2010/09-0256</pub-id>)</mixed-citation></ref>
<ref id="B9"><mixed-citation publication-type="journal"><string-name><surname>Buerkner</surname>, <given-names>P.-C.</given-names></string-name> (<year>2018</year>). <article-title>Advanced Bayesian multilevel modeling with the R package brms</article-title>. <source>The R Journal</source>, <volume>10</volume>(<issue>1</issue>), <fpage>395</fpage>&#8211;<lpage>411</lpage>. <pub-id pub-id-type="doi">10.32614/RJ-2018-017</pub-id></mixed-citation></ref>
<ref id="B10"><mixed-citation publication-type="journal"><string-name><surname>Byrd</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Krivokapi&#263;</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Lee</surname>, <given-names>S.</given-names></string-name> (<year>2006</year>). <article-title>How far, how long: On the temporal scope of prosodic boundary effects</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>120</volume>(<issue>3</issue>), <fpage>1589</fpage>&#8211;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1121/1.2217135</pub-id></mixed-citation></ref>
<ref id="B11"><mixed-citation publication-type="journal"><string-name><surname>Cho</surname>, <given-names>T.</given-names></string-name> (<year>2016</year>). <article-title>Prosodic boundary strengthening in the phonetics&#8211;prosody interface</article-title>. <source>Language and Linguistics Compass</source>, <volume>10</volume>(<issue>3</issue>), <fpage>120</fpage>&#8211;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1111/lnc3.12178</pub-id></mixed-citation></ref>
<ref id="B12"><mixed-citation publication-type="journal"><string-name><surname>Cho</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>McQueen</surname>, <given-names>J. M.</given-names></string-name>, &amp; <string-name><surname>Cox</surname>, <given-names>E. A.</given-names></string-name> (<year>2007</year>). <article-title>Prosodically driven phonetic detail in speech processing: The case of domain-initial strengthening in English</article-title>. <source>Journal of Phonetics</source>, <volume>35</volume>(<issue>2</issue>), <fpage>210</fpage>&#8211;<lpage>243</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2006.03.003</pub-id></mixed-citation></ref>
<ref id="B13"><mixed-citation publication-type="journal"><string-name><surname>Ciaccio</surname>, <given-names>L. A.</given-names></string-name>, &amp; <string-name><surname>Ver&#237;ssimo</surname>, <given-names>J.</given-names></string-name> (<year>2022</year>). <article-title>Investigating variability in morphological processing with Bayesian distributional models</article-title>. <source>Psychonomic Bulletin &amp; Review</source>, <volume>29</volume>(<issue>6</issue>), <fpage>2264</fpage>&#8211;<lpage>2274</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-022-02109-w</pub-id></mixed-citation></ref>
<ref id="B14"><mixed-citation publication-type="journal"><string-name><surname>Civier</surname>, <given-names>O.</given-names></string-name>, <string-name><surname>Bullock</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Max</surname>, <given-names>L.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2013</year>). <article-title>Computational modeling of stuttering caused by impairments in a basal ganglia thalamo-cortical circuit involved in syllable selection and initiation</article-title>. <source>Brain and Language</source>, <volume>126</volume>(<issue>3</issue>), <fpage>263</fpage>&#8211;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1016/j.bandl.2013.05.016</pub-id></mixed-citation></ref>
<ref id="B15"><mixed-citation publication-type="journal"><string-name><surname>Cole</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name> (<year>2016</year>). <article-title>New methods for prosodic transcription: Capturing variability as a source of information</article-title>. <source>Laboratory Phonology</source>, <volume>7</volume>(<issue>1</issue>), <elocation-id>8</elocation-id>, pp. <fpage>1</fpage>&#8211;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.5334/labphon.29</pub-id></mixed-citation></ref>
<ref id="B16"><mixed-citation publication-type="journal"><string-name><surname>Dahl</surname>, <given-names>K. L.</given-names></string-name>, <string-name><surname>C&#225;diz</surname>, <given-names>M. D.</given-names></string-name>, <string-name><surname>Zuk</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, &amp; <string-name><surname>Stepp</surname>, <given-names>C. E.</given-names></string-name> (<year>2024</year>). <article-title>Controlling pitch for prosody: Sensorimotor adaptation in linguistically meaningful contexts</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>67</volume>(<issue>2</issue>), <fpage>440</fpage>&#8211;<lpage>454</lpage>. <pub-id pub-id-type="doi">10.1044/2023_JSLHR-23-00460</pub-id></mixed-citation></ref>
<ref id="B17"><mixed-citation publication-type="journal"><string-name><surname>de Beer</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Hofmann</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Regenbrecht</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Huttenlauch</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Obrig</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Hanne</surname>, <given-names>S.</given-names></string-name> (<year>2022</year>). <article-title>Production and comprehension of prosodic boundary marking in persons with unilateral brain lesions</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>65</volume>(<issue>12</issue>), <fpage>4774</fpage>&#8211;<lpage>4796</lpage>. <pub-id pub-id-type="doi">10.1044/2022_JSLHR-22-00258</pub-id></mixed-citation></ref>
<ref id="B18"><mixed-citation publication-type="journal"><string-name><surname>Dickey</surname>, <given-names>J. M.</given-names></string-name>, &amp; <string-name><surname>Lientz</surname>, <given-names>B. P.</given-names></string-name> (<year>1970</year>). <article-title>The weighted likelihood ratio, sharp hypotheses about chances, the order of a Markov chain</article-title>. <source>The Annals of Mathematical Statistics</source>, <volume>41</volume>(<issue>1</issue>), <fpage>214</fpage>&#8211;<lpage>226</lpage>. <pub-id pub-id-type="doi">10.1214/aoms/1177697203</pub-id></mixed-citation></ref>
<ref id="B19"><mixed-citation publication-type="journal"><string-name><surname>Elman</surname>, <given-names>J. L.</given-names></string-name> (<year>1981</year>). <article-title>Effects of frequency-shifted feedback on the pitch of vocal productions</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>70</volume>(<issue>1</issue>), <fpage>45</fpage>&#8211;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1121/1.386580</pub-id></mixed-citation></ref>
<ref id="B20"><mixed-citation publication-type="journal"><string-name><surname>Ferreira</surname>, <given-names>F.</given-names></string-name>, &amp; <string-name><surname>Karimi</surname>, <given-names>H.</given-names></string-name> (<year>2015</year>). <article-title>Prosody, performance, and cognitive skill: Evidence from individual differences</article-title>. <source>Explicit and Implicit Prosody in Sentence Processing</source>, <volume>46</volume>, <fpage>119</fpage>&#8211;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-12961-7_7</pub-id></mixed-citation></ref>
<ref id="B21"><mixed-citation publication-type="book"><string-name><surname>Flege</surname>, <given-names>J. E.</given-names></string-name> (<year>1995</year>). <chapter-title>Second language speech learning theory, findings, and problems</chapter-title>. In <string-name><given-names>W.</given-names> <surname>Strange</surname></string-name> (Ed.), <source>Speech perception and linguistic experience</source> (pp. <fpage>233</fpage>&#8211;<lpage>277</lpage>). <publisher-name>York Press</publisher-name>.</mixed-citation></ref>
<ref id="B22"><mixed-citation publication-type="journal"><string-name><surname>Floegel</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Fuchs</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Kell</surname>, <given-names>C. A.</given-names></string-name> (<year>2020</year>). <article-title>Differential contributions of the two cerebral hemispheres to temporal and spectral speech feedback control</article-title>. <source>Nature Communications</source>, <volume>11</volume>(<issue>1</issue>), <elocation-id>2839</elocation-id>. <pub-id pub-id-type="doi">10.1038/s41467-020-16743-2</pub-id></mixed-citation></ref>
<ref id="B23"><mixed-citation publication-type="journal"><string-name><surname>Frazier</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Carlson</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Clifton</surname>, <given-names>C.</given-names>, <suffix>JR</suffix></string-name>. (<year>2006</year>). <article-title>Prosodic phrasing is central to language comprehension</article-title>. <source>Trends in Cognitive Sciences</source>, <volume>10</volume>(<issue>6</issue>), <fpage>244</fpage>&#8211;<lpage>249</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2006.04.002</pub-id></mixed-citation></ref>
<ref id="B24"><mixed-citation publication-type="journal"><string-name><surname>Gelman</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Jakulin</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Pittau</surname>, <given-names>M. G.</given-names></string-name>, &amp; <string-name><surname>Su</surname>, <given-names>Y.-S.</given-names></string-name> (<year>2008</year>). <article-title>A weakly informative default prior distribution for logistic and other regression models</article-title>. <source>The Annals of Applied Statistics</source>, <volume>2</volume>(<issue>4</issue>). <pub-id pub-id-type="doi">10.1214/08-AOAS191</pub-id></mixed-citation></ref>
<ref id="B25"><mixed-citation publication-type="journal"><string-name><surname>Ghaffarvand Mokari</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Gafos</surname>, <given-names>A. I.</given-names></string-name>, &amp; <string-name><surname>Williams</surname>, <given-names>D.</given-names></string-name> (<year>2020</year>). <article-title>Perceptuomotor compatibility effects in vowels: Beyond phonemic identity</article-title>. <source>Attention, Perception &amp; Psychophysics</source>, <volume>82</volume>(<issue>5</issue>), <fpage>2751</fpage>&#8211;<lpage>2764</lpage>. <pub-id pub-id-type="doi">10.3758/s13414-020-02014-1</pub-id></mixed-citation></ref>
<ref id="B26"><mixed-citation publication-type="journal"><string-name><surname>Ghosh</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>Mitra</surname>, <given-names>R.</given-names></string-name> (<year>2018</year>). <article-title>On the use of Cauchy prior distributions for Bayesian logistic regression</article-title>. <source>Bayesian Analysis</source>, <volume>13</volume>(<issue>2</issue>). <pub-id pub-id-type="doi">10.1214/17-BA1051</pub-id></mixed-citation></ref>
<ref id="B27"><mixed-citation publication-type="journal"><string-name><surname>Ghosh</surname>, <given-names>S. S.</given-names></string-name>, <string-name><surname>Matthies</surname>, <given-names>M. L.</given-names></string-name>, <string-name><surname>Maas</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Hanson</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>M&#233;nard</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Lane</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name> (<year>2010</year>). <article-title>An investigation of the relation between sibilant production and somatosensory and auditory acuity</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>128</volume>(<issue>5</issue>), <fpage>3079</fpage>&#8211;<lpage>3087</lpage>. <pub-id pub-id-type="doi">10.1121/1.3493430</pub-id></mixed-citation></ref>
<ref id="B28"><mixed-citation publication-type="journal"><string-name><surname>Gollrad</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sommerfeld</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>K&#252;gler</surname>, <given-names>F.</given-names></string-name> (<year>2010</year>). <article-title>Prosodic cue weighting in disambiguation: Case ambiguity in German</article-title>. <source>Speech Prosody 2010</source>, paper 165&#8211;0. <pub-id pub-id-type="doi">10.21437/SpeechProsody.2010-178</pub-id></mixed-citation></ref>
<ref id="B29"><mixed-citation publication-type="journal"><string-name><surname>Grice</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Savino</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Schumacher</surname>, <given-names>P. B.</given-names></string-name>, <string-name><surname>R&#246;hr</surname>, <given-names>C. T.</given-names></string-name>, &amp; <string-name><surname>Ellison</surname>, <given-names>T. M.</given-names></string-name> (<year>2024</year>). <article-title>Rises on pitch accents and edge tones affect serial recall performance at item and domain levels</article-title>. <source>Laboratory Phonology</source>, <volume>15</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.16995/labphon.10473</pub-id></mixed-citation></ref>
<ref id="B30"><mixed-citation publication-type="journal"><string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>1995</year>). <article-title>Speech sound acquisition, coarticulation, and rate effects in a neural network model of speech production</article-title>. <source>Psychological Review</source>, <volume>102</volume>(<issue>3</issue>), <fpage>594</fpage>&#8211;<lpage>621</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295x.102.3.594</pub-id></mixed-citation></ref>
<ref id="B31"><mixed-citation publication-type="book"><string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2016</year>). <source>Neural control of speech</source>. <publisher-name>The MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/10471.001.0001</pub-id></mixed-citation></ref>
<ref id="B32"><mixed-citation publication-type="book"><string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Ghosh</surname>, <given-names>S. S.</given-names></string-name>, <string-name><surname>Nieto-Castanon</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Tourville</surname>, <given-names>J. A.</given-names></string-name> (<year>2013</year>). <chapter-title>A neural model of speech production</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Harrington</surname></string-name> &amp; <string-name><given-names>M.</given-names> <surname>Tabain</surname></string-name> (Eds.), <source>Speech production</source>. <publisher-name>Taylor and Francis</publisher-name>.</mixed-citation></ref>
<ref id="B33"><mixed-citation publication-type="journal"><string-name><surname>Hansen</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Huttenlauch</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>de Beer</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Hanne</surname>, <given-names>S.</given-names></string-name> (<year>2023</year>). <article-title>Individual differences in early disambiguation of prosodic grouping</article-title>. <source>Language and Speech</source>, <volume>66</volume>(<issue>3</issue>), <fpage>706</fpage>&#8211;<lpage>733</lpage>. <pub-id pub-id-type="doi">10.1177/00238309221127374</pub-id></mixed-citation></ref>
<ref id="B34"><mixed-citation publication-type="journal"><string-name><surname>Harmon</surname>, <given-names>T. G.</given-names></string-name>, <string-name><surname>Jacks</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Haley</surname>, <given-names>K. L.</given-names></string-name> (<year>2019</year>). <article-title>Speech fluency in acquired apraxia of speech during narrative discourse: Group comparisons and dual-task effects</article-title>. <source>American Journal of Speech-Language Pathology</source>, <volume>28</volume>(<issue>2S</issue>), <fpage>905</fpage>&#8211;<lpage>914</lpage>. <pub-id pub-id-type="doi">10.1044/2018_AJSLP-MSC18-18-0107</pub-id></mixed-citation></ref>
<ref id="B35"><mixed-citation publication-type="journal"><string-name><surname>Holzgrefe-Lang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wellmann</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>H&#246;hle</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name> (<year>2018</year>). <article-title>Infants&#8217; processing of prosodic cues: Electrophysiological evidence for boundary perception beyond pause detection</article-title>. <source>Language and Speech</source>, <volume>61</volume>(<issue>1</issue>), <fpage>153</fpage>&#8211;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1177/0023830917730590</pub-id></mixed-citation></ref>
<ref id="B36"><mixed-citation publication-type="journal"><string-name><surname>Holzgrefe-Lang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wellmann</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Petrone</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>R&#228;ling</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Truckenbrodt</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>H&#246;hle</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name> (<year>2016</year>). <article-title>How pitch change and final lengthening cue boundary perception in German: Converging evidence from ERPs and prosodic judgements</article-title>. <source>Language, Cognition and Neuroscience</source>, <volume>31</volume>(<issue>7</issue>), <fpage>904</fpage>&#8211;<lpage>920</lpage>. <pub-id pub-id-type="doi">10.1080/23273798.2016.1157195</pub-id></mixed-citation></ref>
<ref id="B37"><mixed-citation publication-type="journal"><string-name><surname>Houde</surname>, <given-names>J. F.</given-names></string-name>, &amp; <string-name><surname>Jordan</surname>, <given-names>M. I.</given-names></string-name> (<year>1998</year>). <article-title>Sensorimotor adaptation in speech production</article-title>. <source>Science</source>, <volume>279</volume>(<issue>5354</issue>), <fpage>1213</fpage>&#8211;<lpage>1216</lpage>. <pub-id pub-id-type="doi">10.1126/science.279.5354.1213</pub-id></mixed-citation></ref>
<ref id="B38"><mixed-citation publication-type="journal"><string-name><surname>Houde</surname>, <given-names>J. F.</given-names></string-name>, &amp; <string-name><surname>Jordan</surname>, <given-names>M. I.</given-names></string-name> (<year>2002</year>). <article-title>Sensorimotor adaptation of speech</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>45</volume>(<issue>2</issue>), <fpage>295</fpage>&#8211;<lpage>310</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2002/023</pub-id>)</mixed-citation></ref>
<ref id="B39"><mixed-citation publication-type="journal"><string-name><surname>Huttenlauch</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>de Beer</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Hanne</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name> (<year>2021</year>). <article-title>Production of prosodic cues in coordinate name sequences addressing varying interlocutors</article-title>. <source>Laboratory Phonology: Journal of the Association for Laboratory Phonology</source>, <volume>12</volume>(<issue>1</issue>), <elocation-id>1</elocation-id>. <pub-id pub-id-type="doi">10.5334/labphon.221</pub-id></mixed-citation></ref>
<ref id="B40"><mixed-citation publication-type="book"><string-name><surname>Jeffreys</surname>, <given-names>H.</given-names></string-name> (<year>1991</year>). <source>Theory of probability</source> (<edition>2nd</edition> ed.). <publisher-name>Clarendon Press</publisher-name>.</mixed-citation></ref>
<ref id="B41"><mixed-citation publication-type="journal"><string-name><surname>Kass</surname>, <given-names>R. E.</given-names></string-name>, &amp; <string-name><surname>Raftery</surname>, <given-names>A. E.</given-names></string-name> (<year>1995</year>). <article-title>Bayes factors</article-title>. <source>Journal of the American Statistical Association</source>, <volume>90</volume>(<issue>430</issue>), <fpage>773</fpage>&#8211;<lpage>795</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1995.10476572</pub-id></mixed-citation></ref>
<ref id="B42"><mixed-citation publication-type="journal"><string-name><surname>Kentner</surname>, <given-names>G.</given-names></string-name>, &amp; <string-name><surname>F&#233;ry</surname>, <given-names>C.</given-names></string-name> (<year>2013</year>). <article-title>A new approach to prosodic grouping</article-title>. <source>The Linguistic Review</source>, <volume>30</volume>(<issue>2</issue>). <pub-id pub-id-type="doi">10.1515/tlr-2013-0009</pub-id></mixed-citation></ref>
<ref id="B43"><mixed-citation publication-type="journal"><string-name><surname>Kentner</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Franz</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Knoop</surname>, <given-names>C. A.</given-names></string-name>, &amp; <string-name><surname>Menninghaus</surname>, <given-names>W.</given-names></string-name> (<year>2023</year>). <article-title>The final lengthening of pre-boundary syllables turns into final shortening as boundary strength levels increase</article-title>. <source>Journal of Phonetics</source>, <volume>97</volume>, <elocation-id>101225</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.wocn.2023.101225</pub-id></mixed-citation></ref>
<ref id="B44"><mixed-citation publication-type="journal"><string-name><surname>Levitt</surname>, <given-names>H.</given-names></string-name> (<year>1971</year>). <article-title>Transformed up-down methods in psychoacoustics</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>49</volume>(<issue>2B</issue>), <fpage>467</fpage>&#8211;<lpage>477</lpage>. <pub-id pub-id-type="doi">10.1121/1.1912375</pub-id></mixed-citation></ref>
<ref id="B45"><mixed-citation publication-type="journal"><string-name><surname>Lialiou</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Grice</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Schumacher</surname>, <given-names>P. B.</given-names></string-name> (in press). <article-title>A test battery for measuring individual cognitive ability: A brief practical tutorial [author accepted manuscript]</article-title>. <source>Europe&#8217;s Journal of Psychology</source>. <pub-id pub-id-type="doi">10.23668/psycharchives.21626</pub-id></mixed-citation></ref>
<ref id="B46"><mixed-citation publication-type="book"><string-name><surname>Lindblom</surname>, <given-names>B.</given-names></string-name> (<year>1990</year>). <chapter-title>Explaining phonetic variation: A sketch of the H&amp;H theory</chapter-title>. In <string-name><given-names>W. J.</given-names> <surname>Hardcastle</surname></string-name> &amp; <string-name><given-names>A.</given-names> <surname>Marchal</surname></string-name> (Eds.), <source>Speech production and speech modelling</source> (pp. <fpage>403</fpage>&#8211;<lpage>439</lpage>). <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-94-009-2037-8_16</pub-id></mixed-citation></ref>
<ref id="B47"><mixed-citation publication-type="journal"><string-name><surname>Lively</surname>, <given-names>S. E.</given-names></string-name>, <string-name><surname>Pisoni</surname>, <given-names>D. B.</given-names></string-name>, <string-name><surname>van Summers</surname>, <given-names>W.</given-names></string-name>, &amp; <string-name><surname>Bernacki</surname>, <given-names>R. H.</given-names></string-name> (<year>1993</year>). <article-title>Effects of cognitive workload on speech production: Acoustic analyses and perceptual consequences</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>93</volume>(<issue>5</issue>), <fpage>2962</fpage>&#8211;<lpage>2973</lpage>. <pub-id pub-id-type="doi">10.1121/1.405815</pub-id></mixed-citation></ref>
<ref id="B48"><mixed-citation publication-type="journal"><string-name><surname>M&#228;nnel</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Friederici</surname>, <given-names>A. D.</given-names></string-name> (<year>2016</year>). <article-title>Neural correlates of prosodic boundary perception in German preschoolers: If pause is present, pitch can go</article-title>. <source>Brain Research</source>, <volume>1632</volume>, <fpage>27</fpage>&#8211;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1016/j.brainres.2015.12.009</pub-id></mixed-citation></ref>
<ref id="B49"><mixed-citation publication-type="journal"><string-name><surname>Martin</surname>, <given-names>C. D.</given-names></string-name>, <string-name><surname>Niziolek</surname>, <given-names>C. A.</given-names></string-name>, <string-name><surname>Du&#241;abeitia</surname>, <given-names>J. A.</given-names></string-name>, <string-name><surname>Perez</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Hernandez</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Carreiras</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Houde</surname>, <given-names>J. F.</given-names></string-name> (<year>2018</year>). <article-title>Online adaptation to altered auditory feedback is predicted by auditory acuity and not by domain-general executive control resources</article-title>. <source>Frontiers in Human Neuroscience</source>, <volume>12</volume>, <elocation-id>91</elocation-id>. <pub-id pub-id-type="doi">10.3389/fnhum.2018.00091</pub-id></mixed-citation></ref>
<ref id="B50"><mixed-citation publication-type="journal"><string-name><surname>McAuliffe</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Socolof</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Mihuc</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Wagner</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Sonderegger</surname>, <given-names>M.</given-names></string-name> (<year>2017</year>). <article-title>Montreal Forced Aligner: Trainable text-speech alignment using Kaldi</article-title>. <source>Proceedings of Interspeech 2017</source>, <fpage>498</fpage>&#8211;<lpage>502</lpage>. <pub-id pub-id-type="doi">10.21437/Interspeech.2017-1386</pub-id></mixed-citation></ref>
<ref id="B51"><mixed-citation publication-type="book"><string-name><surname>McElreath</surname>, <given-names>R.</given-names></string-name> (<year>2020</year>). <source>Statistical rethinking</source>. <publisher-name>Chapman and Hall/CRC</publisher-name>. <pub-id pub-id-type="doi">10.1201/9780429029608</pub-id></mixed-citation></ref>
<ref id="B52"><mixed-citation publication-type="journal"><string-name><surname>Meier</surname>, <given-names>A. M.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2023</year>). <article-title>Neurocomputational modeling of speech motor development</article-title>. <source>Journal of Child Language</source>, <volume>50</volume>(<issue>6</issue>), <fpage>1318</fpage>&#8211;<lpage>1335</lpage>. <pub-id pub-id-type="doi">10.1017/S0305000923000260</pub-id></mixed-citation></ref>
<ref id="B53"><mixed-citation publication-type="journal"><string-name><surname>Miller</surname>, <given-names>H. E.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2021</year>). <article-title>Modelling speech motor programming and apraxia of speech in the DIVA/GODIVA neurocomputational framework</article-title>. <source>Aphasiology</source>, <volume>35</volume>(<issue>4</issue>), <fpage>424</fpage>&#8211;<lpage>441</lpage>. <pub-id pub-id-type="doi">10.1080/02687038.2020.1765307</pub-id></mixed-citation></ref>
<ref id="B54"><mixed-citation publication-type="journal"><string-name><surname>Newman</surname>, <given-names>R. S.</given-names></string-name> (<year>2003</year>). <article-title>Using links between speech perception and speech production to evaluate different acoustic metrics: A preliminary report</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>113</volume>(<issue>5</issue>), <fpage>2850</fpage>&#8211;<lpage>2860</lpage>. <pub-id pub-id-type="doi">10.1121/1.1567280</pub-id></mixed-citation></ref>
<ref id="B55"><mixed-citation publication-type="journal"><string-name><surname>Newsome</surname>, <given-names>W. T.</given-names></string-name>, <string-name><surname>Britten</surname>, <given-names>K. H.</given-names></string-name>, &amp; <string-name><surname>Movshon</surname>, <given-names>J. A.</given-names></string-name> (<year>1989</year>). <article-title>Neuronal correlates of a perceptual decision</article-title>. <source>Nature</source>, <volume>341</volume>(<issue>6237</issue>), <fpage>52</fpage>&#8211;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1038/341052a0</pub-id></mixed-citation></ref>
<ref id="B56"><mixed-citation publication-type="journal"><string-name><surname>Ortega-Llebaria</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Nagao</surname>, <given-names>J.</given-names></string-name> (<year>2025</year>). <article-title>When pitch falls short: Reinforcing prosodic boundaries to signal focus in Japanese</article-title>. <source>Languages</source>, <volume>10</volume>(<issue>9</issue>), <elocation-id>242</elocation-id>. <pub-id pub-id-type="doi">10.3390/languages10090242</pub-id></mixed-citation></ref>
<ref id="B57"><mixed-citation publication-type="journal"><string-name><surname>Oschkinat</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Hoole</surname>, <given-names>P.</given-names></string-name> (<year>2022</year>). <article-title>Reactive feedback control and adaptation to perturbed speech timing in stressed and unstressed syllables</article-title>. <source>Journal of Phonetics</source>, <volume>91</volume>, <elocation-id>101133</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.wocn.2022.101133</pub-id></mixed-citation></ref>
<ref id="B58"><mixed-citation publication-type="journal"><string-name><surname>Oschkinat</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Hoole</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Falk</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Dalla Bella</surname>, <given-names>S.</given-names></string-name> (<year>2022</year>). <article-title>Temporal malleability to auditory feedback perturbation is modulated by rhythmic abilities and auditory acuity</article-title>. <source>Frontiers in Human Neuroscience</source>, <volume>16</volume>, <elocation-id>885074</elocation-id>. <pub-id pub-id-type="doi">10.3389/fnhum.2022.885074</pub-id></mixed-citation></ref>
<ref id="B59"><mixed-citation publication-type="journal"><string-name><surname>Patel</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Niziolek</surname>, <given-names>C. A.</given-names></string-name>, <string-name><surname>Reilly</surname>, <given-names>K. J.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2011</year>). <article-title>Prosodic adaptations to pitch perturbation in running speech</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>54</volume>(<issue>4</issue>), <fpage>1051</fpage>&#8211;<lpage>1059</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2010/10-0162)</pub-id></mixed-citation></ref>
<ref id="B60"><mixed-citation publication-type="journal"><string-name><surname>Patel</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Schell</surname>, <given-names>K. W.</given-names></string-name> (<year>2008</year>). <article-title>The influence of linguistic content on the Lombard effect</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>51</volume>(<issue>1</issue>), <fpage>209</fpage>&#8211;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2008/016)</pub-id></mixed-citation></ref>
<ref id="B61"><mixed-citation publication-type="journal"><string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name> (<year>2012</year>). <article-title>Movement goals and feedback and feedforward control mechanisms in speech production</article-title>. <source>Journal of Neurolinguistics</source>, <volume>25</volume>(<issue>5</issue>), <fpage>382</fpage>&#8211;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1016/j.jneuroling.2010.02.011</pub-id></mixed-citation></ref>
<ref id="B62"><mixed-citation publication-type="journal"><string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name>, <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Lane</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Matthies</surname>, <given-names>M. L.</given-names></string-name>, <string-name><surname>Stockmann</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Zandipour</surname>, <given-names>M.</given-names></string-name> (<year>2004</year>). <article-title>The distinctness of speakers&#8217; productions of vowel contrasts is related to their discrimination of the contrasts</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>116</volume>(<issue>4</issue>), <fpage>2338</fpage>&#8211;<lpage>2344</lpage>. <pub-id pub-id-type="doi">10.1121/1.1787524</pub-id></mixed-citation></ref>
<ref id="B63"><mixed-citation publication-type="journal"><string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name>, <string-name><surname>Matthies</surname>, <given-names>M. L.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Lane</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zandipour</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Marrone</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Stockmann</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2004</year>). <article-title>The distinctness of speakers&#8217; /s/-/s/ contrast is related to their auditory discrimination and use of an articulatory saturation effect</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>47</volume>(<issue>6</issue>), <fpage>1259</fpage>&#8211;<lpage>1269</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2004/095</pub-id>)</mixed-citation></ref>
<ref id="B64"><mixed-citation publication-type="journal"><string-name><surname>Petrone</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Truckenbrodt</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wellmann</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Holzgrefe-Lang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>H&#246;hle</surname>, <given-names>B.</given-names></string-name> (<year>2017</year>). <article-title>Prosodic boundary cues in German: Evidence from the production and perception of bracketed lists</article-title>. <source>Journal of Phonetics</source>, <volume>61</volume>, <fpage>71</fpage>&#8211;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2017.01.002</pub-id></mixed-citation></ref>
<ref id="B65"><mixed-citation publication-type="journal"><string-name><surname>Pijper</surname>, <given-names>J. R. de</given-names></string-name>, &amp; <string-name><surname>Sanderman</surname>, <given-names>A. A.</given-names></string-name> (<year>1994</year>). <article-title>On the perceptual strength of prosodic boundaries and its relation to suprasegmental cues</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>96</volume>(<issue>4</issue>), <fpage>2037</fpage>&#8211;<lpage>2047</lpage>. <pub-id pub-id-type="doi">10.1121/1.410145</pub-id></mixed-citation></ref>
<ref id="B66"><mixed-citation publication-type="journal"><string-name><surname>Poeppel</surname>, <given-names>D.</given-names></string-name> (<year>2003</year>). <article-title>The analysis of speech in different temporal integration windows: Cerebral lateralization as &#8220;asymmetric sampling in time.&#8221;</article-title> <source>Speech Communication</source>, <volume>41</volume>(<issue>1</issue>), <fpage>245</fpage>&#8211;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1016/S0167-6393(02)00107-3</pub-id></mixed-citation></ref>
<ref id="B67"><mixed-citation publication-type="journal"><string-name><surname>R&#233;v&#233;sz</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Michel</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Gilabert</surname>, <given-names>R.</given-names></string-name> (<year>2016</year>). <article-title>Measuring cognitive task demands using dual task methodology, subjective self-ratings, and expert judgments: A validation study</article-title>. <source>Studies in Second Language Acquisition</source>, <volume>38</volume>(<issue>4</issue>), <fpage>703</fpage>&#8211;<lpage>737</lpage>. <pub-id pub-id-type="doi">10.1017/S0272263115000339</pub-id></mixed-citation></ref>
<ref id="B68"><mixed-citation publication-type="journal"><string-name><surname>Schad</surname>, <given-names>D. J.</given-names></string-name>, <string-name><surname>Vasishth</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Hohenstein</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Kliegl</surname>, <given-names>R.</given-names></string-name> (<year>2020</year>). <article-title>How to capitalize on a priori contrasts in linear (mixed) models: A tutorial</article-title>. <source>Journal of Memory and Language</source>, <volume>110</volume>, <elocation-id>104038</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.jml.2019.104038</pub-id></mixed-citation></ref>
<ref id="B69"><mixed-citation publication-type="book"><string-name><surname>Schub&#246;</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Zerbian</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Hanne</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Wartenburger</surname>, <given-names>I.</given-names></string-name> (<year>2023</year>). <source>Prosodic boundary phenomena</source>. <publisher-name>Zenodo</publisher-name>. <pub-id pub-id-type="doi">10.5281/zenodo.7777469</pub-id></mixed-citation></ref>
<ref id="B70"><mixed-citation publication-type="journal"><string-name><surname>Smith</surname>, <given-names>D. J.</given-names></string-name>, <string-name><surname>Stepp</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, &amp; <string-name><surname>Kearney</surname>, <given-names>E.</given-names></string-name> (<year>2020</year>). <article-title>Contributions of auditory and somatosensory feedback to vocal motor control</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>63</volume>(<issue>7</issue>), <fpage>2039</fpage>&#8211;<lpage>2053</lpage>. <pub-id pub-id-type="doi">10.1044/2020_JSLHR-19-00296</pub-id></mixed-citation></ref>
<ref id="B71"><mixed-citation publication-type="webpage"><collab>Stan Development Team</collab>. (<year>2020</year>). <source>RStan: The R interface to Stan</source>. <uri>https://mc-stan.org/</uri></mixed-citation></ref>
<ref id="B72"><mixed-citation publication-type="journal"><string-name><surname>Tomassi</surname>, <given-names>N. E.</given-names></string-name>, <string-name><surname>Weerathunge</surname>, <given-names>H. R.</given-names></string-name>, <string-name><surname>Cushman</surname>, <given-names>M. R.</given-names></string-name>, <string-name><surname>Bohland</surname>, <given-names>J. W.</given-names></string-name>, &amp; <string-name><surname>Stepp</surname>, <given-names>C. E.</given-names></string-name> (<year>2022</year>). <article-title>Assessing ecologically valid methods of auditory feedback measurement in individuals with typical speech</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>65</volume>(<issue>1</issue>), <fpage>121</fpage>&#8211;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1044/2021_JSLHR-21-00377</pub-id></mixed-citation></ref>
<ref id="B73"><mixed-citation publication-type="journal"><string-name><surname>Tourville</surname>, <given-names>J. A.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2011</year>). <article-title>The DIVA model: A neural theory of speech acquisition and production</article-title>. <source>Language and Cognitive Processes</source>, <volume>26</volume>(<issue>7</issue>), <fpage>952</fpage>&#8211;<lpage>981</lpage>. <pub-id pub-id-type="doi">10.1080/01690960903498424</pub-id></mixed-citation></ref>
<ref id="B74"><mixed-citation publication-type="journal"><string-name><surname>Tourville</surname>, <given-names>J. A.</given-names></string-name>, <string-name><surname>Reilly</surname>, <given-names>K. J.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2008</year>). <article-title>Neural mechanisms underlying auditory feedback control of speech</article-title>. <source>NeuroImage</source>, <volume>39</volume>(<issue>3</issue>), <fpage>1429</fpage>&#8211;<lpage>1443</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.09.054</pub-id></mixed-citation></ref>
<ref id="B75"><mixed-citation publication-type="journal"><string-name><surname>Vasishth</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Nicenboim</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Beckman</surname>, <given-names>M. E.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>F.</given-names></string-name>, &amp; <string-name><surname>Kong</surname>, <given-names>E. J.</given-names></string-name> (<year>2018</year>). <article-title>Bayesian data analysis in the phonetic sciences: A tutorial introduction</article-title>. <source>Journal of Phonetics</source>, <volume>71</volume>, <fpage>147</fpage>&#8211;<lpage>161</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2018.07.008</pub-id></mixed-citation></ref>
<ref id="B76"><mixed-citation publication-type="journal"><string-name><surname>Ver&#237;ssimo</surname>, <given-names>J.</given-names></string-name> (<year>2024</year>). <source>A gentle introduction to Bayesian statistics, with applications to bilingualism research</source>. <pub-id pub-id-type="doi">10.31234/osf.io/7wfus_v1</pub-id></mixed-citation></ref>
<ref id="B77"><mixed-citation publication-type="journal"><string-name><surname>Villacorta</surname>, <given-names>V. M.</given-names></string-name>, <string-name><surname>Perkell</surname>, <given-names>J. S.</given-names></string-name>, &amp; <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name> (<year>2007</year>). <article-title>Sensorimotor adaptation to feedback perturbations of vowel acoustics and its relation to perception</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>122</volume>(<issue>4</issue>), <fpage>2306</fpage>&#8211;<lpage>2319</lpage>. <pub-id pub-id-type="doi">10.1121/1.2773966</pub-id></mixed-citation></ref>
<ref id="B78"><mixed-citation publication-type="journal"><string-name><surname>Wagenmakers</surname>, <given-names>E.-J.</given-names></string-name>, <string-name><surname>Lodewyckx</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Kuriyal</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Grasman</surname>, <given-names>R.</given-names></string-name> (<year>2010</year>). <article-title>Bayesian hypothesis testing for psychologists: A tutorial on the Savage-Dickey method</article-title>. <source>Cognitive Psychology</source>, <volume>60</volume>(<issue>3</issue>), <fpage>158</fpage>&#8211;<lpage>189</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogpsych.2009.12.001</pub-id></mixed-citation></ref>
<ref id="B79"><mixed-citation publication-type="journal"><string-name><surname>Weerathunge</surname>, <given-names>H. R.</given-names></string-name>, <string-name><surname>Alzamendi</surname>, <given-names>G. A.</given-names></string-name>, <string-name><surname>Cler</surname>, <given-names>G. J.</given-names></string-name>, <string-name><surname>Guenther</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Stepp</surname>, <given-names>C. E.</given-names></string-name>, &amp; <string-name><surname>Za&#241;artu</surname>, <given-names>M.</given-names></string-name> (<year>2022</year>). <article-title>LaDIVA: A neurocomputational model providing laryngeal motor control for speech acquisition and production</article-title>. <source>PLoS Computational Biology</source>, <volume>18</volume>(<issue>6</issue>), <elocation-id>e1010159</elocation-id>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1010159</pub-id></mixed-citation></ref>
<ref id="B80"><mixed-citation publication-type="book"><string-name><surname>Whalen</surname>, <given-names>D. H.</given-names></string-name> (<year>2020</year>). <chapter-title>The motor theory of speech perception</chapter-title>. In <string-name><given-names>M.</given-names> <surname>Pouplier</surname></string-name> (Ed.), <source>Oxford research encyclopedia of linguistics</source>. <publisher-name>Oxford University Press</publisher-name>. <pub-id pub-id-type="doi">10.1093/acrefore/9780199384655.013.404</pub-id></mixed-citation></ref>
<ref id="B81"><mixed-citation publication-type="journal"><string-name><surname>Wickens</surname>, <given-names>C. D.</given-names></string-name> (<year>2008</year>). <article-title>Multiple resources and mental workload</article-title>. <source>Human Factors</source>, <volume>50</volume>(<issue>3</issue>), <fpage>449</fpage>&#8211;<lpage>455</lpage>. <pub-id pub-id-type="doi">10.1518/001872008X288394</pub-id></mixed-citation></ref>
<ref id="B82"><mixed-citation publication-type="journal"><string-name><surname>Wightman</surname>, <given-names>C. W.</given-names></string-name>, <string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Ostendorf</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Price</surname>, <given-names>P. J.</given-names></string-name> (<year>1992</year>). <article-title>Segmental durations in the vicinity of prosodic phrase boundaries</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>91</volume>(<issue>3</issue>), <fpage>1707</fpage>&#8211;<lpage>1717</lpage>. <pub-id pub-id-type="doi">10.1121/1.402450</pub-id></mixed-citation></ref>
<ref id="B83"><mixed-citation publication-type="journal"><string-name><surname>Winkworth</surname>, <given-names>A. L.</given-names></string-name>, &amp; <string-name><surname>Davis</surname>, <given-names>P. J.</given-names></string-name> (<year>1997</year>). <article-title>Speech breathing and the Lombard effect</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>40</volume>(<issue>1</issue>), <fpage>159</fpage>&#8211;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1044/jslhr.4001.159</pub-id></mixed-citation></ref>
<ref id="B84"><mixed-citation publication-type="journal"><string-name><surname>Xie</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Bux&#243;-Lugo</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Kurumada</surname>, <given-names>C.</given-names></string-name> (<year>2021</year>). <article-title>Encoding and decoding of meaning through structured variability in intonational speech prosody</article-title>. <source>Cognition</source>, <volume>211</volume>, <elocation-id>104619</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.cognition.2021.104619</pub-id></mixed-citation></ref>
<ref id="B85"><mixed-citation publication-type="journal"><string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>W.</given-names></string-name>, &amp; <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name> (<year>2014</year>). <article-title>How listeners weight acoustic cues to intonational phrase boundaries</article-title>. <source>PloS One</source>, <volume>9</volume>(<issue>7</issue>). <pub-id pub-id-type="doi">10.1371/journal.pone.0102166.g001</pub-id></mixed-citation></ref>
<ref id="B86"><mixed-citation publication-type="journal"><string-name><surname>Zhang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Ji</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>He</surname>, <given-names>L.</given-names></string-name> (<year>2015</year>). <article-title>Research on the mechanism for phonating stressed English syllables based on DIVA model</article-title>. <source>Neurocomputing</source>, <volume>152</volume>, <fpage>11</fpage>&#8211;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2014.11.032</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>