<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="Style/article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">2767-0279</journal-id>
<journal-title-group>
<journal-title>Glossa Psycholinguistics</journal-title>
</journal-title-group>
<issn pub-type="epub">2767-0279</issn>
<publisher>
<publisher-name>eScholarship Publishing</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5070/G6011.49062</article-id>
<article-categories>
<subj-group>
<subject>Regular article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Too many tokens: Modeling lexical effects in corpora</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-1293-7899</contrib-id>
<name>
<surname>Gahl</surname>
<given-names>Susanne</given-names>
</name>
<email>gahl@berkeley.edu</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-3178-3944</contrib-id>
<name>
<surname>Baayen</surname>
<given-names>R. Harald</given-names>
</name>
<email>harald.baayen@uni-tuebingen.de</email>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>University of California at Berkeley</aff>
<aff id="aff-2"><label>2</label>Eberhard Karls University, T&#252;bingen</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-08-06">
<day>06</day>
<month>08</month>
<year>2026</year>
</pub-date>
<pub-date pub-type="collection">
<year>2026</year>
</pub-date>
<volume>5</volume>
<issue>1</issue>
<elocation-id>13</elocation-id>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2026 The Author(s)</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://glossapsycholinguistics.journalpub.escholarship.org/articles/10.5070/G6011.49062/"/>
<abstract>
<p>Many statistical models of properties of words in texts and in running speech use as their outcome variables properties of word tokens, rather than properties aggregated over word types. Here, we describe methodological and conceptual problems with token-level (or <italic>token-based</italic>) models, including assumption violations of the models typically used, model redundancy, and distortions of parameters estimating lexical effects &#8211; i.e. the very target of what such models seek to capture. We show that these problems persist &#8211; and in some ways, grow more extreme &#8211; when we follow widely-used practices for specifying random effects structure. We do so by comparing models of spoken word duration of homophones in the Switchboard corpus, a dataset that has been discussed and modeled extensively. We discuss possible remedies and recommendations.</p>
</abstract>
</article-meta>
</front>
<body>
<sec>
<title>1. Introduction</title>
<p>Statistical models of variation in the realization of words have generally taken one of two approaches. The first approach is to aggregate information over word types: In these models, the dependent variables are averages, such as mean latencies, word duration, vowel formant values, letter stroke duration, typing speed, or other properties of words as articulatory, visual, and acoustic events, averaged by word. This is the approach taken in numerous analyses of controlled, balanced data from experiments (e.g. <xref ref-type="bibr" rid="B69">Wright, 2004</xref>), as well as in some corpus-based studies, such as Gahl (<xref ref-type="bibr" rid="B20">2008</xref>) and Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>). The second approach models properties of individual tokens, such as trial-level response latencies, token duration, and so on. This is the approach taken in most corpus-based analyses (e.g. <xref ref-type="bibr" rid="B10">Bell et al., 2003</xref>, <xref ref-type="bibr" rid="B9">2009</xref>; <xref ref-type="bibr" rid="B18">Fosler-Lussier &amp; Morgan, 1999</xref>; <xref ref-type="bibr" rid="B23">Gahl et al., 2012</xref>; <xref ref-type="bibr" rid="B26">Gries, 2015</xref>; <xref ref-type="bibr" rid="B27">Hay, 2007</xref>; <xref ref-type="bibr" rid="B28">Hay et al., 2015</xref>; <xref ref-type="bibr" rid="B33">Kilbourn-Ceron et al., 2020</xref>; <xref ref-type="bibr" rid="B41">Lohmann, 2018b</xref>; <xref ref-type="bibr" rid="B50">Seyfarth, 2014</xref>; <xref ref-type="bibr" rid="B55">Tanner et al., 2020</xref>).</p>
<p>The obvious appeal of token-level analyses lies in the richness of the information that can be considered, such as local speaking rate, the words preceding or following the target, indexical information about talkers and listeners, and so on. Token-level models conform to the sound methodological principle of &#8220;not throwing away data&#8221;, i.e. avoiding information loss due to aggregation. Another alluring feature of such analyses lies in the sheer size of data sets, which can support nuanced statistical models. On top of these advantages, the influence of token-level analyses is, to some degree, self-perpetuating, in that new investigations tend to include token-level models, so as to enable comparisons to previous work. Token-level models have effectively acquired the status of a gold standard, with type-based models looking like a less sophisticated alternative.</p>
<p>The aim of the current study is to point out a set of methodological problems arising with token-level models of unbalanced data sets, such as naturalistic corpus data. We argue that such models introduce biases that, ironically, can get in the way of understanding lexical effects. We are emphatically not claiming that token-level models are to be avoided under all circumstances. Rather, we wish to point out some consequences of token-level modeling that restrict the usefulness of such models to certain types of inquiry. Simplifying our point somewhat: token-level models turn out to be surprisingly ill-suited for modeling what are often simply referred to as &#8220;lexical effects&#8221; or &#8220;word-specific&#8221; effects. The optimal statistical analysis ultimately depends on one&#8217;s theory of the domain being modeled, as well as on the goals of the analysis.</p>
<p>In what follows, we argue that the problems with token-level models (or models sometimes termed <italic>token-based</italic>, in which the observations being modeled constitute tokens, but whose ultimate analytical target are types) include the following:</p>
<list list-type="bullet">
<list-item><p><bold>Ballot-box stuffing:</bold> Because token-based regression models will be penalized for the residual of every token, such models will be driven by high-frequency target words. Because deviations of model predictions are punished for every single observation, token-based models can only achieve excellent fit if they work well for high-frequency words.</p></list-item>
<list-item><p><bold>Model redundancy:</bold> Ballot-box stuffing happens not just once, but for every lexical variable included in a token-based model: Tokens of any given high-frequency target word contribute identical sets of values for all type-level properties of the target. These replicates of co-occurring properties (&#8220;clones&#8221;) of each type can render variables redundant, i.e. predictable from other variables in the model, or from the model as a whole, making it difficult or impossible to evaluate the contributions of any individual predictors.</p></list-item>
<list-item><p><bold>Distorted parameter estimates:</bold> Ballot-box stuffing, while harmful for the reasons just outlined, might, in principle, apply evenly across the ranges of lexical variables. However, at the type level, word frequency is correlated with other variables (<xref ref-type="bibr" rid="B1">Baayen, 2011</xref>; <xref ref-type="bibr" rid="B19">Frauenfelder et al., 1993</xref>; <xref ref-type="bibr" rid="B35">K&#246;hler, 1986</xref>; <xref ref-type="bibr" rid="B37">Landauer &amp; Streeter, 1973</xref>), and high-frequency words tend to have certain phonological and semantic properties in common. The resulting similarities across high-frequency types (i.e. similarities across &#8220;voters&#8221;, rather than sets of identical ballots cast by each voter) has the potential to inflate the predictiveness of some variables and underestimate that of others.</p></list-item>
<list-item><p><bold>Assumption violations:</bold> The uneven distribution of observations and the correlation of frequency with other variables lead to violations of modeling assumptions, such as normality, independence, and homogeneity of residuals.</p></list-item>
</list>
<p>Many researchers trust that the problems mentioned so far can be avoided through judicious model specification, specifically, by-word random effects in mixed-effects regression models. Random effects model the distribution of token-level observations within type-based clusters (<xref ref-type="bibr" rid="B2">Baayen et al., 2008</xref>; <xref ref-type="bibr" rid="B8">Bates, 2005</xref>; <xref ref-type="bibr" rid="B43">Pinheiro &amp; Bates, 2000</xref>; <xref ref-type="bibr" rid="B47">Quen&#233; &amp; van den Bergh, 2008</xref>). This prevents clusters of tokens from gaining undue influence over the model as a whole &#8211; or so the assumption goes. Here, we demonstrate that the problems just outlined persist when the random effects structure of the models include by-word random effects.</p>
<p>The consequences of analyzing token-level vs. type-level information have not been obvious, in part because any given study typically reports only token-level or only type-level results, but not both. In the current study, we report on both type-level and token-level models of spoken word duration, in order to draw attention to problems with token-based models. We analyze a dataset that has previously been analyzed and discussed extensively, viz. homophone duration in the Switchboard corpus, i.e. the duration of words like <italic>time</italic> and <italic>thyme</italic>. By focusing on an existing data set, we wish to highlight consequences of methodological choices, rather than properties of specific variables or item lists. We illustrate the problems with token-based analyses and discuss possible remedies, concluding with a set of recommendations.</p>
</sec>
<sec>
<title>2. Methods</title>
<sec>
<title>2.1 Data</title>
<p>We analyzed the same publicly available data set described in Gahl (<xref ref-type="bibr" rid="B20">2008</xref>, <xref ref-type="bibr" rid="B21">2009</xref>), Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>), and Lohmann (<xref ref-type="bibr" rid="B40">2018a</xref>). With the exception of Gahl (<xref ref-type="bibr" rid="B21">2009</xref>), these previous analyses of the data set all used type-level models, i.e. modeling word duration averaged over word types. We refrained from collecting or proposing any novel variables or making other changes to the dataset, for the sake of comparability to earlier studies and so as to maintain the focus on methodological issues.</p>
<p>According to Gahl (<xref ref-type="bibr" rid="B20">2008</xref>), the initial list of target words for the data set contained all English word forms with at least one non-homographic homophone, based on the CELEX database (<xref ref-type="bibr" rid="B6">Baayen et al., 1995</xref>). Word forms with identical spelling and pronunciation (e.g. nouns and verbs spelled <italic>time</italic>) were treated as tokens of the same type. Gahl (<xref ref-type="bibr" rid="B20">2008</xref>)&#8217;s exclusion criteria excluded the following items: (1) spellings associated with more than one pronunciation, e.g. <italic>tear</italic> (homophonous with <italic>tier</italic> and <italic>tare</italic>); (2) pairs involving function words, such as <italic>in/inn</italic> and <italic>or/ore</italic>, and interjections, such as <italic>whoa/woe</italic>; (3) pairs, such as <italic>source/sauce</italic>, that are homophones in the (Received Pronunciation of British English) CELEX transcriptions, but unlikely to be homophones in the varieties of American English represented in the Switchboard corpus; (4) items containing transcription errors in CELEX; and (5) names of letters of the alphabet. The resulting list contains 409 target types, represented by 79,219 tokens.</p>
<p>The spoken duration of all tokens of these items was extracted from the time-aligned orthographic transcript (<xref ref-type="bibr" rid="B17">Deshmukh et al., 1998</xref>) of the Switchboard corpus (<xref ref-type="bibr" rid="B25">Godfrey et al., 1992</xref>), a corpus of 240 hours of telephone conversations between strangers. We initially retained tokens that were immediately followed by an unfilled pause (defined as a period of silence of 0.5 seconds or longer) or by a filled pause (e.g. a hesitation marker such as <italic>um, uh, err</italic>) and allowed Fluency, coded as a binary factor, to interact with talker age and with the relative frequency of target and homophone. However, these disfluent tokens were too unevenly distributed to allow meaningful models of these interactions. Therefore, we excluded them from all further analyses. Excluding the disfluent tokens left 56,024 tokens of 403 distinct types for analysis.</p>
</sec>
<sec>
<title>2.2 Variables considered</title>
<sec>
<title>2.2.1 Type-level predictors</title>
<p>The type-level variables considered here are those that were included in the multiple linear regression model in Gahl (<xref ref-type="bibr" rid="B20">2008</xref>), the first analysis of the dataset, and/or subsequent analyses using GAMMs (<xref ref-type="bibr" rid="B22">Gahl &amp; Baayen, 2024</xref>; <xref ref-type="bibr" rid="B40">Lohmann, 2018a</xref>). The variables for these analyses were based on prior research on spoken word duration, including Bell et al. (<xref ref-type="bibr" rid="B10">2003</xref>, <xref ref-type="bibr" rid="B9">2009</xref>), Gahl et al. (<xref ref-type="bibr" rid="B23">2012</xref>), Lieberman (<xref ref-type="bibr" rid="B39">1963</xref>), Pluymaekers et al. (<xref ref-type="bibr" rid="B46">2006</xref>), Shields and Balota (<xref ref-type="bibr" rid="B51">1991</xref>), Sorensen et al. (<xref ref-type="bibr" rid="B52">1978</xref>), Walsh and Parker (<xref ref-type="bibr" rid="B59">1983</xref>), and Warner et al. (<xref ref-type="bibr" rid="B60">2004</xref>).</p>
<disp-quote>
<p><italic>Residualized baseline duration</italic> The residuals of a GAM predicting the target&#8217;s baseline duration (as the sum of the average phone durations) from properties of word forms, following Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>); model summary and visualization of smooths for that model are in Appendix A. The rationale for residualizing baseline duration, rather than using the baseline duration itself, is that the distribution of phonological segments in the English lexicon is, to some extent, predictable from lexical variables, such as Phonological Neighborhood Density and lexical frequency (as shown for Dutch, English, German and French in <xref ref-type="bibr" rid="B15">Dautriche et al., 2017</xref>); therefore, estimates of word duration based on by-segment averages do not accomplish what the baseline duration measure is intended to accomplish, which is to control for the inherent duration of the segments (e.g. the duration of a voiceless alveolar stop [t], followed by [a&#1110;], followed by [m]) as distinct from word-level properties, such as Phonological Neighborhood Density, lexical frequency, and so on.</p>
<p><italic>Biphone probability</italic> The average of the target&#8217;s position-specific biphone probabilities, based on Vitevitch and Luce (<xref ref-type="bibr" rid="B58">2004</xref>).</p>
<p><italic>Lemma frequency</italic> The log-transformed frequency of the target (e.g. <italic>time</italic> vs. <italic>thyme</italic>), based on the CELEX database (<xref ref-type="bibr" rid="B6">Baayen et al., 1995</xref>).</p>
<p><italic>Relative frequency</italic> The (log-transformed) CELEX frequency of the target, divided by the frequency of its homophone (e.g. the frequency of <italic>time</italic>, divided by the frequency of <italic>thyme</italic>). The same variable was used in Lohmann (<xref ref-type="bibr" rid="B40">2018a</xref>). The relative frequency is greater than zero for the higher-frequency member of each pair of homophones, and smaller than zero for the lower-frequency member of the pair. The relative frequency entered the models in the form of a tensor product to model the interaction between <monospace>Lemma frequency</monospace> and <monospace>Relative frequency</monospace>. The rationale for including that interaction is that the higher the target lemma frequency, the more it can exceed its homophone twin in frequency.</p>
<p><italic>Word form</italic> A factor identifying the phonological form that homophones have in common, e.g. /ta&#1110;m/ for <italic>time</italic> and <italic>thyme</italic> (termed <italic>lexemes</italic> in <xref ref-type="bibr" rid="B20">Gahl, 2008</xref>, following <xref ref-type="bibr" rid="B38">Levelt et al., 1999</xref>).</p>
<p><italic>Morphological complexity</italic> A binary factor distinguishing morphological simple vs. complex target words, typically third person singular, <italic>-s</italic> plural, or past tense forms, such as <italic>lacks</italic> (vs. <italic>lax</italic>) or <italic>allowed</italic> (vs. <italic>aloud</italic>).</p>
<p><italic>Noun quotient</italic> The estimated proportion of nouns among the tokens of a given form (noun vs. verb uses of <italic>time</italic>), based on CELEX (<xref ref-type="bibr" rid="B6">Baayen et al., 1995</xref>). Nouns occur in phrase-final position more often than words of other syntactic categories in English and, therefore, are more likely to undergo phrase-final lengthening. For example, tokens of <italic>thyme</italic> are more likely to be phrase-final than tokens of <italic>time</italic> (which may represent verbs and, therefore, less likely to be phrase-final). As this variable was bimodally distributed, it was dichotomized, indicating whether the proportion of noun uses of a given form was above vs. below 0.5.</p>
<p><italic>Orthographic regularity</italic> According to Gahl (<xref ref-type="bibr" rid="B20">2008</xref>), a measure indexing the average probability of a word&#8217;s graphemes, normalized by the probability of the most probable pronunciation of each grapheme, based on American English grapheme-to-phoneme probabilities (<xref ref-type="bibr" rid="B11">Berndt et al., 1987</xref>).</p>
<p><italic>Pause quotient</italic> The proportion of tokens of a given target that were immediately followed by a pause.</p>
<p><italic>Phonological Neighborhood Density (PND)</italic> The number of words differing from the target word by addition, deletion, or substitution, based on the English Lexicon Project (<xref ref-type="bibr" rid="B7">Balota et al., 2007</xref>).</p>
</disp-quote>
</sec>
<sec>
<title>2.2.2 Token-specific predictors</title>
<p>Our token-based models contain several additional variables pertaining to specific tokens of words, as follows:</p>
<disp-quote>
<p><italic>Age</italic> The talker&#8217;s age.</p>
<p><italic>Bigram probability</italic> The word-based bigram probability of the target, given the word following it.</p>
<p><italic>Sex</italic> The talker&#8217;s sex (male or female, as coded in <xref ref-type="bibr" rid="B25">Godfrey et al., 1992</xref>).</p>
</disp-quote>
</sec>
</sec>
<sec>
<title>2.3 Statistical tools and strategies</title>
<sec>
<title>2.3.1 GAM(M)s</title>
<p>We made use of the Gaussian Location Scale Additive Model, fitting Generalized Additive Models (<sc>gam</sc>s) and Generalized Additive Mixed Models (<sc>gamm</sc>s), using the packages <monospace>mgcv</monospace> (<xref ref-type="bibr" rid="B64">Wood, 2003</xref>, <xref ref-type="bibr" rid="B65">2004</xref>, <xref ref-type="bibr" rid="B67">2011</xref>, <xref ref-type="bibr" rid="B68">2017</xref>) and <monospace>itsadug</monospace> (<xref ref-type="bibr" rid="B57">van Rij et al., 2020</xref>) in R (<xref ref-type="bibr" rid="B48">R Core Team, 2022</xref>). Tutorials and applications of <sc>gam</sc>s to linguistic data can be found, for example, in Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>), Baayen et al. (<xref ref-type="bibr" rid="B4">2010</xref>), Chuang et al. (<xref ref-type="bibr" rid="B14">2021</xref>), S&#243;skuthy (<xref ref-type="bibr" rid="B53">2021</xref>), Wieling (<xref ref-type="bibr" rid="B61">2018</xref>), and Wieling et al. (<xref ref-type="bibr" rid="B63">2011</xref>, <xref ref-type="bibr" rid="B62">2014</xref>). Here, we summarize properties of GA(M)Ms that are critically important to the current study.</p>
<p>The Generalized Additive (Mixed) Model (henceforth, <sc>gam</sc>, except when specifically referring to models with random effects) is a regression model that relaxes two of the assumptions that underlie Linear (Mixed Effects) Regression models, which it otherwise resembles. The first is the linearity assumption, i.e. the assumption that predicted values change at a constant rate across values of the predictor variables. In <sc>gam</sc>s, the relationship between predictors and outcome is not assumed to be linear, but is modeled as the combination of two sets of functions. The second is the assumption of equal variance. Gaussian Location-Scale <sc>gam</sc>s, which are the types of models we use here, allow the variance, which is assumed to be Gaussian, to change (non-linearly, if necessary) with the predicted mean. That is to say, both mean and variance are modeled as (potentially non-linear) functions of the predictors. The first (<italic>parametric</italic>) set of functions resemble linear predictors in linear mixed models (LMMs); this set includes any continuous predictors specified as linear, as well as fixed-effect factors. The second (<italic>non-parametric</italic>) set of terms estimates the relationship between predictors and outcome as a potentially nonlinear function, specifically, the sum of successively more complex (more &#8220;wiggly&#8221;) basis functions. The coefficients of these functions are determined by a procedure balancing model fit and parsimony. For example, if a predictor is truly linear, a <sc>gam</sc> will penalize any nonlinear components to the point of setting their coefficients to zero.</p>
<p>The estimated number and coefficients of the basis functions are not specified ahead of time, but they can be constrained by the researcher. We did so here; in order to counteract artifactual effects of high-frequency words, we limited the number of basis functions in our initial token-based models to 3 for univariate smooths (i.e. k = 4) and to 4 (i.e. k = 5) for the tensor product of lemma frequency and frequency ratio. We consider several alternative numbers of basis functions in our discussion of the models in 3.8.1.</p>
<p><sc>gam</sc>s can model interactions among continuous variables, yielding model estimates of surfaces, which can be visualized by means of contour plots. These plots may be read in the manner of a topographical map, with elevation indicating predicted target duration. In our plots, warm colors (orange to yellow) indicate &#8220;elevated&#8221;, i.e. longer, predicted duration. Cool colors (green to blue) indicate shorter predicted duration.</p>
<p><sc>gam</sc>s can include Gaussian random effects, analogous to random intercepts and slopes in LMM. Mixed models, i.e. models including fixed and random effects, are very common in many areas of linguistics and psycholinguistics, and we believe most readers are likely to be very familiar with such models. We comment on random effects explicitly here, because of their importance in the methodological issue we wish to point out. Typical uses of random effects in research on lexical processing include by-talker and by-item effects, for clustering observations by talker or by item (e.g. words or experimental stimuli), to allow inferences about entire populations of talkers or items. Rather than estimating a predicted mean for each level of a factor, the idea is to estimate the variance around an overall mean. Most studies of lexical factors in pronunciation variation mentioned above employ by-word random intercepts, estimating the variance of a distribution around the model intercept. Many models additionally include random slopes, estimating the variance of the slope of one or more fixed effects, i.e. the variability in the degree to which, for example, different words are lengthened, depending on their position in an utterance or depending on the age of the talker, or the degree to which different talkers&#8217; pronunciations vary, depending on some experimental manipulation or lexical factor.</p>
<p>As mentioned above, GAMs relax the assumption of equal variance that underlies linear regression models. Instead of relying on an assumption of the variance being equal along the range of the predicted outcome, GAMs can include terms for predicting the variance (along with predicting the mean). We capitalize on this property of GAMs here, by modeling the variance in duration as a function of target frequency (a variable at the center of previous analyses of the data) in several of our models.</p>
<p>In light of the fact that our dataset has been studied previously, particularly in type-based models, and given the exploratory nature of our models, we set <italic>&#945;</italic> = .0001.</p>
</sec>
<sec>
<title>2.3.2 Model criticism and evaluation</title>
<p>GAMs, like Linear Mixed Models (LMMs), rely on the assumption that the model residuals do not co-vary with the predictors: Model fit should be about equally good along the full ranges of the predictors.</p>
<p>A second assumption underlying GAMs (and LMMs) is that the by-group adjustments (the random effects) are Gaussian noise that is supposed to be independent and identically distributed; among other things, this means that the random effects are assumed to be uncorrelated with the other predictors.</p>
<p>We used several tools for comparing models to one another and for model criticism.</p>
<disp-quote>
<p><italic>AIC</italic> The Akaike Information Criterion is a measure of model goodness-of-fit penalized for model complexity, i.e. the number of predictors in the model. Put differently: This measure specifies which of two models fitted to the same datapoints is more likely to have generated the data. Decreasing values of AIC indicate better model fit (measured as 2*Log Likelihood), taking into account the number of predictors.</p>
<p><italic>Concurvity</italic> of each predictor in the model, i.e. a measure of the degree to which each component of the model can be approximated by a weighted sum of the other model components. If the smooth term can be predicted to a large extent by the other terms, then there is a concurvity problem: High concurvity renders estimates unstable, i.e. liable to change drastically in response to seemingly minor changes in the model or data. We assessed the concurvity of each term with all remaining predictors in each model, i.e. x<sub>1</sub> = f(x<sub>2</sub>) + f(x<sub>3</sub>) + &#8230;. + f(x<sub>j</sub>). In some cases, we followed up by assessing the pairwise concurvity between the term in question and each of the other smooth terms, i.e. x<sub>1</sub> = f(x<sub>2</sub>), x<sub>1</sub> = f(x<sub>3</sub>), &#8230;, x<sub>1</sub> = f(x<sub>j</sub>). We used the function <monospace>concurvity</monospace> in <monospace>mgcv</monospace> (<xref ref-type="bibr" rid="B68">Wood, 2017</xref>); the output of that function returns several measures based on the ratio of the squared Euclidean norms of vectors representing the smooth term&#8217;s own space vs. the other terms. Here, we report the measure labeled <italic>estimate</italic> in the output of <monospace>concurvity</monospace>. Concurvity estimates range from 0 to 1, with 1 indicating that a term is fully predictable from others. A fairly common practice is to consider values greater than 0.8 to be unacceptably high (for examples of other studies of language and cognition adopting that strategy and tutorials recommending it, see e.g. <xref ref-type="bibr" rid="B5">Baayen &amp; Linke, 2021</xref>; <xref ref-type="bibr" rid="B14">Chuang et al., 2021</xref>; <xref ref-type="bibr" rid="B36">Lammer et al., 2025</xref>; <xref ref-type="bibr" rid="B49">Saxena et al., 2022</xref>; <xref ref-type="bibr" rid="B56">Tomaschek &amp; Ramscar, 2022</xref>), but we do not consider that value as a strict cutoff point; in fact, in the discussion to follow, we also scrutinize much lower concurvities.</p>
<p><italic>Number of basis functions</italic> The number of basis functions constrains the degree of estimated &#8220;wiggliness&#8221; of the prediction curves (and surfaces). We used the function <monospace>k.check</monospace> in the <monospace>mgcv</monospace> package (<xref ref-type="bibr" rid="B68">Wood, 2017</xref>) to check the number of basis functions.</p>
</disp-quote>
</sec>
</sec></sec>
<sec>
<title>3. Results</title>
<sec>
<title>3.1 Type-based model</title>
<p>We use a published, type-based model of the data set as a baseline. The model summary, replicating Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>), appears in <xref ref-type="table" rid="T1">Table 1</xref>, and the model&#8217;s smooth terms are visualized in <xref ref-type="fig" rid="F1">Figure 1</xref>. In the type-based model, predicted (average) duration increased linearly with residualized baseline duration, and in a nearly linear fashion with the proportion of prepausal tokens. The effects of <sc>noun</sc>&#160;<sc>bias</sc> and O<sc>rthographic</sc>&#160;<sc>regularity</sc> were non-significant. Increasing phonological neighborhood density (PND) was associated with decreasing duration in a nonlinear fashion, with the steepest predicted decrease in the lower range of PND. As for the interaction of lemma frequency and frequency ratio, the contour plot in <xref ref-type="fig" rid="F1">Figure 1</xref> shows the general gradient in the regression surface to be negative. That pattern reflects the oft-observed relationship between increasing lemma frequency and decreasing predicted duration. Increasing lemma frequency was associated with decreasing variance in duration (bottom right panel of <xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
<table-wrap id="T1">
<caption>
<p><bold>Table 1:</bold> Gaussian Location-Scale <sc>gam</sc> fitted to the average duration of the homophone word types, retracing Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>). te(Lemma Freq., Rel. Freq) = Tensor product of (log-transformed) target lemma frequency and ratio of (log-transformed) frequencies of the target and its homophone. AIC = &#8211;230.8.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top"><monospace>Intercept</monospace> [mean]</td>
<td align="left" valign="top">&#8211;1.0387</td>
<td align="left" valign="top">0.0139</td>
<td align="left" valign="top">&#8211;74.7519</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top"><monospace>Noun Bias</monospace> = yes</td>
<td align="left" valign="top">0.0598</td>
<td align="left" valign="top">0.0169</td>
<td align="left" valign="top">3.5387</td>
<td align="left" valign="top">.0004</td>
</tr>
<tr>
<td align="left" valign="top"><monospace>Intercept</monospace> [variance]</td>
<td align="left" valign="top">&#8211;1.8293</td>
<td align="left" valign="top">0.0378</td>
<td align="left" valign="top">&#8211;48.4416</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(<monospace>Proportion with Following Pauses</monospace>)</td>
<td align="left" valign="top">1.8335</td>
<td align="left" valign="top">2.2863</td>
<td align="left" valign="top">32.2159</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(<monospace>Phonological Neighborhood Density</monospace>)</td>
<td align="left" valign="top">4.7174</td>
<td align="left" valign="top">5.7485</td>
<td align="left" valign="top">71.3616</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(<monospace>Orthographic regularity</monospace>)</td>
<td align="left" valign="top">1.0000</td>
<td align="left" valign="top">1.0001</td>
<td align="left" valign="top">5.2430</td>
<td align="left" valign="top">.0220</td>
</tr>
<tr>
<td align="left" valign="top">te(<monospace>Lemma Freq., Rel. Freq.</monospace>)</td>
<td align="left" valign="top">7.2717</td>
<td align="left" valign="top">9.4159</td>
<td align="left" valign="top">108.3451</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(<monospace>Residual Baseline Duration</monospace>)</td>
<td align="left" valign="top">1.1915</td>
<td align="left" valign="top">1.3567</td>
<td align="left" valign="top">122.3611</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(<monospace>Lemma Frequency</monospace>) [variance]</td>
<td align="left" valign="top">2.6120</td>
<td align="left" valign="top">3.2966</td>
<td align="left" valign="top">75.3180</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F1">
<caption>
<p><bold>Figure 1:</bold> Partial effects according to the Gaussian Location-Scale <sc>gam</sc> summarized in <xref ref-type="table" rid="T1">Table 1</xref>, i.e. fitted to the average duration of the homophone word types, following Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>). Y-axis scales are fixed across panels for the partial effects on the mean, to facilitate comparison of effect sizes. Rugs indicate the unique values of predictors. In the contour plot (bottom center panel), word types, i.e. points whose coordinates reflect target lemma frequency (on the x-axis) and relative frequency (on the y-axis), are plotted as black dots.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g1.png"/>
</fig>
<p>Further scrutiny of the model revealed that the deviance residuals of the model were approximately Gaussian and did not co-vary with lemma frequency or with the frequency ratio (all &#124;z&#124; &lt; 1, all p values &gt; .7), suggesting that, for all other predictors, the assumptions of normality and constant variance were met to a satisfactory degree. The two variables that did not give rise to significant effects were either not well established as predictors of spoken word duration (<sc>orthographic</sc>&#160;<sc>regularity</sc>) or estimated in a manner that was likely to be highly imprecise (<sc>noun</sc>&#160;<sc>bias</sc>). In sum, the type-based model recovered a number of well established effects and was silent on variables that independently appear questionable as predictors.</p>
<p>The type-based model could undoubtedly be improved, but we refrain from doing so here: In the context of the present study, the model simply serves as a point of comparison for type-based vs. token-based models.</p>
</sec>
<sec>
<title>3.2 Token-based model retracing the type-based model</title>
<p>We begin our discussion of token-based models with a model that is similar to the type-based model in the choice of lexical predictors, but adds token-level information. Similar to the method usually adopted in token-based analyses using Linear Mixed Effects Regression, this first token-based model includes a by-word-form random intercept. The remaining variables are identical to those in the type-based model, with the following exceptions: First, we did not include the P<sc>ause</sc>&#160;<sc>quotient</sc>, i.e. the proportion of tokens of each type that were followed by a pause; secondly, we included three variables coding token-specific information: the talker&#8217;s sex and age, and the target&#8217;s contextual probability, given the word immediately following the target word in the utterance. The summary of the model is shown in <xref ref-type="table" rid="T2">Table 2</xref>, and visualizations of the regression smooths appear in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<table-wrap id="T2">
<caption>
<p><bold>Table 2:</bold> GAMM fitted to the homophone tokens (n = 56,024), with a by-word-form random effect, but otherwise similar to the predictors in the type-based model (see text). Bigram prob. = Word-based bigram probability of the target, given the following word; PND = Phonological Neighborhood Density; Orthogr. = Orthographic regularity; Resid. baseline = Residualized baseline duration; Freq = target frequency; te(Freq., Freq. Ratio) = tensor product of target frequency and ratio of target and homophone frequency; pron = word form; s.1(Freq) = smooth of variance in the outcome, as predicted by frequency. AIC = 35,075.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.4137</td>
<td align="left" valign="top">0.0216</td>
<td align="left" valign="top">&#8211;65.5462</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0608</td>
<td align="left" valign="top">0.0028</td>
<td align="left" valign="top">&#8211;21.8418</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.1152</td>
<td align="left" valign="top">0.0089</td>
<td align="left" valign="top">12.9958</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">(Intercept).1</td>
<td align="left" valign="top">&#8211;1.1408</td>
<td align="left" valign="top">0.0031</td>
<td align="left" valign="top">&#8211;369.9333</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.9179</td>
<td align="left" valign="top">2.9945</td>
<td align="left" valign="top">1191.9521</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">2.3914</td>
<td align="left" valign="top">2.4514</td>
<td align="left" valign="top">75.7883</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">2.7376</td>
<td align="left" valign="top">2.8591</td>
<td align="left" valign="top">6.8679</td>
<td align="left" valign="top">.0512</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">2.4779</td>
<td align="left" valign="top">2.5612</td>
<td align="left" valign="top">75.3640</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">14.5066</td>
<td align="left" valign="top">15.6688</td>
<td align="left" valign="top">150.6620</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9515</td>
<td align="left" valign="top">2.9982</td>
<td align="left" valign="top">255.2859</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(pron)</td>
<td align="left" valign="top">154.2022</td>
<td align="left" valign="top">206.0000</td>
<td align="left" valign="top">4280.8009</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s.1(Freq)</td>
<td align="left" valign="top">3.4536</td>
<td align="left" valign="top">3.7788</td>
<td align="left" valign="top">1504.2063</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2">
<caption>
<p><bold>Figure 2:</bold> Partial effects according to the Gaussian Location-Scale GAMM fitted to spoken word duration of tokens summarized in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g2.png"/>
</fig>
<p>All of the predictors specific to token-based models reached significance in the expected direction, given prior corpus-based models not restricted to homophones (e.g. <xref ref-type="bibr" rid="B9">Bell et al., 2009</xref>; <xref ref-type="bibr" rid="B30">Horton et al., 2010</xref>; <xref ref-type="bibr" rid="B70">Yuan et al., 2006</xref>), as follows: Increasing word bigram probability was associated with shorter duration. Tokens produced by male talkers tended to be shorter than those produced by female talkers, other things being equal. Increasing talker age was associated with increasing token duration up to about age 30 and above age 55; there was no clearly discernible effect of talker age in the middle range of talker age.</p>
<p>With respect to the predictors that also appeared in the type-based models, the token-based model behaved similarly to the type-based model in some respects: The overall direction of the effect of PND was negative, and the effect was steepest in the lower ranges of that variable. Predicted duration increased with residualized baseline duration, as expected, although the shape of this effect was nonlinear, unlike in the type-based model. The effect of orthographic regularity was associated with a p-value exceeding our preset alpha level, as in the type-based model. Unlike in the type-based model, the effect of noun bias reached significance, such that target forms more likely representing nouns than some other part of speech had longer predicted duration. That pattern is consistent with the idea that nouns are more often phrase-final in English than other parts of speech, i.e. the rationale for including this variable in the original (type-based) analysis of the homophone data (<xref ref-type="bibr" rid="B20">Gahl, 2008</xref>).</p>
<p>The token-based model also recovered the overall association of increasing frequency with shorter predicted duration, but its predictions about the interaction of target frequency with frequency ratio were more complex than in the type-based model. In the type-based model, increasing lemma frequency was associated with decreasing predicted duration across almost the entire range of the frequency ratio. In the token-based model, the effect of lemma frequency was nearly absent for words with relative frequencies of about 2, reasserting itself for words in the highest range of the frequency ratio variable. There was also some evidence for an effect of frequency ratio, such that low-frequency targets had shorter predicted durations if they had very low frequency ratios, i.e. high frequency homophone twins. To the extent that this pattern turns out to be robust, it may reflect an effect known as <italic>frequency inheritance</italic>: Low-frequency homophones of high-frequency words have sometimes been found to behave in some respects like high-frequency words, as if &#8220;inheriting&#8221; the frequency of their high-frequency twins. Such inheritance effects, which have been found in data based on speech errors (<xref ref-type="bibr" rid="B16">Dell, 1990</xref>) and latencies in semantic decisions and translation (<xref ref-type="bibr" rid="B31">Jescheniak &amp; Levelt, 1994</xref>; <xref ref-type="bibr" rid="B32">Jescheniak et al., 2003</xref>), may coexist with lemma-specific effects of frequency on duration (cf. <xref ref-type="bibr" rid="B20">Gahl, 2008</xref>, for discussion).</p>
<p>A striking prediction of the token-based model concerns the shape of the predicted variance in duration as a function of target frequency. Recall that in the type-based model, the predicted variance decreased as target frequency increased. As seen in <xref ref-type="fig" rid="F2">Figure 2</xref>, the prediction seemingly goes in the opposite direction for the token-based model. The predicted variance in duration for a given target frequency increases with target frequency. That pattern, of higher variability of high-frequency targets, is one that we take up in Section 4. The confidence region around the predicted variance is wide for low values of frequency and narrows with increasing frequency, as one would expect, given that uncertainty about an estimate (here, the estimated variance) decreases as sample size increases.</p>
<p>The distribution of the by-word-form random effect suggests departures from normality near the extremes of the by-word-form values. Model diagnostics further revealed that the number of basis functions was overly constrained in the case of the smooth terms for bigram probability, suggesting that the shape of the smooths, as depicted in <xref ref-type="fig" rid="F2">Figure 2</xref>, is underfitting the data; further exploration revealed that this issue persisted when the number of basis functions was increased to as high as 19. We return to this issue in 3.8.1 below.</p>
<p>Setting aside the overly constrained estimate for the effect of bigram probability, the token-based model recovers all of the expected effects of lexical factors on word duration, and it additionally reflects plausible effects of contextual probability, talker sex, and age. The by-word-form random effect did not suggest systematic departures from normality. In sum, the token-based model looks successful at first glance.</p>
</sec>
<sec>
<title>3.3 A token-based model including additional predictors</title>
<p>The data set for the token-based model may well support additional predictors &#8211; and omitting such predictors may distort some of the predictions of the token-based model. Anticipating that objection, we therefore fitted a token-based model including two predictors that did not reach significance in the type-based model, viz. morphological complexity and biphone probability. Model summary and partial effects plots for that model are in Appendix B. Neither morphological complexity nor biphone probability reached significance based on our preset alpha level. The effect of orthographic regularity continued to be non-significant. The shape of the relationship of the remaining terms was similar to the model without the additional terms, and the predicted variance as a function of target frequency once again increased with increasing frequency. We take these results to alleviate concerns one might have about the omission of these variables distorting the initial token-based model (<xref ref-type="fig" rid="F2">Figure 2</xref>).</p>
</sec>
<sec>
<title>3.4 The troubles begin: Concurvity</title>
<p>We have seen two token-based models that appear to be successful at capturing both type-level and token-level information. However, a problem with these models becomes apparent when we consider model concurvity. As noted above, concurvity is a measure of the degree to which each component of the model can be approximated by the other components. The consequences of high concurvity are analogous to those of high collinearity in linear models: Both interfere with model interpretability, by making it impossible to identify each predictor&#8217;s relationship to the outcome, and by rendering individual estimates uninterpretable, i.e. liable to change substantially in response even to minor changes in modeling procedure.</p>
<p>Concurvity values for the type-based model (shown in <xref ref-type="table" rid="T3">Table 3</xref>) indicate high concurvity for the parametric part of the model, which suggests that combinations of the factorial predictors (i.e. talker sex and/or noun bias) are not independent of the smooths, as well as for the predicted variance as a function of frequency. The remaining lexical variables in the type-based model do not give rise to high concurvity.</p>
<table-wrap id="T3">
<caption>
<p><bold>Table 3:</bold> Concurvity of the type-based GAM in <xref ref-type="table" rid="T1">Table 1</xref>. Parm. = Parametric terms; PND = Phonological Neighborhood Density; Ortho = Orthographic regularity; Baseline = Residualized baseline duration; Freq. tensor = Tensor product of target frequency and frequency ratio; Variance = variance predicted from frequency.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Parm</bold>.</td>
<td align="left" valign="top"><bold>Pause quotient</bold></td>
<td align="left" valign="top"><bold>PND</bold></td>
<td align="left" valign="top"><bold>Ortho</bold></td>
<td align="left" valign="top"><bold>Freq. tensor</bold></td>
<td align="left" valign="top"><bold>Baseline</bold></td>
<td align="left" valign="top"><bold>Variance</bold></td>
</tr>
<tr>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.15</td>
<td align="left" valign="top">.31</td>
<td align="left" valign="top">.17</td>
<td align="left" valign="top">.42</td>
<td align="left" valign="top">.21</td>
<td align="left" valign="top">.97</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Matters are different for the two token-based models discussed so far: As shown in <xref ref-type="table" rid="T4">Table 4</xref>, most of the terms representing type-level variables in the token-level models are highly predictable from other smooth terms in the models. The concurvity levels for PND, orthographic regularity, baseline duration, target frequency, and (in the model containing this term) biphone probability are all near .89 or even higher, rendering these terms uninterpretable. This is in stark contrast to variables whose values can vary within lexical types, i.e. local context (word-based bigram probability, for which concurvity values range from .28 to .46) and talker age (concurvity between .01 and 02, suggesting that talker age is independent of lexical variables in our sample).</p>
<table-wrap id="T4">
<caption>
<p><bold>Table 4:</bold> Concurvity of token-based models of word token duration: Mirroring type-based = 2; With biphone probability = <xref ref-type="table" rid="TB1">Table B1</xref> (appendix); No RE = <xref ref-type="table" rid="TC1">Table C1</xref> (appendix) RE, equal variance = <xref ref-type="table" rid="TD1">Table D1</xref> (appendix); No RE, equal variance = <xref ref-type="table" rid="TD1">Table D2</xref> (appendix); No frequency = <xref ref-type="table" rid="TD1">Table D1</xref> (appendix); Parm. = Parametric terms; Bigram = Target probability, given the previous word; PND = Phonological Neighborhood Density; Ortho = Orthographic regularity; Base = Residualized baseline duration; Freq. = Tensor product of target frequency and frequency ratio; Age = talker age; RE = by-word-form random effect; Var = variance predicted from frequency; Biphone = positional biphone probability.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Model</bold></td>
<td align="left" valign="top"><bold>Parm</bold>.</td>
<td align="left" valign="top"><bold>Bigram</bold></td>
<td align="left" valign="top"><bold>PND</bold></td>
<td align="left" valign="top"><bold>Ortho</bold></td>
<td align="left" valign="top"><bold>Base</bold></td>
<td align="left" valign="top"><bold>Freq</bold></td>
<td align="left" valign="top"><bold>Age</bold></td>
<td align="left" valign="top"><bold>RE</bold></td>
<td align="left" valign="top"><bold>Var</bold></td>
<td align="left" valign="top"><bold>Biphone</bold></td>
</tr>
<tr>
<td align="left" valign="top">Mirroring type-based (<xref ref-type="table" rid="T2">Table 2</xref>)</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.44</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.97</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.97</td>
<td align="left" valign="top">.02</td>
<td align="left" valign="top">.53</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">n/a</td>
</tr>
<tr>
<td align="left" valign="top">With Biphone (<xref ref-type="table" rid="TB1">Table B1</xref>)</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.46</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.98</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.97</td>
<td align="left" valign="top">.02</td>
<td align="left" valign="top">.57</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">1</td>
</tr>
<tr>
<td align="left" valign="top">No RE (<xref ref-type="table" rid="TC1">Table C1</xref>))</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.30</td>
<td align="left" valign="top">.70</td>
<td align="left" valign="top">.38</td>
<td align="left" valign="top">.52</td>
<td align="left" valign="top">.48</td>
<td align="left" valign="top">.01</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">n/a</td>
</tr>
<tr>
<td align="left" valign="top">RE, equal variance (<xref ref-type="table" rid="TD1">Table D1</xref>)</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.42</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.97</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.89</td>
<td align="left" valign="top">.02</td>
<td align="left" valign="top">.51</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">n/a</td>
</tr>
<tr>
<td align="left" valign="top">No RE, equal variance (<xref ref-type="table" rid="TD2">Table D2</xref>)</td>
<td align="left" valign="top">.70</td>
<td align="left" valign="top">.28</td>
<td align="left" valign="top">.64</td>
<td align="left" valign="top">.33</td>
<td align="left" valign="top">.45</td>
<td align="left" valign="top">.30</td>
<td align="left" valign="top">.01</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">n/a</td>
</tr>
<tr>
<td align="left" valign="top">No frequency (<xref ref-type="table" rid="TE1">Table E1</xref>)</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.38</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.97</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">.02</td>
<td align="left" valign="top">.37</td>
<td align="left" valign="top">n/a</td>
<td align="left" valign="top">n/a</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="T4">Table 4</xref> further suggests a pattern in the interplay of the by-word-form random effect with other type-level variables: When the random effect is included, PND, orthographic regularity, baseline duration, and target frequency all give rise to high concurvity. When it is not, concurvity levels for these variables go down considerably. The interplay of the by-word-form random effect with target frequency is also worth noting. Frequency is highly predictable (concurvity &gt; .89 or higher) in all models containing the by-word-form random effect, and much less so (concurvity &lt; .5) in models without that term. Conversely, the random effect is <italic>least</italic> predictable (&lt;.4 <italic>vs. &gt;.5</italic>) from other terms when frequency is not a predictor in the model. The by-word-form random intercept, where present, gives rise to concurvity values ranging from .37 to .57, i.e. noticeably higher than those for bigram probability and talker age, but far lower than those for variables pertaining to phonological form (PND and residualized baseline duration). The intermediate degree of concurvity of the random intercept is plausible, in light of the fact that each word form (e.g. /ta&#1110;m/) in the data set receives a single posterior prediction shared by two lexical items (e.g. <italic>time</italic> and <italic>thyme</italic>) with different characteristics.</p>
<p>To understand whether the high model concurvity might be pinned down to specific pairs of predictors, we also inspected pairwise concurvity values. The pairwise concurvity estimates for the model in <xref ref-type="table" rid="T2">Table 2</xref> are shown in <xref ref-type="table" rid="T5">Table 5</xref>. These estimates suggest that the by-word-form random effect is by far the worst offender, giving rise to concurvity values at or near 1 nearly across the board; the only variables that are not fully predictable from the random effect are the two token-specific variables (bigram probability of the target word, given the next word, and talker age). The redundancy of the by-word-form random effect, and its apparent ability to render all other lexical variables in the model redundant, suggest that including that term may be a mistake. A logical response to that problem might to remove that term from the model. That is what we do next.</p>
<table-wrap id="T5">
<caption>
<p><bold>Table 5:</bold> Pairwise concurvity of predictors of the token-based model in <xref ref-type="table" rid="T2">Table 2</xref>. Parm. = Parametric terms; Bigram = Target probability, given the previous word; PND = Phonological Neighborhood Density; Ortho = Orthographic regularity; Baseln = Residualized baseline duration; Freq. = Tensor product of target frequency and frequency ratio; Age = talker age; RE = by-word-form random effect; Variance = variance predicted from frequency.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Variable</bold></td>
<td align="left" valign="top"><bold>Parm</bold>.</td>
<td align="left" valign="top"><bold>Bigram</bold></td>
<td align="left" valign="top"><bold>PND</bold></td>
<td align="left" valign="top"><bold>Ortho</bold></td>
<td align="left" valign="top"><bold>Base</bold></td>
<td align="left" valign="top"><bold>Freq</bold>.</td>
<td align="left" valign="top"><bold>Age</bold></td>
<td align="left" valign="top"><bold>RE</bold></td>
<td align="left" valign="top"><bold>Variance</bold></td>
</tr>
<tr>
<td align="left" valign="top">Parm.</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.05</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">Bigram</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.05</td>
<td align="left" valign="top">.03</td>
<td align="left" valign="top">.01</td>
<td align="left" valign="top">.04</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.03</td>
<td align="left" valign="top">.30</td>
</tr>
<tr>
<td align="left" valign="top">PND</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.03</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.09</td>
<td align="left" valign="top">.06</td>
<td align="left" valign="top">.10</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.12</td>
<td align="left" valign="top">.09</td>
</tr>
<tr>
<td align="left" valign="top">Ortho</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.03</td>
<td align="left" valign="top">.18</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.05</td>
<td align="left" valign="top">.06</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.05</td>
<td align="left" valign="top">.11</td>
</tr>
<tr>
<td align="left" valign="top">Base</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.04</td>
<td align="left" valign="top">.20</td>
<td align="left" valign="top">.15</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.09</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.12</td>
<td align="left" valign="top">.01</td>
</tr>
<tr>
<td align="left" valign="top">Freq.</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.22</td>
<td align="left" valign="top">.56</td>
<td align="left" valign="top">.23</td>
<td align="left" valign="top">.40</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.25</td>
<td align="left" valign="top">1</td>
</tr>
<tr>
<td align="left" valign="top">Age</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">RE</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.35</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.96</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.86</td>
<td align="left" valign="top">.01</td>
<td align="left" valign="top">1</td>
<td align="left" valign="top">.91</td>
</tr>
<tr>
<td align="left" valign="top">Variance</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.15</td>
<td align="left" valign="top">.07</td>
<td align="left" valign="top">.13</td>
<td align="left" valign="top">.01</td>
<td align="left" valign="top">.28</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">.06</td>
<td align="left" valign="top">1</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.5 Token-based model, without the by-word-form random effect</title>
<p>A token-based model without the by-word-form random effect, but otherwise identical to our initial token-based model 2, is summarized in <xref ref-type="table" rid="TC1">Table C.1</xref> in Appendix C. The smooth terms are visualized in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3">
<caption>
<p><bold>Figure 3:</bold> Partial effects according to the Gaussian Location-Scale GAM without a by-word-form random effect, shown in <xref ref-type="table" rid="TC1">Table C.1</xref> in Appendix C.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g3.png"/>
</fig>
<p>The model estimates resemble those of the models with the random effect in some ways: Tokens produced by male talkers tend to be shorter than those produced by female talkers. The parametric effect of noun bias was once again significant (unlike in the type-based model, but like in the token-based models with the random effect), in the same direction as before. The smooth term estimates were also similar to the model with the by-word-form random effects in several respects: Increasing word bigram probability and <sc>pnd</sc> were associated with shorter duration. Increasing talker age was associated with increasing token duration in the lower age range (up to about age 30) and, with greater uncertainty, above age 55. Another point of similarity to the model with the random effect is the increasing predicted variance as a function of target frequency: The tensor product of target lemma frequency and relative frequency yielded a prediction surface that was, overall, similar to that of the models with the random effect.</p>
<p>Removing the random effect also resulted in some differences, however: As one might expect, there was a noticeable narrowing of the confidence regions for the estimates of the lexical variables, i.e. <sc>pnd</sc>, orthographic regularity (which reached significance in this model), and residual baseline duration. Increasing residual baseline duration, which had shown a non-linear effect (flattening in the upper ranges of the predictor) in the token-based models with the random effect, was associated with a nearly linear effect in the model without the random effect, resembling that variable&#8217;s behavior in the type-based model. Removing the random effect resulted in a higher AIC (38,888.81, compared to 35,075.65 in the corresponding model with the random effect). By that criterion, the model with the random effect is preferable. As an aside, we note that the substantial change in the AIC suggests that the lexical variables, although considered well-established as predictors of lexical processing, are far from perfect as predictors of word duration in conversational speech. We do not consider this outcome to be cause for concern. The notion that the lexical variables are significant predictors at all is hardly threatened by the observation that the predictors are imperfect.</p>
<p>The rationale for removing the random effect was the high concurvity seemingly associated with that predictor. Did removing the random effect solve that problem? <xref ref-type="table" rid="T4">Table 4</xref> suggests that the answer may be yes: The lexical predictors (PND, orthographic regularity, baseline duration, and frequency) were associated with much lower concurvity in the model without the random effect.</p>
<p>The only model component that continued to give rise to high concurvity after removing the random effect was the term predicting the variance as a function of target frequency. We therefore remove the variance estimate next, to give a random-effects model the chance to incur minimal redundancy. It is conceivable, after all, that the redundancy of the random effect was exacerbated by estimating the variance in duration as a function of frequency along with the mean. Indeed, the concurvity associated with the term predicting the variance may result from the fact that frequency appears to be predictive of duration mean and variance. This might explain why mean and variance vary in similar ways. Therefore, although the concurvity of the variance may not be the main cause for concern, a random effects model assuming equal variance might conceivably reduce concurvity to acceptable levels.</p>
</sec>
<sec>
<title>3.6 Token-based model assuming equal variance</title>
<p>To understand the interplay of the by-word-form random effect with estimating the variance in duration as a function of frequency, we compare models assuming equal variance with and without the random effects term. Full model summaries and partial effects plots for all predictors in the models can be found in <xref ref-type="table" rid="TD1">Tables D.1</xref> and <xref ref-type="table" rid="TD1">D.2</xref> and <xref ref-type="fig" rid="FD1">Figures D.1</xref> and <xref ref-type="fig" rid="FD2">D.2</xref> in Appendix D. Here, we focus on the patterns of concurvity, and on some salient differences in the smooth terms.</p>
<p><xref ref-type="table" rid="T4">Table 4</xref> shows the effect on model concurvity of including the random effect in the two models assuming equal variance: The model including the random effect incurs high concurvity for all lexical variables; as before, the only variables that are unaffected by the problem are those capturing token-level properties, i.e. word-based bigram probability of the target, and talker age. The model without the random effect shows much more acceptable levels of concurvity. In fact, in that model, even the parametric terms are not fully predictable from the rest of the model, unlike in all other models considered so far. On the minus side, not modeling the variance resulted in a higher AIC (36,467.36, up from 35,075 for the otherwise identical model in <xref ref-type="table" rid="T2">Table 2</xref>), suggesting that modeling the variance is well justified from the point of view of balancing parsimony and goodness-of-fit.</p>
<p>How does dropping the assumption of equal variance affect the predicted relationship between the various predictors and duration? <xref ref-type="fig" rid="F4">Figure 4</xref> shows the model predictions for three variables for the two models without the variance term. For the sake of comparison, the top row of the figure repeats the corresponding partial effects plots of the initial token-based model, i.e. the model including the by-word-form random effect and modeling the variance (<xref ref-type="fig" rid="F2">Figure 2</xref>).</p>
<fig id="F4">
<caption>
<p><bold>Figure 4:</bold> Partial effects of residualized baseline duration, PND, and a tensor product of target frequency and frequency ratio, according to three token-based models. Top row: A model with a by-word-form random effect and modeling variance as a function of frequency (cf. <xref ref-type="table" rid="T2">Table 2</xref>). Middle row: A model without a by-word-form random effect, also modeling variance in duration (cf. <xref ref-type="table" rid="TC1">Table C.1</xref> in Appendix C); Bottom row: A model without a by-word-form random effect, without modeling variance in duration (cf. <xref ref-type="table" rid="TD2">Table D.2</xref> in Appendix D).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g4.png"/>
</fig>
<p>As <xref ref-type="fig" rid="F4">Figure 4</xref> shows, the models differ in the degree of uncertainty around those estimates, as well as in the shapes of the estimated relationships, as follows: The confidence regions around the estimates for the smooth terms are noticeably wider for the two models that include by-word-form random effects, compared to the model without that term (the bottom row in the figure). This is to be expected, given that dropping the random effect leaves it to the remaining terms to capture more of the differences across target words.</p>
<p>As for the shape of the relationship between predictors and predicted token duration, the effect of residualized baseline duration is non-linear in the two models with the random effect, increasing in the lower ranges of residual baseline duration and then flattening or even changing direction (though with high uncertainty) in the upper ranges. The relationship is more closely linear in the model without the random effect and without modeling the variance &#8211; resembling the linear relationship between residual baseline duration and predicted duration in the type-based model (<xref ref-type="fig" rid="F1">Figure 1</xref>). The predicted effect of PND likewise changes shape: That effect is nearly linear in the model with the random effect assuming equal variance (the middle row of <xref ref-type="fig" rid="F4">Figure 4</xref>), unlike in either of the other token-based models, and unlike in the type-based model. Finally, the shape of the prediction surface for the tensor product is similarly enigmatic for the models including the random effect, regardless of the inclusion of the variance term. In the model without the random effect and assuming equal variance, the prediction surface shows twin summits in the upper left quadrant (i.e. the highest values of frequency ratio in the lower ranges of target frequency) towering over the expanse of the remaining surface, which is almost entirely flat.</p>
<p>In sum, the high concurvity we observed for the models with the by-word-form random effect cannot be blamed on modeling the variance as a function of target frequency, as concurvity is high for both models that include the random effect. Moreover, assuming equal variance hurts the model&#8217;s AIC, i.e. its goodness-of-fit balanced against parsimony.</p>
</sec>
<sec>
<title>3.7 A token-based model without frequency as a predictor</title>
<p>There is another possible culprit for the problems associated with the random effect, and that is the decision to include target frequency as a predictor (in the tensor product of target frequency and frequency ratio). In including frequency as a predictor, we followed common practice in corpus-based research, including some of our own. But including frequency as a predictor in token-based models might well be problematic, given that the token count itself reflects lexical frequency (albeit frequency in the corpus under analysis, rather than the CELEX frequency estimates used here). To give the random effects model another chance to account for grouping structure in the data without compromising the independence of the lexical predictors, we remove frequency from the set of predictors in the model.</p>
<p>The model summary and partial effects plots can be found in <xref ref-type="table" rid="TD1">Table D.1</xref> in Appendix D. The pattern of significant effects is essentially unchanged in the model without frequency vs. the earlier models. The visualization of the smooths in <xref ref-type="fig" rid="FD1">Figure D.1</xref> suggests that the model without frequency most closely resembles the initial token-based model, i.e. the model including the by-word-form random effect. Both PND and residual baseline duration give rise to non-linear effects (generally decreasing with increasing PND, and generally increasing with increasing residual baseline duration). As before, the distribution of the by-word-form random effect suggests some departures near the extremes.</p>
<p>Does removing frequency as a predictor solve the problem it was intended to solve? <xref ref-type="table" rid="T4">Table 4</xref> suggests that the answer is no: As in each previous model that included a by-word-form random effect, all lexical predictors (PND, orthographic regularity, and residual baseline duration) are nearly perfectly predictable from the model as a whole; the only non-redundant predictors are the token-level properties, i.e. target bigram probability and talker age. What this suggests is that, while including frequency as a predictor may be problematic, removing it does not solve the problem of high concurvity.</p>
</sec>
<sec>
<title>3.8 Model criticism</title>
<p>So far, we have said very little about whether the various models we considered balanced accuracy and parsimony acceptably well or satisfied the modeling assumptions. These are the questions to which we now turn.</p>
<sec>
<title>3.8.1 Checking the number of basis functions</title>
<p>As mentioned above, we limited the number of basis functions for the smooth terms in our token-based models, in an effort to prevent overfitting. In our initial token-based model, we commented that the estimate of the effect of bigram probability, i.e. a token-level property, was likely over-smoothed. For the token-based model without the by-word-form random effect (cf. <xref ref-type="table" rid="TC1">Table C.1</xref>), <monospace>k.check</monospace> suggested that the number of basis functions was likely too low for all of the predictors, with the possible exception of talker age (see <xref ref-type="table" rid="T6">Table 6</xref>). Given this pattern, we increased k to 20 for all smooth terms. For the resulting model, <monospace>k.check</monospace> suggested that k was acceptably high for all predictors, with the exception of bigram probability, for which k = 20 continued to be likely too low. We return to this issue in 4.1.1 below.</p>
<table-wrap id="T6">
<caption>
<p><bold>Table 6:</bold> Checking the number of basis functions for the models with (left-hand columns) and without (right-hand columns) a by-word-form random intercept. RE = Random effect; s(x) = smooth of x; s.1(Freq) = smooth of the predicted variance, given target frequency.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top" colspan="4"><bold>Include RE, cf. <xref ref-type="table" rid="T2">Table 2</xref></bold></td>
<td align="left" valign="top" colspan="4"><bold>Exclude RE, cf. <xref ref-type="table" rid="TC1">Table C1</xref></bold></td>
</tr>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="left" valign="top"><bold>k</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>k.index</bold></td>
<td align="left" valign="top"><bold>p.value</bold></td>
<td align="left" valign="top"><bold>k</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>k.index</bold></td>
<td align="left" valign="top"><bold>p.value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.92</td>
<td align="left" valign="top">0.90</td>
<td align="left" valign="top">.00</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.90</td>
<td align="left" valign="top">0.86</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.39</td>
<td align="left" valign="top">1.00</td>
<td align="left" valign="top">0.41</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.98</td>
<td align="left" valign="top">0.96</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.74</td>
<td align="left" valign="top">1.02</td>
<td align="left" valign="top">0.83</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.13</td>
<td align="left" valign="top">0.95</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.48</td>
<td align="left" valign="top">1.02</td>
<td align="left" valign="top">0.91</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.96</td>
<td align="left" valign="top">0.93</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">24.00</td>
<td align="left" valign="top">14.51</td>
<td align="left" valign="top">1.00</td>
<td align="left" valign="top">0.46</td>
<td align="left" valign="top">24.00</td>
<td align="left" valign="top">18.38</td>
<td align="left" valign="top">0.93</td>
<td align="left" valign="top">.00</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.95</td>
<td align="left" valign="top">1.00</td>
<td align="left" valign="top">0.37</td>
<td align="left" valign="top">3.00</td>
<td align="left" valign="top">2.95</td>
<td align="left" valign="top">0.98</td>
<td align="left" valign="top">.05</td>
</tr>
<tr>
<td align="left" valign="top">s(pron)</td>
<td align="left" valign="top">206.00</td>
<td align="left" valign="top">154.20</td>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">s.1(Freq)</td>
<td align="left" valign="top">4.00</td>
<td align="left" valign="top">3.45</td>
<td align="left" valign="top">1.01</td>
<td align="left" valign="top">0.76</td>
<td align="left" valign="top">4.00</td>
<td align="left" valign="top">3.50</td>
<td align="left" valign="top">0.92</td>
<td align="left" valign="top">.00</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.8.2 Normality and homogeneity of residuals</title>
<p>Diagnostic plots of the models (included in the Supplementary Materials) with and without by-word-form random effects suggested departures from normality in the upper ranges. In particular, residuals varied with frequency.</p>
</sec>
<sec>
<title>3.8.3 The random-effects assumption</title>
<p>The random-effects assumption is the assumption that the by-group adjustments (the random effects) are uncorrelated with the fixed effects predictors. As shown in <xref ref-type="table" rid="T7">Table 7</xref>, that assumption is violated for the token-based model in <xref ref-type="table" rid="T2">Table 2</xref>, such that all of the lexical predictors are predictive of the random intercepts. The only fixed effect whose association with the random intercept is non-significant at our alpha level is talker age.</p>
<table-wrap id="T7">
<caption>
<p><bold>Table 7:</bold> A model of by-word-form random intercepts in the token-based model shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">0.0134</td>
<td align="left" valign="top">0.0003</td>
<td align="left" valign="top">45.2863</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">8.7738</td>
<td align="left" valign="top">8.9838</td>
<td align="left" valign="top">67.3250</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">8.9947</td>
<td align="left" valign="top">9.0000</td>
<td align="left" valign="top">871.9884</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">8.9867</td>
<td align="left" valign="top">8.9999</td>
<td align="left" valign="top">161.1192</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">8.9962</td>
<td align="left" valign="top">9.0000</td>
<td align="left" valign="top">4536.1619</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">23.9901</td>
<td align="left" valign="top">23.9999</td>
<td align="left" valign="top">2871.6014</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">3.4431</td>
<td align="left" valign="top">4.3121</td>
<td align="left" valign="top">2.1522</td>
<td align="left" valign="top">.0651</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec>
<title>4. Discussion</title>
<p>The accessibility of corpora of naturalistic data and flexible statistical tools have fueled enormous interest in modeling linguistic experience at the token level, i.e. the level of individual utterances. Such models promise major advances over data sets of elicited speech aggregated over many utterances. &#8220;Lab-grown&#8221; data sets restrict not only the ecological validity of conclusions, but also the kinds of questions that can be asked. The advantages of naturalistic data notwithstanding, we are sounding a cautionary note about token-based models.</p>
<p>We began this study by laying out four potential problems with token-based models: Ballot-box stuffing, model redundancy, distorted parameter estimates, and, most concerningly, assumption violations rendering model estimates valid only in a counterfactual world in which the assumptions are met. We now consider the scope of these problems. Are the issues we observed specific to the models we discussed, or are they inherent in some general property of imbalanced speech corpora, or indeed linguistic experience itself?</p>
<p>Assumption violations have the potential to render moot any advantages or disadvantages of token-based models: Estimates of counterfactual worlds hardly suit the kind of inquiry that makes naturalistic data attractive in the first place. We therefore ask whether we can trace the violations to any systematic property of our data. To preview our argument: There are two properties of the data that may be to blame. The first is the imbalanced nature of the sample, i.e. the fact that there are many more observations for high-frequency words than low-frequency ones. We point out some consequences of that imbalance, echoing similar observations in previous research on models with random effects. We argue that the problems extend to models without random effects. The second property of the data that may be to blame for the problems we found concerns the nature of high-frequency words. We argue that some of the problems persist in &#8220;balanced&#8221; corpora with equal sample sizes for all items.</p>
<sec>
<title>4.1 Four problems with token-based models</title>
<p>We begin by summarizing how our models fared vis-&#224;-vis the four potential problems we identified in Section 1.</p>
<sec>
<title>4.1.1 Ballot-box stuffing</title>
<p>Do high-frequency items exert undue influence over token-based models? Interestingly, the by-token punishment for residuals did not make residuals systematically smaller for high-frequency words. In fact, as shown in the Supplementary Materials, there is some indication that residuals actually increased with increasing target frequency. The non-independence of target frequency and model residuals points to an assumption violation, a fact we discuss in 4.1.4 below. Here, we focus on the issue we dubbed <italic>ballot-box stuffing</italic>.</p>
<p>In 3.8.1, we noted that some of the estimates in the token-based models appeared to be overly constrained in their choice of basis functions, i.e. the overall degree of wiggliness of the estimated composite function. Here, we follow up on that observation, demonstrating that this weakness of the models reflects the undue influence of high-frequency types.</p>
<p>In our analysis, we set <italic>k</italic> to 4 by default. With this choice of <italic>k</italic>, we allow for a moderate degree of nonlinearity. The resulting smooth for the effect of bigram probability is shown in the upper left panel of <xref ref-type="fig" rid="F5">Figure 5</xref>. As mentioned earlier, the <monospace>gam.check</monospace> function of the <bold>mgcv</bold> package suggests that <italic>k</italic> = 4 is too low. Setting <italic>k</italic> to 20, the partial effect is much more wiggly (upper right panel of <xref ref-type="fig" rid="F5">Figure 5</xref>), as expected. However, <monospace>gam.check</monospace> suggests that <italic>k</italic> = 20 is still too low. Setting <italic>k</italic> to 40 results in an even wigglier smooth (not shown), but does not change <monospace>gam.check</monospace>&#8217;s verdict that even more basis functions should be invested. Further exposing the flaws of the model where <italic>k</italic> = 4, <xref ref-type="fig" rid="F5">Figure 5</xref> (lower left panel) reveals that the residuals of that model are not equal in all ranges of (log) bigram probability, violating a modeling assumption. (Recall that, while GAMs allow the assumption of equal variances to be relaxed, this particular model only drops that assumption for one of the predictors, viz. target frequency.) Setting <italic>k</italic> = 20 removes the assumption violation: the lower right panel of <xref ref-type="fig" rid="F5">Figure 5</xref> suggests equal variance of the residuals for the entire range of (log) bigram probability. So far, everything speaks in favor of setting <italic>k</italic> to a much higher value than 4.</p>
<fig id="F5">
<caption>
<p><bold>Figure 5:</bold> Smooths (upper panels) and residuals (lower panels) for <italic>k</italic> = 4 (left) and <italic>k</italic> = 20 (right).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g5.png"/>
</fig>
<p>However, further scrutiny reveals that the values of (log) bigram probability are not evenly spread along the horizontal axis: a relatively small number of types contribute large numbers of tokens. This is illustrated by the vertical lines in the upper right and lower left panels, which are histograms visualizing the counts of observations along the range of (log) bigram probability. A mere 3.7% of values occurs more than 100 times, and these jointly account for no less than 43.2% of the total number of datapoints. For example, there are 2,388 tokens for which the bigram probability is &#8211;1.97. The uneven distribution of observations creates a problem for the model: Including all tokens where the bigram probability equals &#8211;1.97 means that this particular value has strong empirical support and will strongly influence the shape of the smooth: Indeed, the smooth in the upper right panel (<italic>k</italic> = 20) shows a very pronounced peak just below <italic>x</italic> = &#8211;&#8211;2. That peak is entirely absent when <italic>k</italic> is set to 4 (upper left panel). The result of increasing <italic>k</italic> is a smooth that essentially connects predictions for strongly supported predictor values. In the absence of a theory predicting that high ranges of bigram probability are associated with extremely variable outcomes (here, spoken word duration), the wiggly curve is uninterpretable and uninformative. We also note that the problem is not a problem of overfitting <italic>per se</italic> and is not resolved by setting the <monospace>gamma</monospace> parameter of the <monospace>gam</monospace> function to 1.4, a recommendation given by Wood (<xref ref-type="bibr" rid="B66">2006, p. 231</xref>) as a measure that may alleviate overfitting: Setting the <monospace>gamma</monospace> parameter to 1.4 for the present data leads to a visually identical wiggly smooth (not shown here), while increasing the AIC of the model. The reason is that the issue is not so much a problem of model specification, but rather of the uneven distribution of observations.</p>
<p>We have discussed this particular predictor at some length. What this example illustrates more generally is that high-frequency words have the ability to influence the model by sheer force of numbers, without actually clarifying the effects of lexical predictors.</p>
</sec>
<sec>
<title>4.1.2 Model redundancy</title>
<p>By-item random effects have been the mechanism of choice for many researchers wishing to prevent high-frequency items from being overly influential in token-based models. As we just pointed out, this strategy appears to be unsuccessful. In addition, it appears that the by-word-form random effects exacerbate an additional problem with token-based models. We found that all of our token-based models that included the by-word-form random effect incurred high concurvity (cf. <xref ref-type="table" rid="T4">Table 4</xref>), unlike the type-based point of comparison (<xref ref-type="table" rid="T3">Table 3</xref>).</p>
<p>A comment on the choice of word form (e.g. /ta&#1110;m/ for <italic>time</italic> and <italic>thyme</italic>) as the grouping variable for the random effect is in order here. One might alternatively group tokens by lemma, i.e. identifying <italic>time</italic> and <italic>thyme</italic> as separate types. We opted for by-word-form random effects expressly to minimize a problem that often arises with corpus data, as pointed out in Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>): Many low-frequency lemmas are represented by so few tokens that a lemma-specific random effects structure is bound to over-fit the data. Pooling low-frequency lemmas with higher-frequency homophone twins mitigates this problem somewhat. Relatedly, we opted for a word form-specific (rather than lemma-specific) residualized baseline duration. In Gahl and Baayen (<xref ref-type="bibr" rid="B22">2024</xref>), whose type-based model served as our starting point, the decision to use a word form-based residualized baseline duration measure was driven by the research question. In the current methodological context, what motivated our decision to adopt this measure was the realization that lemma-specific residualization would increase concurvity yet more.</p>
</sec>
<sec>
<title>4.1.3 Distorted parameter estimates</title>
<p>We commented that the token-by-token influence of high-frequency words has the potential to distort model estimates of the effects of variables other than frequency. Our comments on the wiggliness of the smooth of bigram probability in 4.1.1 illustrate one aspect of this problem. One might limit the wiggliness of that estimate by restricting the number of basis functions, but imposing such a constraint will not elucidate the effect of contextual probability, nor will it help researchers capitalize on the richness of token-level information. An additional way in which high-frequency types can distort model estimates arises due to characteristics of high-frequency words: There are similarities among high-frequency types (see e.g. <xref ref-type="bibr" rid="B1">Baayen, 2011</xref>; <xref ref-type="bibr" rid="B19">Frauenfelder et al., 1993</xref>; <xref ref-type="bibr" rid="B35">K&#246;hler, 1986</xref>; <xref ref-type="bibr" rid="B37">Landauer &amp; Streeter, 1973</xref>). For example, while word-based bigram probability can, in principle, vary within type &#8211; for example, a given low-frequency word might be highly probable in a particular context &#8211; in actual fact, the upper ranges of bigram probability are dominated by high-frequency words. More generally, there are clusters of high-frequency word types along ranges of phonological neighborhood density, orthographic regularity, and degree of polysemy (or vagueness), to name just a few. High frequency items, thus, have the potential to distort estimates of the effects of these other lexical variables.</p>
<p>At this point, one might also ask what type of distortions or divergences across models should be considered meaningful. We believe that, ultimately, that question can only be answered in the context of specific hypotheses and research questions. Directionality changes, such as a given variable being associated with increases in the outcome in one model and with decreases in a different model, would almost certainly qualify: Theory-driven analyses often test predictions about the direction of an effect vs. simply checking whether a given variable significantly predicts an outcome, regardless of the direction of the effect. But directionality changes are not the only type of divergence that might matter: Non-linear effects differ in wiggliness as much as in directionality; the degree of wiggliness may well be central to one analysis (e.g. if one is modeling the specific trajectory of the tip of the tongue during the production of a diphthong as produced by different talkers), but irrelevant to the next (e.g. if one is interested in asking whether some variable is predictive of tongue movement at all).</p>
<p>The phenomenon we have dubbed <italic>ballot-box stuffing</italic> has theoretical implications beyond the issue of wiggliness: There is a long-standing debate about the extent to which apparent effects of lexical frequency are due to frequency vs. other, correlated variables (see e.g. <xref ref-type="bibr" rid="B22">Gahl &amp; Baayen, 2024</xref>; <xref ref-type="bibr" rid="B24">Gardner et al., 1987</xref>). Baayen (<xref ref-type="bibr" rid="B1">2011</xref>), for example, argues that, if other variables are allowed to do their work first, there is not much left for frequency to do. Token-based models may effectively prevent other variables from doing their work &#8220;first&#8221;, because the model is obligated to minimize the residuals for the many tokens of high-frequency words.</p>
</sec>
<sec>
<title>4.1.4 Assumption violations</title>
<p>As mentioned in 3.8.2 above, we found that model residuals in our token-based models increased with lexical frequency, violating the homogeneity assumption. In addition, we saw (in 3.8.3) that the by-word-form posterior modes correlated with the fixed effects predictors. We note that the correlations of frequency with other lexical characteristics in turn make the residuals (and the random effects in models containing those) predictable from other variables in the model, as we demonstrate in the Supplementary Materials.</p>
<p>Given our comments about problems traceable to the by-word-form random effect, one may ask why the residuals of the model <italic>without</italic> the random effect are also problematic. One reason may be that the identity of the word form is also known to the models without random effects: The model specification includes the residualized baseline duration, i.e. the variable reflecting the duration one would expect, given the phonemic content, setting aside what one can predict from the other lexical characteristics, such as PND. Although residualized baseline duration is a continuous variable, for the modestly-sized set of target words under discussion, it has as many distinct values as there are individual word forms, effectively identifying word forms. Removing this variable is, of course, possible, but would entail giving up on identifying what homophone pairs have in common, i.e. the natural experiment enabled by homophones. Indeed, removing the residualized baseline duration from the model would be appropriate if one were to take the (highly implausible) view that segmental content &#8211; e.g. the duration of a long vowel vs. a flap &#8211; was immaterial to word duration.</p>
<p>Some of the problems we just described have been noted before. Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>) warn against including a by-word random intercept in a model of data from the Buckeye corpus of conversational American English (<xref ref-type="bibr" rid="B44">Pitt et al., 2007</xref>). As is typical with such data sets, the data are highly imbalanced, containing many more tokens of high-frequency words than low-frequency ones. In fact, in Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>)&#8217;s data, nearly half of the word types occur only once. As Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>) point out, this means that there are too few observations per predictor, which, in turn, leads to high concurvity, and, in fact, total unidentifiability, as any given observation is predicted by multiple variables. Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>) point out a second problem with the random intercepts of their model: The by-word random intercepts can be predicted from lexical variables in their data. As Baayen and Linke (<xref ref-type="bibr" rid="B5">2021</xref>) put it: &#8220;what should be random noise is in fact structured variation.&#8221; This problem, too, is also apparent in our models, as pointed out in 3.8.3.</p>
</sec>
</sec>
<sec>
<title>4.2 Beyond homophones: Token-level models of imbalanced data</title>
<p>We have been utilizing a data set that has seen multiple previous analyses, in order to keep the focus on the methodological points we are making, rather than the properties of particular estimates of (type-level and token-level) variables. It might be asked whether the problems we are pointing out are specific to English homophones &#8211; or, for that matter, to the analysis of spoken word durations or other acoustic properties. We believe that this is not the case. Our points apply in principle to any analysis of lexical variables in unbalanced data sets, especially when the variables involved are correlated with type frequency. Such might be the case for analyses of reading times in running text, or lexical effects in typing speed in naturalistic samples. One example where similar issues arise is Chuang et al. (<xref ref-type="bibr" rid="B13">2026</xref>), who use GAMs to predict f0 contours, in an analysis of Mandarin tones. In that study, the research question and the nature of the data necessitate a token-level analysis: The research question concerns the existence of word-specific pitch contours. In order to address that question, it is, first of all, necessary to control for non-lexical local determinants of pitch, such as neighboring tones, position within an utterance, and syllable duration. Averages at the type level would make it impossible to control for these token-by-token properties. Secondly, Chuang et al. (<xref ref-type="bibr" rid="B13">2026</xref>) link token-specific pitch contours to their corresponding contextualized embeddings, i.e. another token-specific set of (semantic and form-based) properties of tokens in the context of utterances. In order to keep high-frequency words from dominating the resulting models, while ensuring that estimates of lexical (i.e. type-level) pitch signatures are based on satisfying sample sizes, Chuang et al. (<xref ref-type="bibr" rid="B13">2026</xref>) set an upper limit to the number of tokens per word included in the analysis, as well as a lower limit. While this solution is imperfect, for reasons we are about to discuss, it does allow lexically-specific effects to assert themselves. Importantly, the empirical domain Chuang et al. (<xref ref-type="bibr" rid="B13">2026</xref>) concerns neither English, nor homophones, nor word duration, illustrating that the problems to which we are drawing attention extend beyond the particular dataset under discussion in the current article.</p>
</sec>
<sec>
<title>4.3 Are equal samples the answer?</title>
<p>The problems we described are bound to arise whenever per-item samples are very unequal, even in large samples. When items have very different frequencies, the random intercepts may diverge considerably from normality (Douglas Bates, p.c.). Furthermore, when item-bound (here, by-word-form) predictors are of interest, the random intercepts are confounded with the item predictors. When predictors have values that are instantiated across multiple tokens of the same word form, whether a predictor will reach significance in the presence of a by-item random effect will depend on which values of the predictor happen to recur across these tokens. For predictors with non-repeating values, on the other hand, there is a one-to-one relation between predictor values and random intercepts. As a consequence, concurvity values will be near 1 when smoothing splines are used.</p>
<p>Sampling a set number of tokens for any given word type (see, for an example, <xref ref-type="bibr" rid="B45">Pluymaekers et al., 2005</xref>), might seem to be one strategy for addressing some or all of the problems we noted: In this way, one might keep the data from being overwhelmed by high-frequency words and remove the confound of item frequency and by-item sample size. One might take this strategy one step further and sample repeatedly, so as to avoid relying on a single sampling operation. However, repeating the sampling procedure introduces a new problem: Any given token of a high-frequency word will have a smaller chance of being drawn, and tokens of infrequent words will be included with near certainty (or complete certainty, in the case of hapax legomena). The results, over many runs, still reflect type-level properties very unevenly.</p>
<p>In addition, creating equal sample sizes does not remove inherent properties of high-frequency words and the confounds with other lexical properties that they give rise to. We argue that this problem cannot be overcome by using larger or more evenly sampled data sets. These are strong claims, and we suspect that analyses of additional data sets will be needed to explore these claims fully. But we see reason to believe that the properties of high-frequency words systematically affect statistical models of lexical properties, in ways that have not been fully appreciated and that may necessitate complementing token-based analyses with type-based ones. One reason for this is that lexical frequency is correlated with multiple other lexical variables (see e.g. <xref ref-type="bibr" rid="B19">Frauenfelder et al., 1993</xref>). Therefore, if model residuals are correlated with lexical frequency, then they will also be correlated with other lexical variables.</p>
<p>It might be asked why random effects, whose job it is to account for grouping structure in the data, are themselves correlated with item frequency. We think that one reason for this is that high-frequency word forms tend to be massively ambiguous (or vague), i.e. to be associated with many different meanings (and/or senses: the distinction between ambiguity and vagueness unfortunately does not help here). High-frequency word forms effectively represent many different words, as noted, for example, in Baayen et al. (<xref ref-type="bibr" rid="B3">2006</xref>). Along similar lines, Pimentel et al. (<xref ref-type="bibr" rid="B42">2020</xref>), cited in Hermalin (<xref ref-type="bibr" rid="B29">2025</xref>), found that more contextually predictable words tended to have more meanings/senses. As a consequence, form-level estimates for high-frequency words represent attempts to characterize wildly heterogeneous groups, and increasingly so with increasing frequency. If one believes, as we do, that word meanings and word senses matter for phonetic realization, this heterogeneity will make itself felt in the model predictions: The heterogeneity of the meanings of high-frequency forms may be partly to blame for the relationship between frequency and by-item adjustments. In this connection, it is interesting to note that the predicted variance went up as frequency increased (cf. <xref ref-type="table" rid="T2">Table 2</xref>). This might seem unexpected under the usual assumption that uncertainty decreases with increasing sample size. But it is expected, given that high frequency forms are associated with many different meanings.</p>
<p>If this reasoning is correct, then drawing equal numbers of tokens (say, n = 5 per word form) again will not fix the issue: The five tokens of <italic>hypotenuse</italic> will still mean similar things; whereas the five tokens of <italic>corner</italic> will probably mean different things, having been drawn from such different contexts as <italic>in our corner of the world, they backed him into a corner, the store is right around the corner, the referee awarded a corner</italic>, and so on. The heterogeneity of high-frequency forms will still be present, even if one creates equally sized samples. Careful sense disambiguation based on context-specific information about meaning may mitigate this issue. It is possible to use sense-disambiguation algorithms from Natural Language Processing to assign a discrete sense to the individual word tokens. Such a procedure does not, however, address the question of &#8220;where one sense of a word ends and the next begins&#8221; (<xref ref-type="bibr" rid="B34">Kilgarriff, 2006</xref>, p.29). Whether this approach solves the technical issues due to semantic ambiguity and vagueness remains an open empirical question. Alternatively, contextualized embeddings can be calculated for every single token, using Large Language Models. Such contextualized embeddings estimate words&#8217; meanings in context, but the full richness of an individual speaker&#8217;s communicative intentions for a given word token can only be approximated (for a comparison of these methods in a corpus-based study of Mandarin tone, see <xref ref-type="bibr" rid="B13">Chuang et al., 2026</xref>).</p>
</sec>
<sec>
<title>4.4 Recommendations</title>
<p>As stated at the outset, we are emphatically not claiming that token-level models are to be avoided under all circumstances. Setting aside the issue of model assumptions for a moment, what might be considered ballot-box stuffing in one context may amount to desirable weighting of evidence in the next. We believe caution is in order, particularly when the interest of an analysis concerns types, rather than tokens. Token-level models estimate token-level behavior, which can make for a mismatch between the statistical unit of analysis and the inferential target. This invites the question of how to choose between type-level and token-level models, or perhaps how to combine the two.</p>
<p>Fundamentally, the choice of analytical target depends on one&#8217;s theory of the processes that generated the data. Testing predictions about elements of a (type-based) lexicon vs. acoustic, perceptual, or situational constellations not tied to word types calls for different solutions (see e.g. <xref ref-type="bibr" rid="B22">Gahl &amp; Baayen, 2024</xref>, for discussion). Even when the research question ultimately concerns type-level properties, token-level analyses may be a necessary element of the analysis, precisely because the process that generated the data is hypothesized to involve both type-level and token-level forces. For example, as mentioned above, the research question in Chuang et al. (<xref ref-type="bibr" rid="B13">2026</xref>) concerns the existence of word-specific f0 contours in Mandarin tones, a type-level property of words. In that study, the research question and the nature of the data necessitate a token-level analysis: Averages at the type level make it impossible to control for these token-by-token properties. Superficially, the token-level analysis is necessary to control for utterance-level local determinants of pitch. More fundamentally, the authors&#8217; theory of the process that generated the data involves both type-level forces (the hypothesized lexically-specific tone contours) and utterance-level forces (such as local speaking rate and shades of meaning in specific contexts of use).</p>
<p>Again, the choice between analyses based on tokens and analyses based on types requires reflecting on the goal of the analysis. For example, if the goal is to predict word duration as accurately as possible with as few predictors as possible, then token-level models taking into account pre-pausal lengthening and local speech rate may be perfectly satisfactory. But these predictions will be blind to many processes that may, nevertheless, be operative in human language production. Token-level models can also help address certain questions about learning and development, where one may want to embrace the idea that high-frequency words are represented in the data proportionally to their frequency. During learning, word types are encountered proportional to their frequency in the learner&#8217;s experience; as a consequence, the cognitive system receives its fine-tuning predominantly from the higher-frequency words.</p>
<p>We conclude with a set of recommendations. Again, which of these are optimal depends on the specific (empirical and theoretical) goals of any given research project.</p>
<list list-type="bullet">
<list-item><p>If the goal is to predict properties of tokens, then, despite the problems we noted, token-level models may be the tool of choice. But the prediction accuracy of these models comes at the cost of disentangling contributions of individual predictors: Predictors in such models are inevitably collinear (or concurve), because high-frequency types will be represented by many replicates.</p></list-item>
<list-item><p>If the goal is to evaluate models of lexical memory and retrieval of words conceived as types, then type-based models are often called for. Predictions about how many milliseconds will elapse during the production (or comprehension) of tokens is not the goal of such models, but a means to an end.</p></list-item>
<list-item><p>If the goal is to understand utterance-specific factors and their interplay with lexical properties, we recommend fitting both type-based and token-based models and comparing the resulting parameter estimates and predicted values of each type of model. Consider including interactions of frequency with the other predictors in type-based models (depending on the research question) and model the variance of the outcome (<xref ref-type="bibr" rid="B68">Wood, 2017</xref>), keeping in mind that data sparseness near the extremes of the frequency distribution means different things in type-level vs. token-level models: In any naturalistic corpus, there are many low-frequency word types, represented by few tokens. On the other hand, there are few high-frequency types, so there are few observations informing type-based models in the higher ranges of frequency. The consequences of aggregating data over types are difficult to identify and diagnose. Comparing type-based and token-based models can help identify those consequences.</p></list-item>
<list-item><p>As a further check of the relationship between target frequency and other lexical variables, as well as in order to guard against clusters of observations (at the type level or the token level), examine the distribution of observations along the ranges of the predictors.</p></list-item>
</list>
<p>It is beyond the scope of the present study to evaluate these recommendations. Rather, our goal is to point out that the ability to model large sets of observations has led to a proliferation of models that may actually run counter to some of the goals of psycholinguistic research.</p>
<p>Type-based analyses need not forego information about contextual variables altogether. In fact, when used in combination with token-level information, they can shed light on the consequences of the accumulation of utterance-specific properties for lexical processing. For example, Seyfarth (<xref ref-type="bibr" rid="B50">2014</xref>) found that &#8220;Words that usually appear in predictable contexts are reduced in all contexts, even those in which they are unpredictable&#8221; (<xref ref-type="bibr" rid="B50">Seyfarth, 2014, p.140</xref>). Similar cumulative effects, e.g. of the positions a word most often occurs in, on its production, even in other contexts, have also been reported in Brown et al. (<xref ref-type="bibr" rid="B12">2021</xref>) and S&#243;skuthy and Hay (<xref ref-type="bibr" rid="B54">2017</xref>).</p>
<p>Ultimately, the decision whether to treat any variable, including frequency, as a lexical property depends on hypotheses about the processes one is interested in. Whether form frequency should be considered a property of tokens of, say, <italic>thyme</italic> depends on whether one believes that the frequency of the form [ta&#618;m] was relevant to the processing of tokens of <italic>thyme</italic>.</p>
</sec>
</sec>
<sec>
<title>5. Conclusion</title>
<p>The title of this paper alludes to Austrian Emperor Joseph II&#8217;s purported characterization of Mozart&#8217;s <italic>Abduction from the Seraglio</italic> as containing &#8220;too many notes&#8221;. Mozart&#8217;s riposte, according to anecdote, was that there were just as many notes as needed. We hope readers will consider our critical remarks to be more nuanced than the emperor&#8217;s: The &#8220;needed&#8221; number of observations may equal the token count, depending on the goals of the analysis.</p>
<p>Our observations have potentially broad implications. Any model based on individual tokens of items of unequal frequency is, in principle, subject to the issues we pointed out. These issues apply to other measures besides word duration, and to other domains besides spoken word production.</p>
</sec>
</body>
<back>
<sec>
<title>Appendix A GAM of baseline duration</title>
<p>A target word&#8217;s <italic>baseline duration</italic> is a measure of the expected duration of the word, given its segments. The rationale for controlling baseline duration is that phones have different inherent durations; therefore, spoken word duration is predictable in part based on the phones that are present in a word. For example, tokens of [m] tend to be longer than tokens of [&#638;]. Homophones have identical baseline duration, as a consequence of sharing the same phonemes. The baseline duration was the (log-transformed) sum of the average duration of the target&#8217;s phones, based on the Buckeye corpus (<xref ref-type="bibr" rid="B44">Pitt et al., 2007</xref>). The baseline duration is itself partially predictable from lexical variables. The model residuals entered the models discussed in the main text.</p>
<table-wrap id="TA1">
<caption>
<p><bold>Table A.1:</bold> GAM predicting word form baseline duration. lgPronCelFq = (log-transformed) form frequency; PND = Phonological Neighborhood Density; logAvgBiProb = log-transformed biphone probability.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.1832</td>
<td align="left" valign="top">0.0095</td>
<td align="left" valign="top">&#8211;124.3936</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(lgPronCelFq)</td>
<td align="left" valign="top">1.8174</td>
<td align="left" valign="top">1.9659</td>
<td align="left" valign="top">6.9396</td>
<td align="left" valign="top">0.0007</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">4.6696</td>
<td align="left" valign="top">5.7029</td>
<td align="left" valign="top">37.5923</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(logAvgBiProb)</td>
<td align="left" valign="top">1.0000</td>
<td align="left" valign="top">1.0000</td>
<td align="left" valign="top">24.5515</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="FA1">
<caption>
<p><bold>Figure A.1:</bold> Partial effects of GAM predicting baseline word duration, and qq-plot of model residuals.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g6.png"/>
</fig>
</sec>
<sec>
<title>Appendix B Token-based model including additional predictors</title>
<table-wrap id="TB1">
<caption>
<p><bold>Table B.1:</bold> GAM fitted to the homophone tokens, with a by-word-form random effect and additional predictors that were non-significant in the first token-based model. is_complexTRUE = morphologically complex (i.e. not mononmorphemic).</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.4240</td>
<td align="left" valign="top">0.0216</td>
<td align="left" valign="top">&#8211;65.8158</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0608</td>
<td align="left" valign="top">0.0028</td>
<td align="left" valign="top">&#8211;21.8404</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.1231</td>
<td align="left" valign="top">0.0093</td>
<td align="left" valign="top">13.2535</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">is_complexTRUE</td>
<td align="left" valign="top">0.0315</td>
<td align="left" valign="top">0.0117</td>
<td align="left" valign="top">2.6999</td>
<td align="left" valign="top">0.0069</td>
</tr>
<tr>
<td align="left" valign="top">(Intercept).1</td>
<td align="left" valign="top">&#8211;1.1408</td>
<td align="left" valign="top">0.0031</td>
<td align="left" valign="top">&#8211;369.9374</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.9165</td>
<td align="left" valign="top">2.9943</td>
<td align="left" valign="top">1178.0664</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">2.4962</td>
<td align="left" valign="top">2.5516</td>
<td align="left" valign="top">69.7529</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(logAvgBiProb)</td>
<td align="left" valign="top">1.2227</td>
<td align="left" valign="top">1.2438</td>
<td align="left" valign="top">2.3697</td>
<td align="left" valign="top">0.2242</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">2.6471</td>
<td align="left" valign="top">2.7930</td>
<td align="left" valign="top">3.9898</td>
<td align="left" valign="top">0.1739</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">2.4540</td>
<td align="left" valign="top">2.5419</td>
<td align="left" valign="top">63.9407</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">13.8349</td>
<td align="left" valign="top">15.0525</td>
<td align="left" valign="top">135.4672</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9514</td>
<td align="left" valign="top">2.9982</td>
<td align="left" valign="top">255.3254</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(pron)</td>
<td align="left" valign="top">152.9245</td>
<td align="left" valign="top">206.0000</td>
<td align="left" valign="top">3816.2657</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s.1(Freq)</td>
<td align="left" valign="top">3.4569</td>
<td align="left" valign="top">3.7806</td>
<td align="left" valign="top">1503.7435</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="FB1">
<caption>
<p><bold>Figure B.1:</bold> Partial effects according to the Gaussian Location-Scale GAM fitted to spoken word duration of tokens (n = 56,024), summarized in Table B.1.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g7.png"/>
</fig>
</sec>
<sec>
<title>Appendix C Token-based model without a by-word-form random effect</title>
<table-wrap id="TC1">
<caption>
<p><bold>Table C.1:</bold> GAM fitted to the homophone tokens, without a by-word-form random effect. AIC = 38,888.81.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.3619</td>
<td align="left" valign="top">0.0026</td>
<td align="left" valign="top">&#8211;531.2428</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0641</td>
<td align="left" valign="top">0.0029</td>
<td align="left" valign="top">&#8211;22.2226</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.0171</td>
<td align="left" valign="top">0.0040</td>
<td align="left" valign="top">4.2287</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">(Intercept).1</td>
<td align="left" valign="top">&#8211;1.1024</td>
<td align="left" valign="top">0.0031</td>
<td align="left" valign="top">&#8211;358.1427</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.9049</td>
<td align="left" valign="top">2.9934</td>
<td align="left" valign="top">1143.9668</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">2.9830</td>
<td align="left" valign="top">2.9996</td>
<td align="left" valign="top">1155.9271</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">2.1341</td>
<td align="left" valign="top">2.4353</td>
<td align="left" valign="top">25.8636</td>
<td align="left" valign="top">.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">2.9635</td>
<td align="left" valign="top">2.9985</td>
<td align="left" valign="top">6259.2056</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">18.3831</td>
<td align="left" valign="top">19.6160</td>
<td align="left" valign="top">2444.6596</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9549</td>
<td align="left" valign="top">2.9985</td>
<td align="left" valign="top">271.6112</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
<tr>
<td align="left" valign="top">s.1(Freq)</td>
<td align="left" valign="top">3.5012</td>
<td align="left" valign="top">3.8223</td>
<td align="left" valign="top">843.6087</td>
<td align="left" valign="top">&lt;.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Appendix D Token-based models without predicting the variance</title>
<table-wrap id="TD1">
<caption>
<p><bold>Table D.1:</bold> GAM fitted to the homophone tokens, without modeling the variance; AIC = 36,467.36.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.4133</td>
<td align="left" valign="top">0.0243</td>
<td align="left" valign="top">&#8211;58.1363</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0624</td>
<td align="left" valign="top">0.0029</td>
<td align="left" valign="top">&#8211;21.8165</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.1165</td>
<td align="left" valign="top">0.0106</td>
<td align="left" valign="top">11.0194</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.9210</td>
<td align="left" valign="top">2.9949</td>
<td align="left" valign="top">395.7490</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">1.0840</td>
<td align="left" valign="top">1.0985</td>
<td align="left" valign="top">48.8535</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">3.0000</td>
<td align="left" valign="top">3.0000</td>
<td align="left" valign="top">2.6429</td>
<td align="left" valign="top">0.0474</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">3.0000</td>
<td align="left" valign="top">3.0000</td>
<td align="left" valign="top">18.3040</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">14.1367</td>
<td align="left" valign="top">15.1572</td>
<td align="left" valign="top">7.0818</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9548</td>
<td align="left" valign="top">2.9984</td>
<td align="left" valign="top">80.4330</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(pron)</td>
<td align="left" valign="top">151.6925</td>
<td align="left" valign="top">205.0000</td>
<td align="left" valign="top">17.8510</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="FD1">
<caption>
<p><bold>Figure D.1:</bold> Partial effects according to the Gaussian Location-Scale GAM fitted to spoken word duration of tokens (n = 56,024) summarized in Table D.1, i.e. without modeling the variance in the outcome.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g8.png"/>
</fig>
<table-wrap id="TD2">
<caption>
<p><bold>Table D.2:</bold> GAM fitted to the homophone tokens, without modeling the variance, and without the by-lexeme random effect; AIC = 39,676.71.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.3569</td>
<td align="left" valign="top">0.0027</td>
<td align="left" valign="top">&#8211;510.1780</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0652</td>
<td align="left" valign="top">0.0029</td>
<td align="left" valign="top">&#8211;22.2262</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.0055</td>
<td align="left" valign="top">0.0044</td>
<td align="left" valign="top">1.2401</td>
<td align="left" valign="top">0.2149</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.9476</td>
<td align="left" valign="top">2.9980</td>
<td align="left" valign="top">393.3558</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">2.9874</td>
<td align="left" valign="top">2.9999</td>
<td align="left" valign="top">296.4195</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">2.5299</td>
<td align="left" valign="top">2.8110</td>
<td align="left" valign="top">9.3252</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">2.9618</td>
<td align="left" valign="top">2.9987</td>
<td align="left" valign="top">2287.1847</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">te(Freq., Freq.Ratio)</td>
<td align="left" valign="top">23.9116</td>
<td align="left" valign="top">23.9959</td>
<td align="left" valign="top">97.1434</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9740</td>
<td align="left" valign="top">2.9995</td>
<td align="left" valign="top">86.7165</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="FD2">
<caption>
<p><bold>Figure D.2:</bold> Partial effects according to the Gaussian Location-Scale GAM fitted to spoken word duration of tokens (n = 56,024), summarized in Table D.2, i.e. without modeling the variance in the outcome.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g9.png"/>
</fig>
</sec>
<sec>
<title>Appendix E Token-based models without frequency as a predictor</title>
<table-wrap id="TE1">
<caption>
<p><bold>Table E.1:</bold> GAM fitted to the homophone tokens, without frequency as a predictor.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>A. Parametric coefficients</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Std. Error</bold></td>
<td align="left" valign="top"><bold>t-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="left" valign="top">&#8211;1.2933</td>
<td align="left" valign="top">0.0122</td>
<td align="left" valign="top">&#8211;105.7341</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Sex = male</td>
<td align="left" valign="top">&#8211;0.0622</td>
<td align="left" valign="top">0.0029</td>
<td align="left" valign="top">&#8211;21.7489</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">Noun Bias = yes</td>
<td align="left" valign="top">0.1148</td>
<td align="left" valign="top">0.0101</td>
<td align="left" valign="top">11.3905</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top"><bold>B. Smooth terms</bold></td>
<td align="left" valign="top"><bold>edf</bold></td>
<td align="left" valign="top"><bold>Ref.df</bold></td>
<td align="left" valign="top"><bold>F-value</bold></td>
<td align="left" valign="top"><bold>p-value</bold></td>
</tr>
<tr>
<td align="left" valign="top">s(Bigram prob.)</td>
<td align="left" valign="top">2.8903</td>
<td align="left" valign="top">2.9901</td>
<td align="left" valign="top">434.9655</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(PND)</td>
<td align="left" valign="top">2.6197</td>
<td align="left" valign="top">2.6728</td>
<td align="left" valign="top">32.6225</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(Orthogr.)</td>
<td align="left" valign="top">2.9999</td>
<td align="left" valign="top">3.0000</td>
<td align="left" valign="top">4.3449</td>
<td align="left" valign="top">0.0045</td>
</tr>
<tr>
<td align="left" valign="top">s(Resid. baseline)</td>
<td align="left" valign="top">2.9997</td>
<td align="left" valign="top">2.9998</td>
<td align="left" valign="top">24.5924</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(age)</td>
<td align="left" valign="top">2.9881</td>
<td align="left" valign="top">2.9999</td>
<td align="left" valign="top">80.9129</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
<tr>
<td align="left" valign="top">s(pron)</td>
<td align="left" valign="top">148.7266</td>
<td align="left" valign="top">205.0000</td>
<td align="left" valign="top">28.0788</td>
<td align="left" valign="top">&lt;0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="FE1">
<caption>
<p><bold>Figure E.1:</bold> Partial effects according to the Gaussian Location-Scale GAM fitted to spoken word duration of tokens (n = 56,024), summarized in Table E.1, i.e. without frequency as a predictor.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossapx-5-1-49062-g10.png"/>
</fig>
</sec>
<sec>
<title>Data accessibility statement</title>
<p>The data and analysis code for this study can be accessed at <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://osf.io/7dg38/">https://osf.io/7dg38/</ext-link>.</p>
</sec>
<sec>
<title>Ethics and consent</title>
<p>The data used were from publicly available, anonymized databases. Additional consent procedures were unnecessary.</p>
</sec>
<sec>
<title>Acknowledgments</title>
<p>The authors are grateful to the editor and three anonymous reviewers for their thoughtful feedback.</p>
</sec>
<sec>
<title>Competing interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<sec>
<title>Author contributions</title>
<p>SG contributed Conceptualization, Formal Analysis, Investigation, Methodology, Software, Writing, Original Draft Preparation, and Review &amp; Editing. HB contributed Conceptualization, Formal Analysis, Investigation, Methodology, Software, Writing, and Review &amp; Editing.</p>
</sec>
<sec>
<title>ORCiD IDs</title>
<p><bold>Susanne Gahl:</bold>&#160;<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://orcid.org/0000-0003-1293-7899">https://orcid.org/0000-0003-1293-7899</ext-link></p>
<p><bold>R. Harald Baayen:</bold>&#160;<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://orcid.org/0000-0003-3178-3944">https://orcid.org/0000-0003-3178-3944</ext-link></p>
</sec>
<ref-list>
<ref id="B1"><mixed-citation publication-type="journal"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2011</year>). <article-title>Demythologizing the word frequency effect: A discriminative learning perspective</article-title>. <source>The Mental Lexicon</source>, <volume>5</volume>, <fpage>436</fpage>&#8211;<lpage>461</lpage>. <pub-id pub-id-type="doi">10.1075/ml.5.3.10baa</pub-id></mixed-citation></ref>
<ref id="B2"><mixed-citation publication-type="journal"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Davidson</surname>, <given-names>D. J.</given-names></string-name>, &amp; <string-name><surname>Bates</surname>, <given-names>D. M.</given-names></string-name> (<year>2008</year>). <article-title>Mixed-effects modeling with crossed random effects for subjects and items</article-title>. <source>Journal of Memory and Language</source>, <volume>59</volume>, <fpage>390</fpage>&#8211;<lpage>412</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2007.12.005</pub-id></mixed-citation></ref>
<ref id="B3"><mixed-citation publication-type="journal"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Feldman</surname>, <given-names>L.</given-names></string-name>, &amp; <string-name><surname>Schreuder</surname>, <given-names>R.</given-names></string-name> (<year>2006</year>). <article-title>Morphological influences on the recognition of monosyllabic monomorphemic words</article-title>. <source>Journal of Memory and Language</source>, <volume>53</volume>, <fpage>496</fpage>&#8211;<lpage>512</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2006.03.008</pub-id></mixed-citation></ref>
<ref id="B4"><mixed-citation publication-type="book"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Kuperman</surname>, <given-names>V.</given-names></string-name>, &amp; <string-name><surname>Bertram</surname>, <given-names>R.</given-names></string-name> (<year>2010</year>). <chapter-title>Frequency effects in compound processing</chapter-title>. In <string-name><given-names>S.</given-names> <surname>Scalise</surname></string-name> &amp; <string-name><given-names>I.</given-names> <surname>Vogel</surname></string-name> (Eds.), <source>Compounding</source>. <publisher-name>Benjamins</publisher-name>. <pub-id pub-id-type="doi">10.1075/cilt.311.20baa</pub-id></mixed-citation></ref>
<ref id="B5"><mixed-citation publication-type="book"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, &amp; <string-name><surname>Linke</surname>, <given-names>M.</given-names></string-name> (<year>2021</year>). <chapter-title>Generalized Additive Mixed Models</chapter-title>. In <string-name><given-names>M.</given-names> <surname>Paquot</surname></string-name> &amp; <string-name><given-names>S. T.</given-names> <surname>Gries</surname></string-name> (Eds.), <source>Practical handbook of corpus linguistics</source> (pp. <fpage>563</fpage>&#8211;<lpage>592</lpage>). <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-3-030-46216-1_23</pub-id></mixed-citation></ref>
<ref id="B6"><mixed-citation publication-type="webpage"><string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Piepenbrock</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Gulikers</surname>, <given-names>L.</given-names></string-name> (<year>1995</year>). <source>The CELEX lexical database (release 2)</source>. <publisher-name>Distributed by the Linguistic Data Consortium, University of Pennsylvania</publisher-name>. <uri>https://hdl.handle.net/21.11116/0000-0001-91EF-E</uri></mixed-citation></ref>
<ref id="B7"><mixed-citation publication-type="journal"><string-name><surname>Balota</surname>, <given-names>D. A.</given-names></string-name>, <string-name><surname>Yap</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Hutchison</surname>, <given-names>K. A.</given-names></string-name>, <string-name><surname>Cortese</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Kessler</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Loftis</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Neely</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Nelson</surname>, <given-names>D. L.</given-names></string-name>, <string-name><surname>Simpson</surname>, <given-names>G. B.</given-names></string-name>, &amp; <string-name><surname>Treiman</surname>, <given-names>R.</given-names></string-name> (<year>2007</year>). <article-title>The English Lexicon Project</article-title>. <source>Behavior Research Methods</source>, <volume>39</volume>(<issue>3</issue>), <fpage>445</fpage>&#8211;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.3758/bf03193014</pub-id></mixed-citation></ref>
<ref id="B8"><mixed-citation publication-type="webpage"><string-name><surname>Bates</surname>, <given-names>D. M.</given-names></string-name> (<year>2005</year>). <article-title>Fitting linear mixed models in R</article-title>. <source>R News</source>, <volume>5</volume>, <fpage>27</fpage>&#8211;<lpage>30</lpage>. <uri>https://svn.r-project.org/R-project-web/trunk/md/doc/Rnews/Rnews_2005-1.pdf#page=27</uri></mixed-citation></ref>
<ref id="B9"><mixed-citation publication-type="journal"><string-name><surname>Bell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Brenier</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Gregory</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Girand</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Jurafsky</surname>, <given-names>D.</given-names></string-name> (<year>2009</year>). <article-title>Predictability effects on durations of content and function words in conversational English</article-title>. <source>Journal of Memory and Language</source>, <volume>60</volume>(<issue>1</issue>), <fpage>92</fpage>&#8211;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2008.06.003</pub-id></mixed-citation></ref>
<ref id="B10"><mixed-citation publication-type="journal"><string-name><surname>Bell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Jurafsky</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Fosler-Lussier</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Girand</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Gregory</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Gildea</surname>, <given-names>D.</given-names></string-name> (<year>2003</year>). <article-title>Effects of disfluencies, predictability, and utterance position on word form variation in English conversation</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>113</volume>(<issue>2</issue>), <fpage>1001</fpage>&#8211;<lpage>1024</lpage>. <pub-id pub-id-type="doi">10.1121/1.1534836</pub-id></mixed-citation></ref>
<ref id="B11"><mixed-citation publication-type="journal"><string-name><surname>Berndt</surname>, <given-names>R. S.</given-names></string-name>, <string-name><surname>Reggia</surname>, <given-names>J. A.</given-names></string-name>, &amp; <string-name><surname>Mitchum</surname>, <given-names>C. C.</given-names></string-name> (<year>1987</year>). <article-title>Empirically derived probabilities for grapheme-to-phoneme correspondences in English</article-title>. <source>Behavior Research Methods, Instruments, &amp; Computers</source>, <volume>19</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.3758/BF03207663</pub-id></mixed-citation></ref>
<ref id="B12"><mixed-citation publication-type="journal"><string-name><surname>Brown</surname>, <given-names>E. L.</given-names></string-name>, <string-name><surname>Raymond</surname>, <given-names>W. D.</given-names></string-name>, <string-name><surname>Brown</surname>, <given-names>E. K.</given-names></string-name>, &amp; <string-name><surname>File-Muriel</surname>, <given-names>R. J.</given-names></string-name> (<year>2021</year>). <article-title>Lexically specific accumulation in memory of word and segment speech rates</article-title>. <source>Corpus Linguistics and Linguistic Theory</source>. <pub-id pub-id-type="doi">10.1515/cllt-2020-0016</pub-id></mixed-citation></ref>
<ref id="B13"><mixed-citation publication-type="journal"><string-name><surname>Chuang</surname>, <given-names>Y.-Y.</given-names></string-name>, <string-name><surname>Bell</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Tseng</surname>, <given-names>Y.-H.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2026</year>). <article-title>Word-specific tonal realizations in Mandarin</article-title>. <source>Language</source>, <volume>102</volume>, <fpage>1</fpage>&#8211;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1017/S0097850725000001</pub-id></mixed-citation></ref>
<ref id="B14"><mixed-citation publication-type="book"><string-name><surname>Chuang</surname>, <given-names>Y.-Y.</given-names></string-name>, <string-name><surname>Fon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Papakyritsis</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2021</year>). <chapter-title>Analyzing phonetic data with generalized additive mixed models</chapter-title>. In <string-name><given-names>M. J.</given-names> <surname>Ball</surname></string-name> (Ed.), <source>Manual of clinical phonetics</source> (pp. <fpage>108</fpage>&#8211;<lpage>138</lpage>). <publisher-name>Routledge</publisher-name>. <pub-id pub-id-type="doi">10.4324/9780429320903</pub-id></mixed-citation></ref>
<ref id="B15"><mixed-citation publication-type="journal"><string-name><surname>Dautriche</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Mahowald</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Gibson</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Christophe</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Piantadosi</surname>, <given-names>S. T.</given-names></string-name> (<year>2017</year>). <article-title>Words cluster phonetically beyond phonotactic regularities</article-title>. <source>Cognition</source>, <volume>163</volume>, <fpage>128</fpage>&#8211;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2017.02.001</pub-id></mixed-citation></ref>
<ref id="B16"><mixed-citation publication-type="journal"><string-name><surname>Dell</surname>, <given-names>G. S.</given-names></string-name> (<year>1990</year>). <article-title>Effects of frequency and vocabulary type on phonological speech errors</article-title>. <source>Language and cognitive processes</source>, <volume>5</volume>(<issue>4</issue>), <fpage>313</fpage>&#8211;<lpage>349</lpage>. <pub-id pub-id-type="doi">10.1080/01690969008407066</pub-id></mixed-citation></ref>
<ref id="B17"><mixed-citation publication-type="journal"><string-name><surname>Deshmukh</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Ganapathiraju</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Gleeson</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Hamaker</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Picone</surname>, <given-names>J.</given-names></string-name> (<year>1998</year>). <article-title>Resegmentation of SWITCHBOARD</article-title>. <source>Fifth International Conference on Spoken Language Processing</source>. <pub-id pub-id-type="doi">10.21437/ICSLP.1998-588</pub-id></mixed-citation></ref>
<ref id="B18"><mixed-citation publication-type="journal"><string-name><surname>Fosler-Lussier</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Morgan</surname>, <given-names>N.</given-names></string-name> (<year>1999</year>). <article-title>Effects of speaking rate and word frequency on pronunciations in conversational speech</article-title>. <source>Speech Communication</source>, <volume>29</volume>, <fpage>137</fpage>&#8211;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.1016/S0167-6393(99)00035-7</pub-id></mixed-citation></ref>
<ref id="B19"><mixed-citation publication-type="journal"><string-name><surname>Frauenfelder</surname>, <given-names>U. H.</given-names></string-name>, <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, &amp; <string-name><surname>Hellwig</surname>, <given-names>F. M.</given-names></string-name> (<year>1993</year>). <article-title>Neighborhood density and frequency across languages and modalities</article-title>. <source>Journal of Memory and Language</source>, <volume>32</volume>(<issue>6</issue>), <fpage>781</fpage>&#8211;<lpage>804</lpage>. <pub-id pub-id-type="doi">10.1006/jmla.1993.1039</pub-id></mixed-citation></ref>
<ref id="B20"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name> (<year>2008</year>). <article-title>&#8217;Time&#8217; and &#8217;thyme&#8217; are not homophones: The effect of lemma frequency on word durations in spontaneous speech</article-title>. <source>Language</source>, <volume>84</volume>(<issue>3</issue>), <fpage>474</fpage>&#8211;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.1353/lan.0.0035</pub-id></mixed-citation></ref>
<ref id="B21"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name> (<year>2009</year>). <article-title>Homophone duration in spontaneous speech: A mixed-effects model</article-title>. <source>UC Berkeley Phonology Lab Technical Report</source>. <pub-id pub-id-type="doi">10.5070/P784Q8Q0QN</pub-id></mixed-citation></ref>
<ref id="B22"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2024</year>). <article-title><italic>Time</italic> and <italic>thyme</italic> again: Connecting English spoken word duration to models of the mental lexicon</article-title>. <source>Language</source>, <volume>100</volume>, <fpage>623</fpage>&#8211;<lpage>670</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2024.a947037</pub-id></mixed-citation></ref>
<ref id="B23"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Yao</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>Johnson</surname>, <given-names>K.</given-names></string-name> (<year>2012</year>). <article-title>Why reduce? Phonological neighborhood density and phonetic reduction in spontaneous speech</article-title>. <source>Journal of Memory and Language</source>, <volume>66</volume>(<issue>4</issue>), <fpage>789</fpage>&#8211;<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2011.11.006</pub-id></mixed-citation></ref>
<ref id="B24"><mixed-citation publication-type="journal"><string-name><surname>Gardner</surname>, <given-names>M. K.</given-names></string-name>, <string-name><surname>Rothkopf</surname>, <given-names>E. Z.</given-names></string-name>, <string-name><surname>Lapan</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Lafferty</surname>, <given-names>T.</given-names></string-name> (<year>1987</year>). <article-title>The word frequency effect in lexical decision: Finding a frequency-based component</article-title>. <source>Memory &amp; Cognition</source>, <volume>15</volume>, <fpage>24</fpage>&#8211;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.3758/BF03197709</pub-id></mixed-citation></ref>
<ref id="B25"><mixed-citation publication-type="journal"><string-name><surname>Godfrey</surname>, <given-names>J. J.</given-names></string-name>, <string-name><surname>Holliman</surname>, <given-names>E. C.</given-names></string-name>, &amp; <string-name><surname>McDaniel</surname>, <given-names>J.</given-names></string-name> (<year>1992</year>). <article-title>SWITCHBOARD: Telephone speech corpus for research and development</article-title>. <source>IEEE International Conference on Acoustics, Speech, and Signal Processing</source>, <volume>1</volume>, <fpage>517</fpage>&#8211;<lpage>520</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.1992.225858</pub-id></mixed-citation></ref>
<ref id="B26"><mixed-citation publication-type="journal"><string-name><surname>Gries</surname>, <given-names>S. T.</given-names></string-name> (<year>2015</year>). <article-title>The most under-used statistical method in corpus linguistics: Multi-level (and mixed-effects) models</article-title>. <source>Corpora</source>, <volume>10</volume>(<issue>1</issue>), <fpage>95</fpage>&#8211;<lpage>125</lpage>. <pub-id pub-id-type="doi">10.3366/cor.2015.0068</pub-id></mixed-citation></ref>
<ref id="B27"><mixed-citation publication-type="book"><string-name><surname>Hay</surname>, <given-names>J.</given-names></string-name> (<year>2007</year>). <chapter-title>The phonetics of &#8216;un&#8217;</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Munat</surname></string-name> (Ed.), <source>Lexical creativity, texts and contexts</source> (pp. <fpage>39</fpage>&#8211;<lpage>57</lpage>). <publisher-name>John Benjamins Publishing Company</publisher-name>. <pub-id pub-id-type="doi">10.1075/sfsl.58.09hay</pub-id></mixed-citation></ref>
<ref id="B28"><mixed-citation publication-type="journal"><string-name><surname>Hay</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Pierrehumbert</surname>, <given-names>J. B.</given-names></string-name>, <string-name><surname>Walker</surname>, <given-names>A. J.</given-names></string-name>, &amp; <string-name><surname>LaShell</surname>, <given-names>P.</given-names></string-name> (<year>2015</year>). <article-title>Tracking word frequency effects through 130 years of sound change</article-title>. <source>Cognition</source>, <volume>139</volume>, <fpage>83</fpage>&#8211;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2015.02.012</pub-id></mixed-citation></ref>
<ref id="B29"><mixed-citation publication-type="thesis"><string-name><surname>Hermalin</surname>, <given-names>N.</given-names></string-name> (<year>2025</year>). <source>Preliminary investigations into the communicative efficiency of logographic writing systems and written language</source> [Doctoral dissertation, <publisher-name>University of California at Berkeley</publisher-name>]. <uri>https://escholarship.org/uc/item/0hk9n57m</uri></mixed-citation></ref>
<ref id="B30"><mixed-citation publication-type="journal"><string-name><surname>Horton</surname>, <given-names>W. S.</given-names></string-name>, <string-name><surname>Spieler</surname>, <given-names>D. H.</given-names></string-name>, &amp; <string-name><surname>Shriberg</surname>, <given-names>E.</given-names></string-name> (<year>2010</year>). <article-title>A corpus analysis of patterns of age-related change in conversational speech</article-title>. <source>Psychology and Aging</source>, <volume>25</volume>(<issue>3</issue>), <elocation-id>708</elocation-id>. <pub-id pub-id-type="doi">10.1037/a0019424</pub-id></mixed-citation></ref>
<ref id="B31"><mixed-citation publication-type="journal"><string-name><surname>Jescheniak</surname>, <given-names>J. D.</given-names></string-name>, &amp; <string-name><surname>Levelt</surname>, <given-names>W. J. M.</given-names></string-name> (<year>1994</year>). <article-title>Word frequency effects in speech production: Retrieval of syntactic information and of phonological form</article-title>. <source>Journal of Experimental Psychology: Learning, Memory and Cognition</source>, <volume>20</volume>(<issue>4</issue>), <fpage>824</fpage>&#8211;<lpage>843</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.20.4.824</pub-id></mixed-citation></ref>
<ref id="B32"><mixed-citation publication-type="journal"><string-name><surname>Jescheniak</surname>, <given-names>J. D.</given-names></string-name>, <string-name><surname>Meyer</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Levelt</surname>, <given-names>W. J. M.</given-names></string-name> (<year>2003</year>). <article-title>Specific-word frequency is not all that counts in speech production. Evidence from the production of homophones in dutch and german</article-title>. <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>, <volume>29</volume>, <fpage>432</fpage>&#8211;<lpage>438</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.29.3.432</pub-id></mixed-citation></ref>
<ref id="B33"><mixed-citation publication-type="journal"><string-name><surname>Kilbourn-Ceron</surname>, <given-names>O.</given-names></string-name>, <string-name><surname>Clayards</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Wagner</surname>, <given-names>M.</given-names></string-name> (<year>2020</year>). <article-title>Predictability modulates pronunciation variants through speech planning effects: A case study on coronal stop realizations</article-title>. <source>Laboratory Phonology: Journal of the Association for Laboratory Phonology</source>, <volume>11</volume>(<issue>1</issue>), <elocation-id>5</elocation-id>. <pub-id pub-id-type="doi">10.5334/labphon.168</pub-id></mixed-citation></ref>
<ref id="B34"><mixed-citation publication-type="book"><string-name><surname>Kilgarriff</surname>, <given-names>A.</given-names></string-name> (<year>2006</year>). <chapter-title>Word senses</chapter-title>. In <string-name><given-names>E.</given-names> <surname>Agirre</surname></string-name> &amp; <string-name><given-names>P.</given-names> <surname>Edmonds</surname></string-name> (Eds.), <source>Word sense disambiguation</source> (pp. <fpage>29</fpage>&#8211;<lpage>46</lpage>). <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-1-4020-4809-8_2</pub-id></mixed-citation></ref>
<ref id="B35"><mixed-citation publication-type="book"><string-name><surname>K&#246;hler</surname>, <given-names>R.</given-names></string-name> (<year>1986</year>). <source>Zur linguistischen Synergetik: Struktur und Dynamik der Lexik</source>. <publisher-name>Brockmeyer</publisher-name>.</mixed-citation></ref>
<ref id="B36"><mixed-citation publication-type="journal"><string-name><surname>Lammer</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Beyer</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Riedel-Heller</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sacher</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Glaesmer</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Villringer</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Witte</surname>, <given-names>A. V.</given-names></string-name> (<year>2025</year>). <article-title>Generalized additive mixed models to discern data-driven theoretically informed strategies for public brain, cognitive and mental health</article-title>. <source>European Journal of Epidemiology</source>, <fpage>1</fpage>&#8211;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1007/s10654-025-01296-9</pub-id></mixed-citation></ref>
<ref id="B37"><mixed-citation publication-type="journal"><string-name><surname>Landauer</surname>, <given-names>T. K.</given-names></string-name>, &amp; <string-name><surname>Streeter</surname>, <given-names>L. A.</given-names></string-name> (<year>1973</year>). <article-title>Structural differences between common and rare words: Failure of equivalence assumptions for theories of word recognition</article-title>. <source>Journal of Learning and Verbal Behavior</source>, <volume>12</volume>, <fpage>119</fpage>&#8211;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-5371(73)80001-5</pub-id></mixed-citation></ref>
<ref id="B38"><mixed-citation publication-type="journal"><string-name><surname>Levelt</surname>, <given-names>W. J. M.</given-names></string-name>, <string-name><surname>Roelofs</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Meyer</surname>, <given-names>A. S.</given-names></string-name> (<year>1999</year>). <article-title>A theory of lexical access in speech production</article-title>. <source>Behavioral and Brain Sciences</source>, <volume>22</volume>, <fpage>1</fpage>&#8211;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X99001776</pub-id></mixed-citation></ref>
<ref id="B39"><mixed-citation publication-type="journal"><string-name><surname>Lieberman</surname>, <given-names>P.</given-names></string-name> (<year>1963</year>). <article-title>Some effects of semantic and grammatical context on the production and perception of speech</article-title>. <source>Language and Speech</source>, <volume>6</volume>, <fpage>172</fpage>&#8211;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1177/002383096300600306</pub-id></mixed-citation></ref>
<ref id="B40"><mixed-citation publication-type="journal"><string-name><surname>Lohmann</surname>, <given-names>A.</given-names></string-name> (<year>2018a</year>). <article-title>&#8217;Time&#8217; and &#8217;thyme&#8217; are not homophones: A closer look at Gahl&#8217;s work on the lemma-frequency effect, including a reanalysis</article-title>. <source>Language</source>, <volume>94</volume>(<issue>2</issue>), <fpage>e180</fpage>&#8211;<lpage>e190</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2018.0032</pub-id></mixed-citation></ref>
<ref id="B41"><mixed-citation publication-type="journal"><string-name><surname>Lohmann</surname>, <given-names>A.</given-names></string-name> (<year>2018b</year>). <article-title>Cut (n) and cut (v) are not homophones: Lemma frequency affects the duration of noun&#8211;verb conversion pairs</article-title>. <source>Journal of Linguistics</source>, <volume>54</volume>(<issue>4</issue>), <fpage>753</fpage>&#8211;<lpage>777</lpage>. <pub-id pub-id-type="doi">10.1121/1.4987628</pub-id></mixed-citation></ref>
<ref id="B42"><mixed-citation publication-type="webpage"><string-name><surname>Pimentel</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Maudslay</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Blasi</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name><surname>Cotterell</surname>, <given-names>R.</given-names></string-name> (<year>2020</year>). <source>Speakers fill lexical semantic gaps with context</source>. arXiv. <uri>https://arxiv.org/abs/2010.02172</uri></mixed-citation></ref>
<ref id="B43"><mixed-citation publication-type="book"><string-name><surname>Pinheiro</surname>, <given-names>J. C.</given-names></string-name>, &amp; <string-name><surname>Bates</surname>, <given-names>D. M.</given-names></string-name> (<year>2000</year>). <source>Mixed-effects models in S and S-PLUS</source>. <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-1-4419-0318-1</pub-id></mixed-citation></ref>
<ref id="B44"><mixed-citation publication-type="webpage"><string-name><surname>Pitt</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Dilley</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Johnson</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Kiesling</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Raymond</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Hume</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Fosler-Lussier</surname>, <given-names>E.</given-names></string-name> (<year>2007</year>). <source>Buckeye Corpus of Conversational Speech (2nd release)</source>. <publisher-loc>Columbus, OH</publisher-loc>: <publisher-name>Department of Psychology, Ohio State University</publisher-name>. <uri>https://www.buckeyecorpus.osu.edu</uri></mixed-citation></ref>
<ref id="B45"><mixed-citation publication-type="journal"><string-name><surname>Pluymaekers</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Ernestus</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2005</year>). <article-title>Lexical frequency and acoustic reduction in spoken Dutch</article-title>. <source>Journal of the Acoustical Society of America</source>, <volume>118</volume>, <fpage>2561</fpage>&#8211;<lpage>2569</lpage>. <pub-id pub-id-type="doi">10.1121/1.2011150</pub-id></mixed-citation></ref>
<ref id="B46"><mixed-citation publication-type="journal"><string-name><surname>Pluymaekers</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Ernestus</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2006</year>). <article-title>Articulatory planning is continuous and sensitive to informational redundancy</article-title>. <source>Phonetica</source>, <volume>62</volume>(<issue>2&#8211;4</issue>), <fpage>146</fpage>&#8211;<lpage>159</lpage>. <pub-id pub-id-type="doi">10.1159/000090095</pub-id></mixed-citation></ref>
<ref id="B47"><mixed-citation publication-type="journal"><string-name><surname>Quen&#233;</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>van den Bergh</surname>, <given-names>H.</given-names></string-name> (<year>2008</year>). <article-title>Examples of mixed-effects modeling with crossed random effects and with binomial data</article-title>. <source>Journal of Memory and Language</source>, <volume>59</volume>(<issue>4</issue>), <fpage>413</fpage>&#8211;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2008.02.002</pub-id></mixed-citation></ref>
<ref id="B48"><mixed-citation publication-type="webpage"><collab>R Core Team</collab>. (<year>2022</year>). <source>R: A language and environment for statistical computing</source>. R Foundation for Statistical Computing. <uri>https://www.R-project.org/</uri></mixed-citation></ref>
<ref id="B49"><mixed-citation publication-type="journal"><string-name><surname>Saxena</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Mishra</surname>, <given-names>S. K.</given-names></string-name>, <string-name><surname>Rodrigo</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Choudhury</surname>, <given-names>M.</given-names></string-name> (<year>2022</year>). <article-title>Functional consequences of extended high frequency hearing impairment: Evidence from the speech, spatial, and qualities of hearing scale</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>152</volume>(<issue>5</issue>), <fpage>2946</fpage>&#8211;<lpage>2952</lpage>.</mixed-citation></ref>
<ref id="B50"><mixed-citation publication-type="journal"><string-name><surname>Seyfarth</surname>, <given-names>S.</given-names></string-name> (<year>2014</year>). <article-title>Word informativity influences acoustic duration: Effects of contextual predictability on lexical representation</article-title>. <source>Cognition</source>, <volume>133</volume>(<issue>1</issue>), <fpage>140</fpage>&#8211;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2014.06.013</pub-id></mixed-citation></ref>
<ref id="B51"><mixed-citation publication-type="journal"><string-name><surname>Shields</surname>, <given-names>L. W.</given-names></string-name>, &amp; <string-name><surname>Balota</surname>, <given-names>D. A.</given-names></string-name> (<year>1991</year>). <article-title>Repetition and associative context effects in speech production</article-title>. <source>Language and Speech</source>, <volume>34</volume>, <fpage>47</fpage>&#8211;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1177/002383099103400103</pub-id></mixed-citation></ref>
<ref id="B52"><mixed-citation publication-type="journal"><string-name><surname>Sorensen</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Cooper</surname>, <given-names>W. E.</given-names></string-name>, &amp; <string-name><surname>Paccia</surname>, <given-names>J. M.</given-names></string-name> (<year>1978</year>). <article-title>Speech timing of grammatical categories</article-title>. <source>Cognition</source>, <volume>6</volume>(<issue>2</issue>), <fpage>135</fpage>&#8211;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(78)90019-7</pub-id></mixed-citation></ref>
<ref id="B53"><mixed-citation publication-type="journal"><string-name><surname>S&#243;skuthy</surname>, <given-names>M.</given-names></string-name> (<year>2021</year>). <article-title>Evaluating generalised additive mixed modelling strategies for dynamic speech analysis</article-title>. <source>Journal of Phonetics</source>, <volume>84</volume>, <elocation-id>101017</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.wocn.2020.101017</pub-id></mixed-citation></ref>
<ref id="B54"><mixed-citation publication-type="journal"><string-name><surname>S&#243;skuthy</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Hay</surname>, <given-names>J.</given-names></string-name> (<year>2017</year>). <article-title>Changing word usage predicts changing word durations in New Zealand english</article-title>. <source>Cognition</source>, <volume>166</volume>, <fpage>298</fpage>&#8211;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2017.05.032</pub-id></mixed-citation></ref>
<ref id="B55"><mixed-citation publication-type="journal"><string-name><surname>Tanner</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Sonderegger</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Stuart-Smith</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Fruehwald</surname>, <given-names>J.</given-names></string-name> (<year>2020</year>). <article-title>Toward &#8220;English&#8221; phonetics: Variability in the pre-consonantal voicing effect across English dialects and speakers</article-title>. <source>Frontiers in Artificial Intelligence</source>, <volume>3</volume>, <elocation-id>38</elocation-id>. <pub-id pub-id-type="doi">10.3389/frai.2020.00038</pub-id></mixed-citation></ref>
<ref id="B56"><mixed-citation publication-type="journal"><string-name><surname>Tomaschek</surname>, <given-names>F.</given-names></string-name>, &amp; <string-name><surname>Ramscar</surname>, <given-names>M.</given-names></string-name> (<year>2022</year>). <article-title>Understanding the phonetic characteristics of speech under uncertainty &#8211; Implications of the representation of linguistic knowledge in learning and processing</article-title>. <source>Frontiers in Psychology</source>, <volume>13</volume>, <elocation-id>754395</elocation-id>. <pub-id pub-id-type="doi">10.3389/fpsyg.2022.754395</pub-id></mixed-citation></ref>
<ref id="B57"><mixed-citation publication-type="journal"><string-name><surname>van Rij</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wieling</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name>, &amp; <string-name><surname>van Rijn</surname>, <given-names>H.</given-names></string-name> (<year>2020</year>). <article-title>itsadug: Interpreting Time Series and Autocorrelated Data Using GAMMs [R package version 2.4]</article-title>.</mixed-citation></ref>
<ref id="B58"><mixed-citation publication-type="journal"><string-name><surname>Vitevitch</surname>, <given-names>M. S.</given-names></string-name>, &amp; <string-name><surname>Luce</surname>, <given-names>P. A.</given-names></string-name> (<year>2004</year>). <article-title>A web-based interface to calculate phonotactic probability for words and nonwords in English</article-title>. <source>Behavior Research Methods, Instruments, &amp; Computers</source>, <volume>36</volume>(<issue>3</issue>), <fpage>481</fpage>&#8211;<lpage>487</lpage>. <pub-id pub-id-type="doi">10.3758/BF03195594</pub-id></mixed-citation></ref>
<ref id="B59"><mixed-citation publication-type="journal"><string-name><surname>Walsh</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name><surname>Parker</surname>, <given-names>F.</given-names></string-name> (<year>1983</year>). <article-title>The duration of morphemic and non-morphemic [s] in English</article-title>. <source>Journal of Phonetics</source>, <volume>11</volume>(<issue>2</issue>), <fpage>201</fpage>&#8211;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1016/S0095-4470(19)30816-2</pub-id></mixed-citation></ref>
<ref id="B60"><mixed-citation publication-type="journal"><string-name><surname>Warner</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Jongman</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sereno</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Kemps</surname>, <given-names>R.</given-names></string-name> (<year>2004</year>). <article-title>Incomplete neutralization and other sub-phonemic durational differences in production and perception: Evidence from Dutch</article-title>. <source>Journal of Phonetics</source>, <fpage>251</fpage>&#8211;<lpage>276</lpage>. <pub-id pub-id-type="doi">10.1016/S0095-4470(03)00032-9</pub-id></mixed-citation></ref>
<ref id="B61"><mixed-citation publication-type="journal"><string-name><surname>Wieling</surname>, <given-names>M.</given-names></string-name> (<year>2018</year>). <article-title>Analyzing dynamic phonetic data using generalized additive mixed modeling: A tutorial focusing on articulatory differences between L1 and L2 speakers of English</article-title>. <source>Journal of Phonetics</source>, <volume>70</volume>, <fpage>86</fpage>&#8211;<lpage>116</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2018.03.002</pub-id></mixed-citation></ref>
<ref id="B62"><mixed-citation publication-type="journal"><string-name><surname>Wieling</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Montemagni</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Nerbonne</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2014</year>). <article-title>Lexical differences between Tuscan dialects and standard Italian: Accounting for geographic and socio-demographic variation using generalized additive mixed modeling</article-title>. <source>Language</source>, <volume>90</volume>(<issue>3</issue>), <fpage>669</fpage>&#8211;<lpage>692</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2014.0064</pub-id></mixed-citation></ref>
<ref id="B63"><mixed-citation publication-type="journal"><string-name><surname>Wieling</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Nerbonne</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Baayen</surname>, <given-names>R. H.</given-names></string-name> (<year>2011</year>). <article-title>Quantitative social dialectology: Explaining linguistic variation geographically and socially</article-title>. <source>PLoS ONE</source>, <volume>6</volume>(<issue>9</issue>), <elocation-id>e23613</elocation-id>. <pub-id pub-id-type="doi">10.1371/journal.pone.0023613</pub-id></mixed-citation></ref>
<ref id="B64"><mixed-citation publication-type="journal"><string-name><surname>Wood</surname>, <given-names>S. N.</given-names></string-name> (<year>2003</year>). <article-title>Thin-plate regression splines</article-title>. <source>Journal of the Royal Statistical Society (B)</source>, <volume>65</volume>(<issue>1</issue>), <fpage>95</fpage>&#8211;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1111/1467-9868.00374</pub-id></mixed-citation></ref>
<ref id="B65"><mixed-citation publication-type="journal"><string-name><surname>Wood</surname>, <given-names>S. N.</given-names></string-name> (<year>2004</year>). <article-title>Stable and efficient multiple smoothing parameter estimation for generalized additive models</article-title>. <source>Journal of the American Statistical Association</source>, <volume>99</volume>, <fpage>673</fpage>&#8211;<lpage>686</lpage>. <pub-id pub-id-type="doi">10.1198/016214504000000980</pub-id></mixed-citation></ref>
<ref id="B66"><mixed-citation publication-type="book"><string-name><surname>Wood</surname>, <given-names>S. N.</given-names></string-name> (<year>2006</year>). <source>Generalized Additive Models</source>. <publisher-name>Chapman &amp; Hall/CRC</publisher-name>. <pub-id pub-id-type="doi">10.1201/9781420010404</pub-id></mixed-citation></ref>
<ref id="B67"><mixed-citation publication-type="journal"><string-name><surname>Wood</surname>, <given-names>S. N.</given-names></string-name> (<year>2011</year>). <article-title>Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models</article-title>. <source>Journal of the Royal Statistical Society (B)</source>, <volume>73</volume>(<issue>1</issue>), <fpage>3</fpage>&#8211;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-9868.2010.00749.x</pub-id></mixed-citation></ref>
<ref id="B68"><mixed-citation publication-type="book"><string-name><surname>Wood</surname>, <given-names>S. N.</given-names></string-name> (<year>2017</year>). <source>Generalized Additive Models</source> (<edition>2nd</edition>). <publisher-name>Chapman &amp; Hall/CRC</publisher-name>. <pub-id pub-id-type="doi">10.1201/9781315370279</pub-id></mixed-citation></ref>
<ref id="B69"><mixed-citation publication-type="journal"><string-name><surname>Wright</surname>, <given-names>R.</given-names></string-name> (<year>2004</year>). <article-title>Factors of lexical competition in vowel articulation</article-title>. <source>Papers in Laboratory Phonology VI</source>, <fpage>75</fpage>&#8211;<lpage>87</lpage>.</mixed-citation></ref>
<ref id="B70"><mixed-citation publication-type="journal"><string-name><surname>Yuan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Liberman</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Cieri</surname>, <given-names>C.</given-names></string-name> (<year>2006</year>). <article-title>Towards an integrated understanding of speaking rate in conversation</article-title>. <source>Ninth International Conference on Spoken Language Processing</source>. <pub-id pub-id-type="doi">10.21437/Interspeech.2006-204</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>