About
The official proceedings of ELM provides an open access venue for experimental work on linguistic meaning broadly construed, with a focus on theoretical issues in semantics and pragmatics, their interplay with other components of the grammar, their relation to language processing and acquisition, as well as their connections to human cognition and computation. The papers are developed from presentations at the biennial ELM conference, held at the University of Pennsylvania since 2020.
Volume 2, Issue 1, 2023
ELM Volume 2 (2023)
Articles
- On the interpretation of German einige. The effect of tense and cardinality
We present a study investigating the effect of tense (past vs. future) on the computation of scalar implicatures in connection with the German quantifier einige 'some' in an interactive experiment, which included a financial incentive for participants to consider whether another speaker would share their judgment. We tested the hypothesis that scalar implicatures are less frequently drawn in future tense than in past tense. In addition, we studied to what extent sets with various cardinalities are prototypical representatives of einige + N. We hypothesized that larger cardinalities are more prototypical representatives of the quantifier einige than smaller cardinalities (relative to the cardinality of the total set). We analyzed the experimental data with probabilistic Bayesian models with a linking hypothesis between participants' responses and readings based on utility maximization in simple decision problems. In line with the hypotheses, we found that less scalar implicatures are drawn in future tense than in past tense, which replicates the results of previous research on English some, and that with an increase in set size acceptance of statements involving einige also increases.
- Testing the influence of QUDs on the occurrence of Conditional Perfection
In natural language conversations, speakers often communicate 'if and only if' when they say 'if'. The reasons why in some circumstances, yet not all, conditionals receive a biconditional interpretation remain under investigation. Von Fintel (2001) proposed an account where the interpretation of a conditional ("if p, then q") is predicted to depend on the focus of the conversation which may either lie on the conditions that make the consequent, q, true or on the consequences following when the antecedent, p, is true. To test this account, we present two novel behavioral experiments with non-text based stimuli that take advantage of participants' intuitive understanding of physics. We find some supporting evidence for the tested account that is not conclusive but suggests that other aspects, like the nature of potential alternative causes for the consequent to become true (e.g., with or w/o the influence of an external variable), also play a role for the interpretation of the conditional.
- Comparing Global and Local Accommodation: Rating and Response Time Data
This paper addresses the question to what extent global and local accommodation should be viewed as sharing the same underlying mechanism or whether they are distinct processes that only happen to share the same label. We present offline rating data and response times from a mouse-tracking experiment that directly compared global and local accommodation for five different triggers. The results show that globally accommodating a presupposition led to a larger decrease in acceptance than locally accommodating, and that response times for local accommodation were overall faster. While we take the results to be inconclusive with regard to the question about the underlying mechanism, we conjecture that the contexts tested here were more favorable for local accommodation, and that hence investigating how different contexts affect the relative ease of accommodation type is a promising avenue for future research.
- Real-time processing of indexical and generic expressions: Insights from, and implications for, COVID-related public health messages
We used COVID-related health messages to investigate the incremental processing of expressions such as you, we and people in generic contexts, to further our understanding of how these kinds of expressions are processed in real-time and to explore whether the ease of comprehending public health messages related to the COVID pandemic is influenced by type of referring expression. Results from a self-paced reading study point to an increased processing load in messages with the non-indexical form people (relative to we and you), which we suggest is separable from effects of word length and frequency. We interpret this as preliminary support for the Indexicality Hypothesis, which posits that expressions which can in principle receive indexical interpretations are easier to process than non-indexicals, and also emphasize the need for further work on these kinds of questions.
- On a concessive reading of the rise-fall-rise contour: contextual and semantic factors
This paper presents three auditory rating experiments on the rise-fall-rise contour (RFR). Experiment 1 provides experimental evidence that the RFR makes disagreeing with a prior statement more natural than neutral intonation would. Additionally, the data show that the RFR exhibits a valence asymmetry, noted by Göbel (2019): the amelioration of a disagreement is greater when the RFR is used in a positive reply to a negative statement than in a negative reply to a positive statement. Experiments 2 and 3 investigate factors contributing to this asymmetry, showing that it disappears in replies to questions and is weakened when the reply contains an additive particle. Based on these results, we argue that the RFR has a scalar meaning, following Göbel (2019), with the relevant scale being contextually determined and resulting in an ambiguity resembling the Focus-particle at least.
- Proportions vs. cardinalities: Comparative ambiguities and the COVID pandemic
This paper reports two psycholinguistic experiments on quantity comparatives and superlatives that are potentially ambiguous between cardinal and proportional readings. By using statements about COVID cases and vaccination numbers as a naturalistic context with real-world relevance, this work furthers our understanding of what happens in linguistic environments where multiple measure functions are available—what modulates the choice between them? The results provide new evidence that comparatives and superlatives can refer to scales ranging over degrees of proportion, in addition to degrees of cardinality. Furthermore, this experimental evidence points to a preference for cardinal interpretations, but also show that this is not rigid and can be weakened in favor of proportional readings by semantic factors—including considerations potentially related to stage- vs. individual-level differences—and by certain linguistic forms.
- Reading times show effects of contextual complexity and uncertainty in comprehension of German universal quantifiers
We report three experiments, in which we combined self-paced reading with picture-sentence verification to test how reading times are affected by meaning-related processes. In particular, we investigated German sentences containing the universal quantifier alle ("all") and examined how restrictive processes incrementally interact with other aspects of quantifier meaning, comparably to previous studies using other methods. Our results show that reading times were sensitive towards a match between context and sentence meaning and also towards an interaction between picture complexity and task demands. The results also point to the need for integrated processing models that combine refined notions of the relation between memory and expectations, on the one hand, with assumptions about adaptive processes and about representations involved in compositional interpretation, on the other.
- Representing affect information in word embeddings
A growing body of research in natural language processing (NLP) and natural language understanding (NLU) is investigating human-like knowledge learned or encoded in the word embeddings from large language models. This is a step towards understanding what knowledge language models capture that resembles human understanding of language and communication. Here, we investigated whether and how the affect meaning of a word (i.e., valence, arousal, dominance) is encoded in word embeddings pre-trained in large neural networks. We used the human-labeled dataset (Mohammad 2018) as the ground truth and performed various correlational and classification tests on four types of word embeddings. The embeddings varied in being static or contextualized, and how much affect specific information was prioritized during the pre-training and fine-tuning phase. Our analyses show that word embedding from the vanilla BERT model (Devlin et al. 2019) did not saliently encode the affect information of English words. Only when the BERT model was fine-tuned on emotion related tasks or contained extra contextualized information from emotion-rich contexts could the corresponding embedding encode more relevant affect information.
- Transparency in the processing of temporal ambiguity: The case of embedded tense
We report the results of one acceptability rating study and two self-paced reading studies on the form-meaning mismatch in the interpretation of past-under-past in complement clauses in English. Across the three experiments, we find an off-line and on-line preference for the backward-shifted interpretation, in line with predictions of the structural approach to the ambiguity when assuming a processing preference for morphological transparent interpretation.
- Informational content vs. discourse orientation: experimental and computational perspectives
The aim of this study is to investigate how human speakers and computational language models process (i) the informational content and (ii) the discourse orientation of natural language sentences. These two dimensions of meaning have received little attention outside theoretical literature, especially in the computational linguistics domain. To help fill this void, we present the results of four experiments that exploit the specific semantics of two French adverbs, namely presque (≃ ’almost’) and à peine (≃ ’barely’), which put these two dimensions of meaning at odds. Each experiment focuses on one kind of population (humans or language models), and one kind of meaning (informational content or discourse orientation). Our results show that humans are indeed sensitive to informational content and discourse direction, as assumed in the theoretical literature. Language models exhibit a less transparent behavior. Their performances in dealing with the semantics of presque appear in line with predictions based on the way these models are trained, but this does not extend to à peine.
- Beyond Surprising: English Event Structure in the Maze
To what extent can we tease apart semantic representations and processes from other influences on processing such as probabilistic prediction? In this paper I detail two experiments testing the hypothesis that there are semantic complexity in the lexical representations of result verbs that influences reaction times above and beyond probabilistic distributions. This is done by replicating a self-paced reading study from Levinson & Brennan (2016) while also modelling lexical surprisal. Experiment 1 replicates the original result, but only in experiment 2 using the maze task does the effect emerge beyond surprisals. The more focal maze task results suggest that processing costs associated with bieventive result verbs should be accounted for by grammatical factors, in addition to probabilistic prediction.
- Investigating a shared mechanism in the priming of manner and quantity implicature
In the current paper, we investigate the existence of a shared derivation mechanism between manner and quantity implicature. As per the Gricean-inspired perspective, both manner and quantity implicature are derived in a substantially analogous fashion, relying on the consideration of alternative ways in which the speaker could have spoken, but didn't. In contrast, other accounts (e.g., grammatical accounts) of quantity implicature consider manner implicature and quantity implicature to be distinct in their derivational mechanisms.Previous studies have found that quantity implicature can prime the derivation of subsequent quantity implicature both within and between quantity implicature subtypes in a structural priming paradigm, suggesting that ad hoc, numeral and some quantity implicature are governed by the same derivational mechanism. We have applied a structural priming paradigm to the case of manner implicature to investigate 1) whether manner implicature can be primed, 2) whether manner implicature can prime manner implicature and 3) whether manner implicature can be primed by quantity implicature. Through manner-manner priming, the paper addresses the psycholinguistic reality of manner. While quantity-manner priming probes the existence of a shared derivational mechanism between the phenomena.We show that manner implicature can prime manner implicature under certain experimental circumstances and that ad hoc quantity, but not some quantity implicature can also prime manner implicature, whereas some quantity implicature cannot.
- Default biases in the interpretation of English negation, conjunction, and disjunction
Previous research has hypothesized default interpretive biases for three types of ambiguities with English logical words and, or, and not. First, disjunction (A or B) is hypothesized to be biased towards an exclusive interpretation in upward-entailing environments and an inclusive interpretation in downward-entailing environments (Levinson 2000, Chierchia 2004, Breheny et al. 2005). A negated disjunction (not A or B) is claimed to be biased towards a "neither-nor" interpretation (i.e. wide scope negation: ¬[A ^ B]) and a negated conjunction is said to be biased towards an "either-not" interpretation (i.e. wide-scope negation: ¬[A ^ B]) (Szabolcsi 2002, Szabolcsi & Haddican 2004). We tested these hypotheses within the same experimental paradigm with 149 English-speaking participants and found disjunction to be biased towards an inclusive interpretation across three different entailment environments: episodic declaratives, questions, and conditional antecedents. Our results also confirmed that English negated disjunction is biased towards a "neither-nor" (wide scope negation) interpretation but the results did not show an "either-not" bias (wide scope negation) for negated conjunction.
- Effects of instruction on semantic and pragmatic judgment tasks
Sentence judgment tasks are used often in linguistics studies. However, there is no consensus on how significant the effect of instruction is in such tasks: some argue that instruction is trivial, while others argue that they affect the way participants respond. In this study, we investigate different keywords used in sentence judgment tasks and determine which keyword best teases apart speakers' response to semantically and pragmatically licit and illicit sentences. We test this in English and Mandarin, exploring the possibility of cross-linguistic variation on how speakers respond to different keywords. Our results show that the common keywords used in semantic and pragmatic judgment tasks such as 'natural' do distinguish semantic and pragmatic violations for English speakers, but that the common Mandarin translations of these words fail to distinguish between the two types of violations. Our results highlight the need for language- and study-specific norming procedures in sentence judgment tasks.
- Modeling the Role of Polysemy in Verb Categorization
Recent work has indicated that static word embeddings can predict human semantic categories (Majewska et al. 2021). In this paper, we consider the role of polysemy in semantic categorization, by comparing sense-level embeddings with previously studied static embeddings in their prediction of human-produced categories. We find that the polysemy is crucial for predicting human categories; sense-level embeddings dramatically outperform static embeddings in predicting semantic categories. Our findings highlight the role of polysemy in semantic categorization that is exclusively based on linguistic input.
- Nonboolean Conditionals
On standard analyses, indicative conditionals behave in a Boolean fashion when interacting with and and or. We test this prediction by investigating probability judgments about sentences of the form "If A, then B {and, or} if C, then D". Our findings are incompatible with a Boolean picture. This is challenging for standard analyses of ICs, as well as for several nonclassical analyses. Some trivalent theories, conversely, may account for the data.
- Corpus evidence for the role of world knowledge in ambiguity reduction: Using high positive expectations to inform quantifier scope
Every-negation utterances (e.g., Every vote doesn't count) are ambiguous between a surface scope interpretation (e.g., No vote counts) and an inverse scope interpretation (e.g., Not all votes count). Investigations into the interpretation of these utterances have found variation: child and adult interpretations diverge (e.g., Musolino 1999) and adult interpretations of specific constructions show considerable disagreement (Carden 1973, Heringer 1970, Attali et al. 2021). Can we concretely identify factors to explain some of this variation and predict tendencies in individual interpretations? Here we show that a type of expectation about the world (which we call a high positive expectation), which can surface in the linguistic contexts of every-negation utterances, predicts experimental preferences for the inverse scope interpretation of different every-negation utterances. These findings suggest that (1) world knowledge, as set up in a linguistic context, helps to effectively reduce the ambiguity of potentiallyambiguous utterances for listeners, and (2) given that high positive expectations are a kind of affirmative context, negation use is felicitous in affirmative contexts (e.g., Wason 1961).
- The role of relevance, competence, and priors for scalar inferences
Although it is often assumed that the natural language expressions 'some' and 'or' are interpreted according to their first-order logic counterparts, in certain contexts, they receive a narrower interpretation: 'some' is strengthened to 'some, but not all', and 'or' to 'A or B, but not both'. This process is typically explained as an instance of scalar inference. To test this scalar implicature hypothesis, we collect experimental evidence for the effects and interactions of three factors that should affect the robustness of the scalar inferences of 'some' and 'or': the relevance of the stronger alternative, the speaker's competence about the alternative, and the prior probability that the alternative is true. We find that the interpretation of both triggers was affected by speaker competence, but only 'some' was also affected by prior probability, while relevance did not affect either trigger. Ultimately, our results suggest that the interdependence of the three factors is more complex than just the sum of their effects.
- You must worry! The interpretation of mustn't varies with context and verbal complement
We investigate experimentally whether American English adult speakers are influenced in their interpretation of mustn't by pragmatic context (contexts favoring lack of necessity/necessity not to readings) and/or the semantic properties of the verbal complements of the modal (verbs denoting events in the physical realm vs. verbs expressing undesirable mental activities). In an experiment combining a forced choice task and a gradient acceptability task, participants saw sentences containing mustn't and physical events/negative mental activities in lack of necessity/necessity not to contexts (e.g., You mustn't worry. The woman will give you money) They had to choose the most suitable interpretation of mustn't ('it is necessary not to'/'it is not necessary' interpretations). They then had to rate the acceptability of the sentences containing mustn't in context on a Likert scale from 1 to 7. We find that participants split into two groups: an Interdiction Group, which always treated mustn't as expressing interdiction, and a Variation Group, which tended to interpret mustn't as lack of necessity when the context favored such a reading and when the verbal complement the modal combined with was a negative mental activity. We argue that the lack of necessity reading of mustn't is obtained via pragmatic weakening from its primary interdiction reading, and that this process is sensitive to context, as well as to the cognitive difficulty of imposing or forbidding mental (but not physical) activities to others.
- Tracking the activation of scalar alternatives with semantic priming
From an utterance of Mary ate some of the deep dish, hearers frequently infer that Mary didn't eat all of the deep dish. Similarly, an utterance of The movie is good might lead hearers to conclude that the movie isn't excellent. These inferences are instances of scalar implicature (SI). The standard assumption is that SI arises via hearers' reasoning about alternative utterances that the speaker could have said, but did not. In particular, hearers are taken to consider stronger alternatives such as all (or Mary ate all of the deep dish) and excellent (or The movie is excellent) and derive their negation. In this study, we investigate the psycholinguistic reflexes of this inferential process. We use semantic priming with lexical decision to test whether lexical alternatives such as all and excellent are retrieved and activated in the processing of SI-triggering sentences. The results of our experiments indeed suggest that alternatives play a role in the processing of SI, though a number of empirical puzzles remain.
- When Transformer models are more compositional than humans: The case of the depth charge illusion
State-of-the-art Transformer-based language models like GPT-3 are very good at generating syntactically well-formed and semantically plausible text. However, it is unclear to what extent these models encode the compositional rules of human language and to what extent their impressive performance is due to the use of relatively shallow heuristics, which have also been argued to be a factor in human language processing. One example is the so-called depth charge illusion, which occurs when a semantically complex, incongruous sentence like No head injury is too trivial to be ignored is assigned a plausible but not compositionally licensed meaning (Don't ignore head injuries, even if they appear to be trivial). I present an experiment that investigated how depth charge sentences are processed by Transformer models, which are free of many human performance bottlenecks. The results are mixed: Transformers do show evidence of non-compositionality in depth charge contexts, but also appear to be more compositional than humans in some respects.
- Five degrees of (non)sense: Investigating the connection between bullshit receptivity and susceptibility to semantic illusions
Individual differences in people's tendency to see bullshit statements such as Perceptual reality transcends subtle truth as meaningful and possibly profound have become an active topic of research in judgment and decision making in recent years. However, (psycho)linguistics has so far paid little attention to the topic, despite its obvious appeal for language processing research. I present an experiment that investigated possible shared traits contributing to individual bullshit receptivity and susceptibility to semantic illusions, which occur when compositionally incongruous sentences receive plausible but unlicensed interpretations (e.g., More people have been to Russia than I have). The results show relatively little indication of an individual-level tendency to both fall for bullshit and for linguistic illusions. Implications for future psycholinguistic research into bullshit processing are discussed.
- Crosslinguistic differences on the Present Perfect Puzzle: An experimental approach
In this paper, we analyze how different temporal and referential properties of past-referring adverbials—specifically, hodiernality and deixis—are partially responsible for the crosslinguistic distribution of PAST and PERFECT markers across Dutch, Spanish, and English. To that end, we conducted an acceptability judgment task, where 160 subjects per language rated context-sentence pairs that display either a PAST or a PERFECT marker, and a temporal adverbial that is: (i) either temporally close to or temporally far from the speech time, and (ii), either deictic or not deictic. Results show that: (a) Dutch allows for its PERFECT marker to combine with any past-referring temporal adverbial, (b) Spanish only allows its PERFECT marker to combine with adverbials that locate the event temporally close to speech time, regardless of deixis, and (iii) that English prefers its PAST marker in all past-referring situations, but allows its PERFECT to combine with adverbials that are both deictic and temporally close to speech time, particularly when the adverb specifies an interval that is included in the day of utterance (e.g., this morning), as opposed to adverbs that describe an interval that includes it (e.g., this month).
- Semantics of Non-Doxastic Attitude Ascriptions from Experimental Perspective
The paper presents novel experimental data regarding reports of non-doxastic attitudes (expressed by verbs such as "wants", "fear", "is glad", and etc.) As observed by some theorists, non-doxastic attitude ascriptions differ from the ascriptions of doxastic attitudes (e.g., "believes") in that they do not support simple entailments or presuppositions of their complement clause. In particular, an ascription may intuitively change its truth-value if we alter the informational structure of the embedded clause without modifying its truth conditions. We present two experiments whose results support this observation. Experiment 1 shows that the truth-value and acceptability judgements of non-doxastic attitude ascriptions in a context generally depend on the informational structure of the embedded clause. Experiment 2 reveals that the truth-value judgements vary if we manipulate not only the "presupposition-assertion" structure of embedded clause, but also the components related to non-presuppositional entailments of the clause. This conclusion suggests that the contents on which attitude verbs operate should be represented as structured entities.
- The enduring effects of default focus in let alone ellipsis: Evidence from pupillometry.
The study of clausal ellipsis in sentence processing has revealed that comprehenders are sensitive to multiple, sometimes conflicting, pressures when recovering elided content. This paper presents a pupillometry experiment investigating how the human language processing system responds to sentences in which the location of a pitch accent clashes with global preferences for local correlates. The results are discussed in light of existing literature, including the Enduring Focus Principle, in which locations for default pitch accent continue to influence focus-sensitive processes regardless of overt markers of focus.
- The investigation of quantity implicatures during typical development: a systematic review
The present work is a systematic review on the acquisition of quantity implicatures in typically developing children. The references were selected through the PRISMA method. The criteria for eligibility were that the articles should be peer-reviewed, published articles written in English, containing empirical data on the comprehension of quantity implicatures in first language acquisition during typical development. The aim of this review is three-fold. First, to provide a picture of what empirical data tells us about the acquisition of quantity implicatures, based on both lexical and ad-hoc scales, potentially contributing to theoretical accounts of the phenomenon. Second, to analyze the methodologies that have been used to test children and their adequacy. And lastly to evaluate whether or not systematic review is an accurate analysis method for this type of varied and often complicated data. The results suggest that children improve in implicature derivation with age, especially with lexical scales, and that action-based tasks not based on meta-linguistic evaluations might be better suited to test these inferences, especially as opposed to Truth Value Judgment tasks. The fact that the systematic analysis confirms previously individuated trends in the acquisition of implicatures confirms that this is in fact a useful methodology to analyze the data, even with its limitations.