The egocentric constraint governs human language comprehension and perspective-taking
the verdict
CONTESTED PARTIAL
refutedsupported
the weight of evidence
2 sources for · 1 against
While psychological studies acknowledge that language comprehension and perspective-taking can exhibit egocentric tendencies or biases, the extent and general governing role of egocentric constraints remain contested, with some findings showing limitations or counter-evidence during predictive language tasks.
Speech disfluencies play a role in perspective-taking and audience design in human-human communication (HHC), but little is known about their impact in human-machine dialogue (HMD). In an online Namer-Matcher task, sixty-one participants interacted with a speech agent using either fluent or disfluent speech. Participants completed a partner-modelling questionnaire (PMQ) both before and after the task. Post-interaction evaluations indicated that participants perceived the disfluent agent as more competent, despite no significant differences in pre-task ratings. However, no notable differences were observed in assessments of conversational flexibility or human-likeness. Our findings also reveal evidence of egocentric and allocentric language production when participants interact with speech agents. Interaction with disfluent speech agents appears to increase egocentric communication in comparison to fluent agents. Although the wide credibility intervals mean this effect is not clear-cut. We discuss potential interpretations of this finding, focusing on how disfluencies may impact partner models and language production in HMD.
However, no notable differences were observed in assessments of conversational flexibility or human-likeness. Our
CCS Concepts: • Human-centered computing → HCI theory, concepts and models ; Interaction design theory, concepts and paradigms; Natural language interfaces; • Applied computing → Psychology. Additional Key Words and Phrases: Disfluency, Perspective-Taking, Conversational Agents ACM Reference Format: Rhys Jacka, Paola R. Peña, Sophie Leonard, Éva Székely, and Benjamin R. Cowan. 2025. Talking to...uh...um...Machines: The Impact of Disfluent Speech Agents on Partner Models and Perspective Taking. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25), July 8–10, 2025, Waterloo, ON, Canada. ACM, New York, NY, USA, 12 pages.
2 Related Work 2.1 Perspective taking and audience design in human-machine dialogue Dialogue is a collaborative joint activity [12], within which interlocutors seek to adjust the language they produce to support mutual understanding and coordination of meaning - a process termed audience design [ 5], which reflects allocentric communication, where speakers tailor their utterances based on their partner’s perspective or knowledge. This activity relies on perspective-taking, informed by assumptions of shared knowledge (termed common ground- Clark [13], Horton and Keysar [27]), partners’ communicative competence [33, 45], contextual demands [25, 37], and a partner’s cognitive state [30].
In addition to being shaped by the agent’s assigned role [39], the extent to which users adopt allocentric or egocentric language strategies may also depend on how salient the agent’s communicative capabilities appear during interaction. When users perceive a speech agent as less capable or strug- gling, they may engage in more effortful perspective-taking to accommodate their partner’s presumed difficulties[33]. 2.2 Disfluencies in human-machine dialogue Within HHD, disfluencies - a spoken language performance behaviour such as hesitation markers (uh, um), repetitions, and self-corrections - serve a communicative function.
Consequently, disfluencies have been suggested as a tool for collaborative meaning-making, playing a role in perspective-taking, where speakers adapt their language based on listener needs [3, 10]. Conversational agents that mimic human-like traits and communication behaviours may enable people to interact with them in ways similar to HHD. Research has shown that users adapt their communication styles to the behaviours of voice assistants [26, 31], highlighting the need to consider how these stylistic cues influence user engagement and overall satisfaction.
No other statistically significant effects were observed for the Communicative Flexibility or the Human Likeness subscales, which gives support for H1. 5.2.2 Scalar Modifier Use in Perspective-taking Task. There was strong evidence for the main effects of Perspective conditions on scalar modifier use. Participants were more likely to produce scalar modifiers in the Privileged Ground condition (b = 8.89, SE = 0.75, 95% CrI [7.44, 10.38]) and the Common Ground condition (b = 9.58, SE = 0.76, 95% CrI [8.10, 11.09]) compared to the One Target condition.
(a) Competence and Dependability, (b) Human- Likeness and (c ) Communicative Flexibility, before (pre) and after (post) engaging with the perspective-taking task for both fluency conditions, N=56. scalar modifier use remained uncertain, as the main effect of fluency was negative but not credibly different from zero (b = -6.35, SE = 5.28, 95% CrI [-19.39, 0.84]). Similarly, the interaction effects between Speech Agent and Perspective conditions displayed wide credibility intervals.
Cowan 6.3 Limitations & Future Directions While this study provides valuable insights into how speech disfluency and referential perspective influence language production in human–machine dialogue, it does have its contextual constraints. Much like Peña et al. [39] experimental dialogue partners, our agent was simulated using pre-recorded speech, while effective for control this likely differs from naturalistic use of real-time conversational agents, where agents dynamically respond.
Communication and miscommunication: The role of egocentric processes. 4, 1 (March 2007), 71–84. doi:10.1515/IP.2007.004 Publisher: De Gruyter Mouton Section: Intercultural Pragmatics. [30] Boaz Keysar, Dale J. Barr, Jennifer A. Balin, and Jason S. Brauner. 2000. Taking Perspective in Conversation: The Role of Mutual Knowledge in Comprehension. https://journals.sagepub.com/doi/abs/10.1111/1467-9280.00211 [31] Katerina Koleva, Maurizio Vergari, Tanja Kojić, Sebastian Möller, and Jan-Niklas Voigt-Antons. 2024. Influence of Personality and Communication Behavior of a Conversational Agent on User Experience and Social Presence in Augmented Reality.
Although previous research has demonstrated that language comprehension can be egocentric, there is little evidence for egocentricity during prediction. In particular, comprehenders do not appear to predict egocentrically when the context makes it clear what the speaker is likely to refer to. But do comprehenders predict egocentrically when the context does not make it clear? We tested this hypothesis using a visual-world eye-tracking paradigm, in which participants heard sentences containing the gender-neutral pronoun
They
(e.g.
They would like to wear…
) while viewing four objects (e.g. tie, dress, drill, hairdryer). Two of these objects were plausible targets of the verb (tie and dress), and one was stereotypically compatible with the participant's gender (tie if the participant was male; dress if the participant was female). Participants rapidly fixated targets more than distractors, but there was no evidence that participants ever predicted egocentrically, fixating objects stereotypically compatible with their own gender. These findings suggest that participants do not fall back on their own egocentric perspective when predicting, even when they know that context does not make it clear what the speaker is likely to refer to.
pmc R Soc Open Sci R Soc Open Sci 2753 rsos RSOS Royal Society Open Science 2054-5703 The Royal Society PMC10716638 PMC10716638.1 10716638 10716638 38094271 10.1098/rsos.231252 rsos231252 1 1001 42 205 Psychology and Cognitive Neuroscience Research Articles Evidence against egocentric prediction during language comprehension Evidence against egocentric prediction during language comprehension http://orcid.org/0000-0001-6027-8109 Corps Ruth E.
2023 https://creativecommons.org/licenses/by/4.0/ Published by the Royal Society under the terms of the Creative Commons Attribution License http://creativecommons.org/licenses/by/4.0/ , which permits unrestricted use, provided the original author and source are credited. Although previous research has demonstrated that language comprehension can be egocentric, there is little evidence for egocentricity during prediction. In particular, comprehenders do not appear to predict egocentrically when the context makes it clear what the speaker is likely to refer to. But do comprehenders predict egocentrically when the context does not make it clear?
These findings suggest that participants do not fall back on their own egocentric perspective when predicting, even when they know that context does not make it clear what the speaker is likely to refer to.
prediction , perspective-taking , visual-world paradigm , gender identity Leverhulme Trust http://dx.doi.org/10.13039/501100000275 RPG-2018-259 pmc-status-qastatus 0 pmc-status-live yes pmc-status-embargo no pmc-status-released yes pmc-prop-open-access yes pmc-prop-olf no pmc-prop-manuscript no pmc-prop-legally-suppressed no pmc-prop-has-pdf yes pmc-prop-has-supplement no pmc-prop-pdf-only no pmc-prop-suppress-copyright no pmc-prop-is-real-version no pmc-prop-is-scanned-article no pmc-prop-preprint no pmc-prop-in-epmc yes pmc-license-ref CC BY 1 .
Introduction There is much research demonstrating that language comprehension can be egocentric, with listeners initially comprehending from their own perspective (see [ 1 ] for a review). But there is little evidence for egocentricity during prediction [ 2 , 3 ]. However, these studies have used a visual-world paradigm in which the comprehender sees only one entity that the speaker is likely to refer to, and other entities that are implausible. As a result, the comprehender has a strong sense of what type of entity the speaker is likely to refer to.
In this paper, we ask what happens when the comprehender does not know what the speaker is likely to refer to because there is more than one plausible entity. Does the comprehender now predict from their own (egocentric) perspective? There is much evidence that speakers predict what a speaker is likely to say. For example, Altmann & Kamide [ 4 ] found that participants fixated a picture of a cake (rather than other inedible objects) earlier and for longer when they heard a speaker say The boy will eat… compared to when they heard the speaker say The boy will move .
than in Creel [ 6 ] and Borovsky & Creel [ 7 ], there was only one on-screen entity that was compatible with the linguistic context and with gender stereotypes (e.g. a tie if a male speaker said I would like to wear …). But what happens when there is more than one compatible entity? Do comprehenders then predict egocentrically and assume that the speaker is likely to refer to an entity that is compatible with their own perspective (in this case, their own gender)? In fact, there is much evidence for egocentricity during language comprehension.
Even though participants knew that the confederate had no knowledge of the small candle, they often considered it as a potential referent when the confederate said Now put the small candle above it , and in fact reached for these objects nearly one-fourth of the time. These findings suggest that participants' egocentric biases (i.e. the intrusion of their own perception) affect language comprehension. Thus, although Corps et al . [ 2 ] found that participants predicted consistently, in line with the agent's perspective, participants may predict egocentrically when the agent's perspective does not make it clear what to predict.
Second, previous studies have also focused on bottom-up comprehension, while we focused on top-down prediction. The findings therefore suggest that top-down and bottom-up processing draw on different mechanisms. Consistent with this suggestion, Barr [ 16 ] found that participants were four times more likely to fixate objects in common ground (visible to both the participant and their partner) than objects in privileged ground (visible only to the participant) before they started
This study re-evaluates egocentric behaviour in perspective-taking during referential communication, focusing on the role of stimulus selection. The debate centres on whether listeners prioritize their own perspective before considering the speaker's (late integration) or adopt the speaker's perspective from the outset (early integration). Previous early integration accounts have suggested that late integration findings might be due to the use of target expressions where hidden distractors—objects visible only to the listener—are better referential matches than the actual targets, though this claim lacked empirical testing. To investigate this, 33 neurotypical English-speaking adults participated in a questionnaire-based study. They were presented with ambiguous expressions and asked to indicate which object the speaker was likely referring to, choosing among "Target," "Hidden Distractor," or "Both equally likely," and rated their confidence levels. The items used were from prior studies by Keysar et al. (2003) and Hawkins et al. (2021). The results, analysed using chi-square tests, indicated that participants’ responses were not random but favoured specific interpretations for seven out of eight items, and their confidence in their options was over 80%. Listeners do not make errors because the target expression consistently favours the Hidden Distractor over the Target (cf. Brown-Schmidt & Hanna, 2011). These results highlight a more complex picture regarding the source of listeners’ errors. The study proposes a new classification of items in perspective-taking tasks along with its predictions and next steps for verifying whether this classification of items is causally related to listeners’ performance in perspective-taking tasks.
Kastanas Dimitrios1 & Katsos Napoleon1 1Department of Theoretical and Applied Linguistics, University of Cambridge, United Kingdom Towards a re-evaluation of egocentric behaviour in perspective-taking: the role of stimuli selection Introduction • Visual Perspective Taking (VPT) involves inferring visual or spatial properties of a scene relative to another person or position (Keysar et al., 2003). Results • The Target item is visible to both the speaker and the listener. • The Hidden distractor is visible only to the listener. • Occlusions block the speaker’s view, so the items behind the curtains are visible only to the listener. • Analysis: Chi-square test of independence.
• Analysis of error patterns across neuroimaging, eye-tracking and behavioural paradigms is required for verifying whether our proposed classification is causally related to listeners’ performance in VPT tasks. References Figure 1: The Director Task (Keysar et al., 2003). Adapted from Hawkins et al. (2021). Previous work Acknowledgements Figure 3: Results from our experiment. Target refers to the target item, Hidden Distractor refers to the occluded object from the speaker’s perspective but accessible to the listener, Both equally likely refers to both the Target and Hidden Distractor. Occlusion • Brown-Schmidt, S. & Hanna, J. E.
Target: roll of tape Hidden Distractor: cassette tape Occlusion: curtains Director: speaker (male) Matcher: listener (female) • Eye-tracking, neuroimaging and behavioural data have shown that listeners make a substantial number of errors in this task (Brown-Schmidt & Hanna, 2011). • Prominent theories: • Anchor & Adjustment (Keysar et al., 2003): egocentric processing. • Division of labour (Heller et al., 2016): simultaneous activation of one’s own and the other’s perspective. • Limitation: • Why do listeners make errors in the first place? • The Hidden Distractor is consistently a better referent for the target expression (Hanna & Brown-Schmidt, 2011; Heller et al., 2016).
Everything we examined (3)
This check searched the claim as stated. It did not run a separate search for evidence against it.