Certain monophthongs are universally more acoustically distinguishable than others
the verdict
INSUFFICIENT LEANING
refutedsupported
the weight of evidence
2 sources for · 0 against
The retrieved literature notes general perceptual biases toward certain acoustically extreme vowels and universal markers of difficulty for specific vowel contrasts, but does not fully establish a broader claim about all monophthongs being universally distinguishable.
This study aims to investigate the perception and production of the English /ɪ/–/iː/ vowel contrast by Cypriot Greek speakers of English as a second language (L2). The participants completed a classification test in which they classified the L2 vowels in terms of their first language (L1) categories, a discrimination test in which they distinguished the members of the vowel contrast, and a production test in which they produced the target vowels. The results showed that they classified both L2 /ɪ/–/iː/ mostly in terms of L1 /i/, which denotes the formation of a completely overlapping contrast according to the theoretical framework of the Universal Perceptual Model (UPM), and that they could hardly distinguish the vowel pair. In addition, their productions deviated in most acoustic parameters from the corresponding productions of English controls. The findings suggest that /ɪ/–/iː/ may carry a universal marker of difficulty for speakers with L1s that do not possess this contrast. This distinction is difficult even for experienced L2 speakers probably because they had never been exposed to naturalistic L2 stimuli and they do not use the L2 that much in their daily life. Finally, the study verifies UPM’s predictions about the discriminability of the contrast and extends the model’s implications to speech production; when an L2 vowel contrast is perceived as completely overlapping, speakers activate a (near-) unified interlinguistic exemplar in their vowel space, which represents both L2 vowels.
In addition, Llompart and Reinisch [ 39 ] reported that English /ɪ/–/iː/ was an easy contrast for German speakers as it noted robust perception and better differentiation during production compared to /ɛ/–/æ/. Speech perception is linked to speech production and, therefore, any perceptual deficits are expected to also appear in production. Several psycholinguistic theories attempt to define this link. The motor theory [ 40 , 41 ] suggests that speech perception occurs by detecting intended vocal tract gestures rather than spoken speech. These gestures act as motor commands, providing certain instructions to the articulators.
The participants were instructed to sit in front of the monitor and listen to the stimuli through the PC loudspeakers. Then, they were asked to click on the script label that was acoustically the most similar exemplar to the vowel they heard. The labels included the Greek orthographic representation of the 5 CG vowels, namely, “ι”, “ε”, “α”, “ο”, and “ου”. This is because Greek is orthographically transparent and vowels can be reliably presented using orthography, as seen in [ 54 ]. The speakers classified a total number of
The participants listened through the headphones to a triad of vowels from the PC loudspeakers and were asked to choose whether the middle vowel (X) was the same as the first (A) or the second vowel (B) by clicking on the appropriate label. Each vowel pair appeared in four possible configurations, namely, AAB, ABB, BBA, and BAA, and they discriminated a total number of 48 items (4 trials × 3 repetitions × 2 voices × 1 contrast + 24 filler words); we used tokens from two out of three speakers in this test. The X token was always acoustically different from the A and B tokens to avoid a solely auditory decision. The interstimulus interval was 1 s and the intertrial interval 500 m.s.
The end of the vowel (V) and the beginning of the noise of the second consonant /d/ indicated the vowel’s last point. The vowel formants were measured at their midpoint (50%). The vowel durations were extracted manually through the labeling of the initial and final points of each vowel token. The duration of the vowels was estimated from the measurement of the interval between the starting and ending point of the vocalic part. The normalization of F1 and F2 was produced using the Lobanov method in R ( vowels package; [ 60 ]). 4.2. Results The results of the production test indicated that English /iː/ of CG speakers was acoustically close to the same vowel of English speakers.
In contrast, English /ɪ/ of CG speakers was acoustically between /iː/ and /ɪ/ of English speakers. Furthermore, there was an important overlap between the two English vowels produced by CG speakers. In addition, CG speakers produced the English vowels with similar durations to native English speakers. Figure 3 and Figure 4 present the F1 × F2 and duration of English /ɪ/ and /iː/ as produced by CG and English (RP) speakers. Figure 3 F1 × F2 of English /ɪ/ and /iː/ as produced by CG and RP speakers. Figure 4 Duration of English /ɪ/ and /iː/ as produced by CG and RP speakers. We fitted Bayesian regression models in R to analyze the production data.
As Spanish and CG have a similar vowel system, perhaps, the learning of this contrast by the participants of this study falls within these developmental stages. Finally, in the lens of UPM’s framework, L2 English /ɪ/ and /iː/ can be characterized as disoriented since they are not acoustically close to the productions of native speakers. Several factors determine the orientation such as L1–L2 use, exposure to the L2, phonetic training, etc. The results of this study allow for drawing some general conclusions about the perception–production relationship. This relationship is reflected in the following way.
As /ɪ/ and /iː/ comprise two distinct categories in English, one would expect that their main acoustic parameters would maintain a consistent distance from each other in order to be produced distinctly. While this is the case for native /ɪ/ and /iː/, L2 /ɪ/ and /iː/ acoustically overlap in terms of F1 and, in general, their acoustic spectral distance is smaller compared to the corresponding distance of native productions. This practically means that the members of this completely overlapping contrast, which is the most difficult according to UPM, will hardly be produced with native or near-native properties; they may be produced with similar acoustic properties.
Behavioral studies examining vowel perception in infancy indicate that, for many vowel contrasts, the ease of discrimination changes depending on the order of stimulus presentation, regardless of the language from which the contrast is drawn and the ambient language that infants have experienced. By adulthood, linguistic experience has altered vowel perception; analogous asymmetries are observed for non-native contrasts but are mitigated for native contrasts. Although these directional effects are well documented behaviorally, the brain mechanisms underlying them are poorly understood. In the present study we begin to address this gap. We first review recent behavioral work which shows that vowel perception asymmetries derive from phonetic encoding strategies, rather than general auditory processes. Two existing theoretical models-the Natural Referent Vowel framework and the Native Language Magnet model-are invoked as a means of interpreting these findings. Then we present the results of a neurophysiological study which builds on this prior work. Using event-related brain potentials, we first measured and assessed the mismatch negativity response (MMN, a passive neurophysiological index of auditory change detection) in English and French native-speaking adults to synthetic vowels that either spanned two different phonetic categories (/y/vs./u/) or fell within the same category (/u/). Stimulus presentation was organized such that each vowel was presented as standard and as deviant in different blocks. The vowels were presented with a long (1,600-ms) inter-stimulus interval to restrict access to short-term memory traces and tap into a "phonetic mode" of processing. MMN analyses revealed weak asymmetry effects regardless of the (i) vowel contrast, (ii) language group, and (iii) MMN time window. Then, we conducted time-frequency analyses of the standard epochs for each vowel. In contrast to the MMN analysis, time-frequency analysis revealed significant differences in br
Collectively, these findings suggest that early-latency (pre-attentive) mismatch responses may not be a strong neurophysiological correlate of asymmetric behavioral vowel discrimination. Rather, asymmetries may reflect differences in neural processing efficiency for vowels with certain inherent acoustic-phonetic properties, as revealed by theta oscillatory activity.
These generic speech biases are evident in studies showing that young infants exhibit robust listening preferences for some speech sounds over others ( Polka and Bohn, 2011 ; Nam and Polka, 2016 ), and that some phonetic contrasts are poorly distinguished early on ( Polka et al., 2001 ; Best and McRoberts, 2003 ; Larraza et al., 2020 ) or show directional asymmetries in discrimination ( Polka and Bohn, 2003 , 2011 ; Kuhl et al., 2006 ; Pons et al., 2012 ; Nam and Polka, 2016 ). The present research aims to improve our understanding of the neural mechanisms and processes underlying vowel perception biases observed in adults.
It has been known for years that, early in development, infant perception is biased toward articulatorily and acoustically extreme vowels. These findings have been reviewed and discussed extensively by Polka and Bohn (2003 , 2011) , and have also been reinforced in recent meta-analyses ( Tsuji and Cristia, 2017 ; Polka et al., 2019 ). Evidence supporting this view initially emerged from research revealing that infants show robust directional asymmetries in vowel discrimination tasks.
For example, when producing /i/ (the highest front vowel) F 2 , F 3 , and F 4 converge, when producing /y/ (the highest front rounded vowel) F 2 and F 3 converge, when producing /a/ (the lowest back vowel) and /u/ (the highest back vowel) F 1 and F 2 converge. These convergence points have also been referred to as “focal points” ( Boë and Abry, 1986 ). According to the Dispersion-Focalization Theory, the strong tendency for vowel systems to select members found at the extremes of articulatory/acoustic vowel space is driven by two factors. First, dispersion ensures that vowels are acoustically distant from one another within vowel space, which enhances perceptual differentiation.
The authors interpret these findings as suggesting that /ø/-/o/ has a different phonological status than the other vowel contrasts tested, and that this difference in phonological representation might explain the neural processing differences. Specifically, they postulate that the place of articulation feature [coronal] is universally absent or “underspecified” from phonemic representations in the lexicon–for vowels and consonants alike.
Everything we examined (2)
This check searched the claim as stated. It did not run a separate search for evidence against it.