The Differential Effects of Subtitles on the Comprehension of Native English Connected Speech Varying in Types and Word Familiarity

Saved in:
Bibliographic Details
Title: The Differential Effects of Subtitles on the Comprehension of Native English Connected Speech Varying in Types and Word Familiarity
Language: English
Authors: Wong, Simpson W. L. (ORCID 0000-0002-6606-6382), Lin, Cherry C. Y., Wong, Isabella S. Y., Cheung, Anisa
Source: SAGE Open. Apr-Jun 2020 10(2).
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: http://sagepub.com
Peer Reviewed: Y
Page Count: 13
Publication Date: 2020
Document Type: Journal Articles
Reports - Research
Education Level: Secondary Education
Descriptors: Second Language Learning, English (Second Language), Translation, Connected Discourse, Video Technology, Decoding (Reading), Listening Comprehension Tests, Accuracy, Phonology, Secondary School Students, Adolescents, Foreign Countries, Auditory Perception
Geographic Terms: Hong Kong
DOI: 10.1177/2158244020924378
ISSN: 2158-2440
Abstract: Connected speech produced by native speakers poses a challenge to second language learners. Video subtitles have been found to assist the decoding of English connected speech for learners of English as a foreign language (EFL). However, the presence of subtitles may divert the listeners' attention to the visual cues while paying less attention to the speech signals. To test this proposal, we employed a bi-modal audio-visual listening test and examined whether EFL listeners were able to correctly identify the connected speech when misleading subtitles were present. We further tested whether connected speech with words of lower frequency further reduced the accuracy rate. Twenty-eight adolescent EFL learners, all with more than 10 years of experiences in learning English in schools, were tested with three major types of connected speech phonological processes, namely assimilation, elision, and juncture. The results of statistical analyses showed that matched and mismatched subtitles facilitated the comprehension of both familiar and unfamiliar connected speech. Error analyses revealed the degree of item-specific variations across the three types of connected speech processes as well as across the three subtitling conditions. This research provides insights on the immediate and long-term impact of subtitles on the decoding of English connected speech.
Abstractor: As Provided
Entry Date: 2020
Accession Number: EJ1259676
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFMs7G3X1UDtBXB6upZl-IeAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDH22r9zhV6zEmPkk8gIBEICBmxPEHvhspZJmJkhHX3HhFbJNWk4hGWwmoESMYcA1u-RmO9wIrA7XPl-PVRqpGSLqL7Vqc5KJ7Iulfs9UpuV_QjQZOrq_yPgyEmGO28F7fb7wHgTDAdOiirqqCiOBe6tU6peaHsDBz-DYu8E7yNkZyitareSfOTjo8Mq02qCN1kDvX2Ss2umJaGStoMkylOovboXGrNNXFOc1oSl_
Text:
  Availability: 1
  Value: <anid>AN0144336864;[kbz6]01apr.20;2020Jul03.02:08;v2.2.500</anid> <title id="AN0144336864-1">The Differential Effects of Subtitles on the Comprehension of Native English Connected Speech Varying in Types and Word Familiarity </title> <p>Connected speech produced by native speakers poses a challenge to second language learners. Video subtitles have been found to assist the decoding of English connected speech for learners of English as a foreign language (EFL). However, the presence of subtitles may divert the listeners' attention to the visual cues while paying less attention to the speech signals. To test this proposal, we employed a bi-modal audio-visual listening test and examined whether EFL listeners were able to correctly identify the connected speech when misleading subtitles were present. We further tested whether connected speech with words of lower frequency further reduced the accuracy rate. Twenty-eight adolescent EFL learners, all with more than 10 years of experiences in learning English in schools, were tested with three major types of connected speech phonological processes, namely assimilation, elision, and juncture. The results of statistical analyses showed that matched and mismatched subtitles facilitated the comprehension of both familiar and unfamiliar connected speech. Error analyses revealed the degree of item-specific variations across the three types of connected speech processes as well as across the three subtitling conditions. This research provides insights on the immediate and long-term impact of subtitles on the decoding of English connected speech.</p> <p>Keywords: listening; subtitles; connected speech; learning English as a foreign language; error analysis</p> <hd id="AN0144336864-2">Introduction</hd> <p>Learners of English as a foreign language (EFL) have been found to experience difficulties in comprehending English speech uttered by native speakers even after prolonged listening training in schools ([<reflink idref="bib52" id="ref1">52</reflink>]). One reason is that most of the English speech (presented by non-native English-speaking teachers or pedagogically designed audios spoken by native speakers) appears in EFL classrooms is uttered at a slower rate to accommodate the EFL learners' abilities. To ensure English words are properly learned, they are presented to EFL students in the citation form. Citation form, also known as dictionary form, is the way English words are presented in audio dictionary, that is, each word is presented in isolation and the pronunciation of the consonants and vowels are unaltered by surrounding phonological environment. According to [<reflink idref="bib33" id="ref2">33</reflink>], approximately 60% of words in a corpus of 88,000 American English word tokens were spoken in reduced forms (i.e., speech signals are altered or reduced as compared with the utterance of a word in isolation). The more reduction and alternation of speech signal, the harder it is for EFL learners to perceive, as more demanding phonological reconstruction is required ([<reflink idref="bib25" id="ref3">25</reflink>]).</p> <p>Of the numerous ways to improve listening comprehension, one of the most commonly adopted means is through the use of subtitles, which are frequently displayed in movies as closed captions and can be understood as a textual form of visual aid shown on screen. The usefulness of subtitles in video clips has attracted some scholarly attention in recent decades, and it appears researchers have reached the consensus that they can benefit different cohorts of learners, including normal-hearing children and adults, as well as those with hearing impairments ([<reflink idref="bib27" id="ref4">27</reflink>]). In some empirical studies, subtitles are further classified as standard/inter-language (audio in a second language (L2) + subtitles in the first language (L1)), intra-language (audio in L2 + subtitles in L2), or reversed (dubbed audio in L1 + subtitles in L2). Through comparing performances between various subtitled conditions, studies on the effects of subtitling generally report a positive influence of intra-lingual (L2) subtitles in English (or other language) comprehension among second language learners (e.g., [<reflink idref="bib3" id="ref5">3</reflink>]). Similar findings of the positive effect of subtitles on English comprehension in animations, cartoons, movies, and TV series have been reported ([<reflink idref="bib60" id="ref6">60</reflink>]). Empirical studies on subtitling conducted so far mainly focused on the beneficial effects of subtitles on listening comprehension performance (e.g., [<reflink idref="bib31" id="ref7">31</reflink>]), vocabulary acquisition, and recall of materials (e.g., [<reflink idref="bib11" id="ref8">11</reflink>]). Despite the presence of benefits provided by subtitles to EFL learners, subtitles are viewed differently by different EFL learners. Generally, lower-intermediate and intermediate learners do not have a choice by relying more on subtitles to comprehend the connected speech ([<reflink idref="bib56" id="ref9">56</reflink>]). It was revealed in an eye tracking study that the learners fixated on the subtitles 68% of the time. The interview data further suggested that these learners found reading subtitles easier than listening for extracting meaning and that reading also helped them more readily segment words from the stream of speech. These learners did not employ progressively fewer subtitles during the course of movie viewing ([<reflink idref="bib47" id="ref10">47</reflink>]) and their learning goal was more oriented to improving reading rather than listening skills ([<reflink idref="bib7" id="ref11">7</reflink>]; [<reflink idref="bib8" id="ref12">8</reflink>]). In the long-run, without engagement with both the spoken and the subtitles signals, these less proficient EFL learners will hardly develop the ability to map phonological characteristics of sounds and the semantic meaning of words uttered in connected speech. Even worse, these groups of learners are less confident about their listening ability ([<reflink idref="bib54" id="ref13">54</reflink>]).</p> <p>Although general English language proficiency of the listeners plays a critical role in the use of subtitles ([<reflink idref="bib61" id="ref14">61</reflink>]), what awaits to be addressed is the large individual variations within the specific proficiency group. We speculate whether the nature of connected speech has an effect on listener's performances. The two salient features about connected speech are the types of connected speech phonological processes as well as the familiarity of words within the connected speech. An introduction of the two features and relevant research findings is provided below.</p> <hd id="AN0144336864-3">Processing Efficiency of Familiar and Unfamiliar Connected Speech</hd> <p>As individual words are represented differently in connected speech, multiple phonological representations for individual words exist. A large body of research has shown a strong positive correlation between word frequency and the efficiency of auditory lexical access during spoken word recognition (e.g., [<reflink idref="bib13" id="ref15">13</reflink>]; [<reflink idref="bib14" id="ref16">14</reflink>], [<reflink idref="bib16" id="ref17">16</reflink>]; [<reflink idref="bib28" id="ref18">28</reflink>]; [<reflink idref="bib43" id="ref19">43</reflink>]), with the most frequently occurring words (high-production frequency) having the highest activation strength. In particular, the frequency at which the phonological variants are perceived in the context of continuous speech is shown to determine the efficiency of connected speech processing ([<reflink idref="bib15" id="ref20">15</reflink>]; [<reflink idref="bib24" id="ref21">24</reflink>]; [<reflink idref="bib46" id="ref22">46</reflink>]). For example, words that are frequently heard are recognized faster and more accurately in lexical decision tasks, as compared with those that are infrequently heard. Studies have also shown a reduced priming effect for low production frequency, compared with high production frequency, connected speech ([<reflink idref="bib48" id="ref23">48</reflink>]). Similarly, [<reflink idref="bib45" id="ref24">45</reflink>] demonstrated that higher production frequency is beneficial to the comprehension of reduced variants. Although the above studies examined the impact of word frequency on spoken word recognition in a uni-modal listening environment, it remains uncertain whether a similar word frequency effect can be observed in a bi-modal audio-visual listening context. The present study thus aims to fill this research gap.</p> <hd id="AN0144336864-4">The Consistency of Performances Across Phonological Processes of Connected Speech</hd> <p>For EFL learners, their first language (L1) background and L1 phonological system has been found to restrict the acquisition of L2 connected speech ([<reflink idref="bib58" id="ref25">58</reflink>]). Among Chinese learners of English in particular, three connected speech phonological processes, namely assimilation, elision, and juncture, are largely affected by the cross-linguistic differences between Chinese and English sound systems. <emph>Assimilation</emph> is defined as the process in which phonemes assimilate to the place of the neighboring consonant while retaining their original voicing characteristics ([<reflink idref="bib18" id="ref26">18</reflink>]). For example, the pronunciation of the word phrase /ten players/ is changed from [tεn 'pleɪɚz] to [tem 'pleɪɚz] in running speech. <emph>Assimilation</emph> is absent in Chinese utterances and therefore poses a challenge for Chinese EFL learners in English speech processing ([<reflink idref="bib41" id="ref27">41</reflink>]). In [<reflink idref="bib36" id="ref28">36</reflink>] study, 50 Chinese university sophomores majoring in English exhibited below-chance performances when identifying cases of assimilation. Another connected speech process that poses a challenge to Chinese EFL learners is <emph>(contextual) elision</emph>, in which a vowel or consonant occurring within either the body of a word or at a junction of word boundaries is lost ([<reflink idref="bib18" id="ref29">18</reflink>]). This process is often evidenced in word phrases such as /iced tea/ (its citation form is /aɪst 'ti/), which is pronounced as [aɪs 'ti] (similar to /ice tea/). The process of <emph>elision</emph> is also absent in Chinese utterances ([<reflink idref="bib41" id="ref30">41</reflink>]). In a study conducted in Taiwan, Chinese EFL sophomores with low, mid, and high English proficiency levels correctly identified only 44%, 69%, and 77% of elision cases, respectively ([<reflink idref="bib34" id="ref31">34</reflink>]). Finally, <emph>juncture</emph> is the connected speech process referring to the removal of a clear boundary between two syllables. For example, /a name/ is transformed from its citation form [ə neɪm] to its reduced form [ən eɪm] (like /an aim/). The obligatory gap between words to demarcate information is often absent in connected speech. In order to resolve juncture, a combination of contextual information and subtle cues in the speech signal is utilized by the listener to discern individual words ([<reflink idref="bib51" id="ref32">51</reflink>]). Moreover, it is suggested that English connected speech is like "singing in a legato way," whereas Chinese connected speech is "articulated in a staccato way" ([<reflink idref="bib22" id="ref33">22</reflink>], p. 71; [<reflink idref="bib49" id="ref34">49</reflink>], p. 144). Junctures in the Chinese language can be clearly perceived due to the language's syllable-timed pattern and the fact that the majority of words share the same degree of emphasis. As a result, Chinese EFL learners generally find it hard to identify junctures in English. When junctures are located before unstressed words or between words whose onsets or codas can be joined by the preceding or following words, they are almost unperceivable ([<reflink idref="bib36" id="ref35">36</reflink>]). [<reflink idref="bib51" id="ref36">51</reflink>] empirical study indicated that Hong Kong listeners were only able to correctly identify 60% of junctures presented in British English speech. Given the suboptimal performances in Chinese EFL learners when identifying these three connected speech processes, it is therefore crucial to focus on these aspects in the present study.</p> <hd id="AN0144336864-5">The Use of Subtitles in Non-immersive English Environments</hd> <p>When learning English in a non-immersive English environment, EFL learners receive limited exposure to native English ([<reflink idref="bib4" id="ref37">4</reflink>]). Thus, these EFL learners have to rely heavily on listening materials in school settings as well as mass media, such as movies and TV channels ([<reflink idref="bib53" id="ref38">53</reflink>]). As a common practice in Hong Kong, classroom materials are usually presented alongside a transcript, and the English-speaking movies are required to show Chinese subtitles ([<reflink idref="bib10" id="ref39">10</reflink>]). Ironically, even if native English speech is readily available in non-English-speaking countries, listening skill is still not as good as expected ([<reflink idref="bib57" id="ref40">57</reflink>]). This situation has motivated us to question the exact influences of subtitles on the decoding of connected speech.</p> <p>In addition to the availability of listening materials, other factors such as language proficiency (e.g., [<reflink idref="bib40" id="ref41">40</reflink>]), cognitive load of the multimedia (e.g., [<reflink idref="bib56" id="ref42">56</reflink>]), and design of subtitles (e.g., [<reflink idref="bib12" id="ref43">12</reflink>]) have also been investigated to understand the usefulness of subtitles. Still, a beneficial role of subtitles was assumed in these studies and the investigative focus remained on their degree of facilitation. However, as implied by the poor listening comprehension performances in EFL learners with various language backgrounds who can readily access subtitles (e.g., Chinese: [<reflink idref="bib12" id="ref44">12</reflink>]; French: [<reflink idref="bib30" id="ref45">30</reflink>]; Spanish: [<reflink idref="bib42" id="ref46">42</reflink>]; Russian: [<reflink idref="bib56" id="ref47">56</reflink>]), we suspect that subtitles may not facilitate the acquisition of connected speech under all circumstances and reducing any anxieties triggered when listening to non-L1 speech ([<reflink idref="bib2" id="ref48">2</reflink>]), the negative aspects of subtitles require further investigation.</p> <p>In an audio-visual environment, it is well known that the visual modality is dominant, as visual information is processed significantly faster than auditory information; hence, auditory processing is often subdued ([<reflink idref="bib38" id="ref49">38</reflink>]). Moreover, if subtitles are treated as <emph>obligatory</emph> rather than <emph>supplementary</emph> during listening, EFL learners are likely to experience listening difficulties when no subtitles are provided ([<reflink idref="bib29" id="ref50">29</reflink>]; [<reflink idref="bib47" id="ref51">47</reflink>]). Due to the more enduring nature on the screen of the subtitles than the connected speech ([<reflink idref="bib32" id="ref52">32</reflink>]), ESL learners may develop the tendency to read subtitles for extracting meaning and/or segmenting words in the stream of connected speech ([<reflink idref="bib56" id="ref53">56</reflink>]). One negative effect we speculate is "subtitle-dependency," whereby learners are more likely to trust what they read from subtitles over what they hear from the corresponding audio when they are presented with mismatched subtitles, that is, when the content of the subtitles does not match the content of the audio. Hence, we predict that EFL learners are more prone to listening errors when mismatched subtitles are provided.</p> <hd id="AN0144336864-6">The Present Study</hd> <p>In this section, we shift our focus and provide a brief description of the context of the present study. In Hong Kong, English language is a compulsory subject and is introduced in the curriculum as early as 3 years of age. Despite the official status of the language, it is not commonly spoken in daily life conversation. Students are, however, highly motivated to get good results in English, as proficiency in English directly affects the results of public examination and thus their chance of university admission. Apart from academic use, students in Hong Kong are also exposed to spoken native English through various forms of media such as movies and songs, as well as having lessons with their Native English-speaking teachers in school.</p> <p>Despite the rich environment available for English learning, it is surprising to note that Hong Kong undergraduates who have been learning English for over 15 years were still unable to fully decode connected speech spoken by native English speakers ([<reflink idref="bib59" id="ref54">59</reflink>]); furthermore, in an earlier study, [<reflink idref="bib52" id="ref55">52</reflink>] found that Hong Kong Cantonese speakers who were English language teacher trainees made multiple perceptual errors when decoding English connected speech. It is essential to understand whether high school EFL students in Hong Kong are dependent on subtitles during listening comprehension. Although this enhances immediate decoding of speech, the use of subtitles may hinder auditory perceptual learning of connected speech, causing long-term suboptimal performances in connected speech processing. This potential negative effect motivates us to study how connected speech is processed when subtitles are presented. Do listeners trust their ears or their eyes more? In addition, we compare listeners' accuracy in decoding speech across three types of connected speech (assimilation, elision and juncture) in order to reveal their unique characteristics and the level of difficulties they represent to listeners.</p> <p>In light of the aforementioned research gaps, the four research questions of the present study are listed as follows:</p> <p></p> <ulist> <item> <bold> Research Question 1: </bold> How do EFL learners perform in decoding connected speech with and without the presence of subtitles?</item> <p></p> <item> <bold> Research Question 2: </bold> Do matched subtitles facilitate, and mismatched subtitles conversely interfere with, connected speech decoding?</item> <p></p> <item> <bold> Research Question 3: </bold> Does the familiarity of the words affect decoding performance?</item> <p></p> <item> <bold> Research Question 4: </bold> Are performance levels consistent across the three types of connected speech processes (assimilation, elision, and juncture)?</item> </ulist> <hd id="AN0144336864-7">Method</hd> <p></p> <hd id="AN0144336864-8">Participants</hd> <p>A total of 28 Cantonese-speaking EFL learners (16 females; 12 males) aged between 15 and 16 years were recruited for the current study. The students had no reported difficulties in learning or language acquisition. All participants were 10th graders from a local secondary school in which Cantonese is used as the medium-of-instruction for non-English subjects. These students were recruited by the second author, who was their English teacher. Based on the benchmark against the territory-wide English standard specified by the Hong Kong [<reflink idref="bib23" id="ref56">23</reflink>] as well as via a medium-level listening test ([<reflink idref="bib20" id="ref57">20</reflink>]), the participants were rated "average" in terms of English proficiency.</p> <hd id="AN0144336864-9">Procedure</hd> <p>This work was conducted with the formal approval of Human Research Ethics Committee at (The Education University of Hong Kong). After obtaining informed consent from participants and their parents, respectively, the experiment was carried out in a classroom in the participants' school. All students were presented with the same set of testing materials in a group-testing situation. The uni-modal connected speech decoding task (presented to the participants as an unseen dictation task) was administered first, followed by the bi-modal connected speech decoding test. The test administration order was designed in such a way to avoid the influence of prior exposure to the subtitles in the bi-modal task on the uni-modal task.</p> <hd id="AN0144336864-10">Preparation of Audio</hd> <p>We first generated a list of minimal pairs of connected speech while taking the English proficiency levels of the participants into consideration. The list was developed by our research team and was only adopted in the current study. The stimuli were recorded in a soundproof room using a high-quality Roland R-09HR recorder and digitized at a sample rate of 44.1 kHz with a 16-bit amplitude resolution. The set of stimuli was spoken aloud by a 25-year-old native female speaker who was born and raised in New York. This speaker also advised on the suitability of the test items for the use in this study. A General American (GA) accent was chosen for this study as Hollywood movies and TV programs are popular among Hong Kong Chinese adolescents and they are commonly exposed to this accent. Furthermore, as shown in [<reflink idref="bib10" id="ref58">10</reflink>] study, young adults in Hong Kong regard the GA accent as native accent.</p> <hd id="AN0144336864-11">Measures</hd> <p></p> <hd id="AN0144336864-12">Bi-modal audio-visual connected speech comprehension test</hd> <p>This test assesses listeners' ability to decode minimal pairs of connected speech, distinguished by one differing segment within the whole phrase. By manipulating two parameters, namely the link between the content of the spoken word phrases and the subtitles (matched vs. mismatched) and the degree of word phrase frequency or "familiarity" (familiar vs. unfamiliar), four listening conditions were created: (a) a familiar phrase with matched subtitles, (b) a familiar phrase with mismatched subtitles, (c) an unfamiliar phrase with matched subtitles, and (d) an unfamiliar phrase with mismatched subtitles. The familiar items were connected speech phrases comprising high-frequency words or formulaic phrases (e.g., not at all). Conversely, the resulting unfamiliar items in the minimal pairs contained unfamiliar lexical items and/or non-formulaic phrases. It should be noted that "unfamiliar" phrases such as /hot takes/ may not make sense under normal circumstances. Yet, we cannot rule out the possibility that the phrase may be articulated in a playful situation. Thus, these phrases are regarded as unfamiliar instead of illegitimate.</p> <p>The recordings were deliberately designed to include connected speech processes that have been shown to be difficult for Chinese EFL learners, namely assimilation, elision, and juncture ([<reflink idref="bib9" id="ref59">9</reflink>]; [<reflink idref="bib36" id="ref60">36</reflink>]; [<reflink idref="bib51" id="ref61">51</reflink>]; [<reflink idref="bib60" id="ref62">60</reflink>]). The stimuli were short phrases without filler sentences. As shown in Table 1, the number of words in canonical/citation form was limited to 5 to minimize the working memory load. They also did not carry contextual cues for top-down, meaning-driven predictions. The same set of stimuli was presented in the subtitles-matched and subtitles-mismatched conditions. The number of words was therefore the same across the two conditions. There were altogether 36 minimal pairs of target connected speech patterns (see Table 1 for details). As the whole battery of test can be potentially be used for language assessment, scale reliability was examined. We computed Cronbach's α that inform the degree of relatedness among a set of test items in a battery of tests. Reliability of the whole test (<emph>k</emph> = 72) was.73 (Cronbach's α) which was commonly recognized as acceptable to good reliability ([<reflink idref="bib17" id="ref63">17</reflink>]).</p> <p>Graph</p> <p>Table 1. The Connected Speech Minimal Pairs Speech Stimuli and Their IPA Transcription.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /></colgroup><thead><tr><th align="left" colspan="2">Minimal pairs of connected speech<xref ref-type="table-fn" rid="tfn2">a</xref></th><th align="center">IPA transcription of the citation forms</th><th align="center">IPA transcription of the reduced forms<xref ref-type="table-fn" rid="tfn2">b</xref></th></tr></thead><tbody><tr><td colspan="4">Assimilation</td></tr><tr><td rowspan="2">1.</td><td>A. Ten coins.</td><td>/ten ˌkɔɪnz/</td><td rowspan="2">[ˈtæŋ ˌkɔɪnz]</td></tr><tr><td>B. Tank coins.</td><td>/tæŋk ˌkɔɪnz/</td></tr><tr><td rowspan="2">2.</td><td>A. Good plan.</td><td>/ˈɡʊd ˈplæn/</td><td rowspan="2">[ˈɡʊp ˈplæn]</td></tr><tr><td>B. Goods plan.</td><td>/ˈɡʊds ˈplæn/</td></tr><tr><td rowspan="2">3.</td><td>A. Hand ball.</td><td>/hænd bɔɫ/</td><td rowspan="2">[hæm bɔl]</td></tr><tr><td>B. Han doll.</td><td>/hæn dɒɫ/</td></tr><tr><td rowspan="2">4.</td><td>A. Hot cakes.</td><td>/hɑːt ˈkeɪks/</td><td rowspan="2">[hɑːk ˈkeɪks]</td></tr><tr><td>B. Hot takes.</td><td>/hɑːt ˈteɪks/</td></tr><tr><td rowspan="2">5.</td><td>A. I don't know.</td><td>/ˈaɪ ˈdoʊnt ˈnoʊ/</td><td rowspan="2">[ˈaɪ ˈdoʊn ˈnoʊ]</td></tr><tr><td>B. I dome know.</td><td>/ˈaɪ doʊm ˈnoʊ/</td></tr><tr><td rowspan="2">6.</td><td>A. Batman.</td><td>/ˈbæt ˈmæn/</td><td rowspan="2">[ˈbæp ˈmæn]</td></tr><tr><td>B. Bats man.</td><td>/ˈbæts ˈmæn/</td></tr><tr><td colspan="4">Elision</td></tr><tr><td rowspan="2">7.</td><td>A. Kate and Annie.</td><td>/ˈkeɪt ənd ˈæni/</td><td rowspan="2">[ˈkeɪt ən ˈæni]</td></tr><tr><td>B. Kate and Tanny.</td><td>/ˈkeɪt ənd ˈtæni/</td></tr><tr><td rowspan="2">8.</td><td>A. Next please.</td><td>/ˈnekst ˈpliːz/</td><td rowspan="2">[ˈneks ˈpliːz]</td></tr><tr><td>B. Ness please.</td><td>/ˈnes ˈpliːz/</td></tr><tr><td rowspan="2">9</td><td>A. Give him all.</td><td>/ˈɡɪv ˈhɪm ɔːɫ/</td><td rowspan="2">[ˈɡɪ vɪm ɔːl]</td></tr><tr><td>B. Give film all.</td><td>/ˈɡɪv ˈfɪlm ɔːɫ/</td></tr><tr><td rowspan="2">10.</td><td>A. The best part.</td><td>/ðə ˈbest ˈpa˞t/</td><td rowspan="2">[ðə ˈbes ˈpa˞t]</td></tr><tr><td>B. The best tart.</td><td>/ðə ˈbest ˈta˞t/</td></tr><tr><td rowspan="2">11.</td><td>A. You and me.</td><td>/ju ænd ˈmiː/</td><td rowspan="2">[ju wən ˈmiː]</td></tr><tr><td>B. You wormy.</td><td>/ju ˈwɚmi/</td></tr><tr><td rowspan="2">12.</td><td>A. At the stop.</td><td>/æt ðə stɑːp/</td><td rowspan="2">[ə ðə stɑːp]</td></tr><tr><td>B. Ever stop.</td><td>/ˈɛvɚ stɑːp/</td></tr><tr><td colspan="4">Juncture</td></tr><tr><td rowspan="2">13.</td><td>A. Not at all.</td><td>/ˈnɑːt æt ɔːɫ/</td><td rowspan="2">[ˈnɑː tə tɔːl]</td></tr><tr><td>B. Not that tall.</td><td>/ˈnɑːt ðæt ˈtɒɫ/</td></tr><tr><td rowspan="2">14.</td><td>A. It wasn't easy.</td><td>/ˈɪt ˈwɑːzənt ˈiːzi/</td><td rowspan="2">[ˈɪt ˈwɑːzən ˈtiːzi]</td></tr><tr><td>B. It was teasy.</td><td>/ˈɪt wəz tiːzi/</td></tr><tr><td rowspan="2">15.</td><td>A. Better off.</td><td>/ˈbetɚ ɒf/</td><td rowspan="2">[ˈbetə ɹɒf]</td></tr><tr><td>B. Better ruff.</td><td>/ˈbetɚ ɹəf/</td></tr><tr><td rowspan="2">16.</td><td>A. Leave it to me.</td><td>/liːv ɪt tu miː/</td><td rowspan="2">[liː vɪt tə miː]</td></tr><tr><td>B. Lean fit to me.</td><td>/liːn fɪt tu miː/</td></tr><tr><td rowspan="2">17.</td><td>A. No apple.</td><td>/noʊ ˈæpəl̩/</td><td rowspan="2">[noʊ ˈwæpl̩]</td></tr><tr><td>B. An old wapple.</td><td>/ən oʊld ˈwæpl̩/</td></tr><tr><td rowspan="2">18.</td><td>A. One wish each.</td><td>/wʌn ˈwɪʃ ˈiːtʃ/</td><td rowspan="2">[wʌn ˈwɪ ˈʃ iːtʃ]</td></tr><tr><td>B. One wish sheet.</td><td>/wʌn ˈwɪʃ ˈʃiːt/</td></tr></tbody></table> </ephtml> </p> <p>1 IPA = International Phonetic Alphabet.</p> <p>2 Familiar and unfamiliar word phrases in each pair are presented as "A" and "B," respectively. <sups>b</sups> This is just one of the possible utterances that confuse the perception of the minimal pairs.</p> <p>Participants were told that the recordings would be played with either matched or mismatched subtitles. They were therefore encouraged to focus on aural, as opposed to visual, information when answering the multiple-choice questions. For each trial, the student participants were presented with audio and subtitles concurrently through a speaker and PowerPoint slides (Figure 1). As the experiment was run in a group setting, the experimenters had to make sure that the participants completed each item before proceeding to the next item. Therefore, the bi-modal listening test was controlled and paced by the experimenter.</p> <p>Graph: Figure 1. An illustration showing the experimental setup. Note. In this scenario, the subtitle presented on the PowerPoint did not match the audio presented through the speaker. The participants were instructed to circle the correct answer out of the two options printed on the answer sheet.</p> <p>For half of the trials, the audio stimuli did not match the content of the subtitles, that is, they were presented in the "mismatched" condition. For example, the audio of [I dome know] was played together with the subtitle /I don't know/. The remaining half of the trials presented matched audio and subtitles, that is, the "matched" condition. For example, /I don't know/ was presented in both auditory and visual forms (Figure 2). Students were given two choices on the answer sheet, for this example the choices would be "I dome know" and 'I don't know.' They were instructed to listen and select the phrase that they heard from the two choices printed on the test paper. These 72 items were randomized for presentation. On the randomization list, 64% of the items were presented with the subtitles-matched condition first and 36% of them were presented with the subtitles-mismatched condition first. In addition, the familiar and unfamiliar items in the phrase pairs were presented first in 47% and 53% of the items. These percentages could minimize the possibility that listeners' decision was conditioned by their decision of the previous items. Each correct response yields one mark, and the number of correct responses in each of the listening conditions was summed to yield a mean score for further statistical analyses.</p> <p>Graph: Figure 2. The two independent variables and the resulting four conditions.</p> <hd id="AN0144336864-13">Uni-modal connected speech comprehension test</hd> <p>In the "no-subtitles" condition, the experimenter played the audio stimuli in a fixed order and asked listeners to dictate what they heard. No subtitles were provided within this condition. The total number of test items was 36. An all-or-none scoring criterion was adopted such that an answer is considered correct only if it matches exactly the same as our answer key (Table 1). The marking was carried out by the first and second authors and a 100% agreement was achieved. Each correct response yields one mark, and summing the correct responses would produce a mean score for comparing the listening performances across conditions.</p> <hd id="AN0144336864-14">Results</hd> <p>A mix of inferential and descriptive statistical analyses were performed to investigate the influence of subtitles on decoding three types of English connected speech in Chinese EFL learners. The data analyses were conducted on data obtained from the 28 participants who finished all the tests. With a 3 (<emph>type of connected speech</emph>: assimilation vs. elision vs. juncture) × 2 (<emph>familiarity</emph>: familiar vs. unfamiliar) × 3 (<emph>subtitling</emph>: without subtitles vs. with matched subtitles vs. with mismatched subtitles) design, yielding a total of eighteen listening conditions. The means and standard deviation of these conditions are listed in Table 2.</p> <p>Graph</p> <p>Table 2. Descriptive Statistics for Perception Accuracy Across the Nine Listening Conditions.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /></colgroup><thead><tr><th align="center">Connected speech process</th><th align="center">Without subtitle</th><th align="center">Matched subtitle</th><th align="center">Mismatched subtitle</th></tr></thead><tbody><tr><td colspan="4">Assimilation</td></tr><tr><td> Familiar</td><td>3.21 (1.31)</td><td>5.50 (0.74)</td><td>4.89 (0.95)</td></tr><tr><td> Unfamiliar</td><td>0.17 (0.39)</td><td>4.96 (0.83)</td><td>4.46 (0.96)</td></tr><tr><td colspan="4">Elision</td></tr><tr><td> Familiar</td><td>1.92 (1.21)</td><td>5.39 (0.68)</td><td>5.17 (0.86)</td></tr><tr><td> Unfamiliar</td><td>0.57 (0.92)</td><td>5.35 (0.63)</td><td>5.53 (0.63)</td></tr><tr><td colspan="4">Juncture</td></tr><tr><td> Familiar</td><td>1.07 (0.85)</td><td>5.35 (0.78)</td><td>5.03 (1.10)</td></tr><tr><td> Unfamiliar</td><td>0.17 (0.47)</td><td>5.42 (0.63)</td><td>5.32 (0.61)</td></tr></tbody></table> </ephtml> </p> <p>Repeated-measures analysis of variance (ANOVA) with two within-subject factors was used to compare the performances across various listening conditions. Before the analysis, we tested the assumption using Mauchly's test of sphericity. If this assumption is violated (<emph>p</emph> <.05), a correction of degrees of freedom by Huynh–Feldt estimates of sphericity was carried out ([<reflink idref="bib26" id="ref64">26</reflink>], p. 474). When a significant interaction was detected, contrast analysis was conducted to examine where the difference lied.</p> <hd id="AN0144336864-15">Connected Speech Perception Performance in the No-Subtitles Condition</hd> <p>First, we assessed the performance of connected speech decoding in the "no-subtitles" condition, by comparing the three types of connected speech and examined the effect of familiarity. A 3 (<emph>types</emph>: assimilation vs. elision vs. juncture) × 2 (<emph>familiarity</emph>: familiar vs. unfamiliar) repeated-measures ANOVA was computed. The main effect of <emph>types</emph> was significant, <emph>F</emph>(<reflink idref="bib2" id="ref65">2</reflink>, 54) = 29.96, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.52. A post hoc test of least significant difference (LSD) indicated that the performance of decoding assimilation was significantly better than decoding elision (<emph>p</emph> =.003) and juncture (<emph>p</emph> <.001). Moreover, the performance of decoding elision was significantly better than decoding juncture (<emph>p</emph> <.001). The main effect of familiarity was significant, <emph>F</emph>(<reflink idref="bib1" id="ref66">1</reflink>, 27) = 112.52, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.80, indicating that familiar items were better decoded than unfamiliar ones. The interaction between <emph>types</emph> and <emph>familiarity</emph> was significant, <emph>F</emph>(<reflink idref="bib2" id="ref67">2</reflink>, 54) = 33.41, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.55. Contrast analysis showed that assimilation was more sensitive to familiarity than elision, <emph>F</emph>(<reflink idref="bib1" id="ref68">1</reflink>, 27) = 28.74, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.51, and juncture, <emph>F</emph>(<reflink idref="bib1" id="ref69">1</reflink>, 27) = 92.74, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.77, thus indicating that the change (improvement) of performances from unfamiliar to familiar items was larger in assimilation than in elision and juncture.</p> <hd id="AN0144336864-16">The Influence of Subtitles on Decoding Assimilation, Elision, and Juncture</hd> <p>Next, we examined the interplay between subtitling and familiarity. We conducted a 3 (<emph>subtitling</emph>: without subtitles vs. with matched subtitles vs. with mismatched subtitles) × 2 (<emph>familiarity</emph>: familiar vs. unfamiliar) repeated-measures ANOVA for each of the three types of connected speech, namely assimilation, elision, and juncture. The results will allow us to evaluate whether the two hypothesized variables have the same effect on different types of connected speech.</p> <hd id="AN0144336864-17">Assimilation</hd> <p>The main effect of <emph>subtitling</emph> was significant, <emph>F</emph>(<reflink idref="bib2" id="ref70">2</reflink>, 54) = 301.717, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.91. Post hoc comparison by LSD revealed that the scores obtained in the matched-subtitles condition were significantly higher than those obtained in the mismatched-subtitles (<emph>p</emph> =.001) and no-subtitles conditions (<emph>p</emph> <.001). In addition, the scores obtained in the mismatched-subtitles condition were significantly higher than those obtained in the no-subtitles condition (<emph>p</emph> <.001). The main effect of <emph>familiarity</emph> was significant, <emph>F</emph>(<reflink idref="bib1" id="ref71">1</reflink>, 27) = 126.00, <emph>p</emph> <.001, =.82, indicating that familiar items were better decoded than unfamiliar items. The <emph>subtitling</emph> × <emph>familiarity</emph> interaction was significant, <emph>F</emph>(<reflink idref="bib2" id="ref72">2</reflink>, 54) = 57.95, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.68. Contrast analysis revealed that performances in the no-subtitles condition were more sensitive to familiarity than the matched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref73">1</reflink>, 27) = 64.72, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.70, and mismatched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref74">1</reflink>, 27) = 87.57, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.76. In other words, the change (improvement) of decoding performance from unfamiliar to familiar items was larger in the no-subtitles condition than in the matched- and mismatched-subtitles conditions.</p> <hd id="AN0144336864-18">Elision</hd> <p>Mauchly's test indicated that the assumption of sphericity for the main effect of <emph>subtitling</emph> had been violated, χ<sups>2</sups>(<reflink idref="bib2" id="ref75">2</reflink>) = 17.06, <emph>p</emph> <.001; therefore, degrees of freedom were corrected using Huynh-Feldt estimate of sphericity (ε =.69). The main effect of <emph>subtitling</emph> was significant, <emph>F</emph>(1.13, 36.45) = 455.55, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.94. Post hoc comparison by LSD revealed that the scores obtained in the matched- and mismatched-subtitles condition were significantly higher than those obtained in the no-subtitles condition (<emph>p</emph>s <.001). In addition, the scores obtained in the matched-subtitles and mismatched-subtitles condition were not significantly different (<emph>p</emph> =.83). On the other hand, the main effect of <emph>familiarity</emph> was significant, <emph>F</emph>(<reflink idref="bib1" id="ref76">1</reflink>, 27) = 12.88, <emph>p</emph> =.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.32, indicating that familiar items were better decoded than unfamiliar items. The <emph>subtitling</emph> × <emph>familiarity</emph> interaction was significant, <emph>F</emph>(<reflink idref="bib2" id="ref77">2</reflink>, 54) = 23.99, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.47. Contrast analysis revealed that performances in the no-subtitles condition were more sensitive to familiarity than those in the matched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref78">1</reflink>, 27) = 22.71, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.45, and mismatched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref79">1</reflink>, 27) = 39.87, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.59. In other words, the change (improvement) of decoding performance from unfamiliar to familiar items was larger in the no-subtitles condition than in the matched- and mismatched-subtitles conditions.</p> <hd id="AN0144336864-19">Juncture</hd> <p>Mauchly's test indicated that the assumption of sphericity for the main effect of <emph>subtitling</emph> and the interaction effect between <emph>subtitling</emph> and <emph>familiarity</emph> had been violated, χ<sups>2</sups>(<reflink idref="bib2" id="ref80">2</reflink>) = 6.52, <emph>p</emph> =.038 and χ<sups>2</sups>(<reflink idref="bib2" id="ref81">2</reflink>) = 11.29, <emph>p</emph> =.004, respectively. Therefore, degrees of freedom were corrected using Huynh-Feldt estimate of sphericity (ε =.86 and ε =.77). The main effect of <emph>subtitling</emph> was significant, <emph>F</emph>(1.72, 46.65) = 780.91, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.96. Post hoc comparison by LSD revealed that the scores obtained in the matched- and mismatched-subtitles conditions were significantly higher than those obtained in the no-subtitles condition (<emph>ps</emph><.001). In addition, the scores obtained in the matched-subtitles and mismatched-subtitles conditions were marginally significantly (<emph>p</emph> =.05). The main effect of <emph>familiarity</emph> was nonsignificant, <emph>F</emph>(<reflink idref="bib1" id="ref82">1</reflink>, 27) = 2.89, <emph>p</emph> =.10, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.09, indicating that performances for familiar items were comparable to the unfamiliar items. The <emph>subtitling</emph> × <emph>familiarity</emph> interaction was significant, <emph>F</emph>(1.54, 41.69) = 20.36, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.43. Contrast analysis revealed that performances in the no-subtitles condition were more sensitive to familiarity than those in the matched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref83">1</reflink>, 27) = 21.32, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.44, and mismatched-subtitles condition, <emph>F</emph>(<reflink idref="bib1" id="ref84">1</reflink>, 27) = 24.93, <emph>p</emph> <.001, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi mathvariant="normal">η</mi><mi mathvariant="normal">p</mi><mn>2</mn></msubsup></mrow></math> </ephtml> =.48. In other words, the change (improvement) of decoding performance from unfamiliar to familiar items was larger in the no-subtitles condition than in the matched- and mismatched-subtitles conditions.</p> <p>In summary, decoding of speech embedded with assimilation, elision, and juncture was poor when no subtitles were provided. Interestingly, the matched subtitles did not always serve as better subtitles than mismatched subtitles as a significant difference was found only among the assimilation items. Accuracy in decoding familiar connected speech was higher than unfamiliar connected speech in assimilation and elision only.</p> <hd id="AN0144336864-20">Item Analysis Within Each of the Three Types of Connected Speech</hd> <p>Subsequently, we performed item analysis to further compare listeners' decoding performances across the three subtitling conditions for each individual test item. Specifically, we examined whether speech tokens under the same category (e.g., assimilation) elicited similar levels of accuracy. As shown in Tables 3 to 5, the range of percentage accuracy in speech identification was large across the three types of connected speech as well as across the three subtitling conditions, suggesting large item-specific variations.</p> <p>Graph</p> <p>Table 3. Error Analysis for Assimilation Test Items (N = 28).</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /></colgroup><thead><tr><th /><th /><th align="center" colspan="3">Accuracy rate (%)</th></tr><tr><th align="center">Familiarity</th><th align="center">Test items</th><th align="center">Without subtitles</th><th align="center">With matched subtitles</th><th align="center">With mismatched subtitles</th></tr></thead><tbody><tr><td>F</td><td>Ten coins</td><td>35.7</td><td>85.7</td><td>25.0</td></tr><tr><td>U</td><td>Tank coins</td><td>0</td><td>75.0</td><td>39.3</td></tr><tr><td>F</td><td>Good plan</td><td>67.9</td><td>96.4</td><td>92.9</td></tr><tr><td>U</td><td>Goods plan</td><td>10.7</td><td>92.9</td><td>96.4</td></tr><tr><td>F</td><td>Hand ball</td><td>71.4</td><td>96.4</td><td>89.3</td></tr><tr><td>U</td><td>Han doll</td><td>3.6</td><td>85.7</td><td>85.7</td></tr><tr><td>F</td><td>Hot cakes</td><td>21.4</td><td>96.4</td><td>100</td></tr><tr><td>U</td><td>Hot takes</td><td>3.6</td><td>96.4</td><td>100</td></tr><tr><td>F</td><td>I don't know</td><td>100</td><td>82.1</td><td>46.4</td></tr><tr><td>U</td><td>I dome know</td><td>0</td><td>46.4</td><td>67.9</td></tr><tr><td>F</td><td>Bat man</td><td>25</td><td>92.9</td><td>92.9</td></tr><tr><td>U</td><td>Bats man</td><td>0</td><td>100</td><td>100</td></tr></tbody></table> </ephtml> </p> <p>3 F = familiar; U = unfamiliar.</p> <p>Graph</p> <p>Table 4. Error Analysis for Elision Test Items (N = 28).</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /></colgroup><thead><tr><th /><th /><th align="center" colspan="3">Accuracy (%)</th></tr><tr><th align="center">Familiarity</th><th align="center">Test items</th><th align="center">Without subtitles</th><th align="center">With matched subtitles</th><th align="center">With mismatched subtitles</th></tr></thead><tbody><tr><td>F</td><td>Kate and Annie</td><td>3.6</td><td>92.9</td><td>96.4</td></tr><tr><td>U</td><td>Kate and Tanny</td><td>0</td><td>89.3</td><td>96.4</td></tr><tr><td>F</td><td>Next please</td><td>46.4</td><td>100</td><td>89.3</td></tr><tr><td>U</td><td>Ness please</td><td>0</td><td>100</td><td>92.9</td></tr><tr><td>F</td><td>The best part</td><td>60.7</td><td>100</td><td>100</td></tr><tr><td>U</td><td>The best tart</td><td>7.1</td><td>100</td><td>100</td></tr><tr><td>F</td><td>You and me</td><td>64.3</td><td>96.4</td><td>96.4</td></tr><tr><td>U</td><td>You wormy</td><td>25.0</td><td>96.4</td><td>92.9</td></tr><tr><td>F</td><td>Give him all</td><td>17.9</td><td>78.6</td><td>82.1</td></tr><tr><td>U</td><td>Give film all</td><td>0</td><td>75.0</td><td>71.4</td></tr><tr><td>F</td><td>At the stop</td><td>0</td><td>71.4</td><td>89.3</td></tr><tr><td>U</td><td>Ever stop</td><td>25.0</td><td>75.0</td><td>64.3</td></tr></tbody></table> </ephtml> </p> <p>4 F = familiar; U = unfamiliar.</p> <p>Graph</p> <p>Table 5. Error Analysis for Juncture Test Items (N = 28).</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /><col align="char" char="." /></colgroup><thead><tr><th /><th /><th align="center" colspan="3">Accuracy (%)</th></tr><tr><th align="center">Familiarity</th><th align="center">Test items</th><th align="center">Without subtitles</th><th align="center">With matched subtitles</th><th align="center">With mismatched subtitles</th></tr></thead><tbody><tr><td>F</td><td>Not at all</td><td>0</td><td>75.0</td><td>71.4</td></tr><tr><td>U</td><td>Not that tall</td><td>3.6</td><td>75.0</td><td>71.4</td></tr><tr><td>F</td><td>It wasn't easy</td><td>39.3</td><td>92.9</td><td>92.9</td></tr><tr><td>U</td><td>It was teasy</td><td>0</td><td>94.9</td><td>82.1</td></tr><tr><td>F</td><td>Better off</td><td>3.6</td><td>96.4</td><td>92.9</td></tr><tr><td>U</td><td>Better ruff</td><td>0</td><td>100</td><td>92.9</td></tr><tr><td>F</td><td>Leave it to me</td><td>7.1</td><td>85.7</td><td>78.6</td></tr><tr><td>U</td><td>Lean fit to me</td><td>3.6</td><td>89.3</td><td>82.1</td></tr><tr><td>F</td><td>No apple</td><td>57.1</td><td>85.7</td><td>100</td></tr><tr><td>U</td><td>An old wapple</td><td>10.7</td><td>92.9</td><td>85.7</td></tr><tr><td>F</td><td>One wish each</td><td>0</td><td>92.9</td><td>96.4</td></tr><tr><td>U</td><td>One wish sheet</td><td>0</td><td>100</td><td>89.3</td></tr></tbody></table> </ephtml> </p> <p>5 F = familiar; U = unfamiliar.</p> <p>As noted in the no-subtitles condition, the maximum percentage accuracy in speech identification for assimilation was 100% and the minimum was 0%, yielding a range of 100%. The ranges for elision and juncture in the same condition were 64% and 57.1%, respectively. In both matched-subtitles and mismatched-subtitles conditions, a 100% accuracy rate could be obtained for some items across the three types of connected speech. Although the ranges of percentage correct were similar for elision and juncture (about 30%), the range for assimilation was almost double that of the former two types (about 60%). As shown in Table 3, a number of inaccurate regressive assimilation errors was recorded, particularly for unfamiliar items. For instance, participants commonly misinterpret "tank coins" /ˈtæŋk ˌkɔɪnz/ as "ten coins" /ˈtæŋ ˌkɔɪnz/. For speech items embedded with juncture, a constant pattern of mis-segmentation was observed, such as misinterpreting "not at all" as "not that tall."</p> <p>In general, the performances of decoding assimilation, elision, and juncture were enhanced with the presence of either matched or mismatched subtitles. However, it is worth noting one exception. The accuracy rate of perceiving "I don't know" is in fact lower in the conditions with subtitles (82.1%) compared with the no-subtitles condition (100%). Based on the above results, it is suggested that individual items vary in terms of level of difficulties as well as the sensitivity to subtitles.</p> <hd id="AN0144336864-21">Discussion</hd> <p>The present study aims to examine the role of subtitles in connected speech decoding in Chinese EFL learners. We used both matched and mismatched subtitles to test how listeners process visual and audio information in the face of conflicting situations. Our results showed that the performances of decoding connected speech in both matched and mismatch subtitles could significantly facilitate the processing of the three types of connected speech examined, which is generally in line with previous studies about the usefulness of matched subtitles in listening. These findings can strengthen the claim that provision of subtitles is crucial for decoding native English connected speech for Chinese EFL learners. More importantly, the data in the mismatched condition provide new insight to the use of subtitles in listening comprehension among EFL learners, suggesting that EFL adolescent learners who have not yet attained native-like English listening skills still attempted to link visual information with audio information, rather than disregard the later completely in a bi-modal listening environment. Given that the experimental stimuli used in the present study were phrases of connected speech that differed only by one segment, the participants demonstrated their abilities to identify the critical segments that distinguished the minimal pairs of connected speech.</p> <p>The findings of significant facilitation by mismatched subtitles further suggest that the potential detrimental effect of visual-dominance is minimal in a bi-modal listening environment ([<reflink idref="bib38" id="ref85">38</reflink>]). Listeners were found to be able to process the auditory signals even in the presence of misleading visual information. Furthermore, our results suggest that the listeners had sufficient executive functioning to inhibit irrelevant information as well as to selectively attend to audio information ([<reflink idref="bib35" id="ref86">35</reflink>]). However, it is important to note that our participants were only exposed to subtitles and audio. Taking the limited capacity of cognitive resources into consideration ([<reflink idref="bib21" id="ref87">21</reflink>]), a trade-off between auditory and visual information processing occurs in a bi-modal listening environment. If video is present as in movies and TV programs, listeners' cognitive load may be further increased, leaving little attentional resources to process audio information. Therefore, the use of multimedia in training connected speech decoding needs to take into account the nature and amount of visual information presented to the listeners.</p> <p>As evidenced across the three subtitling conditions, familiar word phrases were better recognized than unfamiliar word phrases. This implies that novel connected speech processes and signals are challenging to EFL listeners, whose phonological repertoire for decoding connected speech may not be versatile enough. We speculate that the immediate success of decoding connected speech with the aid of subtitles may hinder long-term perceptual training. As previously described by [<reflink idref="bib50" id="ref88">50</reflink>], <emph>cognitive task performance</emph> and <emph>learning</emph> are different. <emph>Cognitive task performances</emph> are those actions that operate on mental structures in working memory, whereas <emph>learning</emph> operates on mental structures in long-term memory. In other words, learning takes place only if the content in long-term memory is transformed and results in an increase in expertise. If EFL learners are used to processing connected speech in working memory for immediate outcome without activating the phonological repertoire in long-term memory, their listening skills cannot be improved even after constant exposure to native English speech. As a result, when encountering novel connected speech without subtitles, subtitle-dependent EFL listeners are more likely to struggle in decoding the connected speech signal, though this claim has yet to be verified in future studies. The ingrained "performing without learning" situation is further reinforced by the use of (or reliance on) subtitles. Thus, it is important to prevent the development of this vicious circle within connected speech acquisition.</p> <hd id="AN0144336864-22">The Influence of Frequency on Connected Speech Perception</hd> <p>The current study examines perception of speech embedded with three types of connected speech processes (assimilation, elision, and juncture). When considering word frequency within the no-subtitles condition, the identification of high familiarity words was higher than those of low familiarity. In other words, these three connected speech processes are frequency-sensitive: The higher the frequency, the higher the accuracy, and vice versa. However, when subtitling was considered, the three types of connected speech processes display varying degrees of frequency-sensitivity when presented without subtitles. Yet, both assimilation and elision were found to be more sensitive to familiarity effects than juncture. One possible explanation for this finding is the perceptual differences between various connected speech processes. For assimilation and elision, it may be that the neighboring phonemes serve as cues which contribute to the identification of modified segments. This hypothesis is consistent with the view of probabilistic phonotactics (see [<reflink idref="bib55" id="ref89">55</reflink>]). When viewed in conjunction with the theory that these processes are segmental "overlaps" with gradient variations (e.g., [<reflink idref="bib5" id="ref90">5</reflink>]), residual speech sounds between neighboring phonemes' transitions may be perceived. If this is the case, the frequency effect will become more beneficial as listeners become more experienced. This is due to the shorter phonetic distance and less effortful lexical activation that this entails. In contrast, for juncture, word boundaries are ambiguous. For this process, non-native listeners tend to rely more on sentential and lexical context to aid segmentation as opposed to acoustic-phonetic clues ([<reflink idref="bib1" id="ref91">1</reflink>]; [<reflink idref="bib44" id="ref92">44</reflink>]). As only non-contextual minimal pairs were used in the study, the stimuli lacked supporting top-down information. Therefore, participants relied on limited speech decoding abilities for word segmentation. For these cases, frequency would impose a minimal effect on juncture, as acoustic-phonetic cues may not be the major route utilized for comprehension. Furthermore, the phonetic distance may not be as short as that exhibited for assimilation and elision. These assumptions call for further research to investigate whether EFL learners adopt different speech-decoding strategies for different connected speech processes.</p> <p>Last but not least, it should be noted that listening performances were not perfect even when accurate subtitles were provided. It is possible that some EFL learners may not be able to decode the subtitles because of inadequate reading skill. Although we did not measure reading ability in the current study, our speculation is partly supported by several previous studies showing a strong link between overall language proficiency and the efficiency of subtitle use during listening tasks in EFL learners (e.g., [<reflink idref="bib39" id="ref93">39</reflink>]; [<reflink idref="bib40" id="ref94">40</reflink>]). Still, further research is needed to verify the specific links between reading skills and connected speech processing skills.</p> <hd id="AN0144336864-23">Pedagogical Implications</hd> <p>In line with other studies examining connected speech in EFL learners, suboptimal connected speech comprehension skills were observed in the current sample. This result again pinpoints the need for connected speech listening training in EFL classrooms. As implied in the current study, using onscreen subtitles for educational purposes should be promoted but with greater caution. Pedagogically speaking, the ideal situation is that EFL learners continuously attempt to connect the content of the subtitles and audio when being exposed to them. This is key to ensuring that characteristics of English speech are learned, such as phonotactic properties, novel vocabulary, and grammar of English sounds ([<reflink idref="bib6" id="ref95">6</reflink>]). However, multimedia may only be used for entertainment purposes and the motivation to learn connected speech may not be consistently high across EFL learners. In reality, EFL learners may devote most of their attention to the subtitles and very little attention would be allocated to the audio input. Teachers are advised to remind students of both the positive and negative effects of subtitles on listening training. The focus of English listening comprehension should be placed on training students to comprehend native English connected speech without relying on subtitles. Thus, upon the use of subtitles (or other visual cues) to facilitate the learning of English lexical knowledge, teachers should gradually remove the visual cues so that students can comprehend English connected speech eventually in a subtitles-free listening environment. In addition, the significant familiarity effect obtained in the present study suggests that there is a need to consider this linguistic parameter when planning the listening curriculum. Teachers may design a variety of learning activities to consolidate the lexical familiarity of their students, so as to facilitate listening comprehension. It is interesting to note that assimilation was more sensitive to familiarity than elision. Teachers can tell students this characteristic explicitly and should be more alert to the familiarity effect when training different types of connected speech. Teachers should also choose captioned videos that are of appropriate level to the students</p> <p>However, the familiarity effect should not be over-interpreted as having to pre-teach all difficult vocabulary/phrases from a text prior to the introduction of a listening task (see [<reflink idref="bib37" id="ref96">37</reflink>]). Even though this is believed to scaffold their comprehension and to heighten their awareness of new items in the text (e.g., [<reflink idref="bib8" id="ref97">8</reflink>]), doing so removes the chance for learners to practice inferring the meaning of such items from the context. Hence, teachers should strike a balance between giving vocabulary input and nourishing students' skills on drawing inferences from the speech.</p> <hd id="AN0144336864-24">Limitations and Future Research Directions</hd> <p>The first limitation concerns the number of items for each of the listening conditions. In order to create minimal pairs for the selected connected speech processes, the script was written with a focus on phrases with similar pronunciation. Such constraint, together with another parameter we manipulated (i.e., word familiarity) had limited the number of minimal pairs we could generate. Although the number of items could be increased, the effect size obtained from this study was large enough to substantiate our results. In future studies, the inclusion of additional items is recommended.</p> <p>Another limitation is the presentation of subtitles during the test. The participants were provided with a list of minimal pairs and were requested to listen and choose to make the test procedure easier to understand. However, the option of force-choice questions may provide additional visual cues for participants as they were not required to provide the answers themselves. Given the provision of differential visual cues and the difference in response format, a direct comparison between the performances under the subtitled conditions and no-subtitles condition is deemed to be unfair. Future studies could therefore include a dictation test in the subtitled conditions.</p> <hd id="AN0144336864-25">Conclusion</hd> <p>The usefulness of subtitles for the immediate decoding of connected speech in a bi-modal listening setting is well documented in the literature. However, it is worth reconsidering the impact subtitles have on the acquisition of acoustic, phonetic, and phonological aspects of connected speech in the long run. Consistent with previous research on subtitles, subtitles are shown to be beneficial to listening comprehension in the current study. Hence, we do not advocate to abandon them, especially as EFL learners were found to favor the use of subtitles in their daily lives ([<reflink idref="bib19" id="ref98">19</reflink>]). Still, EFL learners are reminded to equip themselves to become visual-aid-free competent listeners. Moreover, instructors or self-taught learners should be mindful about the exact function of subtitles for various listening tasks and be able to shift their focus back and forth between audio and visual information.</p> <p>We would like to thank all the participants.</p> <ref id="AN0144336864-26"> <title> References </title> <blist> <bibl id="bib1" idref="ref66" type="bt">1</bibl> <bibtext> Altenberg E. P. (2005). The perception of word boundaries in a second language. Second Language Research, 21(4), 325–358. https://doi.org/10.1191/0267658305sr250oa</bibtext> </blist> <blist> <bibl id="bib2" idref="ref48" type="bt">2</bibl> <bibtext> Behroozizad S., Majidi S. (2015). The effect of different modes of English captioning on EFL learners' general listening comprehension: Full text vs. keyword captions. Advances in Language & Literary Studies, 6(4), 115–121. https://doi.org/10.7575/aiac.alls.v.6n.4p.115</bibtext> </blist> <blist> <bibl id="bib3" idref="ref5" type="bt">3</bibl> <bibtext> Bird S. A., Williams J. N. (2002). The effect of bimodal input on implicit and explicit memory: An investigation into the benefits of within-language subtitling. Applied Psycholinguistics, 23(4), 509–533. https://doi.org/10.1017/S0142716402004022</bibtext> </blist> <blist> <bibl id="bib4" idref="ref37" type="bt">4</bibl> <bibtext> Bradlow A. R., Bent T. (2002). The clear speech effect for non-native listeners. The Journal of the Acoustical Society of America, 112(1), 272–284. https://doi.org/10.1121/1.1487837</bibtext> </blist> <blist> <bibl id="bib5" idref="ref90" type="bt">5</bibl> <bibtext> Browman C. P., Goldstein L. (1992). Articulatory phonology: An overview. Phonetica, 49(3–4), 155–180. https://doi.org/10.1159/000261913</bibtext> </blist> <blist> <bibl id="bib6" idref="ref95" type="bt">6</bibl> <bibtext> Brown J. D. (Ed.). (2012). New ways in teaching connected speech. Teachers of English to Speakers of Other Languages.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref11" type="bt">7</bibl> <bibtext> Caimi A. (2006). Audiovisual translation and language learning: The promotion of intralingual subtitles. The Journal of Specialised Translation, 6, 85–98.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref12" type="bt">8</bibl> <bibtext> Chai J., Erlam R. (2008). The effect and the influence of the use of video and captions on second language learning. New Zealand Studies in Applied Linguistics, 14, 25–44.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref59" type="bt">9</bibl> <bibtext> Chan A. Y., Li D. C. (2000). English and Cantonese phonology in contrast: Explaining Cantonese ESL learners' English pronunciation problems. Language Culture and Curriculum, 13(1), 67–85. https://doi.org/10.1080/07908310008666590</bibtext> </blist> <blist> <bibtext> Chan J. Y. H. (2013). Contextual variation and Hong Kong English. World Englishes, 32(1), 54–74. https://doi.org/10.1111/weng.12004</bibtext> </blist> <blist> <bibtext> Chun D. M., Plass J. L. (1996). Effects of multimedia annotation on vocabulary acquisition. The Modern Language Journal, 80(2), 183–198. https://doi.org/10.2307/328635</bibtext> </blist> <blist> <bibtext> Chung J. (1999). The effects of using video texts supported with advance organizers and captions on Chinese college students' listening comprehension: An empirical study. Foreign Language Annals, 32(3), 295–308. <ulink href="http://dx.doi.org/10.18806/tesl.v14i1.678">http://dx.doi.org/10.18806/tesl.v14i1.678</ulink></bibtext> </blist> <blist> <bibtext> Cleland A. A., Gaskell M. G., Quinlan P. T., Tamminen J. (2006). Frequency effects in spoken and visual word recognition: Evidence from dual-task methodologies. Journal of Experimental Psychology: Human Perception and Performance, 32(1), 104–119. https://doi.org/10.1037/0096-1523.32.1.104.</bibtext> </blist> <blist> <bibtext> Connine C. M., Mullennix J., Shernoff E., Yelen J. (1990). Word familiarity and frequency in visual and auditory word recognition. Journal of Experimental Psychology: Learning, Memory, and Cognition, 16(6), 1084–1096. https://doi.org/10.1037/0278-7393.16.6.1084</bibtext> </blist> <blist> <bibtext> Connine C. M., Ranbom L. J., Patterson D. J. (2008). Processing variant forms in spoken word recognition: The role of variant frequency. Perception & Psychophysics, 70(3), 403–411. https://doi.org/10.3758/PP.70.3.403</bibtext> </blist> <blist> <bibtext> Connine C. M., Titone D., Wang J. (1993). Auditory word recognition: Extrinsic and intrinsic effects of word frequency. Journal of Experimental Psychology: Learning, Memory, and Cognition, 19(1), 81–94. https://doi.org/10.1037/0278-7393.19.1.81</bibtext> </blist> <blist> <bibtext> Cortina J. M. (1993). What is coefficient alpha? An examination of theory and applications. Journal of Applied Psychology, 78, 98–104.</bibtext> </blist> <blist> <bibtext> Cruttenden A. (2014). Gimson's pronunciation of English. Routledge.</bibtext> </blist> <blist> <bibtext> Dallas B., McCarthy A., Long G. (2016). Examining the educational benefits of and attitudes toward closed-captioning among undergraduate students. Journal of the Scholarship of Teaching and Learning, 16(2), 56–71. <ulink href="http://dx.doi.org/10.14434/josotl.v16i2.19267">http://dx.doi.org/10.14434/josotl.v16i2.19267</ulink></bibtext> </blist> <blist> <bibtext> Davis R. (1998). Running shoes. <ulink href="http://www.esl-lab.com/runningshoes/runningshoesrd1.htm">http://www.esl-lab.com/runningshoes/runningshoesrd1.htm</ulink></bibtext> </blist> <blist> <bibtext> Drijvers L., Mulder K., Ernestus M. (2016). Alpha and gamma band oscillations index differential processing of acoustically reduced and full forms. Brain and Language, 153-154, 27–37. https://doi.org/10.1016/j.bandl.2016.01.003</bibtext> </blist> <blist> <bibtext> Duanmu S. (2007). The phonology of standard Chinese (2nd ed.). Oxford University Press.</bibtext> </blist> <blist> <bibtext> Education Bureau. (2004). CDC English language curriculum guide (primary 1–6). The Government of the Hong Kong Special Administrative Region.</bibtext> </blist> <blist> <bibtext> Ernestus M. (2014). Acoustic reduction and the roles of abstractions and exemplars in speech processing. Lingua, 142, 27–41. https://doi.org/10.1016/j.lingua.2012.12.006</bibtext> </blist> <blist> <bibtext> Ernestus M., Baayen H., Schreuder R. (2002). The recognition of reduced word forms. Brain and Language, 81(1), 162–173. https://doi.org/10.1006/brln.2001.2514</bibtext> </blist> <blist> <bibtext> Field A. P. (2013). Discovering statistics using SPSS (4th ed.). Sage.</bibtext> </blist> <blist> <bibtext> Gernsbacher M. A. (2015). Video captions benefit everyone. Policy Insights from the Behavioral and Brain Sciences, 2, 195–202. https://doi.org/10.1177/2372732215602130</bibtext> </blist> <blist> <bibtext> Goldinger S. D. (1998). Echoes of echoes? An episodic theory of lexical access. Psychological Review, 105(2), 251–279. https://doi.org/10.1037/0033-295X.105.2.251</bibtext> </blist> <blist> <bibtext> Grgurović M., Hegelheimer V. (2007). Help options and multimedia listening: Students' use of subtitles and the transcript. Language Learning & Technology, 11(1), 45–66.</bibtext> </blist> <blist> <bibtext> Guillory H. G. (1998). The effects of keyword captions to authentic French video on learner comprehension. CALICO Journal, 15(1–3), 89–108.</bibtext> </blist> <blist> <bibtext> Huang H.-C., Eskey D. E. (2000). The effects of closed-captioned television on the listening comprehension of intermediate English as a second language (ESL) students. Journal of Educational Technology Systems, 28, 75–96.</bibtext> </blist> <blist> <bibtext> Hulstijn J. H. (2003). Connectionist models of language processing and the training of listening skills with the aid of multimedia software. Computer Assisted Language Learning, 16(5), 413–425.</bibtext> </blist> <blist> <bibtext> Johnson K. (2004). Massive reduction in conversational American English. In Yoneyama K., Maekawa K. (Eds.), Spontaneous speech: Data and analysis. Proceedings of the 1st session of the 10th international symposium (pp. 29–54). Tokyo, Japan: The National International Institute for Japanese Language.</bibtext> </blist> <blist> <bibtext> Kuo F.-L. (2012). Factors affecting Chinese EFL learners' spoken word recognition. NCUE Journal of Humanities, 6, 1–14.</bibtext> </blist> <blist> <bibtext> Leon-Carrion J., García-Orza J., Pérez-Santamaría F. J. (2004). Development of the inhibitory component of the executive functions in children and adolescents. International Journal of Neuroscience, 114(10), 1291–1311. https://doi.org/10.1371/journal.pone.0077770</bibtext> </blist> <blist> <bibtext> Liang D. (2015). Chinese learners' pronunciation problems and listening difficulties in English connected speech. Asian Social Science, 11(16), 98–106. https://doi.org/10.5539/ass.v11n16p98</bibtext> </blist> <blist> <bibtext> Liao C. Y-W., Yeldham M. (2015). Taiwanese high school EFL teachers' perceptions of their listening instruction. The Asian Journal of Applied Linguistics, 2(2), 92–101.</bibtext> </blist> <blist> <bibtext> Lukas S., Philipp A. M., Koch I. (2010). Switching attention between modalities: Further evidence for visual dominance. Psychological Research, 74, 255–267. https://doi.org/10.1007/s00426-009-0246-y</bibtext> </blist> <blist> <bibtext> Lwo L., Lin C.-T. (2012). The effects of captions in teenagers' multimedia L2 learning. Recall, 24(2), 188–208. https://doi.org/10.1017/S0958344012000067</bibtext> </blist> <blist> <bibtext> Maleki A., Rad M. S. (2011). The effect of visual and textual accompaniments to verbal stimuli on the listening comprehension test performance of Iranian high and low proficient EFL learners. Theory and Practice in Language Studies, 1(1), 28–36.</bibtext> </blist> <blist> <bibtext> Mao H.-Z., Chen H.-Y. (2013). Exploring elision of schwa of /ə/ in English utterances by C & U English Majors. International Journal of Applied Linguistics & English Literature, 2(1), 117–125. https://doi.org/10.7575/ijalel.v.2n.1p.117</bibtext> </blist> <blist> <bibtext> Markham P. L., Peter L. (2003). The influence of English language and Spanish language captions on foreign language listening/reading comprehension. Journal of Educational Technology Systems, 31(3), 331–341. https://doi.org/10.2190/BHUH-420B-FE23-ALA0</bibtext> </blist> <blist> <bibtext> Marslen-Wilson W. D. (1990). Activation, competition, and frequency in lexical access. In Altmann G. T. M. (Ed.), Cognitive models of speech processing: Psycholinguistic and computational perspectives (pp. 148–172). MIT Press.</bibtext> </blist> <blist> <bibtext> Mattys S. L., Melhorn J. F. (2007). Sentential, lexical, and acoustic effects on the perception of word boundaries. Journal of the Acoustical Society of America, 122(1), 554–567. https://doi.org/10.1121/1.2735105</bibtext> </blist> <blist> <bibtext> Mitterer H., Russell K. (2013). How phonological reductions sometimes help the listener. Journal of Experimental Psychology: Learning, Memory, and Cognition, 39(3), 977–984. https://doi.org/10.1037/a0029196</bibtext> </blist> <blist> <bibtext> Pitt M. A., Samuel A. G. (1995). Lexical and sublexical feedback in auditory word recognition. Cognitive Psychology, 29(2), 149–188. https://doi.org/10.1006/cogp.1995.1014</bibtext> </blist> <blist> <bibtext> Pujola J.-T. (2002). CALLing for help: Researching language learning strategies using help facilities in a web-based multimedia program. Recall, 14(2), 235–262. https://doi.org/10.1017/S0958344002000423</bibtext> </blist> <blist> <bibtext> Ranbom L. J., Connine C. M. (2007). Lexical representation of phonological variation in spoken word recognition. Journal of Memory and Language, 57(2), 273–298. https://doi.org/10.1016/j.jml.2007.04.001</bibtext> </blist> <blist> <bibtext> Roach P. (2008). English phonetics and phonology: A practical course. Foreign Language Teaching and Research Press.</bibtext> </blist> <blist> <bibtext> Schnotz W., Kurschner C. (2007). A reconsideration of cognitive load theory. Educational Psychology Review, 19, 469–508. https://doi.org/10.1007/s10648-007-9053-4</bibtext> </blist> <blist> <bibtext> Setter J., Mok P., Low E. L., Zuo D., Tan A. (2014). Word juncture characteristics in world Englishes: A research report. World Englishes, 33, 278–291.</bibtext> </blist> <blist> <bibtext> Shockey L. (2003). Sound patterns of spoken English. Blackwell.</bibtext> </blist> <blist> <bibtext> Vandergrift L. (2011). Second language listening: Presage, process, product, and pedagogy. In Hinkel E. (Ed.), Handbook of research in second language teaching and learning (pp. 455–471). Routledge.</bibtext> </blist> <blist> <bibtext> Vanderplank R. (2016). "Effects of" and "effects with" captions: How exactly does watching a TV programme with same-language subtitles make a difference in language learners? Language Teaching, 49(2), 235–250.</bibtext> </blist> <blist> <bibtext> Vitevitch M. S., Luce P. A. (1999). Probabilistic phonotactics and neighborhood activation in spoken word recognition. Journal of Memory and Language, 40(3), 374–408. https://doi.org/10.1006/jmla.1998.2618</bibtext> </blist> <blist> <bibtext> Winke P., Gass S., Sydorenko T. (2010). The effects of captioning videos used for foreign language listening activities. Language Learning & Technology, 14(1), 65–86. <ulink href="http://llt.msu.edu/vol14num1/winkegasssydorenko.pdf">http://llt.msu.edu/vol14num1/winkegasssydorenko.pdf</ulink></bibtext> </blist> <blist> <bibtext> Wong S. W. L., Dealey J., Mok P., Leung V. W. -H. (in press). Production of English connected speech phonological processes: An assessment of Cantonese ESL learners' difficulties in obtaining native-like speech. The Language Learning Journal. https://doi.org/10.1080/09571736.2019.1642372</bibtext> </blist> <blist> <bibtext> Wong S. W. L., Mok P. P. K., Chung K. K. -H., Leung V. W. H., Bishop D. V. M., Chow B. W. -Y. (2017a). Perception of native English reduced forms in Chinese learners: Its role in listening comprehension and its phonological correlates. TESOL Quarterly, 51(1), 7–31. https://doi.org/10.1002/tesq.273</bibtext> </blist> <blist> <bibtext> Wong S. W. L., Tsui J. K. Y., Chow B. W. -Y., Leung V. W. H., Mok P., Chung K. K. -H., (2017b). Perception of native English reduced forms in adverse environments by Chinese undergraduate students. Journal of Psycholinguistic Research, 46(5), 1149–1165. doi: 10.1007/s10936-017-9486-y</bibtext> </blist> <blist> <bibtext> Yang J. C., Chang P. (2014). Captions and reduced forms instruction: The impact on EFL students' listening comprehension. Recall, 26(1), 44–61. https://doi.org/10.1017/S0958344013000219</bibtext> </blist> <blist> <bibtext> Yeldham M. (2018). Viewing L2 captioned videos: What's in it for the listener? Computer Assisted Language Learning, 31(4), 367–389. https://doi.org/10.1080/09588221.2017.1406956</bibtext> </blist> </ref> <ref id="AN0144336864-27"> <title> Footnotes </title> <blist> <bibtext> Data Availability The data that support the findings of this study are available from the corresponding author (S.W.L.W.), upon reasonable request.</bibtext> </blist> <blist> <bibtext> Declaration of Conflicting Interests The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibtext> Funding The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was funded by Early Career Scheme of the Research Grants Council (RGC) of Hong Kong (ECS 846212).</bibtext> </blist> <blist> <bibtext> ORCID iD Simpson W. L. Wong</bibtext> </blist> <blist> <bibtext>Graph https://orcid.org/0000-0002-6606-6382</bibtext> </blist> </ref> <aug> <p>By Simpson W. L. Wong; Cherry C. Y. Lin; Isabella S. Y. Wong and Anisa Cheung</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib52" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib33" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib25" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib27" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib60" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib31" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib11" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib56" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib47" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib54" firstref="ref13"></nolink> <nolink nlid="nl11" bibid="bib61" firstref="ref14"></nolink> <nolink nlid="nl12" bibid="bib13" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib14" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib16" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib28" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib43" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib15" firstref="ref20"></nolink> <nolink nlid="nl18" bibid="bib24" firstref="ref21"></nolink> <nolink nlid="nl19" bibid="bib46" firstref="ref22"></nolink> <nolink nlid="nl20" bibid="bib48" firstref="ref23"></nolink> <nolink nlid="nl21" bibid="bib45" firstref="ref24"></nolink> <nolink nlid="nl22" bibid="bib58" firstref="ref25"></nolink> <nolink nlid="nl23" bibid="bib18" firstref="ref26"></nolink> <nolink nlid="nl24" bibid="bib41" firstref="ref27"></nolink> <nolink nlid="nl25" bibid="bib36" firstref="ref28"></nolink> <nolink nlid="nl26" bibid="bib34" firstref="ref31"></nolink> <nolink nlid="nl27" bibid="bib51" firstref="ref32"></nolink> <nolink nlid="nl28" bibid="bib22" firstref="ref33"></nolink> <nolink nlid="nl29" bibid="bib49" firstref="ref34"></nolink> <nolink nlid="nl30" bibid="bib53" firstref="ref38"></nolink> <nolink nlid="nl31" bibid="bib10" firstref="ref39"></nolink> <nolink nlid="nl32" bibid="bib57" firstref="ref40"></nolink> <nolink nlid="nl33" bibid="bib40" firstref="ref41"></nolink> <nolink nlid="nl34" bibid="bib12" firstref="ref43"></nolink> <nolink nlid="nl35" bibid="bib30" firstref="ref45"></nolink> <nolink nlid="nl36" bibid="bib42" firstref="ref46"></nolink> <nolink nlid="nl37" bibid="bib38" firstref="ref49"></nolink> <nolink nlid="nl38" bibid="bib29" firstref="ref50"></nolink> <nolink nlid="nl39" bibid="bib32" firstref="ref52"></nolink> <nolink nlid="nl40" bibid="bib59" firstref="ref54"></nolink> <nolink nlid="nl41" bibid="bib23" firstref="ref56"></nolink> <nolink nlid="nl42" bibid="bib20" firstref="ref57"></nolink> <nolink nlid="nl43" bibid="bib17" firstref="ref63"></nolink> <nolink nlid="nl44" bibid="bib26" firstref="ref64"></nolink> <nolink nlid="nl45" bibid="bib35" firstref="ref86"></nolink> <nolink nlid="nl46" bibid="bib21" firstref="ref87"></nolink> <nolink nlid="nl47" bibid="bib50" firstref="ref88"></nolink> <nolink nlid="nl48" bibid="bib55" firstref="ref89"></nolink> <nolink nlid="nl49" bibid="bib44" firstref="ref92"></nolink> <nolink nlid="nl50" bibid="bib39" firstref="ref93"></nolink> <nolink nlid="nl51" bibid="bib37" firstref="ref96"></nolink> <nolink nlid="nl52" bibid="bib19" firstref="ref98"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1259676
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: The Differential Effects of Subtitles on the Comprehension of Native English Connected Speech Varying in Types and Word Familiarity
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Wong%2C+Simpson+W%2E+L%2E%22">Wong, Simpson W. L.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6606-6382">0000-0002-6606-6382</externalLink>)<br /><searchLink fieldCode="AR" term="%22Lin%2C+Cherry+C%2E+Y%2E%22">Lin, Cherry C. Y.</searchLink><br /><searchLink fieldCode="AR" term="%22Wong%2C+Isabella+S%2E+Y%2E%22">Wong, Isabella S. Y.</searchLink><br /><searchLink fieldCode="AR" term="%22Cheung%2C+Anisa%22">Cheung, Anisa</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22SAGE+Open%22"><i>SAGE Open</i></searchLink>. Apr-Jun 2020 10(2).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: http://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 13
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2020
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Translation%22">Translation</searchLink><br /><searchLink fieldCode="DE" term="%22Connected+Discourse%22">Connected Discourse</searchLink><br /><searchLink fieldCode="DE" term="%22Video+Technology%22">Video Technology</searchLink><br /><searchLink fieldCode="DE" term="%22Decoding+%28Reading%29%22">Decoding (Reading)</searchLink><br /><searchLink fieldCode="DE" term="%22Listening+Comprehension+Tests%22">Listening Comprehension Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Phonology%22">Phonology</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Adolescents%22">Adolescents</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Auditory+Perception%22">Auditory Perception</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Hong+Kong%22">Hong Kong</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/2158244020924378
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2158-2440
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Connected speech produced by native speakers poses a challenge to second language learners. Video subtitles have been found to assist the decoding of English connected speech for learners of English as a foreign language (EFL). However, the presence of subtitles may divert the listeners' attention to the visual cues while paying less attention to the speech signals. To test this proposal, we employed a bi-modal audio-visual listening test and examined whether EFL listeners were able to correctly identify the connected speech when misleading subtitles were present. We further tested whether connected speech with words of lower frequency further reduced the accuracy rate. Twenty-eight adolescent EFL learners, all with more than 10 years of experiences in learning English in schools, were tested with three major types of connected speech phonological processes, namely assimilation, elision, and juncture. The results of statistical analyses showed that matched and mismatched subtitles facilitated the comprehension of both familiar and unfamiliar connected speech. Error analyses revealed the degree of item-specific variations across the three types of connected speech processes as well as across the three subtitling conditions. This research provides insights on the immediate and long-term impact of subtitles on the decoding of English connected speech.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2020
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1259676
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1259676
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/2158244020924378
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
    Subjects:
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Translation
        Type: general
      – SubjectFull: Connected Discourse
        Type: general
      – SubjectFull: Video Technology
        Type: general
      – SubjectFull: Decoding (Reading)
        Type: general
      – SubjectFull: Listening Comprehension Tests
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Phonology
        Type: general
      – SubjectFull: Secondary School Students
        Type: general
      – SubjectFull: Adolescents
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Auditory Perception
        Type: general
      – SubjectFull: Hong Kong
        Type: general
    Titles:
      – TitleFull: The Differential Effects of Subtitles on the Comprehension of Native English Connected Speech Varying in Types and Word Familiarity
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Wong, Simpson W. L.
      – PersonEntity:
          Name:
            NameFull: Lin, Cherry C. Y.
      – PersonEntity:
          Name:
            NameFull: Wong, Isabella S. Y.
      – PersonEntity:
          Name:
            NameFull: Cheung, Anisa
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2020
          Identifiers:
            – Type: issn-electronic
              Value: 2158-2440
          Numbering:
            – Type: volume
              Value: 10
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: SAGE Open
              Type: main
ResultId 1