Human Ratings and Complexity, Accuracy, and Fluency (CAF) Indices: A Correlational Study of a Standardised Monologic English-Speaking Test in China
Saved in:
| Title: | Human Ratings and Complexity, Accuracy, and Fluency (CAF) Indices: A Correlational Study of a Standardised Monologic English-Speaking Test in China |
|---|---|
| Language: | English |
| Authors: | Hengzhi Hu (ORCID |
| Source: | SAGE Open. 2025 15(2). |
| Availability: | SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com |
| Peer Reviewed: | Y |
| Page Count: | 13 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Higher Education Postsecondary Education |
| Descriptors: | Foreign Countries, Speech Tests, English (Second Language), Accuracy, Language Fluency, Standardized Tests, Difficulty Level, Objective Tests, Examiners, College Students |
| Geographic Terms: | China |
| DOI: | 10.1177/21582440251343944 |
| ISSN: | 2158-2440 |
| Abstract: | Foreign language (L2) learners' speaking proficiency is often quantified using two dimensions: intuitive human ratings and analytical, linguistic complexity, accuracy, and fluency (CAF) indices. While previous research and assessment practices have predominantly focused on either the subjective approach to L2 speaking or the objective one, it is essential to establish an association between these two seemingly contradictory assessment methods to enhance and promote more credible assessment judgements. To this end, 160 recordings from a monologic task of a standardised English test in China were analysed to quantify CAF in the present study, and the scores were then compared with human ratings given by qualified examiners. Correlation and regression analyses demonstrated that human ratings were positively correlated with speaking fluency, with the number of pauses produced by a candidate being the most significant predictor of human-judged scores. Speaking complexity also positively predicted human ratings, with examiners tending to focus more on grammatical complexity than lexical complexity. In contrast, no correlations were found between human ratings and speaking accuracy. The findings of this study reinforce the possibility of "halo" effects on human raters in L2 assessment and suggest that rater training should focus on helping examiners recognise and mitigate such potential effects. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1477239 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGURRRPxo1yM3SypgtZaYFxAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDNyKsWK_zBKkUC8cgIBEICBm2zMLtTBaifqIRMllvJ97DIMYpJ2zOIJwaepTHODY2a4FCGFS_lSZTfKDFigcuHh_n78Nw4DHRLiyGpNKEq6DXRq-CL-3ZDbEMoIPbf4dIawSWp-1wtmCKu1LWP1ln1TtQF7PnluRSUI_zdQNKZJpPTTPNY-ue32HAgHubVevbIMmHCFZ_9H6NCeO3-OXrQ2cSOv5jPvTHkZNH7q Text: Availability: 1 Value: <anid>AN0186372652;[kbz6]01apr.25;2025Jul07.02:57;v2.2.500</anid> <title id="AN0186372652-1">Human Ratings and Complexity, Accuracy, and Fluency (CAF) Indices: A Correlational Study of a Standardised Monologic English-Speaking Test in China </title> <p>Foreign language (L2) learners' speaking proficiency is often quantified using two dimensions: intuitive human ratings and analytical, linguistic complexity, accuracy, and fluency (CAF) indices. While previous research and assessment practices have predominantly focused on either the subjective approach to L2 speaking or the objective one, it is essential to establish an association between these two seemingly contradictory assessment methods to enhance and promote more credible assessment judgements. To this end, 160 recordings from a monologic task of a standardised English test in China were analysed to quantify CAF in the present study, and the scores were then compared with human ratings given by qualified examiners. Correlation and regression analyses demonstrated that human ratings were positively correlated with speaking fluency, with the number of pauses produced by a candidate being the most significant predictor of human-judged scores. Speaking complexity also positively predicted human ratings, with examiners tending to focus more on grammatical complexity than lexical complexity. In contrast, no correlations were found between human ratings and speaking accuracy. The findings of this study reinforce the possibility of "halo" effects on human raters in L2 assessment and suggest that rater training should focus on helping examiners recognise and mitigate such potential effects.</p> <p>Keywords: assessment; English speaking; human ratings; CAF; standardised tests</p> <hd id="AN0186372652-2">Introduction</hd> <p>In the literature on foreign language (L2) assessment and testing, the conceptualisation of L2 proficiency continues to expand and evolve. Accordingly, the notion of L2 speaking proficiency has undergone considerable changes, with researchers and educators continually reflecting on which core language skills should be evaluated and how they can be more effectively assessed. Today, analytical complexity, accuracy, and fluency (CAF) indices of L2 speaking proficiency underpin countless linguistic analyses. However, speaking is a complex process that is largely communication-oriented, and the extent to which a speaker has achieved the communicative goal cannot always be effectively measured by CAF indicators alone ([<reflink idref="bib69" id="ref1">69</reflink>]). This limitation underscores the importance of subjective human ratings in complementing CAF analyses to provide further insights into L2 learners' speaking proficiency ([<reflink idref="bib3" id="ref2">3</reflink>]; [<reflink idref="bib57" id="ref3">57</reflink>]). Although these two measurement methods have contrasting natures and application processes, and have rarely been employed together in previous research, the need to combine them is increasingly recognised, with initiatives to use human ratings to establish the concurrent validity of CAF analyses now being undertaken ([<reflink idref="bib45" id="ref4">45</reflink>]; [<reflink idref="bib72" id="ref5">72</reflink>]; [<reflink idref="bib79" id="ref6">79</reflink>]). However, little is known about how CAF indices and human ratings correlate with each other. While it may seem commonsensical that moderate to strong correlations should exist between different language proficiency measurements of the same assessment task, the inherent subjectivity of human raters often leads them to unconsciously apply their expectations and experiences ([<reflink idref="bib19" id="ref7">19</reflink>]; [<reflink idref="bib25" id="ref8">25</reflink>]), creating potential scoring problems and necessitating efforts to mitigate the impact of subjectivity.</p> <p>To address this, the present study, underpinned by a positivist paradigm, aims to identify possible correlations between objective measurements of speaking CAF and subjective ratings by qualified examiners in the same monologue tasks. Contextualised within China's standardised English-speaking assessment system, namely the College English Test-Spoken English Test (CET-SET), this study responds to the call for local research on the effectiveness and reliability of assessment methods tailored to the Chinese educational context ([<reflink idref="bib31" id="ref9">31</reflink>]; [<reflink idref="bib37" id="ref10">37</reflink>]), aligning with the global academic trend towards greater emphasis on context-specific assessment practices of CAF and subjective ratings ([<reflink idref="bib69" id="ref11">69</reflink>]). The CET-SET, as part of the College English Test (CET) series, plays a crucial role in evaluating the English proficiency of university students across China. It is widely regarded as a benchmark for assessing spoken English skills among Chinese students and has a significant impact on their academic and professional futures ([<reflink idref="bib81" id="ref12">81</reflink>]). Despite its importance, there is limited research exploring how well this assessment captures the complexities of spoken English proficiency ([<reflink idref="bib80" id="ref13">80</reflink>]), particularly when comparing the objective CAF measures with the more subjective human ratings provided by examiners. Given the influence of the CET-SET on educational outcomes in China, understanding the alignment—or lack thereof—between these two assessment approaches is essential. By situating this study within the CET-SET framework, the research not only addresses a gap in the literature but also contributes to the ongoing discourse on how best to assess L2 speaking proficiency in a way that is both scientifically robust and contextually appropriate.</p> <hd id="AN0186372652-3">Literature Review</hd> <p></p> <hd id="AN0186372652-4">CAF Indexes</hd> <p>In the 1980s, the development of second language acquisition (SLA) theories and research sparked a debate about the distinction between accurate and fluent oral language use in classroom contexts. Researchers such as [<reflink idref="bib8" id="ref14">8</reflink>] sought to differentiate accuracy-oriented L2 learning activities, which emphasised linguistic forms and grammatically correct language use, from fluency-centred activities that focused on spontaneous language production and meaning communication. However, it was not until the late 20th century that the original accuracy and fluency dichotomy was supplemented with complexity as another essential indicator of L2 learners' speaking proficiency, leading to the proposal of the CAF triad as a model of L2 proficiency dimensions ([<reflink idref="bib66" id="ref15">66</reflink>]). With the advancement of cognitive psychology and psycholinguistics, CAF has become a crucial variable in L2 research, and it is regarded "as principal epiphenomena of the psycholinguistic mechanism and processes underlying the acquisition, representation, and processing of L2 knowledge" ([<reflink idref="bib28" id="ref16">28</reflink>], p. 462). Complexity and accuracy demonstrate an L2 learner's current level of language proficiency, while fluency denotes their control of target language (TL) use. However, several controversial issues have arisen in CAF research, one of which is the differing views on what CAF imply and how they should be quantified ([<reflink idref="bib10" id="ref17">10</reflink>]; [<reflink idref="bib29" id="ref18">29</reflink>]; [<reflink idref="bib43" id="ref19">43</reflink>]; [<reflink idref="bib54" id="ref20">54</reflink>]). This necessitates a clear identification of CAF definitions and measurement methods in what follows.</p> <hd id="AN0186372652-5">Complexity</hd> <p>Complexity appears to be the most intricate construct within the CAF triad, and this complication arises from the various aspects it encompasses. According to [<reflink idref="bib54" id="ref21">54</reflink>] review, complexity is categorised into developmental complexity (i.e. the order of the emergence and mastery of linguistic forms in SLA), cognitive complexity (i.e. a learner's perception of the difficulty of the TL), and linguistic complexity (i.e. the linguistic properties related to the TL). Developmental and cognitive complexity remain at a relative level, whereas linguistic complexity, at an absolute level, has been the focus of most SLA research, popularising the definition that complexity concerns "the extent to which the language produced in performing a task is elaborate and varied" ([<reflink idref="bib16" id="ref22">16</reflink>], p. 340).</p> <p>The notion that complexity involves "the number of discrete components that a language feature or a language system consists of and the number of connections between the different components" ([<reflink idref="bib9" id="ref23">9</reflink>], p. 24)—that is, the number of word strings in language structures and the degree to which these structures can form word strings—has brought the measurement of lexical and linguistic complexity into focus. Popular methods for measuring lexical complexity include calculating the Type-Token Ratio (TTR) (i.e. dividing the number of word types in a discourse by the number of words in the same sample) or the Mean Segmental TTR (i.e. calculating the mean of TTR in segments of a given size). However, the former is believed to be influenced by corpus size, and the latter may produce invalid results due to inefficient data segmentation ([<reflink idref="bib20" id="ref24">20</reflink>]). In response, a computer-calculated D score has been proposed and utilised in analysing lexical richness, which offers a complex adjustment to overcome the limitations of traditional measurement methods "by taking 100 random samples of 35 to 50 tokens (without replacement), calculating D using the original formula for TTR, and then calculating an average D score" ([<reflink idref="bib65" id="ref25">65</reflink>], p. 30).</p> <p>Regarding grammatical complexity, earlier studies (see, for example, [<reflink idref="bib51" id="ref26">51</reflink>]; [<reflink idref="bib53" id="ref27">53</reflink>]), particularly those conducted over two decades ago, primarily employed T-unit and C-unit as the basis for measuring unit length and clausal density. However, it has been argued that T-units do not capture the multifaceted nature of spoken data, especially when pauses and repairs are involved, and that the C-unit, as an adjustment of the T-unit, remains poorly defined. Therefore, the concept of the AS-unit has been proposed to measure unit length and clausal density in grammatical complexity at three levels: length, sub-clausal, and subordination. The corresponding indicators are Mean Length of AS-units (MLAS), Mean Length of Clauses (MLC), and Ratio of Subordination (RS), respectively ([<reflink idref="bib21" id="ref28">21</reflink>]; [<reflink idref="bib56" id="ref29">56</reflink>]).</p> <hd id="AN0186372652-6">Accuracy</hd> <p>Regarding accuracy, classic interpretations have also made it a controversial issue within the applied linguistics community. According to [<reflink idref="bib17" id="ref30">17</reflink>], accuracy is "the ability to avoid error in performance, possibly reflecting higher levels of control in the language as well as a conservative orientation, that is, avoidance of challenging structures that might provoke error" (p. 475). From this perspective, accuracy seems to be more closely associated with communication strategies—referring to the purposeful and planned use of language to achieve specific information-conveying goals—than with language utilisation itself ([<reflink idref="bib60" id="ref31">60</reflink>]). In contrast to this viewpoint, accuracy is also simply considered to be the error-free state of one's output in the TL, or one's ability to produce error-free speech ([<reflink idref="bib29" id="ref32">29</reflink>]), and this stance appears to be the prevailing definition underlying most CAF research ([<reflink idref="bib24" id="ref33">24</reflink>]).</p> <p>The measurement of accuracy has focused on two dimensions: global accuracy, which accounts for all the errors in speech, and specific accuracy, which examines only the errors of particular linguistic forms ([<reflink idref="bib34" id="ref34">34</reflink>]). However, the latter can be problematic because different research instruments, such as speaking tasks or tests with varying formats and topics, may prompt the use of contrasting linguistic forms, making it difficult to establish a clear understanding of L2 learners' specific accuracy. Consequently, global accuracy, which encompasses errors in syntax, morphology, and lexis, has been adopted in the present study. A common method for quantifying global accuracy is to calculate the number of errors per 100 words ([<reflink idref="bib18" id="ref35">18</reflink>]), although this approach may be simplistic and susceptible to the influence of speech length. In response, researchers have proposed two more sophisticated variables: the percentage of error-free clauses (PEC) and the percentage of error-free AS-units (PEA), suggesting that these could provide a more precise and valid account of an L2 learner's speaking accuracy ([<reflink idref="bib18" id="ref36">18</reflink>]).</p> <hd id="AN0186372652-7">Fluency</hd> <p>Although fluency is a long-established variable, researchers have divergent interpretations of it. Some (e.g. [<reflink idref="bib22" id="ref37">22</reflink>]; [<reflink idref="bib32" id="ref38">32</reflink>]; [<reflink idref="bib63" id="ref39">63</reflink>]) adhere to the broad layman conception described by [<reflink idref="bib44" id="ref40">44</reflink>] as the ability to use the TL smoothly and spontaneously, while others prefer a narrower but more sophisticated understanding that fluency concerns "the production of language in real time without undue pausing or hesitation" ([<reflink idref="bib18" id="ref41">18</reflink>], p. 139). An even broader perspective is proposed by [<reflink idref="bib24" id="ref42">24</reflink>], who suggest that communication fluency involves the interplay of personal factors, prosodic elements, and educational and cultural levels, defining it as "a multifactorial construct which demands cognition, aptitude, motivation, and experience" in TL processing and performance (p. 3). Though this view is reasonable, the classical interpretation that fluency symbolises real-time language processing has underpinned most previous studies and the measurement of fluency.</p> <p>L2 fluency is typically assessed on three dimensions: speed, breakdown, and repair ([<reflink idref="bib68" id="ref43">68</reflink>]). However, the first dimension, represented by speech rate or articulation rate, may be questionable for several reasons. First, speech speed is subject to individual characteristics and even an L2 learner's first language (L1) rate; for example, [<reflink idref="bib7" id="ref44">7</reflink>] and [<reflink idref="bib64" id="ref45">64</reflink>] regression analyses have shown that L1 speech rate can be a strong predictor of L2 rate. Second, a high speech or articulation rate may be accompanied by numerous pauses, as [<reflink idref="bib1" id="ref46">1</reflink>], referencing [<reflink idref="bib27" id="ref47">27</reflink>] work, suggest that these two variables may be negatively correlated. Third, those who are inclined to use simple words may exhibit a higher level of fluency than those who use complex polysyllabic words, an assumption consistent with [<reflink idref="bib74" id="ref48">74</reflink>] inference from their classroom studies that students may display a slower speech rate when reading polysyllabic words. Therefore, only breakdown, which refers to the average length of filled and unfilled pauses over 250 milliseconds, and repairs, which include repetitions, reformulations, replacements, and false starts, are used in the present study. Fluency is thus indicated by the number of pauses (NP) and the number of repairs (NR) ([<reflink idref="bib68" id="ref49">68</reflink>]).</p> <hd id="AN0186372652-8">Human Ratings</hd> <p>CAF measurement has been commonly used in speech data analysis, as it allows researchers to capture an objective and detailed picture of L2 learners' speaking proficiency from various angles. However, its disadvantages—such as time-consuming processes, the inability to determine whether learners have achieved the required communicative goals, and a lack of ecological validity—have discouraged educators and even researchers from employing this approach. Instead, they have often relied on human evaluation, a long-established educational practice in L2 assessment. In most speaking assessment tasks, particularly in standardised tests or those intended for summative purposes, examiners typically follow preset marking rubrics to evaluate a candidate's performance from either a holistic or an analytic perspective ([<reflink idref="bib42" id="ref50">42</reflink>]). Although this practice has been criticised for lacking objectivity and clarity and being susceptible to personal biases, researchers argue that using speaking rubrics not only engages learners in the assessment process ([<reflink idref="bib58" id="ref51">58</reflink>]) but also provides students and teachers with valuable information, such as whether students have achieved the communicative goals of the tasks ([<reflink idref="bib2" id="ref52">2</reflink>]).</p> <p>Research on the individual investigation of human rating and CAF is abundant, but the studies that have combined these two measurement methods warrant particular attention. Although such research is limited, the need to "compare subjective estimates of [L2] proficiency with its objective measures" has been recognised for some time ([<reflink idref="bib4" id="ref53">4</reflink>], p. 135). The available body of research can be broadly divided into two streams. The first stream includes studies in which CAF scores and human ratings have been used simultaneously to complement each other. For example, a study we are currently conducting analyses Chinese university students' oral English proficiency under different pedagogical approaches using CAF measurements to gain linguistic insights, as well as human ratings, which are compared with China's Standards of English Language Ability, a national framework that identifies Chinese English as a foreign language (EFL) learners' proficiency levels. This design has been largely inspired by [<reflink idref="bib19" id="ref54">19</reflink>] research, where CAF analyses were employed to examine EFL learners' language proficiency in detail, in conjunction with human ratings linked to the Common European Framework of Reference for Languages (CEFR). However, this type of research remains rare, and researchers suggest that the CAF triad offers advantages over marking rubrics or established proficiency frameworks "to benchmark and assess foreign language production, as it employs multiple indicators underlying each aspect," the quantification of which allows comparisons of learners' L2 development over time ([<reflink idref="bib73" id="ref55">73</reflink>], 1.3 Measurement of foreign language production, para. 2).</p> <p>The second stream of research focuses on exploring the relationships between CAF scores and human ratings or human-judged proficiency levels, with most previous studies attempting to establish a predictive link between these two constructs. For instance, [<reflink idref="bib77" id="ref56">77</reflink>] study examined data from the Aptis speaking test and compared various CAF features across different CEFR levels, revealing distinct CAF characteristics between lower and higher proficiency levels, as well as a moderate to strong relationship between most CAF indices and CEFR levels. This finding is similar to that of a later study by [<reflink idref="bib78" id="ref57">78</reflink>], which used the same test corpus but focused on fluency and CEFR proficiency levels, as well as [<reflink idref="bib39" id="ref58">39</reflink>] study, where CAF indices varied significantly across certain CEFR levels. In comparison, [<reflink idref="bib33" id="ref59">33</reflink>] study emphasised speaking complexity and accuracy. By analysing data from EFL learners at different CEFR levels, the researcher concluded that grammatical complexity might not necessarily correlate positively with human-judged proficiency levels, in contrast to the moderately high correlation between accuracy and human ratings. [<reflink idref="bib57" id="ref60">57</reflink>] recent research has been particularly influential in this field. By comparing human ratings of Japanese university students' performance in a monologic oral English test with CAF measures, the researcher found that human ratings were primarily correlated with speaking fluency, as measured by the mean length of pauses, duration of syllables, mean length of run, and phonation time ratio. Additionally, speaking complexity (i.e. MLC, RS, and computer-calculated textual lexical diversity) and accuracy (i.e. PEC) indices also contributed to human ratings, although they explained a smaller portion of the variance compared to fluency measures.</p> <p>Although this category of studies is not the mainstream of research and may not be at the core of L2 assessment development, researchers have generally praised the reliability of CAF indices while highlighting the ecological validity of human ratings as an indicator of the impressionistic judgement formed in communication. However, with previous research presenting heterogeneous findings in different educational contexts—a principal reason being the use of divergent CAF indices without a consensus on standardised quantification of CAF ([<reflink idref="bib43" id="ref61">43</reflink>])—the disagreement over how intuitive human ratings by examiners can be correlated with objective CAF measures remains unresolved. This necessitates the gathering of more empirical evidence to enrich this body of knowledge, which was also the primary purpose of the study presented below.</p> <hd id="AN0186372652-9">Methods</hd> <p></p> <hd id="AN0186372652-10">Context and Participants</hd> <p>This study was part of a larger research project on the development trajectory of Chinese EFL learners' speaking proficiency. The present study adopted a quantitative approach in line with the stated research objective. Eighty students from a comprehensive Chinese higher education institution were recruited for the study with informed consent. In the original research, the participants were evenly allocated to a comparison group and an experimental group, but for the purposes of the present study, they were considered as a single cohort. The sample consisted of 42 males and 38 females, with an average age of 21 years. While the participants were enrolled in different programme streams, they all took College English, a compulsory EFL unit for undergraduates at the research site.</p> <hd id="AN0186372652-11">Data Collection</hd> <p>Data were collected from monologic tasks included in a mock CET-SET Band 6 (CET-SET6), a standardised test in China designed to assess university students' proficiency in spoken English with established validity, reliability, and fairness ([<reflink idref="bib48" id="ref62">48</reflink>]). The test comprises two types of tasks: dialogue and monologue. The monologue tasks were specifically chosen for this study because they provide a more controlled and focused environment for assessing individual speaking proficiency, allowing for a clearer analysis of the CAF measures in isolation from the interactive dynamics that are present in dialogue tasks ([<reflink idref="bib57" id="ref63">57</reflink>]). The test booklets used in the study were adapted from authentic test batteries by a group of CET experts ([<reflink idref="bib47" id="ref64">47</reflink>]) and had been used in our previous study, where their validity—particularly face validity—was established. This meant that the booklets had been reviewed by a group of experts who believed the test accurately measured what it was intended to measure. In the test, candidates were required to discuss a specific topic (e.g. the advantages or disadvantages of online learning/public transportation) provided by the examiner for approximately 2 min after 1 min of preparation, with their responses recorded for further analysis. The test was organised by two qualified examiners who adhered to the administration requirements of the CET-SET6. In the original study, the speaking test was administered both before and after the study to examine development patterns. The present study utilised the already collected data, with 160 speech samples from 80 students available for analysis.</p> <hd id="AN0186372652-12">Data Analysis</hd> <p>Based on the authentic marking rubric of CET-SET6 ([<reflink idref="bib13" id="ref65">13</reflink>]) (see Appendix A), the participants' responses were scored by two qualified raters. The rubric evaluates English speaking proficiency from three perspectives: Accuracy and Range, Length and Coherence, and Flexibility and Appropriateness, each of which is weighted with a maximum of five marks. However, only holistic scores, serving as an indicator of global proficiency, were analysed in this study because the original CET-SET6 does not report a candidate's sub-scores for each criterion. Furthermore, operationalising and analysing global proficiency tends to be more straightforward than other measurement methods ([<reflink idref="bib11" id="ref66">11</reflink>]), and it is common practice in research contexts to adopt an analytical approach to scoring but only report holistic scores for research and educational purposes ([<reflink idref="bib75" id="ref67">75</reflink>]).</p> <p>A pilot study was conducted prior to the main study, involving 34 students from the same research context, with the sample size deemed acceptable ([<reflink idref="bib38" id="ref68">38</reflink>]). The raters for the main study also participated in the pilot study. They received professional training provided by the official CET board, familiarised themselves with the assessment criteria, practised rating by viewing taped performances and setting cut scores according to seven levels—A+ (14.5–15), A (13.5–14.4), B+ (12.5–13.4), B (11–12.4), C+ (9.5–10.9), C (8–9.4), and D (below 7.9)—and explained and discussed the rationale behind the scores awarded. This process ensured the validity and reliability of the human ratings ([<reflink idref="bib49" id="ref69">49</reflink>]). Consequently, the pilot study indicated that adequate inter-rater reliability had been achieved for the instrument and each of the aforementioned criteria (see Table 1). Similarly, the statistics used in the subsequent analysis also demonstrated satisfactory reliability of the instrument in the main study.</p> <p>Table 1. Kappa Values of Instrument.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Criteria of CET-SET6 marking rubric&lt;/th&gt;&lt;th align="left"&gt;Pilot study&lt;/th&gt;&lt;th align="left"&gt;Main study&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Whole Instrument&lt;/td&gt;&lt;td&gt;.74&lt;/td&gt;&lt;td&gt;.78&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Accuracy and Range&lt;/td&gt;&lt;td&gt;.89&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Length and Coherence&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.83&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Flexibility and Appropriateness&lt;/td&gt;&lt;td&gt;.80&lt;/td&gt;&lt;td&gt;.81&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>In addition to human ratings, the participants' responses were transcribed and coded with the identification of AS-units, clauses, errors, pauses, and repairs for the CAF analysis. The data was initially transcribed by the first author of this paper and an English expert together. They then coded the transcripts individually, following the CAF measures. To ensure intercoder reliability, they compared and discussed their transcriptions through three rounds of coding. The final agreement for each code exceeded 90%, which is well above the minimum intercoder agreement threshold of 75% ([<reflink idref="bib50" id="ref70">50</reflink>]). Numerous methods exist for measuring CAF (see the compilation by [<reflink idref="bib41" id="ref71">41</reflink>]), but the ones used in this study (see Table 2) were those identified as appropriate in the literature review. After data collection, Pearson correlation and multiple regression analyses were performed using Statistical Package for the Social Sciences 25.0. The assumptions for correlation (i.e. normality, linearity, and homoscedasticity of data) and regression (i.e. a reasonable ratio of cases, normality, linearity, and homoscedasticity of data, no significant outliers or multicollinearity) tests were not violated in this study.</p> <p>Table 2. Computation Methods of CAF Indices.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Variable&lt;/th&gt;&lt;th align="left"&gt;Index&lt;/th&gt;&lt;th align="left"&gt;Justification&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td rowspan="3"&gt;Grammatical Complexity&lt;/td&gt;&lt;td&gt;MLAS = Number of Tokens/Number of AS-units&lt;/td&gt;&lt;td rowspan="3"&gt;&lt;xref ref-type="bibr" rid="bibr21"&gt;Foster et al. (2000)&lt;/xref&gt; and &lt;xref ref-type="bibr" rid="bibr56"&gt;Norris and Ortega (2009)&lt;/xref&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MLC = Number of Tokens/Number of Clauses&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;RS = Number of Clauses/Number of AS-units&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Lexical Complexity&lt;/td&gt;&lt;td&gt;D-score calculated using the VOCD function in the Computerized Language Analysis programme&lt;/td&gt;&lt;td&gt;&lt;xref ref-type="bibr" rid="bibr65"&gt;Siskova (2012)&lt;/xref&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;Accuracy&lt;/td&gt;&lt;td&gt;PEC = Number of Errorless Clauses/Number of Clauses&lt;/td&gt;&lt;td rowspan="2"&gt;&lt;xref ref-type="bibr" rid="bibr18"&gt;Ellis and Barkhuizen (2005)&lt;/xref&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PEA = Number of Errorless AS-units/Number of AS-units&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;Fluency&lt;/td&gt;&lt;td&gt;NP = Number of Pauses/Speaking Time in Seconds&lt;/td&gt;&lt;td rowspan="2"&gt;&lt;xref ref-type="bibr" rid="bibr68"&gt;Skehan (2009)&lt;/xref&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;NR = Number of Repairs/Speaking Time in Seconds&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0186372652-13">Findings</hd> <p>Despite the descriptive statistics presented in Table 3, the correlation statistics between the variables were particularly noteworthy. As shown in Table 4, a bivariate Pearson's product-moment correlation coefficient was calculated to determine the size and direction of the linear relationship between human ratings and CAF indices. Significant correlations were found between human ratings and the MLAS, MLC, RS, and D-score (<emph>p</emph> &lt;.05), which are indices of syntactic and lexical complexity. However, according to the rule of thumb regarding the Pearson correlation coefficient ([<reflink idref="bib59" id="ref72">59</reflink>]), the linear relationship between human ratings and speaking complexity could only be considered weak (0 &lt; <emph>r</emph> &lt;.3), although changes in complexity indices led to changes in human ratings in the same direction. Conversely, there was a strong negative correlation between human ratings and the indicators of speaking fluency (−.5 &lt; <emph>r</emph> &lt;.0), namely the NP and the NR. This suggested that the lower the NP and the NR, the higher the human ratings. Nevertheless, no correlations were found between human ratings and the PEC and PEA (<emph>p</emph> &gt;.05), indicating that changes in speaking accuracy did not result in any change in human ratings.</p> <p>Table 3. Descriptive Statistics of Speaking Tests.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Variable&lt;/th&gt;&lt;th align="left"&gt;Index&lt;/th&gt;&lt;th align="left"&gt;Minimum&lt;/th&gt;&lt;th align="left"&gt;Maximum&lt;/th&gt;&lt;th align="left"&gt;Mean&lt;/th&gt;&lt;th align="left"&gt;Standard deviation&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Human Ratings&lt;/td&gt;&lt;td /&gt;&lt;td&gt;7.000&lt;/td&gt;&lt;td&gt;15.500&lt;/td&gt;&lt;td&gt;11.228&lt;/td&gt;&lt;td&gt;1.521&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="3"&gt;Grammatical Complexity&lt;/td&gt;&lt;td&gt;MLAS&lt;/td&gt;&lt;td&gt;5.136&lt;/td&gt;&lt;td&gt;23.571&lt;/td&gt;&lt;td&gt;10.018&lt;/td&gt;&lt;td&gt;2.974&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MLC&lt;/td&gt;&lt;td&gt;4.125&lt;/td&gt;&lt;td&gt;8.833&lt;/td&gt;&lt;td&gt;5.927&lt;/td&gt;&lt;td&gt;.913&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;RS&lt;/td&gt;&lt;td&gt;1.000&lt;/td&gt;&lt;td&gt;3.571&lt;/td&gt;&lt;td&gt;1.684&lt;/td&gt;&lt;td&gt;.376&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Lexical Complexity&lt;/td&gt;&lt;td&gt;LCD&lt;/td&gt;&lt;td&gt;23.990&lt;/td&gt;&lt;td&gt;94.630&lt;/td&gt;&lt;td&gt;54.523&lt;/td&gt;&lt;td&gt;14.237&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;Accuracy&lt;/td&gt;&lt;td&gt;PEC&lt;/td&gt;&lt;td&gt;.222&lt;/td&gt;&lt;td&gt;2.182&lt;/td&gt;&lt;td&gt;.718&lt;/td&gt;&lt;td&gt;.195&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PEA&lt;/td&gt;&lt;td&gt;.063&lt;/td&gt;&lt;td&gt;1.000&lt;/td&gt;&lt;td&gt;.599&lt;/td&gt;&lt;td&gt;.181&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;Fluency&lt;/td&gt;&lt;td&gt;NP&lt;/td&gt;&lt;td&gt;.071&lt;/td&gt;&lt;td&gt;.742&lt;/td&gt;&lt;td&gt;.394&lt;/td&gt;&lt;td&gt;.125&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;NR&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.195&lt;/td&gt;&lt;td&gt;.062&lt;/td&gt;&lt;td&gt;.039&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 4. Correlations Amongst Human Ratings and CAF Indices.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Variable&lt;/th&gt;&lt;th /&gt;&lt;th align="left"&gt;Human Ratings&lt;/th&gt;&lt;th align="left"&gt;MLAS&lt;/th&gt;&lt;th align="left"&gt;MLC&lt;/th&gt;&lt;th align="left"&gt;RS&lt;/th&gt;&lt;th align="left"&gt;D-score&lt;/th&gt;&lt;th align="left"&gt;PEC&lt;/th&gt;&lt;th align="left"&gt;PEA&lt;/th&gt;&lt;th align="left"&gt;NP&lt;/th&gt;&lt;th align="left"&gt;NR&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Human Ratings&lt;/td&gt;&lt;td /&gt;&lt;td&gt;&amp;#8211;&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;MLAS&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.281&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;MLC&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.279&lt;/td&gt;&lt;td&gt;.621&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;RS&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.158&lt;/td&gt;&lt;td&gt;.833&lt;/td&gt;&lt;td&gt;.100&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.046&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.207&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;D-score&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.210&lt;/td&gt;&lt;td&gt;.444&lt;/td&gt;&lt;td&gt;.408&lt;/td&gt;&lt;td&gt;.289&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.008&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;PEC&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.154&lt;/td&gt;&lt;td&gt;.307&lt;/td&gt;&lt;td&gt;.143&lt;/td&gt;&lt;td&gt;.290&lt;/td&gt;&lt;td&gt;.084&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.052&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.071&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.289&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;PEA&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.043&lt;/td&gt;&lt;td&gt;.246&lt;/td&gt;&lt;td&gt;.211&lt;/td&gt;&lt;td&gt;.181&lt;/td&gt;&lt;td&gt;.193&lt;/td&gt;&lt;td&gt;.575&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.587&lt;/td&gt;&lt;td&gt;.002&lt;/td&gt;&lt;td&gt;.008&lt;/td&gt;&lt;td&gt;.022&lt;/td&gt;&lt;td&gt;.015&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;NP&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;&amp;#8722;.526&lt;/td&gt;&lt;td&gt;&amp;#8722;.339&lt;/td&gt;&lt;td&gt;&amp;#8722;.335&lt;/td&gt;&lt;td&gt;&amp;#8722;.207&lt;/td&gt;&lt;td&gt;&amp;#8722;.349&lt;/td&gt;&lt;td&gt;&amp;#8722;.188&lt;/td&gt;&lt;td&gt;&amp;#8722;.337&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.009&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.017&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td rowspan="2"&gt;NR&lt;/td&gt;&lt;td&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;&amp;#8722;.559&lt;/td&gt;&lt;td&gt;&amp;#8722;.187&lt;/td&gt;&lt;td&gt;&amp;#8722;.09&lt;/td&gt;&lt;td&gt;&amp;#8722;.179&lt;/td&gt;&lt;td&gt;&amp;#8722;.134&lt;/td&gt;&lt;td&gt;&amp;#8722;.079&lt;/td&gt;&lt;td&gt;.018&lt;/td&gt;&lt;td&gt;.342&lt;/td&gt;&lt;td&gt;&amp;#8211;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;.001&lt;/td&gt;&lt;td&gt;.018&lt;/td&gt;&lt;td&gt;.259&lt;/td&gt;&lt;td&gt;.023&lt;/td&gt;&lt;td&gt;.092&lt;/td&gt;&lt;td&gt;.323&lt;/td&gt;&lt;td&gt;.821&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;&amp;#8211;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Then, standard multiple regression analyses were conducted to estimate the proportion of variance in human ratings that could be accounted for by the CAF indices. The multiple regression model, with all eight predictors, produced <emph>R</emph><sups>2</sups> =.25, <emph>F</emph>(<reflink idref="bib8" id="ref73">8</reflink>, 159) = 6.43, <emph>p</emph> &lt;.001. As shown in Table 5, speaking complexity, as measured by the MLAS, MLC, RS, and D-score, had significant positive regression weights (<emph>p</emph> &lt;.05), suggesting that EFL learners with higher scores on these scales were expected to receive higher human ratings, after controlling for the other variables in the model. Conversely, the NP and NR scales, which are typical of speaking fluency, had significant negative weights (<emph>p</emph> &lt;.05), indicating that, after accounting for the other predictors, students with lower NP and NR scores were anticipated to receive higher human ratings. In contrast, the predictor variables related to speaking accuracy did not contribute significantly to the regression model (<emph>p</emph> &gt;.05). Considering the absolute values of the beta coefficients, NP had the largest predictive effect on human ratings, with the other predictors having a less pronounced effect on the dependent variable.</p> <p>Table 5. Multiple Regression of Human Ratings and CAF Indices.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" rowspan="2"&gt;Predicted variable&lt;/th&gt;&lt;th align="left" rowspan="2"&gt;Predictor variable&lt;/th&gt;&lt;th align="left" colspan="2"&gt;Unstandardised coefficients&lt;/th&gt;&lt;th align="left" colspan="2"&gt;Standardised coefficients&lt;/th&gt;&lt;th align="left" rowspan="2"&gt;Effect size&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;italic&gt;B&lt;/italic&gt;&lt;/th&gt;&lt;th align="left"&gt;Std. Error&lt;/th&gt;&lt;th align="left"&gt;Beta&lt;/th&gt;&lt;th align="left"&gt;Sig.&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td rowspan="6"&gt;Human Ratings&lt;/td&gt;&lt;td&gt;MLAS&lt;/td&gt;&lt;td&gt;.195&lt;/td&gt;&lt;td&gt;.310&lt;/td&gt;&lt;td&gt;.281&lt;/td&gt;&lt;td&gt;.031&lt;/td&gt;&lt;td&gt;.086&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MLC&lt;/td&gt;&lt;td&gt;&amp;#8722;0.085&lt;/td&gt;&lt;td&gt;.567&lt;/td&gt;&lt;td&gt;.279&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.085&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;RS&lt;/td&gt;&lt;td&gt;&amp;#8722;1.065&lt;/td&gt;&lt;td&gt;1.937&lt;/td&gt;&lt;td&gt;.158&lt;/td&gt;&lt;td&gt;.001&lt;/td&gt;&lt;td&gt;.026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;D-score&lt;/td&gt;&lt;td&gt;.003&lt;/td&gt;&lt;td&gt;.009&lt;/td&gt;&lt;td&gt;.210&lt;/td&gt;&lt;td&gt;.022&lt;/td&gt;&lt;td&gt;.046&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;NP&lt;/td&gt;&lt;td&gt;&amp;#8722;4.507&lt;/td&gt;&lt;td&gt;1.042&lt;/td&gt;&lt;td&gt;&amp;#8722;.426&lt;/td&gt;&lt;td&gt;.000&lt;/td&gt;&lt;td&gt;.222&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;NR&lt;/td&gt;&lt;td&gt;&amp;#8722;3.645&lt;/td&gt;&lt;td&gt;3.003&lt;/td&gt;&lt;td&gt;&amp;#8722;.259&lt;/td&gt;&lt;td&gt;.027&lt;/td&gt;&lt;td&gt;.072&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0186372652-14">Discussion</hd> <p>Overall, the present study has indicated that human ratings of a standardised test could be associated with certain CAF indices. Notably, speaking fluency emerged as a substantial predictor of subjective human ratings, with higher levels of fluency corresponding to increased scores. This finding aligns with previous studies that have similarly demonstrated a strong tendency for human ratings of L2 speaking to be predominantly influenced by fluency measures ([<reflink idref="bib14" id="ref74">14</reflink>]; [<reflink idref="bib23" id="ref75">23</reflink>]; [<reflink idref="bib36" id="ref76">36</reflink>]; [<reflink idref="bib57" id="ref77">57</reflink>]; [<reflink idref="bib61" id="ref78">61</reflink>]), suggesting that examiners might award higher scores to candidates who can speak an L2 fluently with fewer pauses or repairs. Specifically, human ratings were correlated with speaking fluency as quantified by the NP, which emerged as the most significant predictor variable among the others. This finding is consistent with [<reflink idref="bib71" id="ref79">71</reflink>] research, which noted that "pause clusters appeared to be strong influences on judges' assessments of level" (p. 239). Although there are abundant findings linking subjective scores with objective fluency indices, one should be cautious about drawing robust conclusions on speaking fluency. According to [<reflink idref="bib14" id="ref80">14</reflink>], if examiners are instructed to focus on certain aspects of speaking fluency, their ratings may become closely aligned with the objective measures of those specific fluency aspects, potentially overlooking other important features of fluency. Thus, the finding that the NP was the most significant predictor of human ratings is logical, as the marking rubric used in the present study specifically highlights the NP as a key measure of speaking fluency ([<reflink idref="bib81" id="ref81">81</reflink>]).</p> <p>In addition, speaking complexity, both grammatical and lexical, was correlated with and could predict human ratings, though its effect was less pronounced compared to speaking fluency. This finding is consistent with previous research, which has also shown that candidates' complex utterances and use of extensive vocabulary can positively influence examiners' judgements during the marking process ([<reflink idref="bib36" id="ref82">36</reflink>]; [<reflink idref="bib57" id="ref83">57</reflink>]; [<reflink idref="bib61" id="ref84">61</reflink>]). [<reflink idref="bib71" id="ref85">71</reflink>] study produced similar findings regarding grammatical complexity but suggested that lexical complexity was not correlated with examiners' assessments, particularly for learners of adjacent proficiency levels (see also [<reflink idref="bib33" id="ref86">33</reflink>]). We attempt to explain this by speculating that human raters might consciously or subconsciously emphasise and tally words during a speaking test. This assumption is plausible, as human raters are inevitably "influenced (often unconsciously) by their previous experience, expectations, knowledge, preferences, and subjective interpretation of assessment scales and their categories and descriptors" ([<reflink idref="bib70" id="ref87">70</reflink>], p. 211). Nevertheless, the seemingly unavoidable influence of subjectivity also raises the possibility that human raters could be "heavily influenced by the occurrence of just a few, judiciously placed, infrequent words and expressions" ([<reflink idref="bib55" id="ref88">55</reflink>], p. 93). In this regard, we justify the correlation between grammatical and lexical complexity with human ratings observed in both the present study and previous ones. We speculate that this correlation stands out not only due to the marking descriptors that emphasise the use of rich vocabulary and complex sentence structures but also because of examiners' expectations and preferences for the diversity and density of a candidate's linguistic features.</p> <p>Regarding accuracy, the present study, in line with some others (e.g. [<reflink idref="bib57" id="ref89">57</reflink>]), suggested that the ability to produce error-free utterances might not be as influential as complexity and fluency measures in obtaining higher scores from human raters. This finding contradicts previous research, which indicated that speaking accuracy tends to be a critical factor in judges' assessments, as errors can disrupt speech and make it less comprehensible or coherent ([<reflink idref="bib33" id="ref90">33</reflink>]; [<reflink idref="bib71" id="ref91">71</reflink>]). However, [<reflink idref="bib76" id="ref92">76</reflink>] highlighted an interesting scenario where, if human raters are overwhelmed with the identification and assessment of certain aspects of speaking performance, they might become more tolerant of other aspects that are also considered important. Based on this perspective, we speculate that the examiners in the present study might have paid less attention to speaking accuracy compared to fluency and complexity. Another possible explanation could lie in the contradictory nature of focussing on form versus focussing on meaning, which might lead examiners to concentrate on only one at a time ([<reflink idref="bib52" id="ref93">52</reflink>]). This view aligns with [<reflink idref="bib15" id="ref94">15</reflink>] finding that raters might not simultaneously attend to both fluency and accuracy. Therefore, although the marking rubric used in the present study "covers not only Accuracy and Range at the grammatical (and lexical) level but also Length and Coherence and Flexibility and Appropriateness at the textual and pragmatic level" ([<reflink idref="bib81" id="ref95">81</reflink>], p. 8), the communicative aspect of the rubric and the participants' speaking performance might have received more attention from the examiners, emphasising meaning-focused output and fluency rather than the proper use of linguistic forms.</p> <p>Admittedly, there is a lack of clear compatibility between the findings of the present study and existing research. One plausible explanation could be that different CAF measures have been used by researchers to quantify L2 speaking proficiency, with the long-standing debate on the efficacy of these heterogeneous CAF indices still unresolved. For instance, in [<reflink idref="bib57" id="ref96">57</reflink>] research, speaking fluency was related not only to pauses and repairs made by an L2 speaker but also to the mean length of run, phonation time ratio, and mean duration of syllable. However, measures concerning speech rate were excluded from the present study due to concerns that a speaker's L2 fluency is highly correlated with their L1 fluency ([<reflink idref="bib7" id="ref97">7</reflink>]; [<reflink idref="bib64" id="ref98">64</reflink>]). Although the study has reached the same conclusion as [<reflink idref="bib57" id="ref99">57</reflink>]—that is, analytical measures of fluency tend to be significantly predictive of human ratings—the adoption of the same analytical indices would undoubtedly invite a closer comparison. However, what is clear is that speaking requires the use of different skills, and some tend to be more influential than others in human ratings, potentially leading to a "halo" effect, where one characteristic of L2 production overshadows the judgement of the entire text ([<reflink idref="bib40" id="ref100">40</reflink>]). Various factors (e.g. different interpretations of marking rubrics, a rater's confusion or overemphasis on certain criteria, the high cognitive load of scoring a candidate's performance) may exert this effect, directing a rater's principal attention to complexity, accuracy, or fluency, and thereby influencing the overall rating process ([<reflink idref="bib6" id="ref101">6</reflink>]; [<reflink idref="bib30" id="ref102">30</reflink>]). While the reasons behind these findings remain unclear given the design of the present study, this scenario highlights the importance of incorporating analytical CAF measures when quantifying speaking proficiency, ensuring the validity and reliability of assessment outcomes.</p> <p>For reasons of authenticity and validity, many language testers still prefer to use performance assessments to evaluate L2 learners' productive abilities, which is understandable given the complexity and difficulty of decoding assessment data and conducting quantitative CAF analyses. Even so, such complications do not necessarily equate to the impracticability of CAF measures, which could be adopted through collaboration between L2 researchers and examiners. That means, by working together, researchers can provide the necessary expertise in CAF analysis, while examiners bring practical insights from the testing field. This collaboration could lead to the development of more refined assessment tools that incorporate both subjective human ratings and objective CAF measures, offering a balanced approach to evaluating L2 speaking proficiency.</p> <p>Practically, the gap between human-judged L2 proficiency and analytic measures highlights the importance of acknowledging and integrating rater perceptions into both marking scales and rater training. This integration should focus on deepening raters' understanding of the nuances of CAF measures and their relationship to overall language proficiency. It is crucial that rater training also includes strategies for maintaining objectivity in assessment, ensuring that raters are aware of potential biases and are equipped to mitigate them effectively. With a shared understanding among test designers and examiners of why a specific sample surpasses another and how certain aspects of subjectivity may influence scoring, the construction and interpretation of verbal descriptors in assessment criteria can be greatly improved. In this context, reliability could be considered "not a problem" in L2 speaking assessment ([<reflink idref="bib5" id="ref103">5</reflink>], p. 150).</p> <p>For further research, a broader agenda should be explored. [<reflink idref="bib35" id="ref104">35</reflink>] latest research has revealed human raters' sensitivity to different types of lexical complexity (e.g. diversity, density, sophistication) corresponding to learners' varying levels of linguistic competence. Hence, there is a possibility that human ratings based on relevant scales could better identify changes in L2 proficiency than objective measures. This interesting finding suggests that considering individual differences in comparing human ratings with CAF indices would enhance the understanding of the complex nature of L2 assessment. That being said, it is also necessary to account for learner heterogeneity, particularly differences in language proficiency, when quantifying speaking CAF. It seems to be a rule of thumb that research with well-justified choices of CAF measures is credible in its own right, whereas the output features of L2 learners at different language levels can vary significantly, with nuances that can only be captured by corresponding CAF indices—for example, complexity by coordination and subordination ([<reflink idref="bib26" id="ref105">26</reflink>]). Finally, in light of the idea of full utilisation of research data ([<reflink idref="bib12" id="ref106">12</reflink>]), an extension of the study could focus on the correlations among CAF, the statistics of which have already been tabulated above, to establish evidence for the debate between the Limited Attention Capacity Hypothesis ([<reflink idref="bib67" id="ref107">67</reflink>]) and the Cognition Hypothesis ([<reflink idref="bib62" id="ref108">62</reflink>]). With the former highlighting a trade-off effect on CAF and the latter suggesting that the attentional resources allocated to a specific dimension of L2 output do not clash with those assigned to others, the disagreement remains unresolved and requires more empirical evidence ([<reflink idref="bib46" id="ref109">46</reflink>]).</p> <hd id="AN0186372652-15">Conclusion</hd> <p>This study examined the relationships between analytical CAF indices and intuitive human ratings within China's EFL context. The analyses of data collected from a standardised test demonstrated that speaking fluency, particularly the NP and NR in speech, significantly predicted human ratings. Similarly, speaking complexity, focussing on grammatical and lexical dimensions, was also a meaningful predictor, although its effect was less pronounced than that of fluency. In contrast, human ratings were not associated with speaking accuracy, despite previous research establishing a link between the CAF triad and human-judged L2 proficiency levels.</p> <p>Admittedly, this study is not without limitations. Firstly, the reliance on a positivist paradigm, while valuable, may lack the depth of insight that a mixed-methods approach could offer, such as by exploring examiners' real-time perceptions of assessment. Additionally, concerns may arise regarding the reliability of the human ratings analysed in the study, despite the professionalism and expertise of the examiners involved. This underscores the need for more systematic research, potentially incorporating many-facet Rasch measurement, to explore the various components of rater-mediated assessment. Nevertheless, the findings presented here are significant as they highlight the potential gap between objective CAF measures and subjective human ratings, suggesting that addressing this gap could enhance the contextual and scoring validity of the specific test employed in the study.</p> <p>Appendix A. CET-SET6 Marking Rubric.</p> <p>Graph</p> <p> <ephtml> &lt;table&gt;&lt;colgroup&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;col align="left" /&gt;&lt;/colgroup&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Mark&lt;/th&gt;&lt;th align="left"&gt;Accuracy and range&lt;/th&gt;&lt;th align="left"&gt;Length and coherence&lt;/th&gt;&lt;th align="left"&gt;Flexibility and appropriateness&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;&amp;#8226; Vocabulary and grammar are basically correct.&lt;break /&gt;&amp;#8226; The vocabulary is rich, and the grammatical structure is relatively complex.&lt;break /&gt;&amp;#8226; Pronunciation is good, despite some native accent that does not affect comprehension.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate can speak at length with coherent language organisation.&lt;break /&gt;&amp;#8226; There are occasional pauses when organising answers and choosing vocabulary, but they do not influence communication.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate can respond to different communication situations and topics excellently.&lt;break /&gt;&amp;#8226; The candidate can engage in discussion actively.&lt;break /&gt;&amp;#8226; The overall use of language is appropriate for different situations, functions and purposes.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;&amp;#8226; There are some grammar and vocabulary mistakes, but they do not seriously impair communication.&lt;break /&gt;&amp;#8226; The vocabulary is relatively rich.&lt;break /&gt;&amp;#8226; Pronunciation is acceptable.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate can conduct coherent talks, but most of them are relatively short.&lt;break /&gt;&amp;#8226; There are occasional pauses when organising answers and choosing vocabulary, which sometimes influence communication.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate can respond to different communication situations and topics without much difficulty.&lt;break /&gt;&amp;#8226; The candidate can engage in discussion.&lt;break /&gt;&amp;#8226; The overall use of language is fairly appropriate for different situations, functions and purposes.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;&amp;#8226; There are grammar and vocabulary mistakes, and they sometimes influence communication.&lt;break /&gt;&amp;#8226; The vocabulary is not rich, and the grammatical structure is relatively simple.&lt;break /&gt;&amp;#8226; The pronunciation is flawed, which sometimes interferes with communication.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate's talks are short.&lt;break /&gt;&amp;#8226; There are frequent long pauses when organising answers and choosing vocabulary, which influence communication. But the candidate is able to complete the speaking tasks.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate cannot engage in discussion actively.&lt;break /&gt;&amp;#8226; Sometimes, the candidate cannot respond to the change of communication topics or contents.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;&amp;#8226; There are many grammar and vocabulary mistakes, which interrupt communication.&lt;break /&gt;&amp;#8226; Communication is seriously affected due to a lack of vocabulary and grammar knowledge.&lt;break /&gt;&amp;#8226; The pronunciation is poor.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate's talks are short and incohesive, which demonstrate limited ability to communicate.&lt;/td&gt;&lt;td&gt;&amp;#8226; The candidate cannot engage in discussion.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;1&lt;/td&gt;&lt;td colspan="3"&gt;&amp;#8226; No description.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ref id="AN0186372652-16"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref46" type="bt">1</bibl> <bibtext> Hengzhi Hu</bibtext> </blist> <blist> <bibtext>Graph</bibtext> </blist> <blist> <bibl id="bib2" idref="ref52" type="bt"></bibl> <bibtext>https://orcid.org/0000-0001-5232-913X Nur Ehsan Mohd Said</bibtext> </blist> <blist> <bibl id="bib3" idref="ref2" type="bt"></bibl> <bibtext>Graph</bibtext> </blist> <blist> <bibl id="bib4" idref="ref53" type="bt"></bibl> <bibtext>https://orcid.org/0000-0002-2891-327X Harwati Hashim</bibtext> </blist> <blist> <bibl id="bib5" idref="ref103" type="bt"></bibl> <bibtext>Graph https://orcid.org/0000-0002-8817-427X</bibtext> </blist> <blist> <bibtext> Informed consent was obtained from all participants in the study.</bibtext> </blist> <blist> <bibtext> The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Center for Research and Instrumentation Management (CRIM) and Faculty of Education, Universiti Kebangsaan Malaysia (UKM).</bibtext> </blist> <blist> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibtext> The datasets used during the current study are available from the corresponding author upon reasonable request.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref101" type="bt">6</bibl> <bibtext> Data Availability Statement included at the end of the article</bibtext> </blist> </ref> <ref id="AN0186372652-17"> <title> References </title> <blist> <bibtext> Aas H. L., Nacey S. (2019). Methodological concerns for investigating pause behavior in spoken corpora. In Degand L., Gilquin G., Meurant L., Simon A. C. (Eds.), Fluency and disfluency across languages and language varieties (pp. 41–66). Presses universitaires de Louvain.</bibtext> </blist> <blist> <bibtext> Alaamer R. A. (2021). A theoretical review on the need to use standardized oral assessment rubrics for ESL learners in Saudi Arabia. English Language Teaching, 14(11), 144–150. https://doi.org/10.5539/elt.v14n11p144</bibtext> </blist> <blist> <bibtext> Alshalan R. F. (2024). Assessing EFL speaking based on reconstructed CAF measures. International Journal of Research in Education Humanities and Commerce, 5(1), 9–24. https://doi.org/10.37602/IJREHC.2024.5102</bibtext> </blist> <blist> <bibtext> Aydoğan H., Akbarov A. (2018). Subjective vs. objective measures of English proficieny in a sample of Turkish students. Acta Didactica Napocensia, 11(2), 135–142. https://doi.org/10.24193/adn.11.2.11</bibtext> </blist> <blist> <bibtext> Bachman L. F. (1988). Problems in examining the validity of the ACTFL oral proficiency interview. Studies in Second Language Acquisition, 10(2), 149–164. https://doi.org/10.1017/S0272263100007282</bibtext> </blist> <blist> <bibtext> Bijani H., Hashempour B., Orabah S. S. B. (2022). Development and validation of a training-embedded speaking assessment rating scale: A multifaceted Rasch analysis in speaking assessment. International Journal of Research in English Education, 7(3), 32–45. https://doi.org/10.52547/ijree.7.3.32</bibtext> </blist> <blist> <bibl id="bib7" idref="ref44" type="bt">7</bibl> <bibtext> Bradlow A. R., Kim M., Blasingame M. (2017). Language-independent talker-specificity in first-language and second-language speech production by bilingual talkers: L1 speaking rate predicts L2 speaking rate. The Journal of the Acoustical Society of America, 141(2), 886–899. https://doi.org/10.1121/1.4976044</bibtext> </blist> <blist> <bibl id="bib8" idref="ref14" type="bt">8</bibl> <bibtext> Brumfit C. (1984). Communicative methodology in language teaching. Cambridge UniversityPress.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref23" type="bt">9</bibl> <bibtext> Bulté B., Housen A. (2012). Defining and operationalising L2 complexity. In Housen A., Kuiken F., Vedder I. (Eds.), Dimensions of L2 performance and proficiency: Complexity, accuracy and fluency in SLA (pp. 23–46). John Benjamins. https://doi.org/10.1075/lllt.32.02bul</bibtext> </blist> <blist> <bibtext> Byrnes H. (2020). Advanced-level grammatical development in instructed SLA. In Malovrh P. A., Benati A. G. (Eds.), The handbook of advanced proficiency in second language acquisition (pp. 133–156). John Wiley &amp; Sons.</bibtext> </blist> <blist> <bibtext> Colantoni L., Steele J., Escudero P. (2015). Second language speech: Theory and practice. Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Chen X., Quan Z., Ma Z., Liu Z. (2022). 我国科学数据开放共享的实现路径研究 [Research on the realization of path of scientific data open sharing in China]. 社会科学前沿 [Advances in Social Sciences], 11(9), 3791–3802. https://doi.org/10.12677/ASS.2022.119519</bibtext> </blist> <blist> <bibtext> Committee of CET Band 4 and Band 6. (2016). 全国大学英语四、六级考试大纲 [Assessment syllabus of CET-4 and CET-6]. Shanghai Jiaotong University Press.</bibtext> </blist> <blist> <bibtext> de Jong N. H. (2016). Fluency in second language assessment. In Tsagari D., Banerjee J. (Eds.), Handbook of second language assessment (pp. 203–218). De Gruyter Mouton. https://doi.org/10.1515/9781614513827-015</bibtext> </blist> <blist> <bibtext> Duijm K., Schoonen R., Hulstijn J. H. (2018). Professional and non-professional raters' responsiveness to fluency and accuracy in L2 speech: An experimental approach. Language Testing, 35(4), 501–527. https://doi.org/10.1177/0265532217712553</bibtext> </blist> <blist> <bibtext> Ellis R. (2003). Task-based language learning and teaching. Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Ellis R. (2009). The differential effects of three types of task planning on the fluency, complexity, and accuracy in L2 oral production. Applied Linguistics, 30(4), 474–509. https://doi.org/10.17507/jltr.0606.17</bibtext> </blist> <blist> <bibtext> Ellis R., Barkhuizen G. (2005). Analysing learner language. Oxford University Press.</bibtext> </blist> <blist> <bibtext> Fernández D. M., González A. B., Padilla J. R. (2021). Investigation the effects of CLIL on language attainment: Instrument design and validation. In Cañado M. L. P. (Ed.), Content and Language Integrated Learning in monolingual settings (pp. 71–102). Springer. https://doi.org/10.1007/978-3-030-68329-0_5</bibtext> </blist> <blist> <bibtext> Fidler M., Cvrček V. (2018). Morphological richness of text. In Fidler M., Cvrček V. (Eds.), Taming the corpus: From inflection and lexis to interpretation (pp. 63–78). Springer. https://doi.org/10.1007/978-3-319-98017-1_4</bibtext> </blist> <blist> <bibtext> Foster P., Tonkyn A., Wigglesworth G. (2000). Measuring spoken language: A unit for all reasons. Applied Linguistics, 21(3), 354–375. https://doi.org/10.1093/applin/21.3.354</bibtext> </blist> <blist> <bibtext> Ghufron M. A. (2017). Language learning strategies used by EFL fluent speakers: A case in Indonesian context. Indonesian Journal of English Teaching, 6(2), 184–202. https://doi.org/10.15642/ijet.2017.6.2.184-202</bibtext> </blist> <blist> <bibtext> Ginther A., Dimova S., Yang R. (2010). Conceptual and empirical relationships between temporal measures of fluency and oral English proficiency with implications for automated scoring. Language Testing, 27(3), 379–399. https://doi.org/10.1177/0265532210364407</bibtext> </blist> <blist> <bibtext> González-Robaina Y., Larenas C. D. (2020). Proposing a theoretical and state-of-the-art didactic model to balance oral communication fluency and accuracy in English as a foreign language. Revista Educación [Journal of Education], 44(1), 549–565. https://doi.org/10.15517/revedu.v44i1.36873</bibtext> </blist> <blist> <bibtext> Green A. (2021). Exploring language assessment and testing: Language in action (2nd ed.). Routledge.</bibtext> </blist> <blist> <bibtext> Hasnain S., Halder S. (2022). Intricacies of the multifaceted triad-complexity, accuracy, and fluency: A review of studies on measures of oral production. Journal of Education. Advance online publication. https://doi.org/10.1177/00220574221101377</bibtext> </blist> <blist> <bibtext> Hieke A. E., Kowal S., O'Connell D. C. (1983). The trouble with "articulatory" pauses. Language and Speech, 26(3), 203–214. https://doi.org/10.1177/002383098302600302</bibtext> </blist> <blist> <bibtext> Housen A., Kuiken F. (2009). Complexity, accuracy and fluency in second language acquisition. Applied Linguistics, 30(4), 461–473. https://doi.org/10.1093/applin/amp048</bibtext> </blist> <blist> <bibtext> Housen A., Kuiken F., Vedder I. (2012). Complexity, accuracy and fluency: Definitions, measurement and research. In Housen A., Kuiken F., Vedder I. (Eds.), Dimensions of L2 performance and proficiency: Complexity, accuracy and fluency (pp. 1–20). John Benjamins Publishing Company. https://doi.org/10.1075/lllt.32.01hou</bibtext> </blist> <blist> <bibtext> Hsieh M. C. (2020). 從多層面 [Rasch] 模式來檢視不同 的評分者等化連結設計對參數估 計的影響 [Investigating the effects of rater equating designs on parameter estimates in the context of preservice principal oral performance]. 教育心理學報 [Bulletin of Educational Psychology], 55(2), 415–436. <ulink href="http://doi.org/10.6251/BEP.202012%5f52(2).0008">http://doi.org/10.6251/BEP.202012%5f52(2).0008</ulink></bibtext> </blist> <blist> <bibtext> Hu H. (2022). Computer-delivered English listening and speaking test in Zhongkao: Test-taker perception, motivation and performance. In Uslu F., Güçlü T., Özdemir M., Altan K., Aslan S. (Eds.), Proceedings of SOCIOINT 2022—9th International conference on education &amp; education of social sciences (pp. 59–75). Ocerint International Organization Center of Academic Research. https://doi.org/10.46529/socioint.202209</bibtext> </blist> <blist> <bibtext> Hussein N. M. H. H., Mohamed F. S., Zaza M. S. (2020). Developing EFL fluency skills among faculty of education students using the multimodal approach. Journal of Faculty of Education, 31(122), 31–56. https://doi.org/10.21608/jfeb.2020.142945</bibtext> </blist> <blist> <bibtext> Inoue C. (2016). A comparative study of the variables used to measure syntactic complexity and accuracy in task-based research. The Language Learning Journal, 44(4), 487–505. https://doi.org/10.1080/09571736.2015.1130079</bibtext> </blist> <blist> <bibtext> Iwashita N., Brown A., McNamara T., O'Hagan S. (2008). Assessed levels of second language speaking proficiency: How distinct? Applied Linguistics, 29(1), 24–49. https://doi.org/10.1093/applin/amm017</bibtext> </blist> <blist> <bibtext> Jeon E. H., In'nami Y., Koizumi R. (2022). Discussion, limitations, and future research. In Jeon E. H., In'nami Y. (Eds.), Understanding L2 proficiency: Theoretical and meta-analytic investigations (pp. 369–385). John Benjamins Publishing Company. https://doi.org/10.1075/bpa.13.12jeo</bibtext> </blist> <blist> <bibtext> Jin T., Mak B. (2013). Distinguishing features in scoring L2 Chinese speaking performance: How do they work? Language Testing, 30(1), 23–47. https://doi.org/10.1177/0265532212442637</bibtext> </blist> <blist> <bibtext> Jin Y., Wang W., Zhang X., Zhao Y. (2020). 大学英语四级口语考试自动评分效度初探 [A preliminary investigation of the scoring validity of the CET-SET automated scoring system]. 中国考试 [China Examinations], 2020(7), 25–33. https://doi.org/10.19360/j.cnki.11-3303/g4.2020.07.004</bibtext> </blist> <blist> <bibtext> Julious S. A. (2005). Sample size of 12 per group rule of thumb for a pilot study. Pharmaceutical Statistics, 4(4), 287–291. https://doi.org/10.1002/pst.185</bibtext> </blist> <blist> <bibtext> Kang O., Yan X. (2018). Linguistic features distinguishing examinees' speaking performances at different proficiency levels. Journal of Language Testing &amp; Assessment, 1, 24–39. https://doi.org/10.23977/langta.2018.11003</bibtext> </blist> <blist> <bibtext> Knoch U. (2009). Diagnostic writing assessment: The development and validation of a rating scale. Peter Lang.</bibtext> </blist> <blist> <bibtext> Koizumi R. (2005). Speaking performance measures of fluency, accuracy, syntactic complexity, and lexical complexity. JABAET (Japan-Britain Association for English Teaching) Journal, 9, 5–33.</bibtext> </blist> <blist> <bibtext> Koizumi R., In'nam Y., Fukazawa M. (2020). Comparison between holistic and analytic rubrics of a paired oral test. JALT Journal, 23, 57–77. https://doi.org/10.20622/jltajournal.23.0_57</bibtext> </blist> <blist> <bibtext> Leńko-Szymańska A., Götz S. (Eds.). (2022). Complexity, accuracy and fluency in learner corpus research. John Benjamins.</bibtext> </blist> <blist> <bibtext> Lennon P. (1990). Investigating fluency in EFL: A quantitative approach. Language Learning, 40(3), 387–417. https://doi.org/10.1111/j.1467-1770.1990.tb00669.x</bibtext> </blist> <blist> <bibtext> Li H., Zhang S. (2021). Exploratory and confirmatory factor analyses of L2 linguistic complexity measures. International Journal of English Linguistics, 11(1), 192–205. https://doi.org/10.5539/ijel.v11n1p192</bibtext> </blist> <blist> <bibtext> Li S. (2022). A dynamic systems theory perspective on L2 writing development. Routledge.</bibtext> </blist> <blist> <bibtext> Liu X., Zhong Q., Li Q., Chen Y., Jiang H., Peng J., Tan Y. (2020). 大学英语口语考试训练教程:(六级版) 真题、模拟与解析 [CET-SET training course: Past exam papers, mock tests and analysis, Band 6 edition]. Tsinghua University Press.</bibtext> </blist> <blist> <bibtext> Lu Z., Li Z., Hou L. (2016). On the validity and reliability of a computer-assisted English speaking test. In Proceedings of the 2016 international conference on intelligent control and computer application (pp. 187–193). Atlantis Press. https://doi.org/10.2991/icca-16.2016.43</bibtext> </blist> <blist> <bibtext> Luoma S. (2010). Assessing speaking. Cambridge University Press. https://doi.org/10.1017/CBO9780511733017</bibtext> </blist> <blist> <bibtext> Mackey A., Gass S. (2005). Second language research: Methodology and design. Routledge. https://doi.org/10.4324/9781003188414</bibtext> </blist> <blist> <bibtext> Markee N. (2000). Conversation analysis. Talyor &amp; Francis.</bibtext> </blist> <blist> <bibtext> McNamara T. (1996). Measuring second language performance. Addison Wesley Longman. https://doi.org/10.2307/330236</bibtext> </blist> <blist> <bibtext> Mhundwa P. H. (2004). Error treatment in students' written assignments in discourse analysis. Journal for Language Teaching, 37(2), 224–236. https://doi.org/10.4314/jlt.v37i2.6017</bibtext> </blist> <blist> <bibtext> Michel M. (2017). Complexity, accuracy and fluency (CAF). In Loewen S., Sato M. (Eds.), The Routledge handbook of instructed second language acquisition (pp. 2–38). Routledge.</bibtext> </blist> <blist> <bibtext> Milton J., Wade J., Hopkins N. (2010). Aural word recognition and oral competence in English as a foreign language. In Chacón-Beltrán R., Abello-Contesse C., Torreblanca-López M. M. (Eds.), Insights into non-native vocabulary teaching and learning (pp. 83–98). Multilingual Matters. https://doi.org/10.21832/9781847692900-007</bibtext> </blist> <blist> <bibtext> Norris J. M., Ortega L. (2009). Towards an organic approach to investigating CAF in instructed SLA: The case of complexity. Applied Linguistics, 30(4), 555–578. https://doi.org/10.1093/applin/amp044</bibtext> </blist> <blist> <bibtext> Ogawa C. (2022). CAF indices and human ratings of oral performances in an opinion-based monologue task. Language Testing in Asia, 12, Article 4. https://doi.org/10.1186/s40468-022-00154-9</bibtext> </blist> <blist> <bibtext> Özdil Ş., Duran E. (2023). Development of persuasive speaking skills rubrics for primary school fourth grade students. International Journal of Education &amp; Literacy Studies, 11(1), 59–67. https://doi.org/10.7575/aiac.ijels.v.11n.1p.59</bibtext> </blist> <blist> <bibtext> Pallant J. (2016). SPSS survival manual (6th ed.). Open University Press.</bibtext> </blist> <blist> <bibtext> Pratama V. M., Zainil Y. (2020). EFL learners' communication strategy on speaking performance of interpersonal conversation in classroom discussion presentation. In Rosa R. N., Al Hafizh M., Ardi H., Arianto M. A. (Eds.),Proceedings of the 7th international conference on english language and teaching (pp. 29–36). Atlantis Press. https://doi.org/10.2991/assehr.k.200306.006</bibtext> </blist> <blist> <bibtext> Révész A., Ekiert M., Torgersen E. N. (2016). The effects of complexity, accuracy, and fluency on communicative adequacy in oral task performance. Applied Linguistics, 37(6), 828–848. https://doi.org/10.1093/applin/amu069</bibtext> </blist> <blist> <bibtext> Robinson P. (2001). Task complexity, cognitive resources, and syllabus design: A triadic framework for examining task influences on SLA. In Robinson P. (Ed.), Cognition and second language instruction (pp. 287–318). Cambridge University press.</bibtext> </blist> <blist> <bibtext> Samifanni F. (2020). The fluency way: A functional method for oral communication. English Language Teaching, 13(3), 100–114. https://doi.org/10.5539/elt.v13n3p100</bibtext> </blist> <blist> <bibtext> Shrosbree M. (2020). The relationship between L1 fluency and L2 fluency among Japanese advanced early learners of English. 外国語教育研究ジャーナル [Journal of Foreign Language Education and Research], 1, 111–117.</bibtext> </blist> <blist> <bibtext> Siskova Z. (2012). Lexical richness in EFL students' narratives. Language Studies Working Papers, 4(2016), 26–36.</bibtext> </blist> <blist> <bibtext> Skehan P. (1989). Individual differences in second language learning. Edward Arnold.</bibtext> </blist> <blist> <bibtext> Skehan P. (1998). A cognitive approach to language learning. Oxford University Press.</bibtext> </blist> <blist> <bibtext> Skehan P. (2009). Modelling second language performance: Integrating complexity, accuracy, fluency and lexis. Applied Linguistics, 30(4), 510–532. https://doi.org/10.1093/applin/amp047</bibtext> </blist> <blist> <bibtext> Sundqvist P., Sandlund E. (2024). Testing talk: Ways to assess second language oral proficiency. Bloomsbury Publishing.</bibtext> </blist> <blist> <bibtext> Taylor L., Galaczi E. D. (2011). Scoring validity. In Taylor L. (Ed.), Examining speaking: Research and practice in assessing second language speaking (pp. 171–233). Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Tonkyn A. P. (2012). Measuring and perceiving changes in oral complexity, accuracy and fluency Examining instructed learners' short-term gains. In Housen A., Kuiken F., Vedder I. (Eds.), Dimensions of L2 performance and proficiency: Complexity, accuracy and fluency in SLA (pp. 221–244). John Benjamins Publishing Company. https://doi.org/10.1075/lllt.32.10ton</bibtext> </blist> <blist> <bibtext> Tran M. N., Saito K. (2021). Effects of the 4/3/2 activity revisited: Extending Boers (2014) and Thai &amp; Boers (2016). Language Teaching Research, 28(2), 326–345. https://doi.org/10.1177/1362168821994136</bibtext> </blist> <blist> <bibtext> Wang Z., Han F. (2021). Developing English language learners' oral production with a digital game-based mobile application. PLoS One, 16(1), Article e0232671. https://doi.org/10.1371/journal.pone.0232671</bibtext> </blist> <blist> <bibtext> Wanzek J., Otaiba S. A., McMaster K. L. (2019). Intensive reading interventions for the elementary grades. Guilford Press.</bibtext> </blist> <blist> <bibtext> Xiao J., Wang J. (2017). 交际语言测试理论视野下的英语专业口语测试 [Oral English testing of English-majors: From the perspective of communicative language testing theory]. In Du X., Huang C., Zhong Y. (Eds.), Proceedings of the 2017 3rd international conference on humanities and social science research (pp. 160–163). Atlantis Press. https://doi.org/10.2991/ichssr-17.2017.32</bibtext> </blist> <blist> <bibtext> Yan X., Ginther A. (2017). Listeners and raters: Similarities and differences in evaluation of accented speech. In Kang O., Ginther A. (Eds.), Assessment in L2 pronunciation (pp. 67–88). Routledge. https://doi.org/10.4324/9781315170756</bibtext> </blist> <blist> <bibtext> Yan X., Kim H. R., Kim J. Y. (2018). Complexity, accuracy and fluency (CAF) features of speaking performances on Aptis across different levels on The Common European Framework of Reference for Languages (CEFR). British Council.</bibtext> </blist> <blist> <bibtext> Yan X., Kim H. R., Kim J. Y. (2020). Dimensionality of speech fluency: Examining the relationships among complexity, accuracy, and fluency (CAF) features of speaking performances on the Aptis test. Language Testing, 38(4), 1–26. https://doi.org/10.1177/0265532220951508</bibtext> </blist> <blist> <bibtext> Zeng Y. (2021). 复杂度,准确度和流利度对于高考英语和托福写作分数占比影响分析 [How complexity, accuracy, and fluency shape scores of English writing in National College Entrance Examination and TOEFL]. 国外英语考试教学与研究 [Overseas English Testing: Pedagogy and Research], 3(4), 158–168. https://doi.org/10.12677/OETPR.2021.34018</bibtext> </blist> <blist> <bibtext> Zhang L. (2022). College English Test–Spoken English Test (CET-SET). Studies in Language Assessment, 11(2), 164–180.</bibtext> </blist> <blist> <bibtext> Zhang Q. (2021). Impacts of world Englishes on local standardized language proficiency testing in the expanding circle. English Today, 38(4), 254–270. https://doi.org/10.1017/S0266078421000158</bibtext> </blist> </ref> <aug> <p>By Hengzhi Hu; Nur Ehsan Mohd Said and Harwati Hashim</p> <p>Reported by Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib69" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib57" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib45" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib72" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib79" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib19" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib25" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib31" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib37" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib81" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib80" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib66" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib28" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib10" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib29" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib43" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib54" firstref="ref20"></nolink> <nolink nlid="nl18" bibid="bib16" firstref="ref22"></nolink> <nolink nlid="nl19" bibid="bib20" firstref="ref24"></nolink> <nolink nlid="nl20" bibid="bib65" firstref="ref25"></nolink> <nolink nlid="nl21" bibid="bib51" firstref="ref26"></nolink> <nolink nlid="nl22" bibid="bib53" firstref="ref27"></nolink> <nolink nlid="nl23" bibid="bib21" firstref="ref28"></nolink> <nolink nlid="nl24" bibid="bib56" firstref="ref29"></nolink> <nolink nlid="nl25" bibid="bib17" firstref="ref30"></nolink> <nolink nlid="nl26" bibid="bib60" firstref="ref31"></nolink> <nolink nlid="nl27" bibid="bib24" firstref="ref33"></nolink> <nolink nlid="nl28" bibid="bib34" firstref="ref34"></nolink> <nolink nlid="nl29" bibid="bib18" firstref="ref35"></nolink> <nolink nlid="nl30" bibid="bib22" firstref="ref37"></nolink> <nolink nlid="nl31" bibid="bib32" firstref="ref38"></nolink> <nolink nlid="nl32" bibid="bib63" firstref="ref39"></nolink> <nolink nlid="nl33" bibid="bib44" firstref="ref40"></nolink> <nolink nlid="nl34" bibid="bib68" firstref="ref43"></nolink> <nolink nlid="nl35" bibid="bib64" firstref="ref45"></nolink> <nolink nlid="nl36" bibid="bib27" firstref="ref47"></nolink> <nolink nlid="nl37" bibid="bib74" firstref="ref48"></nolink> <nolink nlid="nl38" bibid="bib42" firstref="ref50"></nolink> <nolink nlid="nl39" bibid="bib58" firstref="ref51"></nolink> <nolink nlid="nl40" bibid="bib73" firstref="ref55"></nolink> <nolink nlid="nl41" bibid="bib77" firstref="ref56"></nolink> <nolink nlid="nl42" bibid="bib78" firstref="ref57"></nolink> <nolink nlid="nl43" bibid="bib39" firstref="ref58"></nolink> <nolink nlid="nl44" bibid="bib33" firstref="ref59"></nolink> <nolink nlid="nl45" bibid="bib48" firstref="ref62"></nolink> <nolink nlid="nl46" bibid="bib47" firstref="ref64"></nolink> <nolink nlid="nl47" bibid="bib13" firstref="ref65"></nolink> <nolink nlid="nl48" bibid="bib11" firstref="ref66"></nolink> <nolink nlid="nl49" bibid="bib75" firstref="ref67"></nolink> <nolink nlid="nl50" bibid="bib38" firstref="ref68"></nolink> <nolink nlid="nl51" bibid="bib49" firstref="ref69"></nolink> <nolink nlid="nl52" bibid="bib50" firstref="ref70"></nolink> <nolink nlid="nl53" bibid="bib41" firstref="ref71"></nolink> <nolink nlid="nl54" bibid="bib59" firstref="ref72"></nolink> <nolink nlid="nl55" bibid="bib14" firstref="ref74"></nolink> <nolink nlid="nl56" bibid="bib23" firstref="ref75"></nolink> <nolink nlid="nl57" bibid="bib36" firstref="ref76"></nolink> <nolink nlid="nl58" bibid="bib61" firstref="ref78"></nolink> <nolink nlid="nl59" bibid="bib71" firstref="ref79"></nolink> <nolink nlid="nl60" bibid="bib70" firstref="ref87"></nolink> <nolink nlid="nl61" bibid="bib55" firstref="ref88"></nolink> <nolink nlid="nl62" bibid="bib76" firstref="ref92"></nolink> <nolink nlid="nl63" bibid="bib52" firstref="ref93"></nolink> <nolink nlid="nl64" bibid="bib15" firstref="ref94"></nolink> <nolink nlid="nl65" bibid="bib40" firstref="ref100"></nolink> <nolink nlid="nl66" bibid="bib30" firstref="ref102"></nolink> <nolink nlid="nl67" bibid="bib35" firstref="ref104"></nolink> <nolink nlid="nl68" bibid="bib26" firstref="ref105"></nolink> <nolink nlid="nl69" bibid="bib12" firstref="ref106"></nolink> <nolink nlid="nl70" bibid="bib67" firstref="ref107"></nolink> <nolink nlid="nl71" bibid="bib62" firstref="ref108"></nolink> <nolink nlid="nl72" bibid="bib46" firstref="ref109"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1477239 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Human Ratings and Complexity, Accuracy, and Fluency (CAF) Indices: A Correlational Study of a Standardised Monologic English-Speaking Test in China – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hengzhi+Hu%22">Hengzhi Hu</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-5232-913X">0000-0001-5232-913X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Nur+Ehsan+Mohd+Said%22">Nur Ehsan Mohd Said</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-2891-327X">0000-0002-2891-327X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Harwati+Hashim%22">Harwati Hashim</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8817-427X">0000-0002-8817-427X</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22SAGE+Open%22"><i>SAGE Open</i></searchLink>. 2025 15(2). – Name: Avail Label: Availability Group: Avail Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 13 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+Tests%22">Speech Tests</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Fluency%22">Language Fluency</searchLink><br /><searchLink fieldCode="DE" term="%22Standardized+Tests%22">Standardized Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Objective+Tests%22">Objective Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Examiners%22">Examiners</searchLink><br /><searchLink fieldCode="DE" term="%22College+Students%22">College Students</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22China%22">China</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1177/21582440251343944 – Name: ISSN Label: ISSN Group: ISSN Data: 2158-2440 – Name: Abstract Label: Abstract Group: Ab Data: Foreign language (L2) learners' speaking proficiency is often quantified using two dimensions: intuitive human ratings and analytical, linguistic complexity, accuracy, and fluency (CAF) indices. While previous research and assessment practices have predominantly focused on either the subjective approach to L2 speaking or the objective one, it is essential to establish an association between these two seemingly contradictory assessment methods to enhance and promote more credible assessment judgements. To this end, 160 recordings from a monologic task of a standardised English test in China were analysed to quantify CAF in the present study, and the scores were then compared with human ratings given by qualified examiners. Correlation and regression analyses demonstrated that human ratings were positively correlated with speaking fluency, with the number of pauses produced by a candidate being the most significant predictor of human-judged scores. Speaking complexity also positively predicted human ratings, with examiners tending to focus more on grammatical complexity than lexical complexity. In contrast, no correlations were found between human ratings and speaking accuracy. The findings of this study reinforce the possibility of "halo" effects on human raters in L2 assessment and suggest that rater training should focus on helping examiners recognise and mitigate such potential effects. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1477239 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1477239 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1177/21582440251343944 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 13 Subjects: – SubjectFull: Foreign Countries Type: general – SubjectFull: Speech Tests Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Language Fluency Type: general – SubjectFull: Standardized Tests Type: general – SubjectFull: Difficulty Level Type: general – SubjectFull: Objective Tests Type: general – SubjectFull: Examiners Type: general – SubjectFull: College Students Type: general – SubjectFull: China Type: general Titles: – TitleFull: Human Ratings and Complexity, Accuracy, and Fluency (CAF) Indices: A Correlational Study of a Standardised Monologic English-Speaking Test in China Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hengzhi Hu – PersonEntity: Name: NameFull: Nur Ehsan Mohd Said – PersonEntity: Name: NameFull: Harwati Hashim IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 04 Type: published Y: 2025 Identifiers: – Type: issn-electronic Value: 2158-2440 Numbering: – Type: volume Value: 15 – Type: issue Value: 2 Titles: – TitleFull: SAGE Open Type: main |
| ResultId | 1 |