Understanding Intermediate-Level Speakers' Strengths and Weaknesses: An Examination of OPIc Tests from Korean Learners of English

Saved in:
Bibliographic Details
Title: Understanding Intermediate-Level Speakers' Strengths and Weaknesses: An Examination of OPIc Tests from Korean Learners of English
Language: English
Authors: Cox, Troy L.
Source: Foreign Language Annals. Spr 2017 50(1):84-113.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 30
Publication Date: 2017
Document Type: Journal Articles
Reports - Research
Descriptors: Foreign Countries, English (Second Language), Second Language Learning, Language Proficiency, Student Characteristics, Interviews, Linguistics, Language Skills, Performance, Student Needs
Geographic Terms: South Korea
DOI: 10.1111/flan.12258
ISSN: 0015-718X
Abstract: This study profiled Intermediate-level learners in terms of their linguistic characteristics and performance on different proficiency tasks. A stratified random sample of 300 Korean learners of English with holistic ratings of Intermediate Low (IL), Intermediate Mid (IM), and Intermediate High (IH) on Oral Proficiency Interviews-computerized (OPIcs)--100 at each level--were analyzed by trained ACTFL raters to determine what was needed for the learners to progress to the next higher sublevel. The findings indicate that while ILs minimally met all the linguistic characteristics required of the Intermediate level, they needed to improve in the quantity and quality of all the linguistic characteristics they employed and improve their mastery of the types and variety of questions they could use when performing Intermediate tasks to move to the IM sublevel. In contrast, IMs demonstrated a pattern of strength when completing Intermediate tasks, but to move to the IH sublevel they needed to improve their ability to perform all Advanced-level tasks, especially in terms of accuracy when using paragraph-length discourse. Similar to the IMs, for the IHs to move to the Advanced Low sublevel, they needed to improve their accuracy with paragraph-length discourse and expand their content mastery to beyond the autobiographical.
Abstractor: As Provided
Entry Date: 2017
Accession Number: EJ1135287
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwG0s4FYUPYJNuEJDk3GFJTwAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDFb1hTzqDMu67rKDWAIBEICBm7IYVIbw2VGm0ncS3Pa_FGVV6ijd-NZBQ1HMagGnHf2YgplGHRts5jPwyOUkZ2cs3JdcyQYj8klNLFq5Xsf1JaxEbONL8fx-OPsGlHva1ukIPFdtvrVUcv_OC5en_ThWTuQY74xgE044-JUUBHb82G6WBg60pJOqxMqMGC69_VqQjk7tqQzA5GS1DSA-kCqm17xqbETTIk4iWpXn
Text:
  Availability: 1
  Value: <anid>AN0122100053;fla01mar.17;2018Jul02.12:17;v2.2.500</anid> <title id="AN0122100053-1">Understanding Intermediate-Level Speakers' Strengths and Weaknesses: An Examination of OPIc Tests From Korean Learners of English. </title> <p>This study profiled Intermediate‐level learners in terms of their linguistic characteristics and performance on different proficiency tasks. A stratified random sample of 300 Korean learners of English with holistic ratings of Intermediate Low (IL), Intermediate Mid (IM), and Intermediate High (IH) on Oral Proficiency Interviews‐computerized (OPIcs)—100 at each level—were analyzed by trained ACTFL raters to determine what was needed for the learners to progress to the next higher sublevel. The findings indicate that while ILs minimally met all the linguistic characteristics required of the Intermediate level, they needed to improve in the quantity and quality of all the linguistic characteristics they employed and improve their mastery of the types and variety of questions they could use when performing Intermediate tasks to move to the IM sublevel. In contrast, IMs demonstrated a pattern of strength when completing Intermediate tasks, but to move to the IH sublevel they needed to improve their ability to perform all Advanced‐level tasks, especially in terms of accuracy when using paragraph‐length discourse. Similar to the IMs, for the IHs to move to the Advanced Low sublevel, they needed to improve their accuracy with paragraph‐length discourse and expand their content mastery to beyond the autobiographical.</p> <p>English as a foreign/second language; oral proficiency</p> <p>In a recent audit of oral proficiency test results from a large university that the author conducted, it was discovered that a single student had taken either the Oral Proficiency Interview (OPI) or the Oral Proficiency Interview‐computerized (OPIc) nine times over a 3‐year period. Further examination revealed that this student, a language teaching minor who needed a rating of Advanced Low for instructor licensure, was languishing at the Intermediate level. After an initial OPI score of Intermediate High (IH) in 2013, the next four tests resulted in ratings of Intermediate Mid (IM), while the final four tests were rated IH. Reaching the Advanced level is critical for those pursuing teaching licensure (Brooks & Darhower, [<reflink idref="bib6" id="ref1">6</reflink>] ; Chambless, [<reflink idref="bib8" id="ref2">8</reflink>] ), and this student's lack of progression toward higher proficiency on the ACTFL scale represented a real‐world example of the importance of understanding the characteristics of speech and the types of tasks that are required to progress through the three sublevels that constitute the Intermediate level and move into the Advanced range. However, since the rating is holistic, information on the specific aspects of a test taker's performance that prevent that person from being rated at the next adjacent level is not documented, nor is it provided in the final rating. Thus, there can be a disconnect between what the examinees see as their rating, the information that instructors provide to students about the assessment and the rating system, and what raters are attending to when assigning ratings. The purpose of this study was to examine information that is not traditionally available to either test takers or instructors so as to provide more detailed information about the specific profiles of speakers who received the same proficiency rating within the Intermediate range and determine how a test taker's skills along four linear axes (function, text type, content, and accuracy) contributed to their final, global rating.</p> <hd id="AN0122100053-2">Background</hd> <p>The ACTFL defines proficiency as the “ability to use a language to communicate meaningful information in a spontaneous interaction, and in a manner acceptable and appropriate to native speakers of the language” (ACTFL, [<reflink idref="bib4" id="ref3">4</reflink>] , p. 4). The proficiency guidelines (ACTFL, [<reflink idref="bib3" id="ref4">3</reflink>] ) have long been represented as an inverted pyramid, which illustrates that language learning is not linear but rather that the progression from one level to the next can be best represented as a pattern of geometric growth. When envisioning the inverse pyramid, the geometric area in the Novice and Intermediate tiers is much smaller than that of the higher levels. However, the skills that are acquired at those levels form the structural foundation upon which the higher levels are built. For example, while the ability to narrate in the past is a critical characteristic of Advanced‐level communication, language learners usually first learn to report events that have taken place in strings of sentences using the simple past. However, the learner who does not develop the ability to use paragraph‐length discourse will not be able to progress beyond the Intermediate level (ACTFL, [<reflink idref="bib1" id="ref5">1</reflink>] ). Communicative habits that seem to appropriately convey meaning but that are not corrected and extended become ingrained and thus impede progress into and beyond the Advanced level. These fossilized errors in essence become faulty girders and beams that are incapable of supporting the increasing communicative weight when learners are required to carry out more sophisticated functions and address more robust and varied content. Thus, understanding the developmental stages through which learners progress is vital in assisting students in their language‐learning journey, both within a particular level but also from one level into the next.</p> <p>Although the ACTFL guidelines were introduced in 1982 (Liskin‐Gasparro, [<reflink idref="bib19" id="ref6">19</reflink>] ) and the ACTFL recently certified the 1,000th OPI tester worldwide (ACTFL, [<reflink idref="bib5" id="ref7">5</reflink>] ), it is quite likely that many foreign language educators may still be unclear about how exactly to use them to improve student learning outcomes. While there are more than 4,000 institutions of higher education and more than 35,000 high schools (U.S. Department of Education, [<reflink idref="bib29" id="ref8">29</reflink>] , [<reflink idref="bib30" id="ref9">30</reflink>] ) across the United States, only a fraction of the secondary and postsecondary institutions (1 in every 390) has certified personnel to assist in assessment and lead proficiency‐oriented curricular revisions. Even though some institutions have instituted curriculum‐wide training in proficiency assessment (Brooks & Darhower, [<reflink idref="bib6" id="ref10">6</reflink>] ; Gouoni & Feyten, [<reflink idref="bib14" id="ref11">14</reflink>] ), many foreign language educators must rely on written descriptions of the scale with little understanding of how the descriptors relate to actual language production. The result is that a huge segment of the foreign language education community is left with an understanding of the guidelines that is cursory at best or reductive to certain grammatical forms at worst.</p> <p>In simple terms, each major proficiency level (Novice, Intermediate, Advanced, Superior, and Distinguished) is defined as a confluence of four domains: function, text type, content, and accuracy. These features are defined in more detail in Table [NaN] .</p> <p>Speech Characteristics Analyzed by Area of Focus</p> <p> <ephtml> <table><tr><th align="left">Area of Focus</th><th align="center">Characteristics of Speech</th><th align="center">Description</th></tr><tr><td align="left">Function</td><td align="left">Focus on topic/task</td><td align="left">The degree to which the examinee completed the task presented as defined by the major level</td></tr><tr><td align="left">Text type</td><td align="left">Text length</td><td align="left">The extent to which the amount of language completed the function of the task (words and phrases, sentences, strings of sentences, or connected paragraphs)</td></tr><tr><td align="left" /><td align="left">Discourse organization</td><td align="left">The extent to which the text was organized appropriately and the use of appropriate cohesive markers to organize speech</td></tr><tr><td align="left">Content</td><td align="left">Vocabulary use</td><td align="left">The quantity and quality of lexicon needed to accomplish the task appropriately</td></tr><tr><td align="left">Accuracy/comprehensibility expectations</td><td align="left">Fluency</td><td align="left">The extent to which the rate of speech, length of runs, pauses, and other timing features affected the comprehensibility of the message for a native listener</td></tr><tr><td align="left" /><td align="left">Pronunciation</td><td align="left">The extent to which individual words and phrases were articulated in a way that was comprehensible to the listener</td></tr><tr><td align="left" /><td align="left">Grammatical/structural accuracy</td><td align="left">The degree of control of the grammar/syntax needed to accomplish the task in a way that was comprehensible to the listener</td></tr></table> </ephtml> </p> <p>While learners may progress in a linear way on each of these characteristics, conjoint mastery of multiple linear characteristics is necessary for movement through one major level and into the next. Obtaining a rating at the next higher level only occurs through sustained performance of the lower levels (ACTFL, [<reflink idref="bib1" id="ref12">1</reflink>] ; Clifford, [<reflink idref="bib9" id="ref13">9</reflink>] ). Before a rating can be awarded, the speaker must demonstrate a sustained level or “floor” of performance across tasks, text type, content, and accuracy (see Table [NaN] ) as well as a breakdown level or “ceiling” in which the examinee can no longer sustain performance in one or more of the four domains (ACTFL, [<reflink idref="bib1" id="ref14">1</reflink>] ). For examinees in the Intermediate range, the floor is the ability to create with language in sentence‐length utterances that demonstrate control over the content that is needed in daily life; the ceiling is the ability to use paragraph‐length discourse to narrate and describe topics of personal and community interest in all major time frames. While an Intermediate‐level speaker may exhibit some characteristics of Advanced‐level proficiency in certain topic domains or with certain linguistic features, he or she is unable to sustain this level of performance across the requisite range of topics, tasks, or linguistic features with the required level of accuracy and thus does not demonstrate Advanced‐level ability.</p> <p>Floor and Ceiling Performance of Intermediate Speakers</p> <p> <ephtml> <table><tr><th align="left">Level</th><th align="center">Floor (or Intermediate Criteria)</th><th align="center">Ceiling (or Advanced Criteria)</th></tr><tr><td align="left">Function</td><td align="left">Create with language Participate in simple conversations Ask and answer questions</td><td align="left">Narrate and describe in major time frames (past, present, and future) Linguistically negotiate situations with complications</td></tr><tr><td align="left">Text Type</td><td align="left">Sentences</td><td align="left">Paragraphs</td></tr><tr><td align="left">Content</td><td align="left">Self Daily life</td><td align="left">Self Daily life Nonautobiographical topics Topics of general interest</td></tr><tr><td align="left">Accuracy</td><td align="left">Understood by people accustomed to speaking to nonnative speakers</td><td align="left">Can be understood without confusion by monolinguals not accustomed to speaking to nonnative speakers</td></tr></table> </ephtml> </p> <p>As shown in Table [NaN] , the fundamental difference between speech that is rated at any one of the three Intermediate sublevels (Low, Mid, High) lies in the quality and quantity of the examinee's language when engaged in at‐level tasks (Clifford, [<reflink idref="bib9" id="ref15">9</reflink>] ). The Low sublevel is indicative of a speaker who just barely demonstrates competence when performing the tasks for the major level. Meanwhile, a rating at the Mid sublevel indicates that the speaker fulfills all the requirements of the major level with sufficient quantity and quality of language across the assessment criteria. There is no doubt that the examinee can perform the functions of that major level; indeed, the response is much more substantial than that of a speaker at the Low sublevel. The High sublevel rating indicates that the speaker demonstrates a robust ability to meet the criteria for the proficiency level in question and that he or she also attempts and executes with success some of the tasks and can often—but not always—meet the related expectations for text type, context, and level of accuracy that are required at the next higher (adjacent) major level—in this case, Advanced. Thus, a rating of IH indicates that the speaker exhibits Advanced‐level performance most, but not all of the time, by either exhibiting all the traits of the Advanced level in certain topic domains and not others, or by exhibiting Advanced features such as text type and fluency but not others such as pronunciation or grammatical accuracy.</p> <p>While a number of studies have looked at the validity of the OPI and the use of its scale in oral proficiency testing (Dandonoli & Henning, [<reflink idref="bib12" id="ref16">12</reflink>] ; Halleck, [<reflink idref="bib16" id="ref17">16</reflink>] ; Surface & Dierdorff, [<reflink idref="bib21" id="ref18">21</reflink>] ; Thompson, [<reflink idref="bib26" id="ref19">26</reflink>] , [<reflink idref="bib27" id="ref20">27</reflink>] ), little empirical research has specifically sought to document examinees’ strengths and weaknesses at each sublevel within a major level, primarily because the single holistic rating of the speech sample as a whole results in a lack of transparency about exactly what such a rating means and on which dimensions a test taker showed strength or weakness. Thus, while the small percentage of instructors who have received formal OPI training can intuit the reason their students may have received a particular score, the large number of instructors who have less familiarity with the scale may:</p> <p>fail to understand the conjunctive nature of a proficiency rating (Clifford, [<reflink idref="bib9" id="ref21">9</reflink>] ),</p> <p>overestimate their own students’ abilities (Levine & Haus, [<reflink idref="bib17" id="ref22">17</reflink>] ), or</p> <p>confound the performance of rehearsed material with proficiency (Cox, Bown, & Burdis, [<reflink idref="bib11" id="ref23">11</reflink>] ).</p> <p>In an attempt to help instructors understand differences in levels, Liskin‐Gasparro ([<reflink idref="bib18" id="ref24">18</reflink>] ) analyzed the communication strategies of IH and Advanced Low (AL) Spanish speakers and found that the AL speakers used a broader range of communicative strategies; however, she did not analyze other aspects of the interview samples and did not look at the differences among the Intermediate sublevels. Apart from this study, most of the research has focused on learners’ expected proficiency outcomes at particular points in their program of study (Carroll, [<reflink idref="bib7" id="ref25">7</reflink>] ; Chambless, [<reflink idref="bib8" id="ref26">8</reflink>] ; Glisan & Foltz, [<reflink idref="bib15" id="ref27">15</reflink>] ; Gouoni & Feyten, [<reflink idref="bib14" id="ref28">14</reflink>] ) or has compared learners’ results on the two alternate forms of the assessment, the OPIc and the OPI (Surface, Poncheri, & Bhavsar, [<reflink idref="bib22" id="ref29">22</reflink>] ; SWA Consulting, [<reflink idref="bib23" id="ref30">23</reflink>] ; Thompson, Cox, & Knapp, [<reflink idref="bib28" id="ref31">28</reflink>] ). In contrast, this study sought to determine how the scale is operationalized. That is, the purpose of the study was to look into the black box, so to speak, of the Intermediate level to find empirical data and identify the patterns of the linguistic strengths and weaknesses of IL, IM, and IH speakers when they carried out different types of tasks. The study addressed the following questions:</p> <p>What are the most common linguistic features of speakers at each Intermediate sublevel (IL, IM, IH)? Which characteristics prevent speakers from being rated at the next higher sublevel?</p> <p>How well do speakers at each Intermediate sublevel (IL, IM, IH) perform on different task types that operationalize the criteria of Intermediate and Advanced proficiency? Which task types prevent speakers from being rated at the next higher sublevel?</p> <hd id="AN0122100053-3">Method</hd> <p>To answer the research questions, the author analyzed data from a research report that CREDU (a subsidiary of Samsung) had commissioned the ACTFL to study and then wrote a technical report (Cox, [<reflink idref="bib10" id="ref32">10</reflink>] ). The data from that report form the basis for the current article. To determine commonly manifested Intermediate‐level speech characteristics and discover what prevents a test taker from being rated at the next higher adjacent level, experienced ACTFL raters were recruited to analyze existing assessment data from an OPIc.</p> <hd id="AN0122100053-4">Raters</hd> <p>Nine raters selected from the ACTFL's certified OPIc rater pool were recruited: seven were certified ACTFL OPI testers, six were certified ACTFL OPI trainers, three were part of the original OPIc development team, and eight were members of the OPIc Quality Assurance team. The strength of using trained raters guaranteed that the feedback provided was from experts who know the scale intimately. Although the approach is susceptible to the criticism of confirmation bias (raters’ feedback could possibly have been based on the language in the descriptors rather than on unbiased observation), this susceptibility was deemed acceptable (<reflink idref="bib1" id="ref33">1</reflink>) due to the lack of research into what trained raters think as they rate speech samples, and (<reflink idref="bib2" id="ref34">2</reflink>) because only highly experienced raters can provide the necessary analysis of the internal mechanisms that result in ratings across the Intermediate range. While raters typically score OPIcs holistically (ACTFL, [<reflink idref="bib2" id="ref35">2</reflink>] ), for this study, the raters scored the tests analytically by examining specific linguistics features and tasks, analyzing in depth the difficulty of the types of tasks associated with the functions of the Intermediate and Advanced levels, and determining the extent to which different speech features were present in those level‐specific tasks at each of the two major levels (Intermediate and Advanced). To gather qualitative data, raters also had the opportunity to share comments on the specific tasks and on the speech samples that they rated.</p> <hd id="AN0122100053-5">Examinee Data</hd> <p>To control for the variance between native and target language learning, the study was limited to Korean‐speaking adults who were learning English, who were taking the English OPIc, and who were at different OPI levels. All exams were chosen from the existing pool of OPIc assessments taken by Korean test takers. To meet the selection criteria, each assessment had to have been previously double or triple‐rated by raters who had been in exact agreement on the sublevel awarded (e.g., all raters independently rated an examinee as IM). From the list of double‐ and triple‐rated exams, stratified random sampling was used to select 100 exams at each sublevel (Low, Mid, and High) for a total of 300 exams.</p> <hd id="AN0122100053-6">Design</hd> <p>A connected design was used, in which all the raters analyzed a subset of examinees and tasks from existing OPIcs as a way to verify that all the raters were applying the same criteria. Just a single rater then analyzed the subsequent examinees/tasks to allow for the broadest possible survey of examinee response types.</p> <p>To answer the first research question, the approach differed for each of the three sublevels (IL, IM, and IH).</p> <p>For IL speakers, five Intermediate tasks and one Advanced task were analyzed. The objective of examining more Intermediate tasks at this level was to determine in which domains the ILs needed to improve both the quantity and quality of their responses and on which functions test takers needed to improve in order to reach IM.</p> <p>For IM speakers, four Intermediate tasks and three Advanced tasks were analyzed. The Intermediate tasks provided a basis of comparison between the IL's threshold performance and the IM's strong performance at the Intermediate level. The Advanced tasks provided direct information on what the examinees needed to do to move up the scale to the IH rating and beyond. It is important to note that one moves from IM to IH primarily by focusing on and improving the ability with Advanced‐level tasks, although improving performance in Intermediate‐level tasks happens as well.</p> <p>For IH speakers, six Advanced tasks were analyzed. A High rating indicates evidence of performance of all Advanced task types most of the time yet an inability to sustain that performance. Therefore, to determine the linguistic features of IH, the most useful information would come from an analysis of learners’ performance with Advanced‐level tasks.</p> <p>To answer the second research question, a subset of task types was selected for detailed analysis from among the 15 items on each form of the OPIc (Novice High to IM or IM to Advanced). This served to reduce the amount of time that was needed to analyze the profile of any individual examinee, thus resulting in a broader sampling of different examinees from which generalizations could be drawn. Since each form was individually tailored to the examinee, it was not possible to analyze items at the question level; however, since each question represented a specific task type (see Table [NaN] ), the speech samples that were represented in all of the interviews were fundamentally equivalent.</p> <p>Descriptions of Task Type Analyzed by Intermediate Sublevel</p> <p> <ephtml> <table><tr><th align="left">Task Level</th><th align="center">Task Type Description</th><th align="center">IL</th><th align="center">IM</th><th align="center">IH</th></tr><tr><td align="left">Intermediate</td><td align="left">Talk about thing or place</td><td align="center">2</td><td align="center">1</td><td align="center" /></tr><tr><td align="left">Intermediate</td><td align="left">Talk about activity or routine</td><td align="center">1</td><td align="center">1</td><td align="center" /></tr><tr><td align="left">Intermediate</td><td align="left">Ask questions</td><td align="center">1</td><td align="center">1</td><td align="center" /></tr><tr><td align="left">Intermediate</td><td align="left">Intermediate role‐play</td><td align="center">1</td><td align="center">1</td><td align="center" /></tr><tr><td align="left">Advanced</td><td align="left">Past description</td><td align="center">1</td><td align="center">1</td><td align="center">1</td></tr><tr><td align="left">Advanced</td><td align="left">Past narration</td><td align="center" /><td align="center">1</td><td align="center">1</td></tr><tr><td align="left">Advanced</td><td align="left">Advanced role‐play (situation with a complication)</td><td align="center" /><td align="center">1</td><td align="center">1</td></tr><tr><td align="left">Advanced</td><td align="left">Role‐play follow‐up—Narration/description</td><td align="center" /><td align="center" /><td align="center">1</td></tr><tr><td align="left">Advanced</td><td align="left">Narration/description—context beyond personal</td><td align="center" /><td align="center" /><td align="center">1</td></tr><tr><td align="left">Advanced</td><td align="left">Report current event</td><td align="center" /><td align="center" /><td align="center">1</td></tr><tr><td align="center">Total Tasks Analyzed</td><td align="center">6</td><td align="center">7</td><td align="center">6</td></tr></table> </ephtml> </p> <hd id="AN0122100053-7">Procedure</hd> <p>The raters were each assigned 30 examinees, including 10 examinees who had been rated at each sublevel (IL, IM, and IH). Raters used a rubric, shown in Figure [NaN] . For each examinee's speech sample, raters were asked to analyze six or seven individual tasks and were instructed to listen to each task twice. When listening the first time, they were to assess the response holistically on a five‐point scale ranging from “does not meet expectations”—1 (e.g., total breakdown, i.e., the examinee could not produce any language at the intended level) to “exceeds expectations”—5 (e.g., produced language far above the intended task difficulty level). When listening for the second time, the raters were asked to identify any characteristics that would help explain their holistic assessment of the response. For example, a task that received a global assessment of “almost meets expectations” might include weakness in pronunciation, strength in vocabulary use, and moderate ability with the other skills. A second rubric was adapted from one that had been previously employed in a heritage language study (Swender, Martin, Rivera‐Martinez, & Kagan, [<reflink idref="bib24" id="ref36">24</reflink>] ) with a comment field so that raters could note issues (e.g., technical difficulties, other speech characteristics, etc.) that were not easily addressed with the five‐point scale and offer any comments that would add more information to the rating they had awarded.</p> <hd id="AN0122100053-8">Data Analysis</hd> <p>This study was guided by two primary research questions. The first investigated the most common linguistic features of examinees at the Intermediate level by task level (Intermediate or Advanced). The second investigated the way in which task type (see Table [NaN] ) affected examinee performance at the Intermediate level by task level (Intermediate or Advanced). These questions were answered by looking at the mean overall rating (see Figure [NaN] ) by task level. For the first question, the means of the different linguistic features (e.g., fluency, pronunciation, etc.; see Table [NaN] ) were compared and 95% confidence intervals (95% CI) were calculated and graphed to determine how the features differed. For the second question, the means of the task types (“talk about thing or place,” “talk about activity or routine,” etc.; see Table [NaN] ) were compared with 95% CIs calculated and graphed as well. The 95% CI is an estimate of population parameter and is generally represented by an error bar (I) with either a dot or line in the middle indicating the population mean. Examining the length of the error bars and the overlap among variables of interest provided a visual representation of the differences among the variables and their effect sizes. Where there was no overlap between error bars, the means were statistically different from one another. Where there was total overlap, there might not be any difference between the variables.</p> <hd id="AN0122100053-9">Findings</hd> <hd id="AN0122100053-10">IL Speakers</hd> <p>By definition, IL speakers at minimum can accomplish Intermediate‐level functions but are not expected to successfully perform the functions that are required to complete Advanced‐level tasks. In the analysis of the samples rated IL, a rating of 3 was indicative of minimally meeting the requirements. For the five Intermediate‐level tasks, the mean rating of overall task performance was 2.78 (see Table [NaN] ).</p> <p>Speech Criteria Rating on IL Sample</p> <p> <ephtml> <table><tr><th align="left" /><th align="center">Intermediate Tasks (n = 5)</th><th align="center">Advanced Tasks (n = 1)</th></tr><tr><th align="left" /><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Overall</td><td align="char" char=".">544</td><td align="char" char=".">2.78</td><td align="char" char=".">0.78</td><td align="center">[2.73, 2.83]</td><td align="char" char=".">107</td><td align="char" char=".">1.17</td><td align="char" char=".">0.48</td><td align="center">[1.08, 1.26]</td></tr><tr><td align="left">Function</td></tr><tr><td align="left">—Focus on topic and task</td><td align="char" char=".">540</td><td align="char" char=".">2.87</td><td align="char" char=".">1.03</td><td align="center">[2.78, 2.96]</td><td align="char" char=".">107</td><td align="char" char=".">1.23</td><td align="char" char=".">0.45</td><td align="center">[1.15, 1.31]</td></tr><tr><td align="left">Text Type</td></tr><tr><td align="left">—Length</td><td align="char" char=".">541</td><td align="char" char=".">3.06</td><td align="char" char=".">0.72</td><td align="center">[3.00, 3.12]</td><td align="char" char=".">107</td><td align="char" char=".">1.34</td><td align="char" char=".">0.57</td><td align="center">[1.23, 1.45]</td></tr><tr><td align="left">—Discourse organization</td><td align="char" char=".">535</td><td align="char" char=".">3.03</td><td align="char" char=".">0.73</td><td align="center">[2.97, 3.09]</td><td align="char" char=".">107</td><td align="char" char=".">1.69</td><td align="char" char=".">0.85</td><td align="center">[1.53, 1.85]</td></tr><tr><td align="left">Content</td></tr><tr><td align="left">—Vocabulary use</td><td align="char" char=".">539</td><td align="char" char=".">3.12</td><td align="char" char=".">0.65</td><td align="center">[3.07, 3.17]</td><td align="char" char=".">105</td><td align="char" char=".">1.50</td><td align="char" char=".">0.72</td><td align="center">[1.36, 1.64]</td></tr><tr><td align="left">Accuracy</td></tr><tr><td align="left">—Fluency</td><td align="char" char=".">541</td><td align="char" char=".">2.97</td><td align="char" char=".">0.65</td><td align="center">[2.92, 3.02]</td><td align="char" char=".">105</td><td align="char" char=".">2.12</td><td align="char" char=".">1.09</td><td align="center">[1.91, 2.33]</td></tr><tr><td align="left">—Pronunciation</td><td align="char" char=".">539</td><td align="char" char=".">3.19</td><td align="char" char=".">0.63</td><td align="center">[3.14, 3.24]</td><td align="char" char=".">107</td><td align="char" char=".">1.30</td><td align="char" char=".">0.57</td><td align="center">[1.19, 1.41]</td></tr><tr><td align="left">—Grammatical/structural</td><td align="char" char=".">541</td><td align="char" char=".">3.02</td><td align="char" char=".">0.71</td><td align="center">[2.96, 3.08]</td><td align="char" char=".">107</td><td align="char" char=".">1.97</td><td align="char" char=".">1.13</td><td align="center">[1.76, 2.18]</td></tr></table> </ephtml> </p> <p>1 Note: Please note that in some instances raters provided an overall rating but when there was evidence of memorized material they did not rate the individual linguistic features. This is further discussed in the qualitative section of this article.</p> <hd id="AN0122100053-11">Linguistic Features</hd> <p>In examining the linguistic features that contributed to overall task performance on the Intermediate tasks, raters found that two of the features did not meet the required threshold (a score of 3): “fluency” and “focus on topic and task.” For the one Advanced‐level task, the overall mean was 1.17—all of the linguistic features were scored between the criteria “does not meet” (a score of 1) and “almost meets” (a score of 2). A MANOVA showed that all of the linguistic features of the Intermediate‐level tasks were statistically different from those of the Advanced‐level tasks, F(<reflink idref="bib1" id="ref37">1</reflink>, 622) = 26.86, p < 0.0001; Wilk's Λ = 0.79.</p> <p>Figure [NaN] presents the means as well as the 95% CIs as represented by error bars (I). When the speakers were performing Intermediate‐level tasks, performances across the seven categories all clustered around the “minimally meets” threshold of 3. With a difference of 0.32 between the highest (pronunciation) and lowest (“focus on topic and task”) categories, the profile was relatively even. With the Advanced‐level task, scores in only one domain (fluency) exceeded the “almost meets” threshold (a score of 2). With a difference of 0.89 between the highest‐rated domain (“fluency”) and the lowest (“focus on task and topic”), the profile was more disparate.</p> <p>This indicates that the profile of an IL speaker was one in which the ratings on the different categories clustered around the “minimally meets” threshold for Intermediate‐level tasks. With the Advanced‐level task, the developmental profile was less equal. “Focus on topic and task” and “pronunciation” were the strongest areas and were statistically equivalent. The weakest areas were “grammatical/structural,” “text type (length),” and “discourse organization,” with “fluency” and “vocabulary use” just slightly higher than the other three.</p> <hd id="AN0122100053-12">Task Types</hd> <p>To determine the effect of the task type on the overall performance and holistic final rating, the mean of examinees’ overall performance was examined by task type (see Table [NaN] ). None of the Intermediate‐level tasks exceeded the “minimally meets” requirement (a score of 3). Among the Intermediate‐level tasks, “intermediate role‐play” had the lowest mean (mean = 2.62, SD = 0.90), and “talk about activity or routine” had the highest mean (mean = 2.99, SD = 0.86). The Advanced‐level task, “past description,” scored substantially lower than the Intermediate‐level tasks (mean = 1.17, SD = 0.45).</p> <p>Overall Mean of IL Speakers on Task Type</p> <p> <ephtml> <table><tr><th align="left">Task Type</th><th align="center">Task Level</th><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Intermediate role‐play</td><td align="left">Intermediate</td><td align="char" char=".">107</td><td align="char" char=".">2.62</td><td align="char" char=".">0.90</td><td align="center">[2.44, 2.80]</td></tr><tr><td align="left">Talk about thing or place (two prompts)</td><td align="left">Intermediate</td><td align="char" char=".">225</td><td align="char" char=".">2.88</td><td align="char" char=".">0.72</td><td align="center">[2.79, 2.97]</td></tr><tr><td align="left">Talk about activity or routine</td><td align="left">Intermediate</td><td align="char" char=".">106</td><td align="char" char=".">2.99</td><td align="char" char=".">0.86</td><td align="center">[2.83, 3.15]</td></tr><tr><td align="left">Ask questions</td><td align="left">Intermediate</td><td align="char" char=".">106</td><td align="char" char=".">2.80</td><td align="char" char=".">0.80</td><td align="center">[2.64, 2.96]</td></tr><tr><td align="left">Past description</td><td align="left">Advanced</td><td align="char" char=".">107</td><td align="char" char=".">1.17</td><td align="char" char=".">0.45</td><td align="center">[1.09, 1.25]</td></tr></table> </ephtml> </p> <p>In Figure [NaN] , the means as well as the 95% CIs as represented by error bars (I) showed some interesting trends. With the Intermediate level, the performances across the tasks were not significantly different from one another as demonstrated by the error bars; however, “intermediate role‐play” and “ask questions” did appear to be more difficult than the other three Intermediate‐level task types.</p> <p>There were approximately 286 comments on the overall performance of the IL speakers. Many of the comments confirmed what was observed with the quantitative analysis (improving fluency, accuracy, etc.); however, one trend that emerged was the role that rehearsed material or canned/memorized responses had on raters’ ability to assess the sample of speech in a valid way. In nine instances, raters gave a holistic rating of “does not meet” but then used the comments field to note why they did not provide numerical ratings for some of the other linguistic features.</p> <p>Approximately 20% of the rater comments noted that the responses to these specific tasks sounded scripted or rehearsed. This could be an artifact of analyzing Korean examinees, where memorization is often employed as a test preparation strategy. While this may be effective for success on tests of content, sheer memorization and then recitation of such responses is a feature of the Novice level of oral proficiency and therefore would not result in sufficient language for an examinee to be rated at a level higher than Novice. Official ACTFL rating protocols require the raters to listen to the entire speech sample and not just individual tasks, as was the case in this study (ACTFL, [<reflink idref="bib1" id="ref38">1</reflink>] ); however, when a single task type is learned and rehearsed, it does not provide evidence of an examinee's spontaneous, productive speech.</p> <hd id="AN0122100053-13">IM Speakers</hd> <p>By definition, IM speakers fully meet the requirements needed to accomplish the functions that are required by the Intermediate‐level tasks but are not able to successfully sustain the functions that are required at the Advanced level. This was found to be the case for the IM speakers. For the four Intermediate‐level tasks, the mean rating for overall performance was 3.59, an indication that test takers exceeded the “minimally meets” threshold and were approaching the “fully meets” level (see Table [NaN] ).</p> <p>Holistic Rating of IM Speakers on Task by Speech Characteristic</p> <p> <ephtml> <table><tr><th align="left" /><th align="center">Intermediate Tasks (n = 4)</th><th align="center">Advanced Tasks (n = 3)</th></tr><tr><th align="left" /><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Overall</td><td align="char" char=".">440</td><td align="char" char=".">3.59</td><td align="char" char=".">0.73</td><td align="center">[3.52, 3.65]</td><td align="char" char=".">334</td><td align="char" char=".">1.57</td><td align="char" char=".">0.65</td><td align="center">[1.50, 1.64]</td></tr><tr><td align="left">Function</td></tr><tr><td align="left">—Focus on topic and task</td><td align="char" char=".">421</td><td align="char" char=".">3.60</td><td align="char" char=".">0.82</td><td align="center">[3.52, 3.68]</td><td align="char" char=".">347</td><td align="char" char=".">2.39</td><td align="char" char=".">1.07</td><td align="center">[2.28, 2.50]</td></tr><tr><td align="left">Text Type</td></tr><tr><td align="left">—Length</td><td align="char" char=".">422</td><td align="char" char=".">3.71</td><td align="char" char=".">0.57</td><td align="center">[3.66, 3.76]</td><td align="char" char=".">345</td><td align="char" char=".">1.73</td><td align="char" char=".">0.74</td><td align="center">[1.65, 1.81]</td></tr><tr><td align="left">—Discourse organization</td><td align="char" char=".">422</td><td align="char" char=".">3.66</td><td align="char" char=".">0.61</td><td align="center">[3.60, 3.72]</td><td align="char" char=".">347</td><td align="char" char=".">1.80</td><td align="char" char=".">0.8</td><td align="center">[1.72, 1.88]</td></tr><tr><td align="left">Content</td></tr><tr><td align="left">—Vocabulary use</td><td align="char" char=".">420</td><td align="char" char=".">3.74</td><td align="char" char=".">0.53</td><td align="center">[3.69, 3.79]</td><td align="char" char=".">343</td><td align="char" char=".">2.25</td><td align="char" char=".">0.9</td><td align="center">[2.15, 2.35]</td></tr><tr><td align="left">Accuracy</td></tr><tr><td align="left">—Fluency</td><td align="char" char=".">421</td><td align="char" char=".">3.52</td><td align="char" char=".">0.60</td><td align="center">[3.46, 3.58]</td><td align="char" char=".">345</td><td align="char" char=".">1.97</td><td align="char" char=".">0.89</td><td align="center">[1.88, 2.06]</td></tr><tr><td align="left">—Pronunciation</td><td align="char" char=".">421</td><td align="char" char=".">3.68</td><td align="char" char=".">0.55</td><td align="center">[3.63, 3.73]</td><td align="char" char=".">346</td><td align="char" char=".">2.37</td><td align="char" char=".">1.05</td><td align="center">[2.26, 2.48]</td></tr><tr><td align="left">—Grammatical/structural</td><td align="char" char=".">422</td><td align="char" char=".">3.59</td><td align="char" char=".">0.62</td><td align="center">[3.53, 3.65]</td><td align="char" char=".">345</td><td align="char" char=".">1.69</td><td align="char" char=".">0.75</td><td align="center">[1.61, 1.77]</td></tr></table> </ephtml> </p> <p>2 Note: Please note that in some instances, raters provided an overall rating but when there was evidence of memorized material, they did not rate the individual linguistic features. This will be further discussed in the qualitative section of this paper.</p> <hd id="AN0122100053-14">Linguistic Features</hd> <p>In examining the linguistic features that contributed to the overall task performance scores on the Intermediate‐level tasks, raters found that all of the features exceeded the “minimally meets” threshold (a score of 3), though none exceeded the “fully meets” threshold (a score of 4). For the three Advanced‐level tasks, the overall mean was 1.57 and with three of the linguistic features (“focus on topic and task,” “vocabulary use,” and “pronunciation”) exceeding the “almost meets” criterion of 2.</p> <p>Figure [NaN] presents the means as well as the 95% CIs as represented by error bars (I). When the speakers were performing the Intermediate‐level tasks, their performance across the seven categories exceeded the “minimally meets” threshold of 3 but did not reach the “fully meets” threshold of 4. With a difference of just 0.22 between the means for the highest (“vocabulary use”) and lowest (“fluency”) characteristics, the profile was relatively even. With the Advanced‐level tasks, none of the categories met the “minimally meets” threshold and with a difference of 0.70 between the means for the highest (“focus on topic and task”) and lowest (“grammatical/structural”) domains, the profile was more disparate.</p> <p>This indicates that the profile of an IM speaker was one in which the different categories easily exceeded the “minimally meets” threshold for Intermediate‐level tasks. With the Advanced‐level tasks, the developmental profile across domains was less equal. “Focus on topic and task,” “pronunciation,” and “vocabulary use” were the strongest areas and were statistically equivalent. The weakest areas were “grammatical/structural,” “length,” and “discourse organization.”</p> <hd id="AN0122100053-15">Task Types</hd> <p>To determine the effect of the task type on performance, the mean of overall performance was examined by task type (see Table [NaN] ) Among the Intermediate‐level tasks, “talk about activity or routine” had the lowest mean (mean = 3.35, SD = 0.85), and “asking questions” had the highest mean (mean = 3.61, SD = 0.71). With the Advanced‐level tasks, all were scored below the “minimally meets” requirement level, with the lowest mean being “past description” (mean = 1.49, SD = 0.71) and the highest being “advanced role‐play” (mean = 2.46, SD = 0.68).</p> <p>Overall Mean of IM Speakers on Task Type</p> <p> <ephtml> <table><tr><th align="left">Task Type</th><th align="center">Task Level</th><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Intermediate role‐play</td><td align="left">Intermediate</td><td align="char" char=".">111</td><td align="char" char=".">3.60</td><td align="char" char=".">0.72</td><td align="center">[3.46, 3.74]</td></tr><tr><td align="left">Talk about activity or routine</td><td align="left">Intermediate</td><td align="char" char=".">112</td><td align="char" char=".">3.35</td><td align="char" char=".">0.85</td><td align="center">[3.19, 3.51]</td></tr><tr><td align="left">Talk about thing or place</td><td align="left">Intermediate</td><td align="char" char=".">111</td><td align="char" char=".">3.52</td><td align="char" char=".">0.76</td><td align="center">[3.38, 3.66]</td></tr><tr><td align="left">Ask questions</td><td align="left">Intermediate</td><td align="char" char=".">111</td><td align="char" char=".">3.61</td><td align="char" char=".">0.71</td><td align="center">[3.47, 3.75]</td></tr><tr><td align="left">Past description</td><td align="left">Advanced</td><td align="char" char=".">111</td><td align="char" char=".">1.49</td><td align="char" char=".">0.62</td><td align="center">[1.37, 1.61]</td></tr><tr><td align="left">Past narration</td><td align="left">Advanced</td><td align="char" char=".">143</td><td align="char" char=".">1.53</td><td align="char" char=".">0.61</td><td align="center">[1.43, 1.63]</td></tr><tr><td align="left">Advanced role‐play</td><td align="left">Advanced</td><td align="char" char=".">105</td><td align="char" char=".">2.46</td><td align="char" char=".">0.68</td><td align="center">[2.32, 2.60]</td></tr></table> </ephtml> </p> <p>In Figure [NaN] , the means as well as the 95% CIs as represented by error bars (I) showed some interesting trends. With the Intermediate‐level tasks, I across the different tasks overlapped, indicating that the performances were not significantly different. With the Advanced‐level tasks, however, test takers’ scores on “advanced role‐play” were significantly higher than their scores on the other two Advanced‐level tasks. This could be due to the fact that resolving situations with complications at the Advanced level can often be accomplished without paragraph‐length discourse.</p> <hd id="AN0122100053-16">Qualitative Analysis of IM Speakers</hd> <p>There were approximately 216 comments on the overall performance of the IMs. While many of the comments confirmed what was observed with the quantitative analysis—test takers provided both good quantity and quality of speech when completing the Intermediate‐level tasks—for Advanced‐level tasks, there was still a need for improvement in all areas (e.g., in accuracy, text type, and discourse organization). One trend that also emerged with the IM speakers was the role that “rehearsed material” or “canned/memorized responses” played. With the IL speakers, approximately 20% of the comments indicated that test takers’ responses to these specific tasks sounded rehearsed; however, with the IM speakers the rate was much lower—only 12% (or 26 total responses) were considered by raters to constitute instances of rehearsed material. As noted in the IL discussion, while memorization may be an effective strategy for tests of content, sheer memorization and then recitation of such responses is a feature of the Novice level.</p> <hd id="AN0122100053-17">IH Speakers</hd> <p>By definition, IH speakers fully meet the requirements that are needed to accomplish the functions that are assessed by Intermediate‐level tasks (research question 1) and are able to successfully meet the functions and other criteria of the Advanced level most of the time (research question 2). Because the High sublevel is primarily defined in terms of performance at the next higher major level, only Advanced‐level tasks were analyzed. While it would have been interesting to analyze performance on some of the Intermediate‐level tasks as a point of comparison, that was beyond the scope of this study. For the six Advanced‐level tasks, the mean for raters’ performance score for the overall task was 2.13 (see Table [NaN] ).</p> <p>Holistic Rating of IH Speakers on Task by Speech Characteristic</p> <p> <ephtml> <table><tr><th align="left" /><th align="center">Advanced Tasks (n = 3)</th></tr><tr><th align="left" /><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Overall</td><td align="char" char=".">622</td><td align="char" char=".">2.13</td><td align="char" char=".">0.65</td><td align="center">[2.08, 2.18]</td></tr><tr><td align="left">Function</td></tr><tr><td align="left">—Focus on topic and task</td><td align="char" char=".">594</td><td align="char" char=".">2.85</td><td align="char" char=".">0.92</td><td align="center">[2.78, 2.92]</td></tr><tr><td align="left">Text Type</td></tr><tr><td align="left">—Length</td><td align="char" char=".">592</td><td align="char" char=".">2.34</td><td align="char" char=".">0.70</td><td align="center">[2.28, 2.40]</td></tr><tr><td align="left">—Discourse organization</td><td align="char" char=".">594</td><td align="char" char=".">2.38</td><td align="char" char=".">0.71</td><td align="center">[2.32, 2.44]</td></tr><tr><td align="left">Content</td></tr><tr><td align="left">—Vocabulary use</td><td align="char" char=".">588</td><td align="char" char=".">2.64</td><td align="char" char=".">0.68</td><td align="center">[2.59, 2.69]</td></tr><tr><td align="left">Accuracy</td></tr><tr><td align="left">—Fluency</td><td align="char" char=".">595</td><td align="char" char=".">2.46</td><td align="char" char=".">0.77</td><td align="center">[2.40, 2.52]</td></tr><tr><td align="left">—Pronunciation</td><td align="char" char=".">594</td><td align="char" char=".">2.76</td><td align="char" char=".">0.83</td><td align="center">[2.69, 2.83]</td></tr><tr><td align="left">—Grammatical/structural</td><td align="char" char=".">592</td><td align="char" char=".">2.38</td><td align="char" char=".">0.71</td><td align="center">[2.32, 2.44]</td></tr></table> </ephtml> </p> <hd id="AN0122100053-18">Linguistic Features</hd> <p>As noted earlier, a holistic assessment of 3 was indicative of minimally meeting the requirements. As shown in Table [NaN] , none of the IHs exceeded that minimum in any of the seven categories, with the lowest score for “length” (mean = 2.34, SD = 0.69) and the highest for “focus on topic and task” (mean = 2.85, SD = 0.92).</p> <p>Figure [NaN] presents the means as well as the 95% CIs as represented by error bars (I). When these speakers were performing the Advanced‐level tasks, their performance across the seven categories all clustered between the thresholds of 2 and 3. With a difference of 0.51 between the means for the highest domain (“focus on topic and task”) and the lowest (“length”), their profile was relatively even, indicating that the profile of an IH develops evenly across the required set of expectations but does not yet meet expectations at the next level. “Focus on topic and task” and “pronunciation” were the strongest areas and were statistically equivalent. The weakest areas were “grammatical/structural,” “text type (length)” and “discourse organization,” with “fluency” and “vocabulary use” just slightly higher than the other three.</p> <hd id="AN0122100053-19">Task Types</hd> <p>To determine the effect of the task type on performance, the mean of the raters’ scores of overall performance was examined by task type (see Table [NaN] ). Scores for all of the Advanced‐level tasks fell below the “minimally meets” requirement level, with the lowest mean for “current event” (mean = 1.90, SD = 0.66) and the highest for “advanced role‐play” (mean = 2.46, SD = 0.68).</p> <p>Overall Mean of IH Speakers on Task Type</p> <p> <ephtml> <table><tr><th align="left">Task Type</th><th align="center">Task Level</th><th align="center">N</th><th align="center">Mean</th><th align="center">SD</th><th align="center">95% CI</th></tr><tr><td align="left">Past description</td><td align="center">Advanced</td><td align="char" char=".">108</td><td align="char" char=".">2.11</td><td align="char" char=".">0.56</td><td align="center">[1.93, 2.29]</td></tr><tr><td align="left">Past narration</td><td align="center">Advanced</td><td align="char" char=".">73</td><td align="char" char=".">2.36</td><td align="char" char=".">0.63</td><td align="center">[2.22, 2.50]</td></tr><tr><td align="left">Advanced role‐play</td><td align="center">Advanced</td><td align="char" char=".">105</td><td align="char" char=".">2.46</td><td align="char" char=".">0.68</td><td align="center">[2.30, 2.62]</td></tr><tr><td align="left">Role‐play follow‐up</td><td align="center">Advanced</td><td align="char" char=".">106</td><td align="char" char=".">2.10</td><td align="char" char=".">0.58</td><td align="center">[1.96, 2.24]</td></tr><tr><td align="left">Narration/description beyond personal</td><td align="center">Advanced</td><td align="char" char=".">105</td><td align="char" char=".">2.01</td><td align="char" char=".">0.64</td><td align="center">[1.85, 2.17]</td></tr><tr><td align="left">Current event</td><td align="center">Advanced</td><td align="char" char=".">105</td><td align="char" char=".">1.90</td><td align="char" char=".">0.66</td><td align="center">[1.82, 1.98]</td></tr></table> </ephtml> </p> <p>In Figure [NaN] , the means as well as the 95% CIs as represented by error bars (I) show some interesting trends. With the Advanced‐level tasks, “advanced role‐play” and “past narration” had the highest scores, indicating that these were the easiest tasks for the examinees. The next easiest were “past description” and “role‐play follow‐up.” The most difficult were “narration,” “description beyond the personal,” and “current events.”</p> <hd id="AN0122100053-20">Qualitative Analysis of IH Speakers</hd> <p>There were approximately 41 comments on the overall performance of the IH speakers. While many of the comments confirmed what has been learned from the quantitative analysis, the comments indicated that for Advanced‐level tasks, there was still a need for improvement in all areas (e.g., improving accuracy, text type, and discourse organization). One point to note on task type is that as with the IM speakers, the raters found that “advanced role‐play” was more easily performed successfully than the other Advanced‐level tasks, probably because it often does not require paragraph‐level speech. This observation was supported by the quantitative analysis as well.</p> <hd id="AN0122100053-21">Discussion</hd> <p>The purpose of this research project was to provide empirical data on the profiles of examinees who were rated at the Intermediate level and in this way to offer an initial road map for helping students to progress through Intermediate to Advanced levels of proficiency. Furthermore, since OPIc data were used, this study is the first to examine the impact of the different task types that are required at the Intermediate and Advanced levels.</p> <hd id="AN0122100053-22">Speech Characteristics</hd> <p>The first research question examined the linguistic characteristics at each of the sublevels. To answer that question, it is necessary to parse the speakers’ performance on Intermediate‐ and Advanced‐level tasks.</p> <hd id="AN0122100053-23">Intermediate‐Level Tasks</hd> <p>The only way to understand what IL speakers need to do to improve to the IM sublevel is to examine the change in linguistic characteristics between IL and IM speakers when they performed Intermediate‐level tasks. Figure [NaN] shows the 95% CI means and error bars of the linguistic characteristics of five Intermediate‐level tasks for IL speakers and four Intermediate‐level tasks for IM speakers that were rated.<sups>1</sups> While both groups of learners could meet the linguistic demands of the Intermediate level, the IM speakers were stronger in all areas. This indicates an ease in performing Intermediate‐level functions and provides empirical evidence that there is an increase in the quantity and quality of the language produced between the sublevels.</p> <p>As would be expected, the IL speakers’ speech samples averaged near the “minimally meets” threshold (a score of 3), while the IM speakers’ speech samples demonstrated their ability to perform all of the functions that are associated with the Intermediate level using both good quantity and quality of language. Even though all speech samples in this study had been originally double‐ or triple‐rated, it is interesting that at the IL sublevel, the samples were rerated just below the “minimally meets” border rating of 3. Some might argue that this is evidence that OPIc scoring has a compensatory element—the failure to minimally meet the requirements of any single task can be compensated for by stronger performance on other tasks. Thus, the whole speech sample could be rated more highly than some of the individual parts, although it may also be an artifact of the use of rehearsed speech.</p> <p>Table [NaN] lists the mean order rank of the characteristics that IL speakers must improve on when moving toward the IM sublevel. Given that rehearsed material could be subsumed under “focus on task and topic,” it is not surprising that this characteristic was the lowest for the IL speakers. Rather than memorizing responses, students must understand that they must be able to spontaneously (<reflink idref="bib1" id="ref39">1</reflink>) create with language, (<reflink idref="bib2" id="ref40">2</reflink>) perform simple transactions, and (<reflink idref="bib3" id="ref41">3</reflink>) ask and answer questions, and that they should practice tailoring responses to different circumstances rather than going into the autopilot of rehearsed material. While most speakers have a collection of anecdotes that they share in conversations, test takers called attention to their inability to create with language when they offered glibly fluent, memorized responses that did not address the question and were not adapted for the audience.</p> <p>Linguistic Characteristic on Intermediate Task From Weakest to Strongest</p> <p> <ephtml> <table><tr><th align="left">Mean Rank Order</th><th align="center">IL</th><th align="center">IM</th></tr><tr><td align="left">7th</td><td align="left">Function: focus on topic and task</td><td align="left">Accuracy: fluency</td></tr><tr><td align="left">6th</td><td align="left">Accuracy: fluency</td><td align="left">Accuracy: grammatical/structural</td></tr><tr><td align="left">5th</td><td align="left">Accuracy: grammatical/structural</td><td align="left">Function: focus on topic and task</td></tr><tr><td align="left">4th</td><td align="left">Text type: discourse organization</td><td align="left">Text type: discourse organization</td></tr><tr><td align="left">3rd</td><td align="left">Text type: length</td><td align="left">Accuracy: pronunciation</td></tr><tr><td align="left">2nd</td><td align="left">Content: vocabulary use</td><td align="left">Text type: length</td></tr><tr><td align="left">1st</td><td align="left">Accuracy: pronunciation</td><td align="left">Content: vocabulary use</td></tr></table> </ephtml> </p> <p>This issue may be more endemic with the OPIc than the OPI in that it is difficult for OPIc raters to investigate whether speech is being spontaneously created or is simply being recited from memory. For example, when an interviewee struggles to create with the language (e.g., “This uh uh question uh about school uh uh very uh interest…”) and then transitions to a more fluid response (e.g., “Built in the 1940s, the school I attended was part of the Art Deco movement in which…”), an OPI interviewer can interrupt the soliloquy by asking follow‐up and clarification questions that guide the conversation in a new direction. With OPIcs, however, raters must listen for telltale signs of rehearsed responses and then exclude that sample as evidence that the examinee can create with the language. The opportunity cost of using rehearsed material is that there are fewer chances for an examinee to show what can be produced spontaneously. These results indicate that, rather than helping examinees to be rated at a higher level, the uneven juxtaposition of rehearsed material with spontaneous language is a hallmark of the IL sublevel.</p> <p>In addition, to be rated at the next higher sublevel, IL speakers also need to reduce disfluencies in sentence‐level discourse to increase both the quantity and quality of their speech. Often disfluencies arise as learners search for words and self‐correct errors in grammar—an indication that the recall of vocabulary and grammatical structures has not yet been automatized. For learners to progress from conceptual control, which often entails conscientious effort to produce forms, to full control, in which production is automatized, learners must engage in ample, abundant, and varied conversational language practice. The benefit of varied conversational practice is that it allows learners to practice recombining and repurposing rehearsed and memorized material from the Novice level as well as appropriately adjusting and adapting it to new circumstances. Engaged conversational practice over a wide range of personal topics will enable IL speakers to develop the ease and fluency that is needed to progress to the IM sublevel.</p> <hd id="AN0122100053-24">Advanced‐Level Tasks</hd> <p>To understand what IM and IH speakers need to do to move up to the next sublevel, one needs to examine the linguistic characteristics of speakers who performed Advanced‐level tasks. Figure [NaN] shows the 95% CI means and error bars of the linguistic characteristics of the Advanced‐level tasks that were rated (one for IL speakers, three for IM speakers, and six for IH speakers). None of the groups successfully met the linguistic demands of the Advanced level; however, the higher the sublevel, the stronger their performance of each linguistic characteristic. Thus, progression toward the Advanced level requires systematic improvement among all linguistic characteristics.</p> <p>To move to the next higher sublevel, IM and IH speakers need to show progress toward accomplishing Advanced‐level functions. The speech characteristic areas that were found to be the weakest for both of these groups and thus in need of the most improvement were “grammatical/structural,” “length,” and “discourse organization” (see Table [NaN] ).</p> <p>Linguistic Characteristic on Advanced Task From Weakest to Strongest</p> <p> <ephtml> <table><tr><th align="left">Mean Rank Order</th><th align="center">IM</th><th align="center">IH</th></tr><tr><td align="left">7th</td><td align="left">Accuracy: grammatical/structural</td><td align="left">Text type: length</td></tr><tr><td align="left">6th</td><td align="left">Text type: length</td><td align="left">Text type: discourse organization</td></tr><tr><td align="left">5th</td><td align="left">Text type: discourse organization</td><td align="left">Accuracy: grammatical/structural</td></tr><tr><td align="left">4th</td><td align="left">Accuracy: fluency</td><td align="left">Accuracy: fluency</td></tr><tr><td align="left">3rd</td><td align="left">Content: vocabulary use</td><td align="left">Content: vocabulary use</td></tr><tr><td align="left">2nd</td><td align="left">Accuracy: pronunciation</td><td align="left">Accuracy: pronunciation</td></tr><tr><td align="left">1st</td><td align="left">Function: focus on topic and task</td><td align="left">Function: focus on topic and task</td></tr></table> </ephtml> </p> <p>Speakers must move beyond simple sentences to perform the functions that are required at the Advanced level. Thus, as IM and IH speakers engage in descriptions and narrations, sentence complexity will naturally increase. In the case of English, it will involve moving toward complex sentences with embedded clauses (e.g., “The girl over there wearing the red sweater is my cousin”) and will also include moving from partial to full control of tense and aspect when narrating or describing in different time frames (e.g., “When I was walking to school this morning, I ran into my cousin”). The text type progresses along the continuum from “sentences” to “strings of sentences” until it develops into paragraphs with discourse markers (e.g., first, next, then, however). The function of detailed description and narration cannot be attained without increasing length and organizational tags. Thus, grammatical/structural accuracy, length, and organization—the three characteristics that IM and IH speakers must work on—are all interrelated.</p> <p>Just as IL speakers need to enlarge their language base as well as adapt and transfer it to new and varied contexts, IM and IH speakers must develop greater breadth and accuracy; in addition, they must fundamentally reconfigure their speech habits—to add another floor on top of the Intermediate‐level girders that were mentioned at the beginning of the article. They need to move beyond conversational exchanges and practice carrying out Advanced‐level functions. While IM speakers would likely benefit from drawing content from familiar, autobiographic domains and adding more complexity and length to their utterances, IH speakers may benefit from moving beyond the autobiographical by acquiring more content domains. Advanced‐level speakers are often compared to news reporters—they can describe the setting and narrate the details of stories over a wide range of topics. Thus, IH speakers would benefit from opportunities to practice sharing content by describing settings and narrating stories in many different domains. The ability to describe or narrate in all time frames requires speakers to use enough language (text type) to paint a verbal picture (discourse organization and vocabulary) with enough precision (accuracy) that a monolingual listener (accuracy) can visualize the scene.</p> <p>Once again, a speaker's inability to fulfill these functions at the Advanced level may result from the failure to use enough language, to organize it meaningfully, or to offer enough precision to communicate without causing confusion or misunderstanding. Thus, concentrating on different aspects of the four construct axes can help learners gain the deliberate practice (Ericsson, [<reflink idref="bib13" id="ref42">13</reflink>] ) needed for incremental growth. While this type of growth often is best achieved during intensive immersion experiences like study abroad (Pearson, Fonseca‐Greber, & Foell, [<reflink idref="bib20" id="ref43">20</reflink>] ), growth can also be facilitated by instructors requiring vocabulary development and grammar learning to be completed out of class and deliberately allocating a very high percentage of class time to extended communication activities. For example, when focusing on past narration (an Advanced‐level function), learners could work on spontaneously producing detailed paragraph‐length discourse using a variety of sentence structures and connecting devices without penalty for grammatical errors. Then, to improve accuracy, learners could work at the sentence level to correct their recorded speech. The inverse (creating a written base text and then spontaneously enhancing and elaborating on it by adding detail and content and varying the sentence structure) would also help learners focus on improving their language along all four of the required dimensions.</p> <hd id="AN0122100053-25">Task Type Difficulty</hd> <p>The second research question examined the difficulty of the Intermediate‐ and Advanced‐level tasks at each of the sublevels. To answer this question, it is necessary to parse performance by major level.</p> <hd id="AN0122100053-26">Intermediate‐Level Tasks</hd> <p>The only way to understand what IL speakers need to do to improve to the IM sublevel is to examine the performance differences between IL and IM speakers on the different Intermediate‐level tasks. Figure [NaN] shows the 95% CI means and error bars of the linguistic characteristics of four different Intermediate‐level task types that were rated. The IL speakers were at or just under the “minimally meets” threshold of 3, while the IM speakers could clearly perform the different Intermediate‐level tasks. As occurred with the linguistic characteristics, it was expected that the IL speakers would reach the threshold; however, this result could have been another instance where the whole was greater than the sum of its parts, or it could have been due to many of the responses being rehearsed and thus not able to be rated.</p> <p>It is important to point out that the ordering of task difficulty was different between the IL and IM speakers (see Table [NaN] ). For IM speakers, their strengths were performing “intermediate role‐play” and “ask questions,” both of which required transactional language. Yet those same tasks were the most difficult for the IL speakers. Engaging in role‐plays is not part of everyday, spontaneous conversation; however, “intermediate role‐play” in the OPIc was designed to allow examinees to demonstrate their ability to handle simple transactions or social situations (e.g., make a purchase, accept or propose an invitation) that are not readily elicited through a conversational format. Because the successful completion of “intermediate role‐play” required the speaker to ask questions, it is not surprising that “ask questions” was the other function most in need of improvement. Thus, for speakers to move to the next sublevel, improving the ability to ask questions spontaneously in interactional contexts must take priority while speakers focus on the linguistic features that have already been discussed.</p> <p>Ordering of Intermediate Tasks From Weakest to Strongest</p> <p> <ephtml> <table><tr><th align="left">Mean Rank Order</th><th align="center">IL</th><th align="center">IM</th></tr><tr><td align="left">4th</td><td align="left">Intermediate role‐play</td><td align="left">Talk about activity or routine</td></tr><tr><td align="left">3rd</td><td align="left">Ask questions</td><td align="left">Talk about thing or place</td></tr><tr><td align="left">2nd</td><td align="left">Talk about thing or place</td><td align="left">Intermediate role‐play</td></tr><tr><td align="left">1st</td><td align="left">Talk about activity or routine</td><td align="left">Ask questions</td></tr></table> </ephtml> </p> <hd id="AN0122100053-27">Advanced‐Level Tasks</hd> <p>To understand what IM and IH speakers need to do to improve up a sublevel, one needs to examine the change in their differences in their performance on Advanced‐level tasks. Figure [NaN] shows the 95% CI means and error bars of Advanced‐level tasks (one for IL speakers, three for IM speakers, and six for IH speakers). None of the groups successfully performed the Advanced‐level tasks; however, the higher the sublevel, the stronger the performance.</p> <p>The easiest tasks for both the IMs and the IHs was the “advanced role‐play” (see Table [NaN] ). While “intermediate role‐play” was designed to elicit the language that is needed for simple conversational exchanges, the role‐play at the Advanced level added a complication and placed the transaction in a more formal setting. This required that examinees use more precise language and actively negotiate with the interlocutor. As noted above, the “advanced role‐play” required test takers to add new girders in their linguistic structure. However, since such encounters are still transactional even at the Advanced level, paragraph‐length discourse may not be needed to accomplish the task. The ordering of the “role‐play follow‐up” for the IH speakers was also somewhat surprising given the relative ease with which they performed the role‐play itself. The purpose of the follow‐up was to provide another opportunity for examinees to describe or narrate a personal instance in which they had experienced something similar to what was in the role‐play, and that does require paragraph‐length discourse. This finding provides empirical evidence that resolving complicated situations may be the first trait that is acquired in the progression toward the Advanced level, but discussion and narration within the same topic domain is more difficult.</p> <p>Ordering of Advanced Tasks From Weakest to Strongest</p> <p> <ephtml> <table><tr><th align="left">Mean Rank Order</th><th align="center">IM</th><th align="center">IH</th></tr><tr><td align="left">6th</td><td align="center">—</td><td align="left">Current event</td></tr><tr><td align="left">5th</td><td align="center">—</td><td align="left">Narration/description beyond personal</td></tr><tr><td align="left">4th</td><td align="center">—</td><td align="left">Role‐play follow‐up</td></tr><tr><td align="left">3rd</td><td align="center">Past description</td><td align="left">Past description</td></tr><tr><td align="left">2nd</td><td align="center">Past narration</td><td align="left">Past narration</td></tr><tr><td align="left">1st</td><td align="center">Advanced role‐play</td><td align="left">Advanced role‐play</td></tr></table> </ephtml> </p> <p>That “past narration” was easier than “past description” was somewhat unexpected. Typically, “past narration” requires greater command of grammatical/structural accuracy, which intuitively would seem to be more difficult than offering a detailed description. It could be that autobiographical narrations “sound” better to raters because the rhetorical structure is different from that of a description. Furthermore, it would have been interesting to examine how IM speakers responded to the other Advanced‐level tasks such as “current event” to see if the ordering of all tasks was the same between IM and IH speakers. Clearly, more research must be conducted to explore this phenomenon.</p> <p>When certified testers attempt to gather evidence of Advanced language proficiency, they often employ a three‐prong strategy in which examinees are (<reflink idref="bib1" id="ref44">1</reflink>) asked to describe a setting or situation; (<reflink idref="bib2" id="ref45">2</reflink>) asked to elaborate, clarify, or expand the same topic; and (<reflink idref="bib3" id="ref46">3</reflink>) prompted to relate the story from the outset to the conclusion (Swender & Vicars, [<reflink idref="bib25" id="ref47">25</reflink>] ). A similar strategy could be used with IM and IH speakers if they try to describe, elaborate, and narrate in all the major time frames, moving from topics that are personal to those that are more general. As the most difficult tasks for IH speakers were “narration/description beyond personal” and “current event,” it is evident that the increased cognitive load of discussing general issues spontaneously could be impacting their linguistic control. Thus, having students read or listen to authentic material and then asking them to describe, elaborate, and recount the story will help examinees have content they can incorporate and use as they gain the language skills and build girders that are needed to move along the continuum of the Intermediate sublevels to the Advanced level.</p> <hd id="AN0122100053-28">Limitations and Future Directions</hd> <p>While this study identified common patterns of language growth, there are some issues that still must be taken into consideration. First, human performance is variable and not every learner will follow the same path through the Intermediate sublevels into the Advanced level. Thus, while this research reports general trends, it would not be surprising for individual exceptions to occur. Second, using trained raters has both strengths and weaknesses. One strength is that it ensures that those doing the rating understand the scale well and know what to look for. However, that familiarity could be a weakness as it may lead to confirmation bias in which those same raters use circular logic to justify their ratings. Thus, a rater listening to an examinee who is already known to be an IL speaker will only be looking for evidence to support that rating rather than simply rating the task on its merits alone; however, this does not seem to invalidate the insights that were gained into test takers’ linguistic characteristics and the way in which task types differentiate among proficiency levels. Finally, since this study was conducted with Korean speakers learning English, it is unknown to what extent these findings would be generalizable to other languages and learners.</p> <hd id="AN0122100053-29">Conclusion</hd> <p>Understanding the stages that learners go through as they progress through the Intermediate range into the Advanced levels using the ACTFL proficiency framework has received very little attention. Fortunately, the OPIc allows a better view into what may be happening as learners progress through that major level as both linguistic characteristics and task types can be analyzed.</p> <p>To return to the student who languished at the Intermediate level for 3 years across nine attempts to demonstrate Advanced‐level proficiency, it might have been helpful if her instructors had intervened and helped her specifically target the different linguistic areas and task types at each developmental sublevel. For example, when she started as an IL speaker, she could have been instructed to work on adapting her memorized language to different circumstances and to work on the spontaneous back‐and‐forth that characterizes transactional language as well as asking and answering questions about personal experiences and daily life contexts. As she developed into an IM speaker, the transactional language that used to be a weakness should now be a strength, and she could practice moving beyond simple sentences to more complex strings of sentences. Recording and transcribing what she said could provide the foundation for learning how to combine simple sentences using subordination and how to enrich the content by adding detail. For example, an instructor could ask her to elaborate on what she was speaking about by telling her that for every person (or object) she mentioned (e.g., a cousin), she needed to think of three traits (e.g., physical description, hobbies, occupation) that she could incorporate into the description. Repeatedly rerecording more structurally complex versions that were also more rich in content could help the learner establish patterns of more complex grammar use, transition from shorter to longer text types, and increasingly incorporate nonautobiographical and general content. Such continued practice across a variety of contexts would force rehearsed material to be adapted and allow her to confirm that she could create with the language as she emerged into the IH sublevel.</p> <p>As an IH speaker, this learner should be using Advanced language most of the time, although she would be unable to sustain it. The learning approach that allowed her to move into the IH sublevel would also allow her to move to the AL sublevel, but there would be a few caveats. Successful communication at the Advanced level requires that learners demonstrate the ability to create with language in longer text types using discourse markers and showing automatized fluency that incorporates more complex grammatical structures. Because Advanced‐level speech requires that entirely new girders be built in the learner's speech paradigm, it is often difficult to reach the level of fluency that is needed without abundant opportunities to produce paragraph‐length discourse. Thus, since classroom time alone is typically insufficient, other opportunities for extensive language practice must be incorporated. This could include study abroad, foreign language housing, speaking partners, or technological solutions that would allow consistent partnering with native speakers.</p> <p>Offering feedback on grammatical and structural errors that cause confusion or misunderstanding is also essential. Perhaps having this learner record and transcribe a response, circle and identify errors that she was aware of, and then rerecord herself would encourage her to notice and correct errors that might otherwise fossilize and thus limit the conjoint progression that is needed to reach the next major level. Once again, the learner should seek to communicate with a level of automaticity that would lead to increased fluency and would help her move beyond the purely autobiographical into narration and description that extends beyond the personal frame to include topics of general interest as well as current events. As both language learners and instructors come to understand the necessity of simultaneous and interrelated, or conjoint, development in function, text type, content, and accuracy, they can structure learning so as to scaffold performance on each of these dimensions to help learners more easily progress through the Intermediate level into the Advanced range.</p> <hd id="AN0122100053-30">Note</hd> <p>IH speakers did not have any Intermediate tasks analyzed.</p> <hd id="AN0122100053-31">Acknowledgments</hd> <p>This article was based on a research report written for the ACTFL, and the original members of that research team (Elvira Swender, Cynthia Martin, and Danielle Tezcan) provided valuable assistance throughout. I am extremely grateful for their generosity and friendship.</p> <ref id="AN0122100053-32"> <title>Footnotes</title> <blist> <bibl id="bib1" idref="ref5" type="bt">1</bibl> <bibtext>Troy L. Cox (PhD, Brigham Young University) is Associate Director of Research and Assessment, Center for Language Studies, Brigham Young University, Provo, Utah. </bibtext> </blist> </ref> <ref id="AN0122100053-33"> <title>References</title> <blist> <bibtext>ACTFL. ( 2012a ). Oral proficiency interview familiarization manual. Alexandria, VA: Author. </bibtext> </blist> <blist> <bibl id="bib2" idref="ref34" type="bt">2</bibl> <bibtext>ACTFL. ( 2012b ). Oral proficiency interview computerized familiarization manual. Alexandria, VA: Author. </bibtext> </blist> <blist> <bibl id="bib3" idref="ref4" type="bt">3</bibl> <bibtext>ACTFL. ( 2012c ). Proficiency guidelines 2012. Alexandria, VA: Author. </bibtext> </blist> <blist> <bibl id="bib4" idref="ref3" type="bt">4</bibl> <bibtext>ACTFL. ( 2012d ). Performance descriptors 2012. Alexandria, VA: Author. </bibtext> </blist> <blist> <bibl id="bib5" idref="ref7" type="bt">5</bibl> <bibtext>ACTFL. ( 2016 ). ACTFL achieves milestone of 1,000 certified ACTFL OPI testers. Retrieved February 21, 2017, from https://<ulink href="http://www.actfl.org/news/press-releases/actfl-achieves-milestone-1000-certified-actfl-opi-testers">www.actfl.org/news/press-releases/actfl-achieves-milestone-1000-certified-actfl-opi-testers</ulink></bibtext> </blist> <blist> <bibl id="bib6" idref="ref1" type="bt">6</bibl> <bibtext>Brooks, F. B., & Darhower, M. A. ( 2014 ). It takes a department! A study of the culture of proficiency in three successful foreign language teacher education programs. Foreign Language Annals, 47, 592 – 613. </bibtext> </blist> <blist> <bibl id="bib7" idref="ref25" type="bt">7</bibl> <bibtext>Carroll, J. B. ( 1967 ). Foreign language proficiency levels attained by language majors near graduation from college. Foreign Language Annals, 1, 131 – 151. </bibtext> </blist> <blist> <bibl id="bib8" idref="ref2" type="bt">8</bibl> <bibtext>Chambless, K. S. ( 2012 ). Teachers’ oral proficiency in the target language: Research on its role in language teaching and learning. Foreign Language Annals [Supplement], 45, s141 – s162. </bibtext> </blist> <blist> <bibl id="bib9" idref="ref13" type="bt">9</bibl> <bibtext>Clifford, R. ( 2016 ). A rationale for criterion‐referenced proficiency testing. Foreign Language Annals, 49, 224 – 234. </bibtext> </blist> <blist> <bibl id="bib10" idref="ref32" type="bt">10</bibl> <bibtext>Cox, T. ( 2015 ). Findings of the ACTFL‐CREDU research project: Linguistic profiles of Korean speakers of English. White paper submitted to ACTFL. </bibtext> </blist> <blist> <bibl id="bib11" idref="ref23" type="bt">11</bibl> <bibtext>Cox, T. L., Bown, J., & Burdis, J. ( 2015 ). Exploring proficiency‐based vs. performance‐based items with elicited imitation assessment. Foreign Language Annals, 48, 350 – 371. </bibtext> </blist> <blist> <bibl id="bib12" idref="ref16" type="bt">12</bibl> <bibtext>Dandonoli, P., & Henning, G. ( 1990 ). An investigation of the construct validity of the ACTFL proficiency guidelines and oral interview procedure. Foreign Language Annals, 23, 11 – 22. </bibtext> </blist> <blist> <bibl id="bib13" idref="ref42" type="bt">13</bibl> <bibtext>Ericsson, K. A. ( 2006 ). The influence of experience and deliberate practice on the development of superior expert performance. The Cambridge Handbook of Expertise and Expert Performance, 38, 685 – 705. </bibtext> </blist> <blist> <bibl id="bib14" idref="ref11" type="bt">14</bibl> <bibtext>Gouoni, J. M., & Feyten, C. M. ( 1999 ). Effects of the ACTFL OPI‐type training on student performance, instructional methods, and classroom materials in the secondary foreign language classroom. Foreign Language Annals, 32, 189 – 200. </bibtext> </blist> <blist> <bibl id="bib15" idref="ref27" type="bt">15</bibl> <bibtext>Glisan, E. W., & Foltz, D. A. ( 1998 ). Assessing students’ oral proficiency in an outcome‐based curriculum: Student performance and teacher intuitions. Modern Language Journal, 82, 1 – 18. </bibtext> </blist> <blist> <bibl id="bib16" idref="ref17" type="bt">16</bibl> <bibtext>Halleck, G. B. ( 1996 ). Interrater reliability of the OPI. Using academic trainee raters. Foreign Language Annals, 29, 223 – 238. </bibtext> </blist> <blist> <bibl id="bib17" idref="ref22" type="bt">17</bibl> <bibtext>Levine, M. G., & Haus, G. J. ( 1987 ). The accuracy of teacher judgment of the oral proficiency of high school foreign language students. Foreign Language Annals, 20, 45 – 50. </bibtext> </blist> <blist> <bibl id="bib18" idref="ref24" type="bt">18</bibl> <bibtext>Liskin‐Gasparro, J. E. ( 1996 ). Circumlocution, communication strategies, and the ACTFL proficiency guidelines: An analysis of student discourse. Foreign Language Annals, 29, 317 – 330. </bibtext> </blist> <blist> <bibl id="bib19" idref="ref6" type="bt">19</bibl> <bibtext>Liskin‐Gasparro, J. E. ( 2003 ). The ACTFL proficiency guidelines and the oral proficiency interview: A brief history and analysis of their survival. Foreign Language Annals, 36, 483 – 490. </bibtext> </blist> <blist> <bibl id="bib20" idref="ref43" type="bt">20</bibl> <bibtext>Pearson, L., Fonseca‐Greber, B., & Foell, K. ( 2006 ). Advanced proficiency for foreign language teacher candidates: What can we do to help them achieve this goal ? Foreign Language Annals, 39, 507 – 519. </bibtext> </blist> <blist> <bibl id="bib21" idref="ref18" type="bt">21</bibl> <bibtext>Surface, E. A., & Dierdorff, E. C. ( 2003 ). Reliability and the ACTFL oral proficiency interview: Reporting indices of interrater consistency and agreement for 19 languages. Foreign Language Annals, 36, 507 – 519. </bibtext> </blist> <blist> <bibl id="bib22" idref="ref29" type="bt">22</bibl> <bibtext>Surface, E., Poncheri, R., & Bhavsar, K. ( 2008 ). Two studies investigating the reliability and validity of the English ACTFL OPIc with Korean test takers: The ACTFL OPIc validation project technical report. Retrieved August 15, 2015, from <ulink href="http://www.languagetesting.com/wp‐content/uploads/2013/08/ACTFL‐OPIc‐English‐Validation‐2008.pdf">http://www.languagetesting.com/wp‐content/uploads/2013/08/ACTFL‐OPIc‐English‐Validation‐2008.pdf</ulink></bibtext> </blist> <blist> <bibl id="bib23" idref="ref30" type="bt">23</bibl> <bibtext>SWA Consulting Inc. ( 2009 ). Brief reliability report 5: Test‐retest reliability and absolute agreement rates of English ACTFL OPIc proficiency ratings for double and single rated tests within a sample of Korean test takers. Raleigh, NC: Author. </bibtext> </blist> <blist> <bibl id="bib24" idref="ref36" type="bt">24</bibl> <bibtext>Swender, E., Martin, C. L., Rivera‐Martinez, M., & Kagan, O. E. ( 2014 ). Exploring oral proficiency profiles of heritage speakers of Russian and Spanish. Foreign Language Annals, 47, 423 – 446. </bibtext> </blist> <blist> <bibl id="bib25" idref="ref47" type="bt">25</bibl> <bibtext>Swender, E., & Vicars, R. ( 2012 ). Oral proficiency interview training manual. Alexandria, VA: ACTFL. </bibtext> </blist> <blist> <bibl id="bib26" idref="ref19" type="bt">26</bibl> <bibtext>Thompson, I. ( 1995 ). A study of interrater reliability of the ACTFL oral proficiency interview in five European languages: Data from ESL, French, German, Russian, and Spanish. Foreign Language Annals, 28, 407 – 422. </bibtext> </blist> <blist> <bibl id="bib27" idref="ref20" type="bt">27</bibl> <bibtext>Thompson, I. ( 1996 ). Assessing foreign language skills. Data from Russian. Modern Language Journal, 80, 47 – 65. </bibtext> </blist> <blist> <bibl id="bib28" idref="ref31" type="bt">28</bibl> <bibtext>Thompson, G. L., Cox, T. L., & Knapp, N. ( 2016 ). Comparing the OPI and the OPIc: The effect of test method on oral proficiency scores and student preference. Foreign Language Annals, 49, 79 – 92. </bibtext> </blist> <blist> <bibl id="bib29" idref="ref8" type="bt">29</bibl> <bibtext>U.S. Department of Education. ( 2016a, December 16). Degree‐granting institutions and branches. Retrieved December 15, 2016, from http://nces.ed.gov//programs/digest/d02/dt244.asp </bibtext> </blist> <blist> <bibl id="bib30" idref="ref9" type="bt">30</bibl> <bibtext>U.S. Department of Education. ( 2016b, December 16). High school facts at a glance. Retrieved December 15, 2016, from <ulink href="http://www2.ed.gov/about/offices/list/ovae/pi/hs/hsfacts.html">http://www2.ed.gov/about/offices/list/ovae/pi/hs/hsfacts.html</ulink></bibtext> </blist> </ref> <p>Graph: Holistic Assessment Grid Example</p> <p>Graph: Holistic Assessment of IL Linguistic Characteristics by Intermediate and Advanced Task Level</p> <p>Graph: Holistic Assessment of IL Speakers on Intermediate and Advanced Task Type— Qualitative Comments</p> <p>Graph: Holistic Assessment of IM Linguistic Characteristics by Intermediate and Advanced Task Level</p> <p>Graph: Holistic Ratings of IM Speakers on Intermediate and Advanced Task Types</p> <p>Graph: Holistic Ratings of IH Speakers on Advanced‐Level Tasks</p> <p>Graph: Holistic Ratings of IH Speakers on Advanced Task Types</p> <p>Graph: Linguistic Characteristic Ratings of IL and IM Speakers on Intermediate‐Level Tasks</p> <p>Graph: Linguistic Characteristic Ratings of IL, IM, and IH Speakers on Advanced‐Level Tasks</p> <p>Graph: Holistic Ratings of IL and IM Speakers on Intermediate‐Level Tasksgr10</p> <p>Graph: Holistic Ratings of IL, IM, and IH Speakers on Advanced‐Level Tasks</p> <aug> <p>By Troy L. Cox</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1135287
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Understanding Intermediate-Level Speakers' Strengths and Weaknesses: An Examination of OPIc Tests from Korean Learners of English
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Cox%2C+Troy+L%2E%22">Cox, Troy L.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Foreign+Language+Annals%22"><i>Foreign Language Annals</i></searchLink>. Spr 2017 50(1):84-113.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 30
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2017
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Proficiency%22">Language Proficiency</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Characteristics%22">Student Characteristics</searchLink><br /><searchLink fieldCode="DE" term="%22Interviews%22">Interviews</searchLink><br /><searchLink fieldCode="DE" term="%22Linguistics%22">Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Skills%22">Language Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Performance%22">Performance</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Needs%22">Student Needs</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22South+Korea%22">South Korea</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/flan.12258
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0015-718X
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study profiled Intermediate-level learners in terms of their linguistic characteristics and performance on different proficiency tasks. A stratified random sample of 300 Korean learners of English with holistic ratings of Intermediate Low (IL), Intermediate Mid (IM), and Intermediate High (IH) on Oral Proficiency Interviews-computerized (OPIcs)--100 at each level--were analyzed by trained ACTFL raters to determine what was needed for the learners to progress to the next higher sublevel. The findings indicate that while ILs minimally met all the linguistic characteristics required of the Intermediate level, they needed to improve in the quantity and quality of all the linguistic characteristics they employed and improve their mastery of the types and variety of questions they could use when performing Intermediate tasks to move to the IM sublevel. In contrast, IMs demonstrated a pattern of strength when completing Intermediate tasks, but to move to the IH sublevel they needed to improve their ability to perform all Advanced-level tasks, especially in terms of accuracy when using paragraph-length discourse. Similar to the IMs, for the IHs to move to the Advanced Low sublevel, they needed to improve their accuracy with paragraph-length discourse and expand their content mastery to beyond the autobiographical.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2017
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1135287
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1135287
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/flan.12258
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 30
        StartPage: 84
    Subjects:
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Language Proficiency
        Type: general
      – SubjectFull: Student Characteristics
        Type: general
      – SubjectFull: Interviews
        Type: general
      – SubjectFull: Linguistics
        Type: general
      – SubjectFull: Language Skills
        Type: general
      – SubjectFull: Performance
        Type: general
      – SubjectFull: Student Needs
        Type: general
      – SubjectFull: South Korea
        Type: general
    Titles:
      – TitleFull: Understanding Intermediate-Level Speakers' Strengths and Weaknesses: An Examination of OPIc Tests from Korean Learners of English
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Cox, Troy L.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-print
              Value: 0015-718X
          Numbering:
            – Type: volume
              Value: 50
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Foreign Language Annals
              Type: main
ResultId 1