Can the Oral Proficiency Interview -- Computer (ACTFL OPIc) Be Used Instead of the Oral Proficiency Interview (ACTFL OPI)? An Aligned Rank Transform (ART) Analysis

Saved in:
Bibliographic Details
Title: Can the Oral Proficiency Interview -- Computer (ACTFL OPIc) Be Used Instead of the Oral Proficiency Interview (ACTFL OPI)? An Aligned Rank Transform (ART) Analysis
Language: English
Authors: Troy L. Cox (ORCID 0000-0001-9379-5102), Gregory L. Thompson (ORCID 0000-0003-4709-4836), Steven S. Stokes
Source: Foreign Language Annals. 2025 58(2):300-325.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 26
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Oral Language, Language Proficiency, Interviews, Computer Uses in Education, College Students, Second Language Learning, Spanish, Gender Differences, Age Differences, Scores, Testing, Test Reliability, Language Tests, Test Bias
Assessment and Survey Identifiers: ACTFL Oral Proficiency Interview
DOI: 10.1111/flan.12804
ISSN: 0015-718X
1944-9720
Abstract: This study investigated the differences between the ACTFL Oral Proficiency Interview (OPI) and the ACTFL Oral Proficiency Interview - Computer (OPIc) among Spanish learners at a U.S. university. Participants (N = 154) were randomly assigned to take both tests in a counterbalanced order to mitigate test order effects. Data were analyzed using an aligned rank transform (ART) analysis, focusing on variables such as gender, age, language courses, missionary experience, and self-assessed Spanish ability. Results showed a strong correlation between ACTFL OPI and ACTFL OPIc ratings ([tau] = 0.79, p < 0.001), though ACTFL OPIc ratings were slightly higher on average. No significant order effects were found, indicating the order of test administration did not influence ratings. The reliability of both tests was confirmed, and no significant biases were detected. The findings suggest that both ACTFL OPI and ACTFL OPIc are effective, holistic measures of Spanish oral proficiency, with ACTFL OPIc offering a slight advantage in rating outcomes. Pedagogically, this suggests flexibility in test choice without compromising assessment integrity.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1477432
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwEmqkykbDCFYX0Qd6CoTKnLAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDBPImCA6fkvaA7ep5wIBEICBmzvy3snpL42pQLN5StYq4cqLDb4111GighvFLboe-4Ga7qFfg_JfHcDojjGghto6jsienR0A3KHPJ18zyg-zOUC4WGPE12QO8eaSYBqF7hVu-N6mBTtPyBqQjiTGkt6dFxS6KFdnhii4YryWWQLXdqln-qwvqcFKq_Tbwmh-wbQobaGvUYqMMoE3YHH91m2M8DVrgYBccEpkousI
Text:
  Availability: 1
  Value: &lt;anid&gt;AN0186745416;fla01jul.25;2025Jul22.02:37;v2.2.500&lt;/anid&gt; &lt;title id=&quot;AN0186745416-1&quot;&gt;Can the Oral Proficiency Interview ‐ Computer (ACTFL OPIc) be used instead of the Oral Proficiency Interview (ACTFL OPI)? An aligned rank transform (ART) analysis&#160;&lt;/title&gt; &lt;p&gt;This study investigated the differences between the ACTFL Oral Proficiency Interview (OPI) and the ACTFL Oral Proficiency Interview ‐ Computer (OPIc) among Spanish learners at a U.S. university. Participants (N = 154) were randomly assigned to take both tests in a counterbalanced order to mitigate test order effects. Data were analyzed using an aligned rank transform (ART) analysis, focusing on variables such as gender, age, language courses, missionary experience, and self‐assessed Spanish ability. Results showed a strong correlation between ACTFL OPI and ACTFL OPIc ratings (τ = 0.79, p &amp;lt;.001), though ACTFL OPIc ratings were slightly higher on average. No significant order effects were found, indicating the order of test administration did not influence ratings. The reliability of both tests was confirmed, and no significant biases were detected. The findings suggest that both ACTFL OPI and ACTFL OPIc are effective, holistic measures of Spanish oral proficiency, with ACTFL OPIc offering a slight advantage in rating outcomes. Pedagogically, this suggests flexibility in test choice without compromising assessment integrity.&lt;/p&gt; &lt;p&gt;Keywords: learner characteristics; oral proficiency (OPI and OPIc); self‐assessment; Spanish; test development and validation&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-gra-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-gra-0001.jpg&quot; title=&quot;.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-3&quot;&gt;INTRODUCTION&lt;/hd&gt; &lt;p&gt;This study examines the comparability of two oral proficiency assessment formats for second language learners: the traditional ACTFL Oral Proficiency Interview (OPI) and the computer‐based ACTFL Oral Proficiency Interview ‐ Computer (OPIc). This study focuses on Spanish learners at a U.S. university, examining the reliability, potential biases, and practical implications of these testing methods.&lt;/p&gt; &lt;p&gt;The ACTFL OPI, long considered the gold standard for evaluating speaking skills, involves a face‐to‐face or telephonic interaction between a tester and an examinee. It offers a structured yet flexible approach, tailoring the interview to the examinee&#39;s proficiency level. The interview protocol for the ACTFL OPI is designed to elicit speech from an examinee that can then be rated against the ACTFL Proficiency Guidelines ([&lt;reflink idref=&quot;bib3&quot; id=&quot;ref1&quot;&gt;3&lt;/reflink&gt;]). These guidelines describe the geometric progression of language needed to perform different functions. As global tasks and functions increase in difficulty, the accuracy expectations increase, the content domains expand, and the text types become longer and more varied.&lt;/p&gt; &lt;p&gt;There are five major proficiency level descriptors that describe the alignment of function, accuracy, content, and text type: Novice, Intermediate, Advanced, Superior and Distinguished. The first three major levels are divided into sublevels of Low, Mid, and High. However, logistical challenges and costs associated with having a trained tester administer the exams have led to interest in alternative methods (Malone &amp;amp; Montee,&#160;[&lt;reflink idref=&quot;bib34&quot; id=&quot;ref2&quot;&gt;34&lt;/reflink&gt;]). The ACTFL OPIc was developed as a computer‐delivered test mimicking the OPI format but eliminating the need for a trained tester to administer each test. Test‐takers respond to prompts from a virtual avatar, allowing for consistent administration and scalability. Proponents argue it offers advantages in cost‐effectiveness, convenience, and the potential for more standardized administration (Isbell &amp;amp; Winke,&#160;[&lt;reflink idref=&quot;bib23&quot; id=&quot;ref3&quot;&gt;23&lt;/reflink&gt;]). However, some concerns remain regarding its ability to accurately replicate the interactive nature of the ACTFL OPI and its effectiveness in assessing higher levels of proficiency (Isbell &amp;amp; Winke,&#160;[&lt;reflink idref=&quot;bib23&quot; id=&quot;ref4&quot;&gt;23&lt;/reflink&gt;]; Thompson et al.,&#160;[&lt;reflink idref=&quot;bib54&quot; id=&quot;ref5&quot;&gt;54&lt;/reflink&gt;]).&lt;/p&gt; &lt;p&gt;Toulmin&#39;s model (1958) of argumentation provides a framework for analyzing and constructing arguments in practical contexts, emphasizing how reasoning is shaped by specific fields and audiences. The model identifies six key components: the claim, which is the main assertion; grounds, the evidence supporting the claim; the warrant, which connects the evidence to the claim; backing, offering additional justification for the warrant; the qualifier, indicating the strength of the claim; and the rebuttal, addressing potential counterarguments. Unlike formal logic, Toulmin&#39;s model is adaptable to real‐world discourse, making it particularly useful for evaluating the validity and practical implications of findings. Messick ([&lt;reflink idref=&quot;bib38&quot; id=&quot;ref6&quot;&gt;38&lt;/reflink&gt;]) proposed the use of this framework for test validation and is applied in the present study to assess the validity of using the ACTFL OPIc as an alternative to the traditional ACTFL OPI. The claim is that the ACTFL OPIc is a valid and reliable alternative to the ACTFL OPI for assessing oral proficiency in Spanish learners, with minimal bias and comparable rating outcomes.&lt;/p&gt; &lt;p&gt;The grounds are the evidence supporting the claim. Previous research comparing the ACTFL OPI and ACTFL OPIc has indicated a high degree of correlation between the two formats, but little has been done to explore potential areas of bias (Cubbellotti,&#160;[&lt;reflink idref=&quot;bib14&quot; id=&quot;ref7&quot;&gt;14&lt;/reflink&gt;]; Surface et al.,&#160;[&lt;reflink idref=&quot;bib52&quot; id=&quot;ref8&quot;&gt;52&lt;/reflink&gt;]; SWA Consulting,&#160;[&lt;reflink idref=&quot;bib53&quot; id=&quot;ref9&quot;&gt;53&lt;/reflink&gt;]; Thompson et al.,&#160;[&lt;reflink idref=&quot;bib54&quot; id=&quot;ref10&quot;&gt;54&lt;/reflink&gt;]). Thus, more research is needed to support the claim that the rating outcomes are comparable with minimal bias. Research on language exams has increasingly focused on the interaction between individual characteristics and test modes, revealing potential biases that can affect test outcomes. Studies indicate that personal characteristics, such as gender and ethnicity, can influence performance, particularly in different test formats. For instance, Baluyan ([&lt;reflink idref=&quot;bib6&quot; id=&quot;ref11&quot;&gt;6&lt;/reflink&gt;]) emphasizes the importance of considering test takers&#39; individual traits during test design to enhance reliability and authenticity. Ozdemir and Alshamrani&#39;s ([&lt;reflink idref=&quot;bib43&quot; id=&quot;ref12&quot;&gt;43&lt;/reflink&gt;]) work highlights the use of differential item functioning (DIF) methods to identify biased items across genders, suggesting that certain test items may favor one gender over another. Joo&#39;s ([&lt;reflink idref=&quot;bib26&quot; id=&quot;ref13&quot;&gt;26&lt;/reflink&gt;]) investigation into computerized versus face‐to‐face testing formats found significant rater biases, although no gender or age biases were detected. These findings underscore the necessity for ongoing research to mitigate bias and ensure fairness in language assessments, particularly as testing modalities evolve. However, it is essential to recognize that while biases exist, advancements in bias detection methodologies, including software tools, are improving the fairness of language tests (Ross &amp;amp; Okabe,&#160;[&lt;reflink idref=&quot;bib47&quot; id=&quot;ref14&quot;&gt;47&lt;/reflink&gt;]).&lt;/p&gt; &lt;p&gt;Research on language exams has increasingly examined the interaction between individual characteristics and test mode as potential sources of bias. Studies have highlighted various factors, including psychological aspects like emotioncy, which can influence test performance. For instance, Karami et al. ([&lt;reflink idref=&quot;bib28&quot; id=&quot;ref15&quot;&gt;28&lt;/reflink&gt;]) found that EFL learners&#39; emotioncy affected their vocabulary test outcomes, indicating a psychological source of bias. In addition, gender differentials were also analyzed in a high‐stakes language proficiency test, revealing no significant bias, suggesting that individual characteristics can interact differently with test modes (Karami,&#160;[&lt;reflink idref=&quot;bib27&quot; id=&quot;ref16&quot;&gt;27&lt;/reflink&gt;]). Furthermore, the Duolingo English Test was analyzed for fairness, showing that subjective perceptions of test validity and access significantly impacted performance, emphasizing the importance of individual perspectives in assessing bias (Yao,&#160;[&lt;reflink idref=&quot;bib60&quot; id=&quot;ref17&quot;&gt;60&lt;/reflink&gt;]). These findings collectively underscore the complexity of bias in language testing, influenced by both individual characteristics and the mode of assessment.&lt;/p&gt; &lt;p&gt;This study aims to address this gap by investigating possible biases related to age and gender of participants, language learning history, and self‐assessed ability. As proficiency tests claim to measure real‐world communication skills regardless of how the language was acquired, they should not favor examinees based on factors like gender, age, or learning background. This study pays particular attention to the role of self‐assessment in the OPIc, as it directly impacts the ratings examinees can earn. Unlike the OPI, the OPIc requires test‐takers to self‐assess their language ability to determine the administered form they will receive, each with a specific rating range (for a detailed description of the OPIc see ACTFL,&#160;[&lt;reflink idref=&quot;bib2&quot; id=&quot;ref18&quot;&gt;2&lt;/reflink&gt;]; Isbell &amp;amp; Winke,&#160;[&lt;reflink idref=&quot;bib23&quot; id=&quot;ref19&quot;&gt;23&lt;/reflink&gt;]). This addition to the OPIc could potentially alter the construct being measured as accurate self‐assessment is necessary to be able to receive the correct test form.&lt;/p&gt; &lt;p&gt;To address these questions, the study used an ART analysis to analyze data from 154 Spanish learners at a U.S. university. Participants were randomly assigned to take both the ACTFL OPI and ACTFL OPIc in a counterbalanced order to control for potential order effects. This experimental design aims to ensure that any observed differences in ratings can be attributed to the testing format rather than external variables. This study seeks to provide insights into the comparability of these two testing formats, their reliability in assessing oral proficiency, and any potential biases that may affect test outcomes. The findings have significant implications for language education, assessment practices, and the ongoing development of computer‐based language testing tools.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-4&quot;&gt;Research questions and study purpose&lt;/hd&gt; &lt;p&gt;The purpose of this paper is to determine whether ACTFL OPI versus ACTFL OPIc ratings achieve comparable outcomes with minimal bias. Sources of rating discrepancy are an important issue, given the potentially high‐stakes nature of many language certification experiences. In particular, we examine the following research questions:&lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;ulist&gt; &lt;item&gt; 1. Do the examinee characteristics of age or gender make a difference between ACTFL OPI and ACTFL OPIc rating outcomes?&lt;/item&gt; &lt;p&gt;&lt;/p&gt; &lt;item&gt; 2. Does the language learning experience of test‐takers in terms of years of language study or immersion experience make a difference between ACTFL OPI and ACTFL OPIc rating outcomes?&lt;/item&gt; &lt;p&gt;&lt;/p&gt; &lt;item&gt; 3. What is the relationship between individual self‐assessment and ACTFL OPI versus ACTFL OPIc rating outcomes?&lt;/item&gt; &lt;/ulist&gt; &lt;p&gt;By exploring these questions, this study aims to provide new insights into the comparability of these assessment methods, contributing to best practices in oral proficiency testing.&lt;/p&gt; &lt;p&gt;The findings have significant pedagogical implications for language teaching. If both the ACTFL OPIc and the ACTFL OPI assess students fairly without biases, the OPIc could be used more frequently, providing a more flexible and scalable assessment option that improves program efficiency. Understanding potential biases and limitations of each oral exam format can inform the development of more equitable assessment practices for language learners. Ultimately, this study seeks to advance language education by providing evidence‐based recommendations for oral proficiency assessment.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-5&quot;&gt;LITERATURE REVIEW&lt;/hd&gt; &lt;p&gt;This literature review focuses on two key areas:&lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;ulist&gt; &lt;item&gt; 1. Differences between direct and semi‐direct testing, relevant to comparing the ACTFL OPI and OPIc.&lt;/item&gt; &lt;p&gt;&lt;/p&gt; &lt;item&gt; 2. Overview of language testing bias research.&lt;/item&gt; &lt;/ulist&gt; &lt;p&gt;These topics help define the study&#39;s scope and contribute to a better understanding of the research results, clarifying the context of oral proficiency assessment methods.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-6&quot;&gt;The use of direct and semi‐direct testing in language exams&lt;/hd&gt; &lt;p&gt;Language testing is a key aspect of language education, essential for assessing proficiency and guiding instructional and high‐stakes decisions, such as university admissions and job placements. Among the various methods of assessment, direct and semi‐direct testing have gained notable attention (Hughes,&#160;[&lt;reflink idref=&quot;bib21&quot; id=&quot;ref20&quot;&gt;21&lt;/reflink&gt;]; Stansfield,&#160;[&lt;reflink idref=&quot;bib51&quot; id=&quot;ref21&quot;&gt;51&lt;/reflink&gt;]). Direct testing involves tasks that replicate real‐life language use, like speaking proficiency tests (e.g., ACTFL OPI). In contrast, semi‐direct testing uses controlled methods to simulate real‐life language use, such as computerized oral tests (e.g., ACTFL OPIc).&lt;/p&gt; &lt;p&gt;Quaid and Barrett ([&lt;reflink idref=&quot;bib44&quot; id=&quot;ref22&quot;&gt;44&lt;/reflink&gt;]) consider four facets in their discussion of direct and semi‐direct testing: practicality, face validity, reliability, and concurrent validity. In terms of practicality, semi‐direct tests offer notable benefits, such as scalability, reduced administrative demands, and improved accessibility for test‐takers in remote locations. Regarding face validity, mixed reactions have been found among test‐takers; some prefer the natural interaction of OPIs, while others feel more at ease with the lower‐pressure environment of computer‐based assessments (Thompson et al.,&#160;[&lt;reflink idref=&quot;bib54&quot; id=&quot;ref23&quot;&gt;54&lt;/reflink&gt;]). On reliability, semi‐direct tests can achieve high reliability in scoring due to standardized prompts and a reduced influence of interlocutor variability. However, there are concerns about raters potentially engaging in selective listening with semi‐direct tests, which could impact scoring consistency (Nakatsuhara,&#160;[&lt;reflink idref=&quot;bib39&quot; id=&quot;ref24&quot;&gt;39&lt;/reflink&gt;]). In terms of concurrent validity, numerous studies show strong correlations between scores from direct and semi‐direct tests, reinforcing their potential interchangeability (O&#39;Loughlin,&#160;[&lt;reflink idref=&quot;bib40&quot; id=&quot;ref25&quot;&gt;40&lt;/reflink&gt;]; Shohamy,&#160;[&lt;reflink idref=&quot;bib49&quot; id=&quot;ref26&quot;&gt;49&lt;/reflink&gt;]). Nevertheless, further research is needed to better understand score alignment, particularly in high‐stakes contexts, where both absolute and adjacent score agreements are crucial.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-7&quot;&gt;Direct testing&lt;/hd&gt; &lt;p&gt;Direct testing tasks require test‐takers to perform actual language use, closely mimicking real‐life activities. These tasks elicit specific language behaviors observable by the examiner. The ACTFL OPI is an example of a speaking test involving a conversation with a trained rater.&lt;/p&gt; &lt;p&gt;Advantages of direct testing include its greater authenticity, as these exams closely replicate real‐world language use, providing a more accurate measure of a test‐taker&#39;s practical language ability. Fulcher ([&lt;reflink idref=&quot;bib20&quot; id=&quot;ref27&quot;&gt;20&lt;/reflink&gt;]) supports this, noting that direct tests validly measure a learner&#39;s real‐world language use abilities, and Pusey and Butler ([&lt;reflink idref=&quot;bib45&quot; id=&quot;ref28&quot;&gt;45&lt;/reflink&gt;]) emphasize the enhanced ecological validity of such tests. Examples include oral interviews, role‐plays, and essay writing, which have increased face validity and positive washback, encouraging learners to engage in meaningful language use (Cheng &amp;amp; Curtis,&#160;[&lt;reflink idref=&quot;bib11&quot; id=&quot;ref29&quot;&gt;11&lt;/reflink&gt;]). This can lead to improved skills and motivation, with direct tests offering rich qualitative data on a learner&#39;s abilities (Brown &amp;amp; Abeywickrama,&#160;[&lt;reflink idref=&quot;bib8&quot; id=&quot;ref30&quot;&gt;8&lt;/reflink&gt;]). In considering the advantages of direct testing, Thompson et al. ([&lt;reflink idref=&quot;bib54&quot; id=&quot;ref31&quot;&gt;54&lt;/reflink&gt;]) found that while roughly 32% of students were rated higher on the ACTFL OPIc, over 71% of the students expressed a preference for the ACTFL OPI. This finding supports the notion of the positive washback from direct tests, especially in the areas of oral proficiency.&lt;/p&gt; &lt;p&gt;However, direct testing has its disadvantages. Scoring, particularly in speaking and writing exams, can be subjective, even with detailed rubrics, potentially affecting reliability (McNamara,&#160;[&lt;reflink idref=&quot;bib37&quot; id=&quot;ref32&quot;&gt;37&lt;/reflink&gt;]). McNamara highlights examiner bias and variability in scoring as significant issues. Furthermore, direct testing is resource‐intensive, requiring trained examiners, suitable environments, and considerable time for administration and scoring, making it less feasible for large‐scale assessments (Weir,&#160;[&lt;reflink idref=&quot;bib57&quot; id=&quot;ref33&quot;&gt;57&lt;/reflink&gt;]). Additionally, test‐taker anxiety or unfamiliarity with the test format can lead to variability in performance, affecting reliability. Examples of direct testing include the IELTS ([&lt;reflink idref=&quot;bib22&quot; id=&quot;ref34&quot;&gt;22&lt;/reflink&gt;].) speaking component, which involves a face‐to‐face interview, and the TOEFL ([&lt;reflink idref=&quot;bib55&quot; id=&quot;ref35&quot;&gt;55&lt;/reflink&gt;].) writing section, which requires essay production.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-8&quot;&gt;Semi‐direct testing&lt;/hd&gt; &lt;p&gt;Semi‐direct testing aims to combine the authenticity of direct testing with greater practicality and reliability. These tests, such as the ACTFL OPIc, simulate real‐life language use through prerecorded prompts. While not used in the ACTFL OPIc, some semi‐direct tests incorporate automated scoring systems to lessen the cost and time needed to score them. Semi‐direct tests are more practical and efficient for large‐scale administration, as noted by Ockey ([&lt;reflink idref=&quot;bib42&quot; id=&quot;ref36&quot;&gt;42&lt;/reflink&gt;]) because they are able to be scored later after the administration of the test. Automated scoring can be used with both direct and semi‐direct testing, and this can enhance consistency and reliability by minimizing human bias and variability (Xi,&#160;[&lt;reflink idref=&quot;bib59&quot; id=&quot;ref37&quot;&gt;59&lt;/reflink&gt;]). Another advantage to semi‐direct tests is that these tests are also more accessible, particularly in remote or resource‐limited settings, as they can be administered online (Chapelle &amp;amp; Douglas,&#160;[&lt;reflink idref=&quot;bib10&quot; id=&quot;ref38&quot;&gt;10&lt;/reflink&gt;]). Semi‐direct testing environments can also reduce test‐taker anxiety, potentially providing a more accurate representation of language abilities (Thompson et al.,&#160;[&lt;reflink idref=&quot;bib54&quot; id=&quot;ref39&quot;&gt;54&lt;/reflink&gt;]).&lt;/p&gt; &lt;p&gt;Despite these advantages, semi‐direct testing has limitations as well. The lack of interactive communication is a primary criticism, with Fulcher ([&lt;reflink idref=&quot;bib19&quot; id=&quot;ref40&quot;&gt;19&lt;/reflink&gt;]) arguing that the absence of a real interlocutor limits the assessment of interactive and pragmatic skills. Additionally, the controlled nature of semi‐direct tasks can impact the naturalness of language production (Bijani &amp;amp; Khabiri,&#160;[&lt;reflink idref=&quot;bib7&quot; id=&quot;ref41&quot;&gt;7&lt;/reflink&gt;]). Technical issues and digital literacy can also pose challenges (Isbell &amp;amp; Kremmel,&#160;[&lt;reflink idref=&quot;bib24&quot; id=&quot;ref42&quot;&gt;24&lt;/reflink&gt;]), and access to the necessary technology may be limited in some areas.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-9&quot;&gt;Direct versus semi‐direct testing&lt;/hd&gt; &lt;p&gt;Comparative studies between direct and semi‐direct testing provide valuable insights into their strengths and weaknesses. Kenyon and Tschirner ([&lt;reflink idref=&quot;bib29&quot; id=&quot;ref43&quot;&gt;29&lt;/reflink&gt;]) found a 90% agreement between direct (ACTFL OPI) and semi‐direct (German Speaking Test) assessments. Kiddle and Kormos ([&lt;reflink idref=&quot;bib30&quot; id=&quot;ref44&quot;&gt;30&lt;/reflink&gt;]) observed no significant rating differences between the two methods but noted that students perceived the direct test as fairer and less stressful. As previously mentioned, Thompson et al. ([&lt;reflink idref=&quot;bib54&quot; id=&quot;ref45&quot;&gt;54&lt;/reflink&gt;]) reported that over 71% of students preferred the ACTFL OPI over the ACTFL OPIc, despite many being rated higher on the latter.&lt;/p&gt; &lt;p&gt;Research indicates that while direct testing offers authenticity and detailed insights into language abilities, it is resource‐intensive and subjective. Semi‐direct testing provides practicality and reliability but may lack the interactive and authentic nature of direct tests (Nakatsuhara,&#160;[&lt;reflink idref=&quot;bib39&quot; id=&quot;ref46&quot;&gt;39&lt;/reflink&gt;]). As language education evolves, ongoing research and innovation in testing methods will be crucial to meet the diverse needs of learners and educators.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-10&quot;&gt;Gender bias in language testing&lt;/hd&gt; &lt;p&gt;Gender bias in language testing has been widely studied with varied results from no difference at all (O&#39;Loughlin,&#160;[&lt;reflink idref=&quot;bib41&quot; id=&quot;ref47&quot;&gt;41&lt;/reflink&gt;]) to some depending on test and item type (Espinosa &amp;amp; Gardeazabal,&#160;[&lt;reflink idref=&quot;bib18&quot; id=&quot;ref48&quot;&gt;18&lt;/reflink&gt;]; Masoumi &amp;amp; Sadeghi,&#160;[&lt;reflink idref=&quot;bib36&quot; id=&quot;ref49&quot;&gt;36&lt;/reflink&gt;]). Chubbuck et al. ([&lt;reflink idref=&quot;bib12&quot; id=&quot;ref50&quot;&gt;12&lt;/reflink&gt;]) observed that language assessments can inadvertently favor one gender due to content that resonates more with their experiences and interests. For instance, they found that passages about sports or certain science articles on the SAT disadvantaged female test‐takers because these topics were less familiar to them, despite increasing female participation in these fields. Arias et al. ([&lt;reflink idref=&quot;bib4&quot; id=&quot;ref51&quot;&gt;4&lt;/reflink&gt;]) found that with high‐stakes exams, stereotypes and risk aversion could contribute to differences in rating outcomes between genders. Several strategies have been proposed to counteract gender bias. One approach is selecting test content carefully to limit gender bias. Kunnan ([&lt;reflink idref=&quot;bib31&quot; id=&quot;ref52&quot;&gt;31&lt;/reflink&gt;]) emphasizes the importance of using diverse and balanced content that reflects a wide range of interests and experiences. Additionally, DIF analysis can identify items that favor one gender, allowing test developers to revise or eliminate biased items (Ozdemir &amp;amp; Alshamrani,&#160;[&lt;reflink idref=&quot;bib43&quot; id=&quot;ref53&quot;&gt;43&lt;/reflink&gt;]). Including multiple‐choice questions with scenarios equally familiar to all genders can also help minimize the influence of gender‐specific knowledge.&lt;/p&gt; &lt;p&gt;How might gender result in different ratings based on the test type between OPI and OPIc? Since the OPI allows the interviewer to explore topics of interest to the examinee based on the warm‐up portion of the interview, the content of the test might naturally move in directions that are comfortable to examinees. The OPIc has a database of questions that are selected based on examinee&#39;s interest from areas chosen during their initial questionnaire. However, since the test is rated holistically, it is impossible to study DIF from actual test rating data on the OPIc because each question is not rated individually. Comparing the two test types, however, could give some indication on whether further study of the OPIc test question database is warranted.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-11&quot;&gt;Age bias in language testing&lt;/hd&gt; &lt;p&gt;The literature investigating age bias for direct and indirect testing methods is sparse, though important to take under consideration as cognitive abilities, language acquisition processes, and life experiences vary across different age groups. Research suggests that language assessments often fail to account for these differences, potentially disadvantaging older or younger test‐takers (Educational Testing Service ETS,&#160;[&lt;reflink idref=&quot;bib16&quot; id=&quot;ref54&quot;&gt;16&lt;/reflink&gt;]). Young children and adolescents are still developing their language skills, while older adults may experience cognitive decline in areas such as working memory and processing speed (Craik &amp;amp; Bialystok,&#160;[&lt;reflink idref=&quot;bib13&quot; id=&quot;ref55&quot;&gt;13&lt;/reflink&gt;]). Fulcher ([&lt;reflink idref=&quot;bib19&quot; id=&quot;ref56&quot;&gt;19&lt;/reflink&gt;]) highlights that many language tests are designed with a one‐size‐fits‐all approach, often catering to the cognitive styles and experiences of younger adults, thus impacting older test‐takers.&lt;/p&gt; &lt;p&gt;To address age bias, test developers have explored several approaches. One method is designing age‐appropriate test content that considers the cognitive and experiential backgrounds of different age groups. Adaptive testing, which adjusts question difficulty based on performance, has shown promise in providing a more equitable assessment across ages. This method ensures that all test‐takers encounter questions matching their proficiency level, reducing bias likelihood.&lt;/p&gt; &lt;p&gt;For the OPIc and OPI, familiarity with technology has a potential impact on ratings in terms of the content that is presented to the examinee. The OPIc relies on examinees to respond to surveys that determine both the form level and topics that will be presented, while the OPI allows an interviewer to elicit this same information. The OPIc, thus, presupposes a familiarity with technology that could be correlated with age. Verifying the extent to which this could result in individual disparate outcomes, therefore is important.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-12&quot;&gt;Acquisition setting bias in language testing&lt;/hd&gt; &lt;p&gt;The setting in which a language is acquired, whether in a formal classroom, through immersion, or via informal learning, can strongly influence one&#39;s language proficiency. Isbell et al. ([&lt;reflink idref=&quot;bib25&quot; id=&quot;ref57&quot;&gt;25&lt;/reflink&gt;]) found that differences in OPIc ratings across languages mainly reflected learner backgrounds and instruction amounts rather than the languages themselves. They noted that differences in student ratings on the OPIc across various languages in the tertiary setting were more related to pre‐tertiary study, university courses, heritage status, motivation, and study abroad experiences than the language studied. This would seem to imply that how the language was learned greatly impacted their ratings, and the language being studied was secondary to that factor. However, Isbell et al. ([&lt;reflink idref=&quot;bib25&quot; id=&quot;ref58&quot;&gt;25&lt;/reflink&gt;]) primarily looked at students enrolled in university courses and did not include those that learned the language in informal contexts such as missionary service or living abroad with no formal schooling.&lt;/p&gt; &lt;p&gt;To mitigate acquisition setting bias, assessments should recognize and value diverse language learning experiences. One approach is including performance‐based assessments that evaluate practical language use in real‐world contexts, as suggested by Bachman and Palmer ([&lt;reflink idref=&quot;bib5&quot; id=&quot;ref59&quot;&gt;5&lt;/reflink&gt;]). These assessments can more accurately reflect the skills of individuals who have learned the language through immersion or informal settings. Additionally, incorporating a variety of question types and tasks catering to different learning backgrounds can help create a more balanced assessment. The ACTFL OPI&#39;s trained tester can adjust their questions to the examinee&#39;s background, while the ACTFL OPIc uses pre‐survey responses and self‐assessed levels to generate questions that may end up being less personalized than the ones from the trained rater in the ACTFL OPI.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-13&quot;&gt;The impact of self‐assessment in achieving comparable outcomes&lt;/hd&gt; &lt;p&gt;Self‐assessment allows test‐takers to evaluate their language abilities, but results can be influenced by preconceptions such as overconfidence, underconfidence, and cultural factors. Ross ([&lt;reflink idref=&quot;bib46&quot; id=&quot;ref60&quot;&gt;46&lt;/reflink&gt;]) found that self‐assessment accuracy varies widely depending on the skill assessed and the context. Li and Zhang ([&lt;reflink idref=&quot;bib32&quot; id=&quot;ref61&quot;&gt;32&lt;/reflink&gt;]) found a significant positive correlation between self‐assessment and language performance, indicating that individuals who confidently rated their skills tended to better understand their abilities. When learners are aware of their abilities, they can set realistic goals and track their progress much more effectively.&lt;/p&gt; &lt;p&gt;To address the role of self‐assessment in language testing, researchers have explored several strategies. One approach is calibrating self‐assessment tools, providing clear criteria and examples to guide evaluations (Brown et al.,&#160;[&lt;reflink idref=&quot;bib9&quot; id=&quot;ref62&quot;&gt;9&lt;/reflink&gt;]). Computer‐adaptive self‐assessment tests can improve accuracy, showing a good correlation between self‐assessment and speaking proficiency (Winke et al.,&#160;[&lt;reflink idref=&quot;bib58&quot; id=&quot;ref63&quot;&gt;58&lt;/reflink&gt;]). Incorporating peer and teacher assessments alongside self‐assessment can provide a more balanced and accurate picture of language proficiency (Li &amp;amp; Zhang,&#160;[&lt;reflink idref=&quot;bib32&quot; id=&quot;ref64&quot;&gt;32&lt;/reflink&gt;]).&lt;/p&gt; &lt;hd id=&quot;AN0186745416-14&quot;&gt;Integrative approaches to avoiding bias&lt;/hd&gt; &lt;p&gt;In addition to addressing specific types of bias, integrative approaches that consider multiple dimensions of bias simultaneously have been proposed. Ongoing research and development are essential to continually refine language assessments and eliminate biases. Modern psychometric techniques, such as item response theory (IRT) and computer‐adaptive testing (CAT), offer sophisticated methods for analyzing test data and identifying bias. Involving diverse groups of stakeholders, including test‐takers, educators, and psychometricians, in the test development process can provide valuable insights and help create more equitable assessments.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-15&quot;&gt;Evaluating bias in language testing&lt;/hd&gt; &lt;p&gt;Many statistical techniques have been used to evaluate bias in language testing including DIF (Ross &amp;amp; Okabe,&#160;[&lt;reflink idref=&quot;bib47&quot; id=&quot;ref65&quot;&gt;47&lt;/reflink&gt;]; Schaap,&#160;[&lt;reflink idref=&quot;bib48&quot; id=&quot;ref66&quot;&gt;48&lt;/reflink&gt;]); and IRT with Rasch Modeling (Elder,&#160;[&lt;reflink idref=&quot;bib17&quot; id=&quot;ref67&quot;&gt;17&lt;/reflink&gt;]); however, the data need to meet certain assumptions to use these techniques. Since the OPI and the OPIc are rated holistically, item‐level data do not exist, which precludes any DIF analysis. Rasch analysis is widely employed and would work well for this study as the test scores are ordinal rankings based on a proficiency scale. However, a Rasch analysis can be problematic if the data overfit the model (indicated by low fit statistics and meaning that it lacks some of the randomness in human performance). If the data overfit the model, then this can result in inflated reliability (Smith,&#160;[&lt;reflink idref=&quot;bib50&quot; id=&quot;ref68&quot;&gt;50&lt;/reflink&gt;]), distort parameter estimates (Linacre,&#160;[&lt;reflink idref=&quot;bib33&quot; id=&quot;ref69&quot;&gt;33&lt;/reflink&gt;]) and, most importantly for this study, impact diagnostic value and test fairness (Eckes,&#160;[&lt;reflink idref=&quot;bib15&quot; id=&quot;ref70&quot;&gt;15&lt;/reflink&gt;]). For instances in which data overfit the Rasch model, aligned rank transform (ART) analysis can be used to analyze non‐normally distributed data in which traditional ANOVA assumptions are violated (Mansouri,&#160;[&lt;reflink idref=&quot;bib35&quot; id=&quot;ref71&quot;&gt;35&lt;/reflink&gt;]).&lt;/p&gt; &lt;hd id=&quot;AN0186745416-16&quot;&gt;PROCEDURES&lt;/hd&gt; &lt;p&gt;The organization of this study involved three main steps. First, prospective participants (Spanish students) completed a background survey collecting demographic information potentially leading to test rating bias. Second, the 154 students who were selected were randomly divided into two groups for a counterbalanced design to mitigate test order effects: Group 1 (&lt;emph&gt;n&lt;/emph&gt; = 77): ACTFL OPI first and Group 2 (&lt;emph&gt;n&lt;/emph&gt; = 77): ACTFL OPIc first. Finally, an ART analysis was performed.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-17&quot;&gt;Participants&lt;/hd&gt; &lt;p&gt;Our sample consisted of 154 students (81 females and 73 males, mean age = 23.4, SD = 5.29) at a large private university in the western United States (see Figure&#160;1). The population for this study was relatively homogenous in regard to their age. Our consideration of age as an area of potential bias is quite limited since 88% of the participants reported being between 17 and 25 years old. While there can exist considerable variation in the types of life experiences had during these years including time abroad, language study, language classes, etc., this study&#39;s consideration of bias based on age is limited due to the small sample size (&lt;emph&gt;n&lt;/emph&gt; = 16) of students who were 26+ years of age. Data on race and ethnicity were not gathered for this study, but given the composition of students at the university where the data were collected, the vast majority of the students were white.&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0001.jpg&quot; title=&quot;1 Participants by age and gender.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;The study included a range of participants, from freshmen to graduate students, with varying levels of Spanish language experience (average # of university courses = 5.35, SD = 4.45, median = 4 courses). Some had no formal language courses, while one had taken at least 20 language courses (see Figure&#160;2). The number of courses was self‐reported and specifically referred to the number of university‐level courses that each participant had taken though some respondents may have included courses they had tested out of via challenge exam. It is likely that the participants in this study had taken even more courses if classes taken in primary and secondary school had been included in the questionnaire. Participants included Spanish majors, minors, and students from other majors in the university&#39;s language certificate program. Lower‐division students mostly fulfilled General Education requirements, while most upper‐division students had significant practical experience through living or studying abroad.&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0002.jpg&quot; title=&quot;2 Participants by number of formal university Spanish courses.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;Many advanced speakers in this study had spent 18‐24 months as proselyting missionaries in the U.S. and abroad (see Figure&#160;3). This Spanish‐speaking missionary service was a significant factor in their language‐learning background due to its intense, immersive nature. Missionaries for the Church of Jesus Christ of Latter‐day Saints receive 3 to 6 weeks of intensive formal language training at the Missionary Training Center before being sent to their assigned areas. In the Missionary Training Center, they are immersed in the native language and culture in which they will be proselyting. However, the immersive aspect of the mission experience may vary depending on the service location. Those serving in English‐dominant countries, who spend their time proselyting to Spanish speakers in those areas, might have less intense language immersion due to the heavy usage of English around them. To account for this variation, the study examined service location as a dependent variable, recognizing its potential impact on language proficiency.&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0003.jpg&quot; title=&quot;3 Participants by missionary Spanish language experience.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;Participation in the study was voluntary but incentivized. Researchers covered all testing costs through a grant and upper‐division students could use their higher rating for program and professional requirements. All participants received a nationally recognized proficiency certificate. These incentives resulted in high participation rates, with few students opting out of either the ACTFL OPI or OPIc exams.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-21&quot;&gt;Instruments&lt;/hd&gt; &lt;p&gt;The instruments used in this study were (a) a pre‐survey, (b) the telephonic version of the ACTFL OPI, (c) the ACTFL OPIc, and (d) a post‐survey. Each of the instruments is described in detail in the following sections.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-22&quot;&gt;Pre‐survey&lt;/hd&gt; &lt;p&gt;The pre‐survey gathered demographic data and language experience information from participants, including the number of years of formal K‐12 experience and the number of university‐level courses of Spanish and Spanish‐speaking missionary service and location. Participants also used a 5‐point Likert scale (1=poor, 5=excellent) to self‐assess their Spanish abilities in eight areas: speaking, listening, grammar knowledge, pronunciation, vocabulary, writing ability, reading ability, and cultural knowledge&lt;/p&gt; &lt;p&gt;The survey aimed to collect data on students&#39; language goals and familiarity with the ACTFL OPI and OPIc exams. A PDF of the complete survey is available in an OSF Repository. https://osf.io/9zkuf/?view_only=b5ad9ffcc2e143a38087256ef9b914d2.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-23&quot;&gt;ACTFL OPI and ACTFL OPIc&lt;/hd&gt; &lt;p&gt;The ACTFL OPI and OPIc are language proficiency assessment tools with similar structures but distinct delivery methods. The OPI is a 15–30‐min recorded interview conducted by a human interviewer, following a four‐stage protocol: warm‐up, floor‐ceiling establishment, role‐play, and wind‐down. The interviewer assesses the examinee&#39;s language abilities across various topics, determining their proficiency level.&lt;/p&gt; &lt;p&gt;The OPIc, a technology‐mediated version, replaces the warm‐up with demographic questions and relies on examinee self‐assessment to determine the test form. Before beginning the OPIc, test takers complete a background survey that allows them to choose topics of interest to them (e.g., entertainment, hobbies, sports, etc.) (ACTFL,&#160;[&lt;reflink idref=&quot;bib1&quot; id=&quot;ref72&quot;&gt;1&lt;/reflink&gt;]). Additionally, examinees are asked to complete a self‐assessment of their own language abilities from five different descriptions derived from the NCSSFL/ACTFL Can‐do statements (ACTFL,&#160;[&lt;reflink idref=&quot;bib2&quot; id=&quot;ref73&quot;&gt;2&lt;/reflink&gt;]). Based on test‐takers responses to the survey and self‐assessment, an individualized test is created that selects items determined in part by the responses in the interest survey as well as on the examinee&#39;s self‐reported proficiency level.&lt;/p&gt; &lt;p&gt;Both tests involve human raters evaluating the examinee&#39;s responses, with a third rater arbitrating in case of disagreement. The OPIc&#39;s reliance on self‐assessment can potentially lead some learners to inaccurately rate themselves if they misjudge their abilities. If the examinee is not successful in their self‐assessment, they will receive a test form that does not match their actual language abilities, and this may keep them from receiving the rating that best reflects their ability in the language.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-24&quot;&gt;Post‐survey&lt;/hd&gt; &lt;p&gt;The post‐survey gathered data on participants&#39; experiences and attitudes toward both test methods. Before receiving final ratings, participants predicted their outcomes based on ACTFL Proficiency Guidelines. They expressed test preference (ACTFL OPI vs. OPIc) using a 9‐point Likert scale from &quot;Dislike extremely&quot; to &quot;Like extremely.&quot; An open‐ended question asked them to explain their preference. The survey aimed to understand participants&#39; perceptions of each test method and their self‐assessment accuracy. Detailed analysis of the post‐survey results can be found in Authors (2016).&lt;/p&gt; &lt;hd id=&quot;AN0186745416-25&quot;&gt;DATA ANALYSIS&lt;/hd&gt; &lt;p&gt;The data were initially going to be analyzed with many‐facet Rasch model (MFRM) using FACETS software, however the data did not fit the model as only 35 of the 154 participants had fit statistics between 0.5 and 1.5. The MFRM control file data and analysis can be found at OSF Repository https://osf.io/9zkuf/?view_only=b5ad9ffcc2e143a38087256ef9b914d2.&lt;/p&gt; &lt;p&gt;Thus, we used the ART analysis to run a mixed model as OPI and OPIc ratings are ordinal in nature. We used the open‐source software package Jamovi (Version 2.3.28.0), and the data analysis can be found at https://osf.io/9zkuf/?view_only=b5ad9ffcc2e143a38087256ef9b914d2.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-26&quot;&gt;RESULTS&lt;/hd&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-27&quot;&gt;ACTFL OPIc and ACTFL OPI differences&lt;/hd&gt; &lt;p&gt;Comparing ACTFL OPI and OPIc ratings for examinees who took both tests (see Table&#160;1):&lt;/p&gt; &lt;olist&gt; &lt;item&gt; TABLE Contingency table.&lt;/item&gt; &lt;/olist&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th&amp;gt;ACTFL OPIc ratings&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th&amp;gt;NM&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;NH&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;IL&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;IM&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;IH&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;AL&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;AM&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;AH&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;S&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;Total&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;ACTFL OPI ratings&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;NM&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;NH&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;&amp;lt;ext-link href=&quot;&amp;amp;#42;&quot; /&amp;gt; 1&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;IL&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;10&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;&amp;lt;ext-link href=&quot;&amp;amp;#42;&quot; /&amp;gt; 1&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;15&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;IM&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;15&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;22&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;IH&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;16&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;&amp;lt;ext-link href=&quot;&amp;amp;#42;&quot; /&amp;gt; 1&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;24&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;AL&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;23&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;23&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;47&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;AM&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;7&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;21&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;30&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;AH&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;&amp;lt;ext-link href=&quot;&amp;amp;#42;&quot; /&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;9&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;S&amp;lt;/td&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td /&amp;gt;&amp;lt;td&amp;gt;Total&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;30&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;24&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;35&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;50&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;154&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt;1 * Ratings were two sublevels different&lt;/p&gt; &lt;ulist&gt; &lt;item&gt;54.5% (&lt;reflink idref=&quot;bib84&quot; id=&quot;ref74&quot;&gt;84&lt;/reflink&gt;) had identical ratings&lt;/item&gt; &lt;item&gt;29.9% (&lt;reflink idref=&quot;bib46&quot; id=&quot;ref75&quot;&gt;46&lt;/reflink&gt;) were rated one sublevel higher on OPIc&lt;/item&gt; &lt;item&gt;12.3% (&lt;reflink idref=&quot;bib19&quot; id=&quot;ref76&quot;&gt;19&lt;/reflink&gt;) were rated one sublevel lower on OPIc&lt;/item&gt; &lt;item&gt;96% (&lt;reflink idref=&quot;bib149&quot; id=&quot;ref77&quot;&gt;149&lt;/reflink&gt;) were within one sublevel difference&lt;/item&gt; &lt;/ulist&gt; &lt;p&gt;Five examinees had a two‐sublevel difference&lt;/p&gt; &lt;p&gt;To assess the relationship between OPIc and OPI ratings, Kendall&#39;s Tau (τ) was calculated. This non‐parametric measure evaluates the strength and direction of association between two variables. Results showed a Kendall&#39;s Tau value of τ = 0.79 (&lt;emph&gt;t&lt;/emph&gt; = 12.2, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001), indicating a strong positive correlation between ACTFL OPIc and OPI ratings. This suggests that students performing well on one test tend to perform similarly on the other. These findings demonstrate high consistency between the two test formats, with most examinees receiving very similar ratings across both assessments.&lt;/p&gt; &lt;p&gt;Examinees were rated slightly higher on the ACTFL OPIc (mean = 6.71, SD = 1.41) compared to the OPI (mean = 6.53, SD = 1.57). This difference was statistically significant [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref78&quot;&gt;1&lt;/reflink&gt;, 152) = 10.44, &lt;emph&gt;p&lt;/emph&gt; = .002, partial eta = 0.064] with a medium effect size. The most significant rating discrepancies occurred among examinees in the OPI&#39;s lower sublevels, particularly those rated Intermediate or Advanced Low on the OPI but Mid sublevel on the OPIc.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-28&quot;&gt;Research Question 1: Gender &amp;amp; age&lt;/hd&gt; &lt;p&gt;Tests should be impartial, with personal characteristics like age and gender not affecting ratings. There were no significant interactions found between gender and test type, and age and test type.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-29&quot;&gt;Gender and test type&lt;/hd&gt; &lt;p&gt;An ART mixed‐model analysis was conducted to assess the effects of gender (male vs. female), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures.&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of gender, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref79&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref80&quot;&gt;150&lt;/reflink&gt;) = 39.40, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001], with males having higher ratings on their tests than females (see Table&#160;2 and Figure&#160;4). A significant main effect of test type was also found, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref81&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref82&quot;&gt;150&lt;/reflink&gt;) = 10.59, &lt;emph&gt;p&lt;/emph&gt; = .003] indicating that participants had higher ratings on the OPIc compared to the OPI. However, the main effect of test order was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref83&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref84&quot;&gt;150&lt;/reflink&gt;) = 2.28, &lt;emph&gt;p&lt;/emph&gt; =. 133]. Interactions between gender and test type (&lt;emph&gt;F&#160;&lt;/emph&gt;(&lt;reflink idref=&quot;bib1&quot; id=&quot;ref85&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref86&quot;&gt;150&lt;/reflink&gt;) = 1.16, &lt;emph&gt;p&lt;/emph&gt; = .283) were not significant.&lt;/p&gt; &lt;p&gt;2 TABLE Estimated marginal means of gender by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;Gender&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Male&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.99&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.282&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.43&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.54&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Female&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.62&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.271&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;276&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.09&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.16&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Male&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.11&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.282&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.55&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.67&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Female&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.86&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.278&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;271&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.31&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.41&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0004.jpg&quot; title=&quot;4 Data visualization of estimated marginal means of gender by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-31&quot;&gt;Age and test type&lt;/hd&gt; &lt;p&gt;An ART mixed‐model analysis was conducted to assess the effects of age (17–19, 20–21, 22–23, 24–25, 26+ and Not Reported (NR)), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures.&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of Age, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib5&quot; id=&quot;ref87&quot;&gt;5&lt;/reflink&gt;,&lt;reflink idref=&quot;bib142&quot; id=&quot;ref88&quot;&gt;142&lt;/reflink&gt;) = 11.58, &lt;emph&gt;p &lt;/emph&gt;&amp;lt; .001], with the younger groups having lower ratings on their tests than the higher groups (see Table&#160;3 and Figure&#160;5). A significant main effect of test type was also found, &lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref89&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib148&quot; id=&quot;ref90&quot;&gt;148&lt;/reflink&gt;) = 13.49, indicating that participants had higher ratings on the OPIc compared to the OPI. However, the main effect of test order was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref91&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib147&quot; id=&quot;ref92&quot;&gt;147&lt;/reflink&gt;) = 3.47, &lt;emph&gt;p&lt;/emph&gt; = .065. Interactions between age and test type [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib5&quot; id=&quot;ref93&quot;&gt;5&lt;/reflink&gt;,&lt;reflink idref=&quot;bib148&quot; id=&quot;ref94&quot;&gt;148&lt;/reflink&gt;) = 1.01, &lt;emph&gt;p&lt;/emph&gt; = .415) were not significant.&lt;/p&gt; &lt;p&gt;3 TABLE Estimated marginal means of age by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;AgeCode&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;17&amp;amp;#8211;19&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.03&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.48&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.57&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;20&amp;amp;#8211;21&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.75&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.298&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.16&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.33&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;22&amp;amp;#8211;23&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.69&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.196&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.07&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;24&amp;amp;#8211;25&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.24&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.185&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.87&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.6&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;26+&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.75&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.324&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.11&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.39&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;NR&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.09&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.651&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.81&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.38&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;17&amp;amp;#8211;19&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.21&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.66&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.76&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;20&amp;amp;#8211;21&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.90&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.298&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.32&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.49&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;22&amp;amp;#8211;23&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.80&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.196&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.41&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.19&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;24&amp;amp;#8211;25&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.38&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.185&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.01&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.74&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;26+&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.19&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.324&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.55&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.83&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;NR&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.84&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.651&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;171&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.56&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;9.13&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0005.jpg&quot; title=&quot;5 Data visualization of estimated marginal means of age by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-33&quot;&gt;Research Question 2: Language learning experience – formal study and missionary service locat...&lt;/hd&gt; &lt;p&gt;In terms of measuring oral proficiency, it should not matter how the language was learned, or whether it was acquired in a more formal environment, informally through an immersion experience, or a combination of both. There were no significant interactions found between formal study and test type, and missionary service location and test type.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-34&quot;&gt;Formal study and test type&lt;/hd&gt; &lt;p&gt;An ART mixed‐model analysis was conducted to assess the effects of formal study as determined by number of classes (0, 1, 2, 3, 4 to 5, 6 to 7, 8 to 9, 10 to 11, 12+), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures.&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of number of Spanish courses, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib8&quot; id=&quot;ref95&quot;&gt;8&lt;/reflink&gt;,&lt;reflink idref=&quot;bib144&quot; id=&quot;ref96&quot;&gt;144&lt;/reflink&gt;) = 6.34, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001], with those that have had more classes with higher ratings on their tests than those without (see Table&#160;4 and Figure&#160;6). A significant main effect of test type was also found, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref97&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib145&quot; id=&quot;ref98&quot;&gt;145&lt;/reflink&gt;) = 10.97, indicating that participants had higher ratings on the OPIc compared to the OPI. However, the main effect of test order was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref99&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib144&quot; id=&quot;ref100&quot;&gt;144&lt;/reflink&gt;) = 1.83, &lt;emph&gt;p&lt;/emph&gt; = .178. Interactions between number of Spanish courses and test type [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib8&quot; id=&quot;ref101&quot;&gt;8&lt;/reflink&gt;,&lt;reflink idref=&quot;bib144&quot; id=&quot;ref102&quot;&gt;144&lt;/reflink&gt;) = 0.548, &lt;emph&gt;p&lt;/emph&gt; = .819) were not significant.&lt;/p&gt; &lt;p&gt;4 TABLE Estimated marginal means of formal courses by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;Spanish&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;Courses&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;0&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.92&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.401&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.13&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.71&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.17&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.62&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.72&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.53&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.421&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.7&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.36&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.64&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.243&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.16&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.12&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;amp;#8211;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.46&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.334&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.81&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.12&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;6&amp;amp;#8211;7&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.08&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.273&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.54&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.62&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;8&amp;amp;#8211;9&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.37&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.47&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.93&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;10&amp;amp;#8211;11&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.26&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.401&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.47&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.05&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;12+&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.48&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.333&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.83&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.14&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;0&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.38&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.401&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.58&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.17&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.43&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.277&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.88&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.98&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.83&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.421&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.66&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.71&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.243&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.23&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.19&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;amp;#8211;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.71&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.334&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.06&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.37&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;6&amp;amp;#8211;7&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.29&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.273&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.75&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.82&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;8&amp;amp;#8211;9&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.37&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.47&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.93&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;10&amp;amp;#8211;11&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.53&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.401&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.74&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.32&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;12+&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.54&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.333&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;167&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.89&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.2&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0006.jpg&quot; title=&quot;6 Data visualization of estimated marginal means of formal courses by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-36&quot;&gt;Missionary service location and test type&lt;/hd&gt; &lt;p&gt;An ART mixed‐model analysis was conducted to assess the effects of missionary service location (none, English dominant environment, and Spanish dominant environment), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures.&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of missionary service location, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib2&quot; id=&quot;ref103&quot;&gt;2&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref104&quot;&gt;150&lt;/reflink&gt;) = 54.27, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001], with those that have had an immersion experience with higher ratings on their tests than those without (see Table&#160;5 and Figure&#160;7). A significant main effect of test type was also found, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref105&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib151&quot; id=&quot;ref106&quot;&gt;151&lt;/reflink&gt;) = 4.84, indicating that participants had higher ratings on the OPIc compared to the OPI. However, the main effect of test order was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref107&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref108&quot;&gt;150&lt;/reflink&gt;) = 3.47, &lt;emph&gt;p&lt;/emph&gt; = .586. Interactions between missionary service location and test type [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib2&quot; id=&quot;ref109&quot;&gt;2&lt;/reflink&gt;,&lt;reflink idref=&quot;bib151&quot; id=&quot;ref110&quot;&gt;151&lt;/reflink&gt;) = 1.68, &lt;emph&gt;p&lt;/emph&gt; = .190) were not significant.&lt;/p&gt; &lt;p&gt;5 TABLE Estimated marginal means of missionary service location by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;Mission service location&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;None&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.52&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.136&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.25&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.79&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;English Dom&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.00&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.261&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.49&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.51&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Spanish Dom&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.61&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.151&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.31&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.91&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;None&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.82&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.136&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.55&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.08&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;English Dom&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.05&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.261&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.54&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.56&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;Spanish Dom&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.71&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.151&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;181&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.41&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.01&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0007.jpg&quot; title=&quot;7 Data visualization of estimated marginal means of missionary service location by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-38&quot;&gt;Research Question 3: Self‐perceived ability – ACTFL OPIc test form and self‐assessment&lt;/hd&gt; &lt;p&gt;Before evaluating what, if any, bias existed, we examined the reliability of the Self‐Report variable based on the examinee&#39;s responses to the pre‐survey questions to determine if we could use it as a grouping variable. We then examined if there was any difference in ratings between the ACTFL OPIc and OPI based on someone&#39;s perception of their ability as measured by two different variables: ACTFL OPIc test form and self‐assessment.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-39&quot;&gt;Reliability of self‐report&lt;/hd&gt; &lt;p&gt;As noted in the instrument&#39;s section, there were eight statements related to speaking, listening, grammar knowledge, pronunciation, vocabulary, writing ability, reading ability, and cultural knowledge of a 5‐point scale to which the examinees responded to self‐assess their language ability on the pre‐survey self‐assessment questionnaire. The self‐assessment rating scale was found to be reliable Cronbach alpha = 0.92 (see Figure&#160;8).&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0008.jpg&quot; title=&quot;8 Correlation heatmap of self‐assessment statements.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-41&quot;&gt;Self‐assessed ability and ACTFL OPIc test form&lt;/hd&gt; &lt;p&gt;In terms of measuring oral proficiency, it should not matter how an examinee perceives their language ability. There were no significant interactions found between self‐assessed ability and test type; and OPIc test form and test type.&lt;/p&gt; &lt;p&gt; &lt;bold&gt;Self‐assessed ability and test type.&lt;/bold&gt; An ART mixed‐model analysis was conducted to assess the effects of self‐assessed ability on a 5‐point rating scale (Poor, Fair, Good, Very Good, Excellent, No Response(NR)), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures.&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of self‐assessed ability, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib5&quot; id=&quot;ref111&quot;&gt;5&lt;/reflink&gt;,&lt;reflink idref=&quot;bib148&quot; id=&quot;ref112&quot;&gt;148&lt;/reflink&gt;) = 17.89, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001], with those that rated themselves higher with higher ratings on their tests than those that rated themselves lower (see Table&#160;6). However, the main effect of test type was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref113&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib148&quot; id=&quot;ref114&quot;&gt;148&lt;/reflink&gt;) = 3.47, &lt;emph&gt;p&lt;/emph&gt; = .065], nor was the main effect of test order [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref115&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib147&quot; id=&quot;ref116&quot;&gt;147&lt;/reflink&gt;) = 1.92, &lt;emph&gt;p&lt;/emph&gt; = .168. Interactions between the number of self‐assessed ability and test type [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib5&quot; id=&quot;ref117&quot;&gt;5&lt;/reflink&gt;,&lt;reflink idref=&quot;bib148&quot; id=&quot;ref118&quot;&gt;148&lt;/reflink&gt;) = 1.59, &lt;emph&gt;p&lt;/emph&gt; = .166) were not significant.&lt;/p&gt; &lt;p&gt;6 TABLE Estimated marginal means of self‐assessed ability by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;SelfRating&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.13&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;1.218&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2.73&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.54&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.33&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.296&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;3.75&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.92&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.24&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.158&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.92&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.55&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.15&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.157&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.84&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.46&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.92&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.351&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;175&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.22&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.61&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;NR&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.07&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.609&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.86&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.27&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.13&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;1.218&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;2.73&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.54&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.80&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.296&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.22&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.39&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.47&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.158&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.16&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.78&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.20&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.157&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.89&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.51&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.00&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.351&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;175&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.31&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.69&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;NR&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.82&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.609&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;174&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.61&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;9.02&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;bold&gt;Self‐assessed ability and test type&lt;/bold&gt;. An ART mixed‐model analysis was conducted to assess the effects of OPIc Test Form (Form 2, Form 3, Form 4 or Form 5), test order (OPI first vs. OPIc first), and test type (OPI vs. OPIc) on ordinal ratings of test scores, with participant ID included as a random effect to account for repeated measures (Figure&#160;9).&lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0009.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0009.jpg&quot; title=&quot;9 Data visualization of estimated marginal means of self‐assessed ability by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;The fixed effects model showed significant main effects of OPIc Test Form, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib3&quot; id=&quot;ref119&quot;&gt;3&lt;/reflink&gt;,&lt;reflink idref=&quot;bib149&quot; id=&quot;ref120&quot;&gt;149&lt;/reflink&gt;) = 71.24, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001], with those that had higher test forms getting higher ratings on their tests than those with lower ones (see Table&#160;7 and Figure&#160;10). However, the main effect of test type was not significant, [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref121&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref122&quot;&gt;150&lt;/reflink&gt;) = 0.40, &lt;emph&gt;p&lt;/emph&gt; = .528], nor was the main effect of test order [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib1&quot; id=&quot;ref123&quot;&gt;1&lt;/reflink&gt;,&lt;reflink idref=&quot;bib149&quot; id=&quot;ref124&quot;&gt;149&lt;/reflink&gt;) = 3.23, &lt;emph&gt;p&lt;/emph&gt; = .074. Interactions between the number of OPIc test form and test type [&lt;emph&gt;F&lt;/emph&gt; (&lt;reflink idref=&quot;bib3&quot; id=&quot;ref125&quot;&gt;3&lt;/reflink&gt;,&lt;reflink idref=&quot;bib150&quot; id=&quot;ref126&quot;&gt;150&lt;/reflink&gt;) = 0.79, &lt;emph&gt;p&lt;/emph&gt; = .500) were not significant.&lt;/p&gt; &lt;p&gt;7 TABLE Estimated marginal means of OPIc test form by test type.&lt;/p&gt; &lt;p&gt; &lt;ephtml&gt; &amp;lt;table&amp;gt;&amp;lt;thead valign=&quot;bottom&quot;&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;OPIc&amp;lt;/th&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th /&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;95% Confidence interval&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr valign=&quot;bottom&quot;&amp;gt;&amp;lt;th&amp;gt;TestForm&amp;lt;/th&amp;gt;&amp;lt;th&amp;gt;TestType&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Mean&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;SE&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;df&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Lower&amp;lt;/th&amp;gt;&amp;lt;th align=&quot;left&quot;&amp;gt;Upper&amp;lt;/th&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/thead&amp;gt;&amp;lt;tbody valign=&quot;top&quot;&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.95&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.580&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;3.81&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.10&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.01&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.145&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.72&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.29&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.06&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.114&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.84&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.29&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPI&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.95&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.201&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.56&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.35&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;2&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.62&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.580&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;3.48&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.77&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;3&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.28&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.145&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;4.99&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;5.56&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;4&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.22&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.114&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;6.99&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.44&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;5&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;OPIc&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.15&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;0.201&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;192&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;7.76&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;8.55&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&amp;lt;/tbody&amp;gt;&amp;lt;/table&amp;gt; &lt;/ephtml&gt; &lt;/p&gt; &lt;p&gt; &lt;img src=&quot;https://imageserver.ebscohost.com/img/embimages/rdk/FLA/01jul25/flan12804-fig-0010.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS&quot; alt=&quot;flan12804-fig-0010.jpg&quot; title=&quot;10 Data visualization of estimated marginal means of OPIc test form by test type.&quot; /&gt; &lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;hd id=&quot;AN0186745416-44&quot;&gt;DISCUSSION&lt;/hd&gt; &lt;p&gt;The findings from this study support the first part of the argument claim (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref127&quot;&gt;56&lt;/reflink&gt;]) framework that the ACTFL OPIc is a valid and reliable alternative to the ACTFL OPI for assessing oral proficiency in Spanish learners. The grounds (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref128&quot;&gt;56&lt;/reflink&gt;]) with which make this claim is that the data demonstrate a strong correlation between ACTFL OPI and ACTFL OPIc ratings (τ = 0.79, &lt;emph&gt;p&lt;/emph&gt; &amp;lt; .001), with 96% of ratings falling within one sublevel of each other based on the warrant that both tests are grounded in the ACTFL proficiency guidelines. These results align with prior research that has established the comparability and reliability of OPI and OPIc assessments.&lt;/p&gt; &lt;p&gt;The consistency between these tests can be attributed to warrant (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref129&quot;&gt;56&lt;/reflink&gt;]) of their shared foundation in the ACTFL Proficiency Guidelines, which ensures alignment in the constructs being measured. The high agreement rates between test ratings further indicate that the OPIc reliably replicates the outcomes of the OPI. Importantly, the lack of significant bias in the results supports the fairness of using either test across diverse populations. This validates the OPIc as a credible option for language educators and institutions seeking flexibility without sacrificing assessment integrity.&lt;/p&gt; &lt;p&gt;For the second part of the claim (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref130&quot;&gt;56&lt;/reflink&gt;]) that there are comparable rating outcomes with minimal bias, we address the results in relation to the research questions posed in the study as the backing (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref131&quot;&gt;56&lt;/reflink&gt;]) and suggest ideas for future research.&lt;/p&gt; &lt;p&gt;First, controlling for the gender and age of participants, no statistical or practical differences were found between ACTFL OPI and OPIc rating outcomes. In our sample, females appeared to perform slightly better than males on the ACTFL OPIc, but this difference was not statistically significant and could be attributed to general variation among participants. Some participants over the age of 35 were outliers compared to the younger majority, but they only exhibited slightly larger standard errors. Future research should include a broader age range to determine if age could be a contributing factor, given the average participant age of 23.4 years. The homogeneity of this population means age cannot be ruled out as a bias, though it was not identified as one in our study.&lt;/p&gt; &lt;p&gt;Second, using ART, we found that the choice of test was not significantly impacted by years of formal Spanish language study or participation in a Spanish immersion experience. Participants with varying levels of language training and experience were neither advantaged nor disadvantaged by either testing modality. This suggests that selecting between the ACTFL OPI and OPIc does not inherently benefit or harm someone based on their formal language training. However, our study did not include enough heritage learners of Spanish to determine the impact of being raised in a Spanish‐speaking home on performance. Future research should specifically consider heritage learners, as Isbell et al. ([&lt;reflink idref=&quot;bib25&quot; id=&quot;ref132&quot;&gt;25&lt;/reflink&gt;]) identified heritage status as an important factor in language proficiency.&lt;/p&gt; &lt;p&gt;Third, ACTFL OPIc test‐takers are rated based on the test form assigned after a self‐assessment. An inaccurate self‐assessment can result in a test form that is too difficult or too easy, potentially leading to inaccurate ratings (Li &amp;amp; Zhang,&#160;[&lt;reflink idref=&quot;bib32&quot; id=&quot;ref133&quot;&gt;32&lt;/reflink&gt;]). In our sample, test form was statistically unrelated to students&#39; self‐assessments, although some test‐takers inaccurately assessed their ability, either above or below their actual level. Overassessment occurred approximately 6% of the time (8 out of 154 participants), invalidating individual ratings but not significantly impacting mean outcomes. It is difficult to determine underassessment due to the fact that participants who underassessed would simply receive the highest possible rating for the exam they had chosen without knowing that they could have been rated higher had they not underassessed. However, given that more students were rated higher on the OPIc than the OPI, this is likely to be a relatively small portion, especially since the chosen form on the OPIc offers a wide range of ratings.&lt;/p&gt; &lt;p&gt;An examinee&#39;s perception of their language ability might influence their test ratings. ACTFL OPIc test forms are designed with rating ranges based on the number of questions at each major level. Form 3 targets speakers from Intermediate Mid to Advanced Low, but ratings can range from Intermediate Low to Advanced Low. Form 4 targets Advanced Low and Advanced Mid, though ratings from Intermediate High to Advanced High may be assigned. Contrast that with Form 5, which is designed for performance at the Advanced High to Superior, but ratings can range from Advanced Low to Superior. If a test‐taker&#39;s performance does not meet the lowest level, a rating of &quot;BR&quot; (below rating) can be awarded. Unlike the ACTFL OPI, which is conducted live and allows interviewers to adjust question levels real‐time based on performance, the OPIc relies on self‐assessment for test form selection. Institutions may preselect test forms for specific groups—for example, requiring teacher candidates to take Form 4 to demonstrate their eligibility to teach. However, this approach may disadvantage individual test‐takers who are unaware of how test form selection affects their rating outcomes. Future research should explore the relationship between ACTFL OPIc test forms, self‐assessment, and outcomes, particularly how training students to assess their proficiency levels could improve test form selection and evaluation accuracy.&lt;/p&gt; &lt;p&gt;The results in this study further support the findings of Surface et al. ([&lt;reflink idref=&quot;bib52&quot; id=&quot;ref134&quot;&gt;52&lt;/reflink&gt;]), SWA Consulting ([&lt;reflink idref=&quot;bib53&quot; id=&quot;ref135&quot;&gt;53&lt;/reflink&gt;]), and Thompson et al. ([&lt;reflink idref=&quot;bib54&quot; id=&quot;ref136&quot;&gt;54&lt;/reflink&gt;]), which indicated no significant difference between participants&#39; ratings on the ACTFL OPI and OPIc even though some differences were found especially at the Advanced Low and Advanced Mid‐levels. The study supports using either test to determine proficiency, given that gender, age, and language learning environment do not produce systematic or statistically significant differences. However, most participants in this study were rated at the Advanced level, so future research should examine how less proficient students perform on both exams to see if these variables affect lower‐level students.&lt;/p&gt; &lt;p&gt;Nevertheless, as qualifiers (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref137&quot;&gt;56&lt;/reflink&gt;]), it is important to note the study&#39;s findings should be understood within the context of its limitations. The data primarily reflect the experiences of Spanish learners at a U.S. university in which 38.9% had served a mission in which Spanish was the dominant language, which may limit generalizability to other languages, proficiency levels, or test‐taker demographics. Future research should explore the applicability of these findings to heritage learners, as well as the impact of self‐assessment accuracy on OPIc outcomes.&lt;/p&gt; &lt;p&gt;Potential counterarguments or rebuttals (Toulmin,&#160;[&lt;reflink idref=&quot;bib56&quot; id=&quot;ref138&quot;&gt;56&lt;/reflink&gt;]), such as concerns about self‐assessment accuracy and the lack of interactive elements in the ACTFL OPIc, warrant consideration. While self‐assessment inaccuracies may influence the ACTFL OPIc&#39;s outcomes, the overall impact on the test&#39;s reliability appears minimal in this study. Additionally, although the ACTFL OPIc lacks the interactive nature of the ACTFL OPI, its consistency and reliability in rating effectively mitigate this limitation. Moreover, future advancements in artificial intelligence may enhance the interactivity of semi‐direct tests like the ACTFL OPIc, further bridging this gap.&lt;/p&gt; &lt;p&gt;However, this study significantly contributes to the limited body of published research on these tests, thereby enhancing the overall understanding and providing essential validity evidence. This localized research is invaluable as it addresses specific contextual factors that may influence test performance and outcomes, which are often overlooked in broader studies. Furthermore, it is imperative for test users to scrutinize the fairness of these tests for their local populations. This ensures that the tests are equitable and just, taking into account the diverse backgrounds and needs of the test‐takers. By doing so, researchers and practitioners can foster a more inclusive and accurate assessment environment, ultimately leading to more reliable and valid test results.&lt;/p&gt; &lt;p&gt;The findings are encouraging for proponents of the ACTFL OPIc, indicating that it produces ratings comparable to the ACTFL OPI. Given the importance of accurately determining participants&#39; ratings without bias, this study found no systematic or statistically significant differences caused by the factors studied. As the ACTFL OPI and OPIc are high‐stakes exams important for employment and education, these findings suggest that the ratings accurately represent participants&#39; performance, providing confidence in the fairness and reliability of both exams.&lt;/p&gt; &lt;hd id=&quot;AN0186745416-45&quot;&gt;CONCLUSIONS&lt;/hd&gt; &lt;p&gt;Avoiding testing bias in language exams is a complex but essential goal to ensure fair and accurate assessment of language proficiency. Previous research mentioned in this study has highlighted the potential influence of gender, age, acquisition setting, and self‐assessment biases on test outcomes, prompting the development of various strategies to mitigate these biases. By employing methods such as careful content selection, DIF analysis, adaptive testing, performance‐based assessments, and integrative approaches, test developers can create more equitable language assessments. Ongoing research and collaboration among stakeholders are crucial to continually improve the fairness and validity of language tests, ultimately benefiting all test‐takers.&lt;/p&gt; &lt;p&gt;In addition, as technology continues to advance, the line between direct and semi‐direct testing may blur further. Innovations in artificial intelligence and natural language processing hold the potential to create more interactive and adaptive semi‐direct tests that closely mimic direct testing&#39;s authenticity while retaining their efficiency and scalability. Future research and development in this area will be necessary in enhancing the effectiveness and fairness of proficiency testing.&lt;/p&gt; &lt;p&gt;GRAPH: Supporting Information&lt;/p&gt; &lt;ref id=&quot;AN0186745416-46&quot;&gt; &lt;title&gt; REFERENCES &lt;/title&gt; &lt;blist&gt; &lt;bibl id=&quot;bib1&quot; idref=&quot;ref72&quot; type=&quot;bt&quot;&gt;1&lt;/bibl&gt; &lt;bibtext&gt; ACTFL. (2018). ACTFL OPIc examinee handbook. Retrieved from https://&lt;ulink href=&quot;http://www.languagetesting.com/pub/media/wysiwyg/PDF/opic-examinee-handbook.pdf&quot;&gt;www.languagetesting.com/pub/media/wysiwyg/PDF/opic-examinee-handbook.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib2&quot; idref=&quot;ref18&quot; type=&quot;bt&quot;&gt;2&lt;/bibl&gt; &lt;bibtext&gt; ACTFL (2024a). ACTFL Oral Proficiency Interview‐Computer familiarization guide. Retrieved from https://&lt;ulink href=&quot;http://www.languagetesting.com/pub/media/wysiwyg/PDF/ACTFL-ACTFLOPIc-familiarization-guide.pdf&quot;&gt;www.languagetesting.com/pub/media/wysiwyg/PDF/ACTFL-ACTFLOPIc-familiarization-guide.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib3&quot; idref=&quot;ref1&quot; type=&quot;bt&quot;&gt;3&lt;/bibl&gt; &lt;bibtext&gt; ACTFL (2024b). ACTFL proficiency guidelines 2024. Retrieved from https://&lt;ulink href=&quot;http://www.actfl.org/uploads/files/general/Resources-Publications/ACTFL%5fProficiency%5fGuidelines%5f2024.pdf&quot;&gt;www.actfl.org/uploads/files/general/Resources-Publications/ACTFL%5fProficiency%5fGuidelines%5f2024.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib4&quot; idref=&quot;ref51&quot; type=&quot;bt&quot;&gt;4&lt;/bibl&gt; &lt;bibtext&gt; Arias, O., Canals, C., Mizala, A., &amp;amp; Meneses, F. (2023). Gender gaps in mathematics and language: The bias of competitive achievement tests. PLoS One, 18 (3), e0283384. https://doi.org/10.1371/journal.pone.0283384&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib5&quot; idref=&quot;ref59&quot; type=&quot;bt&quot;&gt;5&lt;/bibl&gt; &lt;bibtext&gt; Bachman, L., &amp;amp; Palmer, A. (2010). Language assessment in practice. Oxford University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib6&quot; idref=&quot;ref11&quot; type=&quot;bt&quot;&gt;6&lt;/bibl&gt; &lt;bibtext&gt; Baluyan, S. (2019). Taking account of test takers&#39; personal characteristics in language test development process, SHS Web of Conferences (70, p. 04002). EDP Sciences https://doi.org/10.1051/shsconf/20197004002&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib7&quot; idref=&quot;ref41&quot; type=&quot;bt&quot;&gt;7&lt;/bibl&gt; &lt;bibtext&gt; Bijani, H., &amp;amp; Khabiri, M. (2017). Direct and semi‐direct validation: Test takers&#39; perceptions, evaluations and anxiety towards speaking module of an English proficiency test. Journal of Language and Translation, 7 (1), 25 – 41.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib8&quot; idref=&quot;ref30&quot; type=&quot;bt&quot;&gt;8&lt;/bibl&gt; &lt;bibtext&gt; Brown, H. D., &amp;amp; Abeywickrama, P. (2019). Language assessment: Principles and classroom practices. Pearson.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibl id=&quot;bib9&quot; idref=&quot;ref62&quot; type=&quot;bt&quot;&gt;9&lt;/bibl&gt; &lt;bibtext&gt; Brown, N. A., Dewey, D. P., &amp;amp; Cox, T. L. (2014). Assessing the validity of Can‐Do statements in retrospective (Then‐Now) self‐assessment. Foreign Language Annals, 47 (2), 261 – 285. https://doi.org/10.1111/flan.12082&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Chapelle, C. A., &amp;amp; Douglas, D. (2006). Assessing language through computer technology. Cambridge University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Cheng, L., &amp;amp; Curtis, A. (2004). Washback or backwash: A review of the impact of testing on teaching and learning. In L. Cheng, Y. Watanabe, &amp;amp; A. Curtis (Eds.), Washback in language testing: Research contexts and methods (pp. 3 – 18). Lawrence Erlbaum Associates.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Chubbuck, K., Curley, W. E., &amp;amp; King, T. C. (2016). Who&#39;s on first? Gender differences in performance on the SAT &#174; test on critical reading items with sports and science content. ETS Research Report Series, 2016, 1 – 116. https://doi.org/10.1002/ets2.12109&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Craik, F. I. M., &amp;amp; Bialystok, E. (2006). Cognition through the lifespan: Mechanisms of change. Trends in Cognitive Sciences, 10 (3), 131 – 138.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Cubbellotti, S. (2015). Examination evaluation of the ACTFL OPIc&#174; in Arabic, English, and Spanish for the ACE Review. ACTFL. Retrieved from &lt;ulink href=&quot;http://www.languagetesting.com/pub/media/wysiwyg/research/reports/Examination%5fEvaluation%5fof%5fthe%5fACTFL%5fACTFLOPIc%5fin%5fArabic%5fEnglish%5fand%5fSpanish%5ffor%5fthe%5fACE%5fReview.pdf&quot;&gt;www.languagetesting.com/pub/media/wysiwyg/research/reports/Examination%5fEvaluation%5fof%5fthe%5fACTFL%5fACTFLOPIc%5fin%5fArabic%5fEnglish%5fand%5fSpanish%5ffor%5fthe%5fACE%5fReview.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Eckes, T. (2012). Operational rater types in writing assessment: Linking rater cognition to rater behavior. Language Assessment Quarterly, 9 (3), 270 – 292.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Educational Testing Service (ETS) (2022). ETS guidelines for developing fair tests and communications. Retrieved from https://&lt;ulink href=&quot;http://www.ets.org/pdfs/about/fair-tests-and-communications.pdf&quot;&gt;www.ets.org/pdfs/about/fair-tests-and-communications.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Elder, C. (2012). Bias in language assessment. In C. A. Chapelle (Ed.), The Encyclopedia of applied linguistics (pp. 1 – 7). Blackwell https://doi.org/10.1002/9781405198431.wbeal1198&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Espinosa, M. P., &amp;amp; Gardeazabal, J. (2020). The gender‐bias effect of test scoring and framing: A concern for personnel selection and college admission. The B.E. Journal of Economic Analysis &amp;amp; Policy, 20 (4), 1 – 23. https://doi.org/10.1515/bejeap-2019-0316&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Fulcher, G. (2003). Testing second language speaking. Longman.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Fulcher, G. (2010). Practical language testing. Routledge.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Hughes, A. (2003). Testing for language teachers (2nd ed.). Cambridge University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; IELTS. (n.d.). Welcome to IELTS. https://ielts.org/&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Isbell, D., &amp;amp; Winke, P. (2019). ACTFL Oral Proficiency Interview – Computer (OPIc). Language Testing, 36 (3), 467 – 477. https://doi.org/10.1177/0265532219828253&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Isbell, D. R., &amp;amp; Kremmel, B. (2020). Test review: Current options in at‐home language proficiency tests for making high‐stakes decisions. Language Testing, 37 (4), 600 – 619.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Isbell, D. R., Winke, P., &amp;amp; Gass, S. M. (2019). Using the ACTFL OPIc to assess proficiency and monitor progress in a tertiary foreign languages program. Language Testing, 36 (3), 439 – 465. https://doi.org/10.1177/0265532218798139&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Joo, M. J. (2008). Investigation of rater, gender, and age biases to test formats: FTFI and COT. 영어학, 8 (1), 1 – 20.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Karami, H. (2013). An investigation of the gender differential performance on a high‐stakes language proficiency test in Iran. Asia Pacific Education Review, 14, 435 – 444. https://doi.org/10.1007/s12564-013-9272-y&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Karami, M., Pishghadam, R., &amp;amp; Baghaei, P. (2019). A probe into EFL learners&#39; emotioncy as a source of test bias: Insights from differential item functioning analysis. Studies in Educational Evaluation, 60, 170 – 178.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Kenyon, D. M., &amp;amp; Tschirner, E. (2000). The rating of direct and semi‐direct oral proficiency interviews: Comparing performance at lower proficiency levels. The Modern Language Journal, 84 (1), 85 – 101.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Kiddle, T., &amp;amp; Kormos, J. (2011). The effect of mode of response on a semidirect test of oral proficiency. Language Assessment Quarterly, 8 (4), 342 – 360.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Kunnan, A. J. (2014). Fairness and justice in language assessment. In A. J. Kunnan (Ed.), The companion to language assessment (pp. 1098 – 1114). Wiley‐Blackwell.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Li, M., &amp;amp; Zhang, X. (2021). A meta‐analysis of self‐assessment and language performance in language testing and assessment. Language Testing, 38 (2), 189 – 218. https://doi.org/10.1177/0265532220932481&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Linacre, J. M. (1994). Many‐facet Rasch measurement. MESA Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Malone, M. E., &amp;amp; Montee, M. J. (2010). Oral proficiency assessment: Current approaches and applications for post‐secondary foreign language programs. Language and linguistics compass, 4 (10), 972 – 986.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Mansouri, H. (1999). Aligned rank transform tests in linear models. Journal of Statistical Planning and Inference, 79 (1), 141 – 155. https://doi.org/10.1016/S0378-3758(98)00229-8&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Masoumi, G. A., &amp;amp; Sadeghi, K. (2020). Impact of test format on vocabulary test performance of EFL learners: The role of gender. Language Testing in Asia, 10 (1), 1 – 13. https://doi.org/10.1186/s40468-020-00099-x&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; McNamara, T. F. (2000). Language testing. Oxford University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13 – 103). Macmillan.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Nakatsuhara, F. (2008). Measuring spoken fluency in semi‐direct and direct oral tasks: A question of reliability? Language Testing, 25 (1), 55 – 79.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; O&#39;Loughlin, K. J. (2001). The equivalence of direct and semi‐direct speaking tests. Cambridge University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; O&#39;Loughlin, K. (2002). The impact of gender in oral proficiency testing. Language Testing, 19 (2), 169 – 192. https://doi.org/10.1191/0265532202lt226oa&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Ockey, G. J. (2009). The effects of group members&#39; personalities on a test taker&#39;s L2 group oral discussion test scores. Language Testing, 26 (2), 161 – 186.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Ozdemir, B., &amp;amp; Alshamrani, A. H. (2020). Examining the fairness of language test across gender with IRT‐based differential item and test functioning methods. International Journal of Learning, Teaching and Educational Research, 19 (6), 27 – 45. https://doi.org/10.26803/ijlter.19.6.2&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Quaid, E. D., &amp;amp; Barrett, A. (2020). Toward the future of computer‐assisted language testing: Assessing spoken performance through semi‐direct tests, Recent developments in technology‐enhanced and computer‐assisted language learning (pp. 208 – 235). IGI Global.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Pusey, K., &amp;amp; Butler, Y. G. (2023). Investigating the ecological validity of second language writing assessment tasks. System, 119, 103174.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Ross, S. (1998). Self‐assessment in second language testing: A meta‐analysis and analysis of experiential factors. Language Testing, 15 (1), 1 – 20.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Ross, S. J., &amp;amp; Okabe, J. (2006). The subjective and objective interface of bias detection on language tests. International Journal of Testing, 6 (3), 229 – 253. https://doi.org/10.1207/s15327574ijt0603_2&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Schaap, P. (2011). The differential item functioning and structural equivalence of a nonverbal cognitive ability test for five language groups. SA Journal of Industrial Psychology, 37 (1), 1 – 16.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Shohamy, E. (1994). The validity of direct versus semi‐direct oral tests. Language Testing, 11 (2), 99 – 123. https://doi.org/10.1177/026553229401100202&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Smith, R. M. (2000). Fit analysis in latent trait measurement models. Journal of Applied Measurement, 1 (2), 199 – 218.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Stansfield, C. W. (1991). A comparative analysis of simulated and direct oral proficiency interviews. In S. Anivan (Ed.), Current developments in language testing (pp. 199 – 209). Regional Language Centre.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Surface, E., Poncheri, R., &amp;amp; Bhavsar, K. (2008). Two studies investigating the reliability and validity of the English ACTFL OPIc with Korean test takers: The ACTFL OPIc validation project technical report. Retrieved May 11, 2021, from https://aappl.actfl.org/sites/default/files/assessments/acereports/ACTFL-ACTFLOPIc-English-Validation-2008.pdf&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; SWA Consulting. (2009). Brief reliability report 5: Test‐retest reliability and absolute agreement rates of English ACTFL OPIc&#174; proficiency ratings for double and single rated tests within a sample of Korean test takers. SWA Consulting. Retrieved from &lt;ulink href=&quot;http://www.languagetesting.com/pub/media/wysiwyg/research/ACTFL-ACTFLOPIc-Retest-Reliability-Study-2009.pdf&quot;&gt;www.languagetesting.com/pub/media/wysiwyg/research/ACTFL-ACTFLOPIc-Retest-Reliability-Study-2009.pdf&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Thompson, G. L., Cox, T. L., &amp;amp; Knapp, N. (2016). Comparing the OPI and the OPIc: The effect of test method on oral proficiency scores and student preference. Foreign Language Annals, 49 (1), 75 – 92.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; TOEFL. (n.d.). TOEFL. https://&lt;ulink href=&quot;http://www.ets.org/toefl.html&quot;&gt;www.ets.org/toefl.html&lt;/ulink&gt;&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Toulmin, S. E. (1958). The Uses of Argument. Cambridge University Press.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Weir, C. J. (2005). Language testing and validation: An evidence‐based approach. Palgrave Macmillan.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Winke, P., Zhang, X., &amp;amp; Pierce, S. J. (2023). A closer look at a marginalized test method: self‐assessment as a measure of speaking proficiency. Studies in Second Language Acquisition, 45 (2), 416 – 441.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Xi, X. (2010). Automated scoring and feedback systems: Where are we and where are we heading? Language Testing, 27 (3), 291 – 300.&lt;/bibtext&gt; &lt;/blist&gt; &lt;blist&gt; &lt;bibtext&gt; Yao, D. (2023). Examining the subjective fairness of at‐home and online tests: Taking duolingo English test as an example. PLoS One, 18 (9), e0291629. https://doi.org/10.1371/journal.pone.0291629&lt;/bibtext&gt; &lt;/blist&gt; &lt;/ref&gt; &lt;aug&gt; &lt;p&gt;By Troy L. Cox; Gregory L. Thompson and Steven S. Stokes&lt;/p&gt; &lt;p&gt;Reported by Author; Author; Author&lt;/p&gt; &lt;/aug&gt; &lt;nolink nlid=&quot;nl1&quot; bibid=&quot;bib34&quot; firstref=&quot;ref2&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl2&quot; bibid=&quot;bib23&quot; firstref=&quot;ref3&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl3&quot; bibid=&quot;bib54&quot; firstref=&quot;ref5&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl4&quot; bibid=&quot;bib38&quot; firstref=&quot;ref6&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl5&quot; bibid=&quot;bib14&quot; firstref=&quot;ref7&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl6&quot; bibid=&quot;bib52&quot; firstref=&quot;ref8&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl7&quot; bibid=&quot;bib53&quot; firstref=&quot;ref9&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl8&quot; bibid=&quot;bib43&quot; firstref=&quot;ref12&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl9&quot; bibid=&quot;bib26&quot; firstref=&quot;ref13&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl10&quot; bibid=&quot;bib47&quot; firstref=&quot;ref14&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl11&quot; bibid=&quot;bib28&quot; firstref=&quot;ref15&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl12&quot; bibid=&quot;bib27&quot; firstref=&quot;ref16&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl13&quot; bibid=&quot;bib60&quot; firstref=&quot;ref17&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl14&quot; bibid=&quot;bib21&quot; firstref=&quot;ref20&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl15&quot; bibid=&quot;bib51&quot; firstref=&quot;ref21&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl16&quot; bibid=&quot;bib44&quot; firstref=&quot;ref22&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl17&quot; bibid=&quot;bib39&quot; firstref=&quot;ref24&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl18&quot; bibid=&quot;bib40&quot; firstref=&quot;ref25&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl19&quot; bibid=&quot;bib49&quot; firstref=&quot;ref26&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl20&quot; bibid=&quot;bib20&quot; firstref=&quot;ref27&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl21&quot; bibid=&quot;bib45&quot; firstref=&quot;ref28&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl22&quot; bibid=&quot;bib11&quot; firstref=&quot;ref29&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl23&quot; bibid=&quot;bib37&quot; firstref=&quot;ref32&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl24&quot; bibid=&quot;bib57&quot; firstref=&quot;ref33&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl25&quot; bibid=&quot;bib22&quot; firstref=&quot;ref34&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl26&quot; bibid=&quot;bib55&quot; firstref=&quot;ref35&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl27&quot; bibid=&quot;bib42&quot; firstref=&quot;ref36&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl28&quot; bibid=&quot;bib59&quot; firstref=&quot;ref37&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl29&quot; bibid=&quot;bib10&quot; firstref=&quot;ref38&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl30&quot; bibid=&quot;bib19&quot; firstref=&quot;ref40&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl31&quot; bibid=&quot;bib24&quot; firstref=&quot;ref42&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl32&quot; bibid=&quot;bib29&quot; firstref=&quot;ref43&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl33&quot; bibid=&quot;bib30&quot; firstref=&quot;ref44&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl34&quot; bibid=&quot;bib41&quot; firstref=&quot;ref47&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl35&quot; bibid=&quot;bib18&quot; firstref=&quot;ref48&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl36&quot; bibid=&quot;bib36&quot; firstref=&quot;ref49&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl37&quot; bibid=&quot;bib12&quot; firstref=&quot;ref50&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl38&quot; bibid=&quot;bib31&quot; firstref=&quot;ref52&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl39&quot; bibid=&quot;bib16&quot; firstref=&quot;ref54&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl40&quot; bibid=&quot;bib13&quot; firstref=&quot;ref55&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl41&quot; bibid=&quot;bib25&quot; firstref=&quot;ref57&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl42&quot; bibid=&quot;bib46&quot; firstref=&quot;ref60&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl43&quot; bibid=&quot;bib32&quot; firstref=&quot;ref61&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl44&quot; bibid=&quot;bib58&quot; firstref=&quot;ref63&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl45&quot; bibid=&quot;bib48&quot; firstref=&quot;ref66&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl46&quot; bibid=&quot;bib17&quot; firstref=&quot;ref67&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl47&quot; bibid=&quot;bib50&quot; firstref=&quot;ref68&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl48&quot; bibid=&quot;bib33&quot; firstref=&quot;ref69&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl49&quot; bibid=&quot;bib15&quot; firstref=&quot;ref70&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl50&quot; bibid=&quot;bib35&quot; firstref=&quot;ref71&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl51&quot; bibid=&quot;bib84&quot; firstref=&quot;ref74&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl52&quot; bibid=&quot;bib149&quot; firstref=&quot;ref77&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl53&quot; bibid=&quot;bib150&quot; firstref=&quot;ref80&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl54&quot; bibid=&quot;bib142&quot; firstref=&quot;ref88&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl55&quot; bibid=&quot;bib148&quot; firstref=&quot;ref90&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl56&quot; bibid=&quot;bib147&quot; firstref=&quot;ref92&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl57&quot; bibid=&quot;bib144&quot; firstref=&quot;ref96&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl58&quot; bibid=&quot;bib145&quot; firstref=&quot;ref98&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl59&quot; bibid=&quot;bib151&quot; firstref=&quot;ref106&quot;&gt;&lt;/nolink&gt; &lt;nolink nlid=&quot;nl60&quot; bibid=&quot;bib56&quot; firstref=&quot;ref127&quot;&gt;&lt;/nolink&gt;
Header DbId: eric
DbLabel: ERIC
An: EJ1477432
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Can the Oral Proficiency Interview -- Computer (ACTFL OPIc) Be Used Instead of the Oral Proficiency Interview (ACTFL OPI)? An Aligned Rank Transform (ART) Analysis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: &lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Troy+L%2E+Cox%22&quot;&gt;Troy L. Cox&lt;/searchLink&gt; (ORCID &lt;externalLink term=&quot;http://orcid.org/0000-0001-9379-5102&quot;&gt;0000-0001-9379-5102&lt;/externalLink&gt;)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Gregory+L%2E+Thompson%22&quot;&gt;Gregory L. Thompson&lt;/searchLink&gt; (ORCID &lt;externalLink term=&quot;http://orcid.org/0000-0003-4709-4836&quot;&gt;0000-0003-4709-4836&lt;/externalLink&gt;)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Steven+S%2E+Stokes%22&quot;&gt;Steven S. Stokes&lt;/searchLink&gt;
– Name: TitleSource
  Label: Source
  Group: Src
  Data: &lt;searchLink fieldCode=&quot;SO&quot; term=&quot;%22Foreign+Language+Annals%22&quot;&gt;&lt;i&gt;Foreign Language Annals&lt;/i&gt;&lt;/searchLink&gt;. 2025 58(2):300-325.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley &amp; Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 26
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles&lt;br /&gt;Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: &lt;searchLink fieldCode=&quot;EL&quot; term=&quot;%22Higher+Education%22&quot;&gt;Higher Education&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;EL&quot; term=&quot;%22Postsecondary+Education%22&quot;&gt;Postsecondary Education&lt;/searchLink&gt;
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Oral+Language%22&quot;&gt;Oral Language&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+Proficiency%22&quot;&gt;Language Proficiency&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Interviews%22&quot;&gt;Interviews&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Computer+Uses+in+Education%22&quot;&gt;Computer Uses in Education&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22College+Students%22&quot;&gt;College Students&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Second+Language+Learning%22&quot;&gt;Second Language Learning&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Spanish%22&quot;&gt;Spanish&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Gender+Differences%22&quot;&gt;Gender Differences&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Age+Differences%22&quot;&gt;Age Differences&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Scores%22&quot;&gt;Scores&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Testing%22&quot;&gt;Testing&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Test+Reliability%22&quot;&gt;Test Reliability&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+Tests%22&quot;&gt;Language Tests&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Test+Bias%22&quot;&gt;Test Bias&lt;/searchLink&gt;
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;SU&quot; term=&quot;%22ACTFL+Oral+Proficiency+Interview%22&quot;&gt;ACTFL Oral Proficiency Interview&lt;/searchLink&gt;
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/flan.12804
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0015-718X&lt;br /&gt;1944-9720
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study investigated the differences between the ACTFL Oral Proficiency Interview (OPI) and the ACTFL Oral Proficiency Interview - Computer (OPIc) among Spanish learners at a U.S. university. Participants (N = 154) were randomly assigned to take both tests in a counterbalanced order to mitigate test order effects. Data were analyzed using an aligned rank transform (ART) analysis, focusing on variables such as gender, age, language courses, missionary experience, and self-assessed Spanish ability. Results showed a strong correlation between ACTFL OPI and ACTFL OPIc ratings ([tau] = 0.79, p &lt; 0.001), though ACTFL OPIc ratings were slightly higher on average. No significant order effects were found, indicating the order of test administration did not influence ratings. The reliability of both tests was confirmed, and no significant biases were detected. The findings suggest that both ACTFL OPI and ACTFL OPIc are effective, holistic measures of Spanish oral proficiency, with ACTFL OPIc offering a slight advantage in rating outcomes. Pedagogically, this suggests flexibility in test choice without compromising assessment integrity.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1477432
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1477432
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/flan.12804
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 26
        StartPage: 300
    Subjects:
      – SubjectFull: Oral Language
        Type: general
      – SubjectFull: Language Proficiency
        Type: general
      – SubjectFull: Interviews
        Type: general
      – SubjectFull: Computer Uses in Education
        Type: general
      – SubjectFull: College Students
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Spanish
        Type: general
      – SubjectFull: Gender Differences
        Type: general
      – SubjectFull: Age Differences
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Testing
        Type: general
      – SubjectFull: Test Reliability
        Type: general
      – SubjectFull: Language Tests
        Type: general
      – SubjectFull: Test Bias
        Type: general
      – SubjectFull: ACTFL Oral Proficiency Interview
        Type: general
    Titles:
      – TitleFull: Can the Oral Proficiency Interview -- Computer (ACTFL OPIc) Be Used Instead of the Oral Proficiency Interview (ACTFL OPI)? An Aligned Rank Transform (ART) Analysis
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Troy L. Cox
      – PersonEntity:
          Name:
            NameFull: Gregory L. Thompson
      – PersonEntity:
          Name:
            NameFull: Steven S. Stokes
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0015-718X
            – Type: issn-electronic
              Value: 1944-9720
          Numbering:
            – Type: volume
              Value: 58
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Foreign Language Annals
              Type: main
ResultId 1