The Feasibility of Computerized Adaptive Testing of the National Benchmark Test: A Simulation Study

Saved in:
Bibliographic Details
Title: The Feasibility of Computerized Adaptive Testing of the National Benchmark Test: A Simulation Study
Language: English
Authors: Musa Adekunle Ayanwale, Mdutshekelwa Ndlovu
Source: Journal of Pedagogical Research. 2024 8(2):95-112.
Availability: Journal of Pedagogical Research. Duzce University, Faculty of Education, Konuralp Campus, 81620, Duzce, Turkey. e-mail: ijopr.editor@gmail.com; Web site: https://www.ijopr.com/
Peer Reviewed: Y
Page Count: 18
Publication Date: 2024
Document Type: Journal Articles
Reports - Research
Descriptors: Adaptive Testing, Benchmarking, National Competency Tests, Computer Assisted Testing, High Stakes Tests, Foreign Countries, Test Validity, Simulation, Test Items, Computer Software, Monte Carlo Methods
Geographic Terms: South Africa
ISSN: 2602-3717
Abstract: The COVID-19 pandemic has had a significant impact on high-stakes testing, including the national benchmark tests in South Africa. Current linear testing formats have been criticized for their limitations, leading to a shift towards Computerized Adaptive Testing [CAT]. Assessments with CAT are more precise and take less time. Evaluation of CAT programs requires simulation studies. To assess the feasibility of implementing CAT in NBTs, SimulCAT, a simulation tool, was utilized. The SimulCAT simulation involved creating 10,000 examinees with a normal distribution characterized by a mean of 0 and a standard deviation of 1. A pool of 500 test items was employed, and specific parameters were established for the item selection algorithm, CAT administration rules, item exposure control, and termination criteria. The termination criteria required a standard error of less than 0.35 to ensure accurate abilities estimation. The findings from the simulation study demonstrated that fixed-length tests provided higher testing precision without any systematic error, as indicated by measurement statistics like CBIAS, CMAE, and CRMSE. However, fixed-length tests exhibited a higher item exposure rate, which could be mitigated by selecting items with fewer dependencies on specific item parameters (a-parameters). On the other hand, variable-length tests demonstrated increased redundancy. Based on these results, CAT is recommended as an alternative approach for conducting NBTs due to its capability to accurately measure individual abilities and reduce the testing duration. For high-stakes assessments like the NBTs, fixed-length tests are preferred as they offer superior testing precision while minimizing item exposure rates.
Abstractor: As Provided
Entry Date: 2024
Accession Number: EJ1428037
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1428037
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1428037
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: The Feasibility of Computerized Adaptive Testing of the National Benchmark Test: A Simulation Study
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Musa+Adekunle+Ayanwale%22">Musa Adekunle Ayanwale</searchLink><br /><searchLink fieldCode="AR" term="%22Mdutshekelwa+Ndlovu%22">Mdutshekelwa Ndlovu</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Pedagogical+Research%22"><i>Journal of Pedagogical Research</i></searchLink>. 2024 8(2):95-112.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Journal of Pedagogical Research. Duzce University, Faculty of Education, Konuralp Campus, 81620, Duzce, Turkey. e-mail: ijopr.editor@gmail.com; Web site: https://www.ijopr.com/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 18
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Adaptive+Testing%22">Adaptive Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Benchmarking%22">Benchmarking</searchLink><br /><searchLink fieldCode="DE" term="%22National+Competency+Tests%22">National Competency Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22High+Stakes+Tests%22">High Stakes Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Validity%22">Test Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Simulation%22">Simulation</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Software%22">Computer Software</searchLink><br /><searchLink fieldCode="DE" term="%22Monte+Carlo+Methods%22">Monte Carlo Methods</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22South+Africa%22">South Africa</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2602-3717
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The COVID-19 pandemic has had a significant impact on high-stakes testing, including the national benchmark tests in South Africa. Current linear testing formats have been criticized for their limitations, leading to a shift towards Computerized Adaptive Testing [CAT]. Assessments with CAT are more precise and take less time. Evaluation of CAT programs requires simulation studies. To assess the feasibility of implementing CAT in NBTs, SimulCAT, a simulation tool, was utilized. The SimulCAT simulation involved creating 10,000 examinees with a normal distribution characterized by a mean of 0 and a standard deviation of 1. A pool of 500 test items was employed, and specific parameters were established for the item selection algorithm, CAT administration rules, item exposure control, and termination criteria. The termination criteria required a standard error of less than 0.35 to ensure accurate abilities estimation. The findings from the simulation study demonstrated that fixed-length tests provided higher testing precision without any systematic error, as indicated by measurement statistics like CBIAS, CMAE, and CRMSE. However, fixed-length tests exhibited a higher item exposure rate, which could be mitigated by selecting items with fewer dependencies on specific item parameters (a-parameters). On the other hand, variable-length tests demonstrated increased redundancy. Based on these results, CAT is recommended as an alternative approach for conducting NBTs due to its capability to accurately measure individual abilities and reduce the testing duration. For high-stakes assessments like the NBTs, fixed-length tests are preferred as they offer superior testing precision while minimizing item exposure rates.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1428037
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1428037
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 18
        StartPage: 95
    Subjects:
      – SubjectFull: Adaptive Testing
        Type: general
      – SubjectFull: Benchmarking
        Type: general
      – SubjectFull: National Competency Tests
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: High Stakes Tests
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Test Validity
        Type: general
      – SubjectFull: Simulation
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Computer Software
        Type: general
      – SubjectFull: Monte Carlo Methods
        Type: general
      – SubjectFull: South Africa
        Type: general
    Titles:
      – TitleFull: The Feasibility of Computerized Adaptive Testing of the National Benchmark Test: A Simulation Study
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Musa Adekunle Ayanwale
      – PersonEntity:
          Name:
            NameFull: Mdutshekelwa Ndlovu
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-electronic
              Value: 2602-3717
          Numbering:
            – Type: volume
              Value: 8
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Journal of Pedagogical Research
              Type: main
ResultId 1